Overview
Link Scraper works by parsing the HTML content of a web page to extract all the links. It returns the links in a structured format, including the link URL, title, and more.
Live Test Web Link Grounding Data for AI Agents Source →
The tool
Once your client is connected to the VerveContext server, this appears in its tool list as WebLinkGroundingDataforAIAgents. It is read-only and open-world — it fetches and never mutates anything on your side — so most clients call it without asking you to confirm.
{
"name": "WebLinkGroundingDataforAIAgents",
"arguments": {
"url": "https://en.wikipedia.org/wiki/Web_scraping"
}
}You do not name the tool yourself; the model picks it. Asking about https://en.wikipedia.org/wiki/Web_scraping in the terms this source covers is enough for it to reach for WebLinkGroundingDataforAIAgents on its own — naming it explicitly also works, and is the way to force the call.
Connecting
One server URL covers every source in the catalog, including this one. Authorization is OAuth: the client opens a browser once, and there is no key to paste into a config file.
{
"mcpServers": {
"vervecontext": {
"url": "https://api.vervecontext.com/v1/mcp"
}
}
}https://api.vervecontext.com/v1/mcpPer-client setup — Claude, Cursor, VS Code, ChatGPT — is on the MCP setup page.
Arguments
These are the properties on the tool's inputSchema, so a well-behaved client validates them before the call is made. Premium arguments are accepted on every plan but only take effect on plans that include them.
| Argument | Type | Description |
|---|---|---|
urlRequired | string | The URL of the web page to scrape links from url |
maxlinksOptionalPremium | number | Maximum number of links to scrape and return default 50 |
includequeryOptional | boolean | Include query strings in the scraped links |
What the model gets back
The result carries a structuredContent object matching the tool's declared outputSchema, so a client reads fields without parsing prose. status is "ok" and error is null on success; a null field means the value was not available for that input, not that the call failed.
{
"status": "ok",
"error": null,
"data": {
"url": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"linkCount": 16,
"externalLinkCount": 13,
"internalLinkCount": 3,
"links": [
{
"text": "Documentation",
"href": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html",
"external": false
},
{
"text": "Amazon EC2 Instance Types Guide",
"href": "https://docs.aws.amazon.com/ec2/latest/instancetypes/instance-types.html",
"external": true
},
{
"text": "Amazon EC2 Auto Scaling",
"href": "https://docs.aws.amazon.com/autoscaling/",
"external": true
}
],
"uniqueDomains": [
"docs.aws.amazon.com",
"aws.amazon.com"
],
"maxLinksReached": false
}
}
Response fields
Paths are relative to data. Premium fields are absent rather than zeroed on plans that do not include them, so check for presence instead of comparing to 0.
| Field | Type | Example | Description |
|---|---|---|---|
url | string | http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html | The scraped page's URL, normalized to a bare http:// origin with any tabs, newlines and carriage returns stripped |
linkCount | number | 16 | Total number of links returned, after filtering out empty anchors and page-fragment links |
externalLinkCount | number | 13 | Number of returned links that point to a different domain than the scraped page |
internalLinkCount | number | 3 | Number of returned links that point to the same domain as the scraped page |
links | array[3] | Every link found on the page, up to the requested limit | |
links.0.text | string | Documentation | The link's visible anchor text, with tabs, newlines and carriage returns removed |
links.0.href | string | http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html | The link's destination URL, resolved to an absolute URL when the source was root-relative |
links.0.external | boolean | false | Whether the link points to a different domain than the scraped page |
uniqueDomainsPremium | array | ["docs.aws.amazon.com","aws.amazon.com"] | List of unique external domains found |
maxLinksReached | boolean | false | Whether the number of links found met or exceeded the maximum link limit, meaning more links may exist beyond what was returned |
Why ground on it
A model can produce something that looks like this answer from its training data, and be confidently out of date or simply wrong. This source returns the current value in a shape you can check, which is the difference between an answer you can cite and one you have to hedge.
Point an evaluation at linkCount: it is the field most worth pinning a claim to, and it is either present and current or absent — never plausibly invented.
Failure modes
Errors come back as tool errors carrying a sentence the model can act on, not a bare status code. Error handling covers the full list.
| Status | What it means |
|---|---|
400 / 422 | The arguments did not validate. The message names the offending one. |
401 | The OAuth session is invalid or expired — reconnect the server. |
403 | Blocked by a key restriction or an IP allow-list. Never a bad identity. |
404 | This source is not part of VerveContext. Check the catalog. |
429 | Out of credits, or a brief rate limit. The message tells them apart. |
A call costs 10 credits each time the tool actually runs; a model that reasons about the tool without calling it costs nothing.
Use cases
- Site Architecture Audits
- Crawlers inspect internal navigation paths and anchor text across catalog pages to detect orphan URLs and track link equity distribution.
- Dead Link Discovery
- Run published articles through the scraper to catch external links, resolve relative URLs, and forward destination targets to status checkers.
- Competitor Footprint Mapping
- Growth analysts feed rival blog posts into the endpoint to inventory citation targets and detect outbound partner referrals.
- Web Archival Pipelines
- When archiving resource directories, extract all outbound destinations and clean anchor text before caching document references.
Other ways to use Web Link Grounding Data for AI Agents
Set up Web Link Grounding Data for AI Agents on VerveContext, or reach the same source a different way. Your VerveContext account and credits work on all of them — one key, one balance.
Related
More in Data Scraping: