Docs/Sources/Web Link Grounding Data for AI Agents

Web Link Grounding Data for AI Agents

Scrape web page links

OperationalCredits 10 per callp50 1734msData Scraping

Overview

Link Scraper works by parsing the HTML content of a web page to extract all the links. It returns the links in a structured format, including the link URL, title, and more.

Live Test Web Link Grounding Data for AI Agents Source →

The tool

Once your client is connected to the VerveContext server, this appears in its tool list as WebLinkGroundingDataforAIAgents. It is read-only and open-world — it fetches and never mutates anything on your side — so most clients call it without asking you to confirm.

Tool call
{
  "name": "WebLinkGroundingDataforAIAgents",
  "arguments": {
    "url": "https://en.wikipedia.org/wiki/Web_scraping"
  }
}

You do not name the tool yourself; the model picks it. Asking about https://en.wikipedia.org/wiki/Web_scraping in the terms this source covers is enough for it to reach for WebLinkGroundingDataforAIAgents on its own — naming it explicitly also works, and is the way to force the call.

Connecting

One server URL covers every source in the catalog, including this one. Authorization is OAuth: the client opens a browser once, and there is no key to paste into a config file.

{
  "mcpServers": {
    "vervecontext": {
      "url": "https://api.vervecontext.com/v1/mcp"
    }
  }
}

Per-client setup — Claude, Cursor, VS Code, ChatGPT — is on the MCP setup page.

Arguments

These are the properties on the tool's inputSchema, so a well-behaved client validates them before the call is made. Premium arguments are accepted on every plan but only take effect on plans that include them.

ArgumentTypeDescription
urlRequiredstringThe URL of the web page to scrape links from
url
maxlinksOptionalPremiumnumberMaximum number of links to scrape and return
default 50
includequeryOptionalbooleanInclude query strings in the scraped links

What the model gets back

The result carries a structuredContent object matching the tool's declared outputSchema, so a client reads fields without parsing prose. status is "ok" and error is null on success; a null field means the value was not available for that input, not that the call failed.

Result
{
  "status": "ok",
  "error": null,
  "data": {
    "url": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
    "linkCount": 16,
    "externalLinkCount": 13,
    "internalLinkCount": 3,
    "links": [
      {
        "text": "Documentation",
        "href": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html",
        "external": false
      },
      {
        "text": "Amazon EC2 Instance Types Guide",
        "href": "https://docs.aws.amazon.com/ec2/latest/instancetypes/instance-types.html",
        "external": true
      },
      {
        "text": "Amazon EC2 Auto Scaling",
        "href": "https://docs.aws.amazon.com/autoscaling/",
        "external": true
      }
    ],
    "uniqueDomains": [
      "docs.aws.amazon.com",
      "aws.amazon.com"
    ],
    "maxLinksReached": false
  }
}

Response fields

Paths are relative to data. Premium fields are absent rather than zeroed on plans that do not include them, so check for presence instead of comparing to 0.

FieldTypeExampleDescription
urlstringhttp://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.htmlThe scraped page's URL, normalized to a bare http:// origin with any tabs, newlines and carriage returns stripped
linkCountnumber16Total number of links returned, after filtering out empty anchors and page-fragment links
externalLinkCountnumber13Number of returned links that point to a different domain than the scraped page
internalLinkCountnumber3Number of returned links that point to the same domain as the scraped page
linksarray[3]Every link found on the page, up to the requested limit
links.0.textstringDocumentationThe link's visible anchor text, with tabs, newlines and carriage returns removed
links.0.hrefstringhttp://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.htmlThe link's destination URL, resolved to an absolute URL when the source was root-relative
links.0.externalbooleanfalseWhether the link points to a different domain than the scraped page
uniqueDomainsPremiumarray["docs.aws.amazon.com","aws.amazon.com"]List of unique external domains found
maxLinksReachedbooleanfalseWhether the number of links found met or exceeded the maximum link limit, meaning more links may exist beyond what was returned

Why ground on it

A model can produce something that looks like this answer from its training data, and be confidently out of date or simply wrong. This source returns the current value in a shape you can check, which is the difference between an answer you can cite and one you have to hedge.

Point an evaluation at linkCount: it is the field most worth pinning a claim to, and it is either present and current or absent — never plausibly invented.

Failure modes

Errors come back as tool errors carrying a sentence the model can act on, not a bare status code. Error handling covers the full list.

StatusWhat it means
400 / 422The arguments did not validate. The message names the offending one.
401The OAuth session is invalid or expired — reconnect the server.
403Blocked by a key restriction or an IP allow-list. Never a bad identity.
404This source is not part of VerveContext. Check the catalog.
429Out of credits, or a brief rate limit. The message tells them apart.

A call costs 10 credits each time the tool actually runs; a model that reasons about the tool without calling it costs nothing.

Use cases

Site Architecture Audits
Crawlers inspect internal navigation paths and anchor text across catalog pages to detect orphan URLs and track link equity distribution.
Dead Link Discovery
Run published articles through the scraper to catch external links, resolve relative URLs, and forward destination targets to status checkers.
Competitor Footprint Mapping
Growth analysts feed rival blog posts into the endpoint to inventory citation targets and detect outbound partner referrals.
Web Archival Pipelines
When archiving resource directories, extract all outbound destinations and clean anchor text before caching document references.

Other ways to use Web Link Grounding Data for AI Agents

Set up Web Link Grounding Data for AI Agents on VerveContext, or reach the same source a different way. Your VerveContext account and credits work on all of them — one key, one balance.

Call it as a REST APIOne HTTPS endpoint and an x-api-key header, with SDKs for Node, Python and .NET.APIVerve →Reference →
Give it to an AI agentConnect over MCP and your agent calls it as a native tool — Claude, Cursor, ChatGPT.VerveKit →Reference →
Google Sheets or ExcelA =VERVE() formula fills a column — no script, no export, recalculates in place.VerveSheets →Reference →

More in Data Scraping:

Was this page helpful?