Docs/APIs/Link Scraper

Link Scraper

Scrape web page links

OperationalCredits 25 per callp50 1734msData ScrapingStar

Overview

Link Scraper works by parsing the HTML content of a web page to extract all the links. It returns the links in a structured format, including the link URL, title, and more.

Live Test Link Scraper API →

Endpoint

One host, one path per API. The block below shows this call in four languages; every one of them is the same HTTP request. Making requests covers the timeouts, retries and parameter rules that apply to all of them. The SDKs wrap the same call in a typed client.

POSThttps://api.apiverve.com/v1/linkscraper
curl -X POST https://api.apiverve.com/v1/linkscraper \
  -H "x-api-key: your_api_key_here" \
  -H "Content-Type: application/json" \
  -d '{
  "url": "https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
  "maxlinks": 20,
  "includequery": false
}'

Replace your_api_key_here with the key from your dashboard. When the inputs arrive as a list rather than one at a time, batch requests run up to 200 of them through this same API in a single call.

Authentication

Send your key in the x-api-key header. That is the only auth step — there is no token exchange and no per-endpoint scope to configure. Authentication covers creating, rotating and revoking keys.

401 is the only auth verdict

A 401 means the key is missing, invalid or expired. A 403 means the key is valid but not permitted here — blocked by a key restriction or an IP allow-list. Running out of credits is a 429.

Parameters

Sent as JSON in the request body. Premium parameters are accepted on every plan but only take effect on plans that include them.

ParameterTypeDescription
urlRequiredstringThe URL of the web page to scrape links from
url
maxlinksOptionalPremiumnumberMaximum number of links to scrape and return
default 50
includequeryOptionalbooleanInclude query strings in the scraped links

Response

Every API returns the same three top-level keys, so one response handler covers your whole integration: status, error and data. Only data changes shape. Response format covers the envelope, the other output formats and how premium fields are withheld.

Sample response
{
  "status": "ok",
  "error": null,
  "data": {
    "url": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
    "linkCount": 16,
    "externalLinkCount": 13,
    "internalLinkCount": 3,
    "links": [
      {
        "text": "Documentation",
        "href": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html",
        "external": false
      },
      {
        "text": "Amazon EC2 Instance Types Guide",
        "href": "https://docs.aws.amazon.com/ec2/latest/instancetypes/instance-types.html",
        "external": true
      },
      {
        "text": "Amazon EC2 Auto Scaling",
        "href": "https://docs.aws.amazon.com/autoscaling/",
        "external": true
      }
    ],
    "uniqueDomains": [
      "docs.aws.amazon.com",
      "aws.amazon.com"
    ],
    "maxLinksReached": false
  }
}

Response fields

Paths are relative to data. Premium fields are absent rather than zeroed on plans that do not include them, so check for presence instead of comparing to 0.

FieldTypeExampleDescription
urlstring"http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html"The scraped page's URL, normalized to a bare http:// origin with any tabs, newlines and carriage returns stripped
linkCountnumber16Total number of links returned, after filtering out empty anchors and page-fragment links
externalLinkCountnumber13Number of returned links that point to a different domain than the scraped page
internalLinkCountnumber3Number of returned links that point to the same domain as the scraped page
linksarray[3]Every link found on the page, up to the requested limit
textstring"Documentation"The link's visible anchor text, with tabs, newlines and carriage returns removed
hrefstring"http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html"The link's destination URL, resolved to an absolute URL when the source was root-relative
externalbooleanfalseWhether the link points to a different domain than the scraped page
uniqueDomainsPremiumarray[docs.aws.amazon.com, ...]List of unique external domains found
maxLinksReachedbooleanfalseWhether the number of links found met or exceeded the maximum link limit, meaning more links may exist beyond what was returned

Errors

Read the HTTP status first, then error for the specific reason. The body names the parameter that has to change. Error handling covers the full status list and which of them are worth retrying.

StatusMeaningWhat to do
400Input was rejectedRead error; it names the parameter.
401Key missing or invalidCheck the header name and the key value.
403Key valid, but not permittedA key restriction or IP allow-list; see key scoping.
429Rate limited, or out of creditsRead error to tell them apart; see rate limits.

Use cases

Site Architecture Audits
Crawlers inspect internal navigation paths and anchor text across catalog pages to detect orphan URLs and track link equity distribution.
Dead Link Discovery
Run published articles through the scraper to catch external links, resolve relative URLs, and forward destination targets to status checkers.
Competitor Footprint Mapping
Growth analysts feed rival blog posts into the endpoint to inventory citation targets and detect outbound partner referrals.
Web Archival Pipelines
When archiving resource directories, extract all outbound destinations and clean anchor text before caching document references.

Other ways to use Link Scraper

Set up Link Scraper on APIVerve, or reach the same source a different way. Your APIVerve account and credits work on all of them — one key, one balance.

Give it to an AI agentConnect over MCP and your agent calls it as a native tool — Claude, Cursor, ChatGPT.VerveKit →Reference →
Google Sheets or ExcelA =VERVE() formula fills a column — no script, no export, recalculates in place.VerveSheets →Reference →
Ground an agent on itA cited, machine-checkable fact your model can't produce on its own.VerveContext →Reference →

More in Data Scraping:

Was this page helpful?