Overview
Link Scraper works by parsing the HTML content of a web page to extract all the links. It returns the links in a structured format, including the link URL, title, and more.
Endpoint
One host, one path per API. The block below shows this call in four languages; every one of them is the same HTTP request. Making requests covers the timeouts, retries and parameter rules that apply to all of them. The SDKs wrap the same call in a typed client.
curl -X POST https://api.apiverve.com/v1/linkscraper \
-H "x-api-key: your_api_key_here" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"maxlinks": 20,
"includequery": false
}'const res = await fetch('https://api.apiverve.com/v1/linkscraper', {
method: 'POST',
headers: {
'x-api-key': 'your_api_key_here',
'Content-Type': 'application/json',
},
body: JSON.stringify({
"url": "https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"maxlinks": 20,
"includequery": false
}),
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const { data } = await res.json();
console.log(data);import requests
res = requests.post(
"https://api.apiverve.com/v1/linkscraper",
headers={"x-api-key": "your_api_key_here"},
json={
"url": "https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"maxlinks": 20,
"includequery": False
},
timeout=15,
)
res.raise_for_status()
print(res.json()["data"])package main
import (
"fmt"
"io"
"net/http"
"strings"
)
func main() {
body := strings.NewReader(`{
"url": "https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"maxlinks": 20,
"includequery": false
}`)
req, _ := http.NewRequest("POST", "https://api.apiverve.com/v1/linkscraper", body)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("x-api-key", "your_api_key_here")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, _ := io.ReadAll(res.Body)
fmt.Println(string(out))
}Replace your_api_key_here with the key from your dashboard. When the inputs arrive as a list rather than one at a time, batch requests run up to 200 of them through this same API in a single call.
Authentication
Send your key in the x-api-key header. That is the only auth step — there is no token exchange and no per-endpoint scope to configure. Authentication covers creating, rotating and revoking keys.
A 401 means the key is missing, invalid or expired. A 403 means the key is valid but not permitted here — blocked by a key restriction or an IP allow-list. Running out of credits is a 429.
Parameters
Sent as JSON in the request body. Premium parameters are accepted on every plan but only take effect on plans that include them.
| Parameter | Type | Description |
|---|---|---|
urlRequired | string | The URL of the web page to scrape links from url |
maxlinksOptionalPremium | number | Maximum number of links to scrape and return default 50 |
includequeryOptional | boolean | Include query strings in the scraped links |
Response
Every API returns the same three top-level keys, so one response handler covers your whole integration: status, error and data. Only data changes shape. Response format covers the envelope, the other output formats and how premium fields are withheld.
{
"status": "ok",
"error": null,
"data": {
"url": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html",
"linkCount": 16,
"externalLinkCount": 13,
"internalLinkCount": 3,
"links": [
{
"text": "Documentation",
"href": "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html",
"external": false
},
{
"text": "Amazon EC2 Instance Types Guide",
"href": "https://docs.aws.amazon.com/ec2/latest/instancetypes/instance-types.html",
"external": true
},
{
"text": "Amazon EC2 Auto Scaling",
"href": "https://docs.aws.amazon.com/autoscaling/",
"external": true
}
],
"uniqueDomains": [
"docs.aws.amazon.com",
"aws.amazon.com"
],
"maxLinksReached": false
}
}
Response fields
Paths are relative to data. Premium fields are absent rather than zeroed on plans that do not include them, so check for presence instead of comparing to 0.
| Field | Type | Example | Description |
|---|---|---|---|
url | string | "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html" | The scraped page's URL, normalized to a bare http:// origin with any tabs, newlines and carriage returns stripped |
linkCount | number | 16 | Total number of links returned, after filtering out empty anchors and page-fragment links |
externalLinkCount | number | 13 | Number of returned links that point to a different domain than the scraped page |
internalLinkCount | number | 3 | Number of returned links that point to the same domain as the scraped page |
links | array[3] | Every link found on the page, up to the requested limit | |
text | string | "Documentation" | The link's visible anchor text, with tabs, newlines and carriage returns removed |
href | string | "http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/concepts.html/index.html" | The link's destination URL, resolved to an absolute URL when the source was root-relative |
external | boolean | false | Whether the link points to a different domain than the scraped page |
uniqueDomainsPremium | array | [docs.aws.amazon.com, ...] | List of unique external domains found |
maxLinksReached | boolean | false | Whether the number of links found met or exceeded the maximum link limit, meaning more links may exist beyond what was returned |
Errors
Read the HTTP status first, then error for the specific reason. The body names the parameter that has to change. Error handling covers the full status list and which of them are worth retrying.
| Status | Meaning | What to do |
|---|---|---|
400 | Input was rejected | Read error; it names the parameter. |
401 | Key missing or invalid | Check the header name and the key value. |
403 | Key valid, but not permitted | A key restriction or IP allow-list; see key scoping. |
429 | Rate limited, or out of credits | Read error to tell them apart; see rate limits. |
Use cases
- Site Architecture Audits
- Crawlers inspect internal navigation paths and anchor text across catalog pages to detect orphan URLs and track link equity distribution.
- Dead Link Discovery
- Run published articles through the scraper to catch external links, resolve relative URLs, and forward destination targets to status checkers.
- Competitor Footprint Mapping
- Growth analysts feed rival blog posts into the endpoint to inventory citation targets and detect outbound partner referrals.
- Web Archival Pipelines
- When archiving resource directories, extract all outbound destinations and clean anchor text before caching document references.
Other ways to use Link Scraper
Set up Link Scraper on APIVerve, or reach the same source a different way. Your APIVerve account and credits work on all of them — one key, one balance.
Related
More in Data Scraping: