Docs/APIs/HTML to Text

HTML to Text

Convert HTML to text

OperationalCredits 2 per callp50 1767msData ProcessingStar

Overview

HTML to Text works by parsing the HTML content and extracting the text. Text extraction is done by removing the HTML tags and returning the plain text content. Advanced algorithms are used to ensure accurate text extraction.

Endpoint

One host, one path per API. The block below shows this call in four languages; every one of them is the same HTTP request. Making requests covers the timeouts, retries and parameter rules that apply to all of them. The SDKs wrap the same call in a typed client.

POSThttps://api.apiverve.com/v1/htmltotext
curl -X POST https://api.apiverve.com/v1/htmltotext \
  -H "x-api-key: your_api_key_here" \
  -H "Content-Type: application/json" \
  -d '{
  "html": "<!doctype html> <html>  <head> <title>This is the title of the webpage!</title> </head> <body> <p>This is an example paragraph. Anything in the <strong>body</strong> tag will appear on the page, just like this <strong>p</strong> tag and its contents.</p> </body> </html>"
}'

Replace your_api_key_here with the key from your dashboard. When the inputs arrive as a list rather than one at a time, batch requests run up to 200 of them through this same API in a single call.

Authentication

Send your key in the x-api-key header. That is the only auth step — there is no token exchange and no per-endpoint scope to configure. Authentication covers creating, rotating and revoking keys.

401 is the only auth verdict

A 401 means the key is missing, invalid or expired. A 403 means the key is valid but not permitted here — blocked by a key restriction or an IP allow-list. Running out of credits is a 429.

Parameters

Sent as JSON in the request body. Premium parameters are accepted on every plan but only take effect on plans that include them.

ParameterTypeDescription
htmlRequiredstringThe HTML to convert to text

Response

Every API returns the same three top-level keys, so one response handler covers your whole integration: status, error and data. Only data changes shape. Response format covers the envelope, the other output formats and how premium fields are withheld.

Sample response
{
  "status": "ok",
  "error": null,
  "data": {
    "text": "This is an example paragraph. Anything in the body tag will appear on the page, just like this p tag and its contents.",
    "parsed": true,
    "extractionMethod": "article",
    "detectedLanguage": {
      "language": "english",
      "confidence": 0.3507446808510638
    },
    "characterCount": 118,
    "wordCount": 23
  }
}

Response fields

Paths are relative to data. Premium fields are absent rather than zeroed on plans that do not include them, so check for presence instead of comparing to 0.

FieldTypeExampleDescription
textstring"This is an example paragraph. Anything in the body tag will appear on the page, just like this p tag and its contents."The article text extracted from the HTML, or null when no article body could be identified
parsedbooleantrueWhether article text was successfully extracted - false when the HTML held no identifiable article body
extractionMethodstring"article"How the text was obtained: article when article-body extraction succeeded, markup when it fell back to reading all the markup, none when no text was found.
detectedLanguagePremiumobject{...}Detected language with confidence score
languagePremiumstring"english"
confidencePremiumnumber0.3507446808510638
characterCountnumber118Number of characters in extracted text
wordCountnumber23Number of words in extracted text

Errors

Read the HTTP status first, then error for the specific reason. The body names the parameter that has to change. Error handling covers the full status list and which of them are worth retrying.

StatusMeaningWhat to do
400Input was rejectedRead error; it names the parameter.
401Key missing or invalidCheck the header name and the key value.
403Key valid, but not permittedA key restriction or IP allow-list; see key scoping.
429Rate limited, or out of creditsRead error to tell them apart; see rate limits.

Use cases

Inbound Email Parsing
Strip HTML markup from customer support emails before passing unformatted message bodies into ticket triage queues.
Search Engine Indexing
When crawling web pages, remove formatting tags and extract readable text to compute accurate word counts for search indices.
LLM Prompt Preparation
To minimize prompt token counts, data engineers strip raw article markup down to readable text before querying language models.
Feed Snippet Generation
Content aggregators convert rich blog markup into plain text snippets for previews, push notifications, and mobile readers.

Other ways to use HTML to Text

Set up HTML to Text on APIVerve, or reach the same source a different way. Your APIVerve account and credits work on all of them — one key, one balance.

Give it to an AI agentConnect over MCP and your agent calls it as a native tool — Claude, Cursor, ChatGPT.VerveKitReference →
Google Sheets or ExcelA =VERVE() formula fills a column — no script, no export, recalculates in place.VerveSheetsReference →
Ground an agent on itA cited, machine-checkable fact your model can't produce on its own.VerveContextReference →

More in Data Processing:

Was this page helpful?