NEWScrapingAnt MCP for Claude Code, Cursor & Windsurf — try it free →
Skip to main content
Version: v2

Markdown Transformation Endpoint

Welcome to the documentation for ScrapingAnt's new endpoint which extracts data from websites and automatically converts the HTML output into Markdown format. This feature is particularly useful for leveraging results in Language Learning Models (LLMs) and Retrieval-Augmented Generation (RAG).

Endpoint Description​

The Markdown Transformation Endpoint provides a seamless way to scrape web content, converting it directly from HTML to Markdown. This simplifies the process of integrating scraped data into text-based models and applications.

Features​

  • HTML to Markdown Conversion: Automatically converts the extracted HTML content to Markdown, maintaining the essential structure and style in a simpler text format.
  • Easy Integration with LLMs and RAG: The output in Markdown format is ready to be used with various language models and retrieval systems without additional processing.

Request examples​

The fixture page used below is https://scrapingant.github.io/scrapingant-examples/fixtures/markdown-article.html (an article with navigation, a cookie banner, a table, a code block and a footer). The three snippets are the executed scripts from the evidence packet.

curl 'https://api.scrapingant.com/v2/markdown?url=https://scrapingant.github.io/scrapingant-examples/fixtures/markdown-article.html' \
-H 'x-api-key: YOUR_API_KEY'
import requests

r = requests.get(
"https://api.scrapingant.com/v2/markdown",
params={"url": "https://scrapingant.github.io/scrapingant-examples/fixtures/markdown-article.html"},
headers={"x-api-key": "YOUR_API_KEY"},
timeout=120,
)
print(r.status_code, r.headers.get("Ant-credits-cost"))
doc = r.json()
print(doc["url"])
print(doc["markdown"])
const url = "https://scrapingant.github.io/scrapingant-examples/fixtures/markdown-article.html";
const res = await fetch(
"https://api.scrapingant.com/v2/markdown?" + new URLSearchParams({ url }),
{ headers: { "x-api-key": "YOUR_API_KEY" } },
);
const { url: finalUrl, markdown } = (await res.json()) as { url: string; markdown: string };
console.log(res.status, res.headers.get("ant-credits-cost"), finalUrl);
console.log(markdown);

Parameters​

The endpoint takes the same parameters as the general endpoint: url and x-api-key (required), browser (default true), proxy_type, proxy_country, timeout (5–60 seconds, default 60), return_page_source, cookies, js_snippet, wait_for_selector and block_resource. js_snippet, wait_for_selector, block_resource and return_page_source work only with browser=true.

The endpoint supports GET, POST, PUT and DELETE; the method and body are forwarded to the target page as described in POST, PUT and DELETE requests.

Response format​

The response is a JSON object with two properties:

{
"url": "https://example.com",
"markdown": "# Heading 1\n\nThis is a paragraph of text.\n\n## Heading 2\n\nAnother paragraph of text."
}

Every successful response carries the Ant-credits-cost header with the credits spent on the request. Note that this endpoint's responses do not carry the Ant-page-status-code header; a target page that answers with an error status (for example 404) is still converted and returned with HTTP 200.

What the conversion does​

The service removes <script> and <noscript> elements and converts the rest of the page to Markdown with html2text (version 2020.1.16). It does not remove navigation, footers, cookie banners or advertising blocks: the whole page comes back as Markdown. In the output, headings become # lines, links stay inline as [text](url), lists keep their markers, tables become pipe-separated rows, <pre> blocks become 4-space-indented blocks without a language marker, images become ![alt](src), and lines are wrapped at 78 characters where a break is possible.

To drop parts of a page before conversion, remove them in the browser with a base64-encoded js_snippet (requires browser=true), for example:

document.querySelectorAll('nav, footer, .cookie, aside').forEach(e => e.remove());

Pricing​

A request with browser=true (the default) through a datacenter proxy costs 10 API credits; with browser=false it costs 1 API credit. Residential proxies cost 25 credits without a browser and 125 with JavaScript rendering. Only successful responses are charged; the exact cost is in the Ant-credits-cost header of each response. See API credits cost.

Errors​

Errors come back as JSON with a detail string; the status codes are listed on the Errors page. If the Markdown conversion itself fails, the API answers 500 with Markdown extraction error. Please try again later or contact us via support@scrapingant.com.