Scrape a Page
Extract content from any single URL as markdown, HTML, or structured JSON.
Scrape a Page
Extract clean content from any URL. Returns markdown by default — perfect for feeding into LLMs.
Basic Usage
curl -X POST https://scrapeforllm.com/api/app/scrapes \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"url": "https://example.com/blog/post",
"type": "scrape"
}'const response = await fetch("https://scrapeforllm.com/api/app/scrapes", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: "Bearer YOUR_API_KEY",
},
body: JSON.stringify({
url: "https://example.com/blog/post",
type: "scrape",
}),
});
const data = await response.json();
console.log(data.scrape.result.data.markdown);import requests
response = requests.post(
"https://scrapeforllm.com/api/app/scrapes",
headers={
"Content-Type": "application/json",
"Authorization": "Bearer YOUR_API_KEY",
},
json={
"url": "https://example.com/blog/post",
"type": "scrape",
},
)
data = response.json()
print(data["scrape"]["result"]["data"]["markdown"])Response:
{
"scrape": {
"id": "550e8400-e29b-41d4-a716-446655440000",
"url": "https://example.com/blog/post",
"type": "scrape",
"status": "completed",
"creditsUsed": 1,
"format": "markdown",
"result": {
"data": {
"markdown": "# Blog Post Title\n\nThe full content of the page...",
"metadata": {
"title": "Blog Post Title",
"description": "Meta description",
"sourceURL": "https://example.com/blog/post",
"statusCode": 200
}
}
},
"createdAt": "2025-01-15T10:30:00.000Z",
"completedAt": "2025-01-15T10:30:02.000Z"
}
}PDFs and Documents
Point scrape at a PDF, Word, Excel or PowerPoint URL and you get the same clean
markdown you'd get from a web page — no separate endpoint, no extra parameters,
still 1 credit. Multi-page documents are extracted in full.
curl -X POST https://scrapeforllm.com/api/app/scrapes \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_API_KEY" \
-d '{
"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"type": "scrape"
}'const response = await fetch("https://scrapeforllm.com/api/app/scrapes", {
method: "POST",
headers: {
"Content-Type": "application/json",
Authorization: "Bearer YOUR_API_KEY",
},
body: JSON.stringify({
url: "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
type: "scrape",
}),
});
const data = await response.json();
console.log(data.scrape.result.data.markdown);import requests
response = requests.post(
"https://scrapeforllm.com/api/app/scrapes",
headers={
"Content-Type": "application/json",
"Authorization": "Bearer YOUR_API_KEY",
},
json={
"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"type": "scrape",
},
)
print(response.json()["scrape"]["result"]["data"]["markdown"])| Type | Extension | What you get |
|---|---|---|
.pdf | Full text, including long multi-page documents | |
| Word | .docx | Text with headings and structure preserved |
| Excel | .xlsx | Sheets converted to markdown tables |
| PowerPoint | .pptx | Slide titles as headings, bullets as lists |
Legacy and open formats (.doc, .xls, .ppt, .odt, .ods, .odp, .rtf,
.epub, .csv) are accepted too.
A spreadsheet comes back ready to feed straight to a model:
| Company | Revenue | Employees |
| --------- | -------- | --------- |
| Acme Corp | 42000000 | 1234 |
| Globex | 7500000 | 321 |Text is extracted from the document's own text layer. Scanned or image-only PDFs
have no text layer and fail with a 422 rather than returning an empty document —
OCR is not applied.
Uploading a file
When a document isn't reachable by URL — a local report, a file behind a login, a
contract on disk — upload it directly to /api/app/scrapes/upload as
multipart/form-data. Same markdown out, same 1 credit.
curl -X POST https://scrapeforllm.com/api/app/scrapes/upload \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "file=@/path/to/report.pdf" \
-F 'formats=["markdown"]'const form = new FormData();
form.append("file", fileInput.files[0]);
form.append("formats", JSON.stringify(["markdown"]));
const response = await fetch(
"https://scrapeforllm.com/api/app/scrapes/upload",
{
method: "POST",
headers: { Authorization: "Bearer YOUR_API_KEY" },
body: form, // do not set Content-Type — the browser adds the boundary
},
);
const data = await response.json();
console.log(data.scrape.result.data.markdown);import requests
with open("report.pdf", "rb") as f:
response = requests.post(
"https://scrapeforllm.com/api/app/scrapes/upload",
headers={"Authorization": "Bearer YOUR_API_KEY"},
files={"file": f},
data={"formats": '["markdown"]'},
)
print(response.json()["scrape"]["result"]["data"]["markdown"])| Field | Required | Notes |
|---|---|---|
file | Yes | The document. Max 25 MB. |
formats | No | JSON array — markdown (default), html, rawHtml |
A PDF with no extractable text layer — a scan, or a page of images — returns
HTTP 422 Parsing failed rather than an empty document, so a failed extraction
is never mistaken for an empty one. OCR is not applied.
Useful for
Research papers, SEC and regulatory filings, tax and government forms, investor decks, whitepapers, and any report published as a PDF — all converted to LLM-ready markdown in one call.
Anti-bot Scraping
Some sites block automated requests — either by rejecting datacenter IP ranges
(Reddit, many company databases) or with a browser challenge like Cloudflare
Turnstile (RocketReach, review sites). Set the proxy option to get past them.
It runs a waterfall: the cheap path first, escalating only when a site
actually blocks you — so you never overpay for pages that were easy.
{
"url": "https://www.reddit.com/r/programming/",
"type": "scrape",
"options": { "proxy": "auto" }
}proxy | What it does | Credits (by tier reached) |
|---|---|---|
basic | Fast datacenter fetch (default; omit proxy for this) | 1 |
auto | Datacenter → residential IP → stealth browser, escalating only on a block | 1 / 2 / 5 |
stealth | Skips straight to residential + a stealth browser, for sites you know are hard | 2 / 5 |
You are billed for the tier that succeeds, reported as result.data.metadata.proxyTier
(datacenter, residential-fallback, or camoufox-stealth). An auto scrape of
an unprotected page costs the same 1 credit as a basic scrape.
Blocked scrapes escalate automatically
You don't have to ask for this. If a normal scrape is blocked, it's retried
through the waterfall automatically and billed by the tier that succeeds —
pages that aren't blocked are untouched and still cost 1 credit. Setting
proxy explicitly just skips straight to that tier. Turn the behaviour off per
account in Profile → Scraping if you want blocked scrapes to fail for free
instead (they are never charged when they fail).
What it does and doesn't beat
Anti-bot mode beats IP-reputation blocks (Reddit-class), fingerprint
challenges like Cloudflare Turnstile (RocketReach-class), and Akamai Bot
Manager. It does not unlock Google search results — use the search type
for web search — or PerimeterX "Press & Hold" walls.
Credits
Each scrape costs 1 credit. If you include json format, it costs 5
credits (uses LLM extraction). Anti-bot scrapes (proxy: "auto" or
"stealth") are billed 1–5 credits by the tier that succeeded — see above.