ScrapeforLLM Docs
ScrapeforLLM Docs
Getting StartedScrape a PageScreenshotCrawl a SiteMap, Search & ExtractList & Get ScrapesError Codes

Scrape a Page

Extract content from any single URL as markdown, HTML, or structured JSON.

Scrape a Page

Extract clean content from any URL. Returns markdown by default — perfect for feeding into LLMs.

Basic Usage

curl -X POST https://scrapeforllm.com/api/app/scrapes \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "url": "https://example.com/blog/post",
    "type": "scrape"
  }'
const response = await fetch("https://scrapeforllm.com/api/app/scrapes", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    Authorization: "Bearer YOUR_API_KEY",
  },
  body: JSON.stringify({
    url: "https://example.com/blog/post",
    type: "scrape",
  }),
});

const data = await response.json();
console.log(data.scrape.result.data.markdown);
import requests

response = requests.post(
    "https://scrapeforllm.com/api/app/scrapes",
    headers={
        "Content-Type": "application/json",
        "Authorization": "Bearer YOUR_API_KEY",
    },
    json={
        "url": "https://example.com/blog/post",
        "type": "scrape",
    },
)

data = response.json()
print(data["scrape"]["result"]["data"]["markdown"])

Response:

{
  "scrape": {
    "id": "550e8400-e29b-41d4-a716-446655440000",
    "url": "https://example.com/blog/post",
    "type": "scrape",
    "status": "completed",
    "creditsUsed": 1,
    "format": "markdown",
    "result": {
      "data": {
        "markdown": "# Blog Post Title\n\nThe full content of the page...",
        "metadata": {
          "title": "Blog Post Title",
          "description": "Meta description",
          "sourceURL": "https://example.com/blog/post",
          "statusCode": 200
        }
      }
    },
    "createdAt": "2025-01-15T10:30:00.000Z",
    "completedAt": "2025-01-15T10:30:02.000Z"
  }
}

PDFs and Documents

Point scrape at a PDF, Word, Excel or PowerPoint URL and you get the same clean markdown you'd get from a web page — no separate endpoint, no extra parameters, still 1 credit. Multi-page documents are extracted in full.

curl -X POST https://scrapeforllm.com/api/app/scrapes \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -d '{
    "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    "type": "scrape"
  }'
const response = await fetch("https://scrapeforllm.com/api/app/scrapes", {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    Authorization: "Bearer YOUR_API_KEY",
  },
  body: JSON.stringify({
    url: "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    type: "scrape",
  }),
});

const data = await response.json();
console.log(data.scrape.result.data.markdown);
import requests

response = requests.post(
    "https://scrapeforllm.com/api/app/scrapes",
    headers={
        "Content-Type": "application/json",
        "Authorization": "Bearer YOUR_API_KEY",
    },
    json={
        "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
        "type": "scrape",
    },
)

print(response.json()["scrape"]["result"]["data"]["markdown"])
TypeExtensionWhat you get
PDF.pdfFull text, including long multi-page documents
Word.docxText with headings and structure preserved
Excel.xlsxSheets converted to markdown tables
PowerPoint.pptxSlide titles as headings, bullets as lists

Legacy and open formats (.doc, .xls, .ppt, .odt, .ods, .odp, .rtf, .epub, .csv) are accepted too.

A spreadsheet comes back ready to feed straight to a model:

| Company   | Revenue  | Employees |
| --------- | -------- | --------- |
| Acme Corp | 42000000 | 1234      |
| Globex    | 7500000  | 321       |

Text is extracted from the document's own text layer. Scanned or image-only PDFs have no text layer and fail with a 422 rather than returning an empty document — OCR is not applied.

Uploading a file

When a document isn't reachable by URL — a local report, a file behind a login, a contract on disk — upload it directly to /api/app/scrapes/upload as multipart/form-data. Same markdown out, same 1 credit.

curl -X POST https://scrapeforllm.com/api/app/scrapes/upload \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "file=@/path/to/report.pdf" \
  -F 'formats=["markdown"]'
const form = new FormData();
form.append("file", fileInput.files[0]);
form.append("formats", JSON.stringify(["markdown"]));

const response = await fetch(
  "https://scrapeforllm.com/api/app/scrapes/upload",
  {
    method: "POST",
    headers: { Authorization: "Bearer YOUR_API_KEY" },
    body: form, // do not set Content-Type — the browser adds the boundary
  },
);

const data = await response.json();
console.log(data.scrape.result.data.markdown);
import requests

with open("report.pdf", "rb") as f:
    response = requests.post(
        "https://scrapeforllm.com/api/app/scrapes/upload",
        headers={"Authorization": "Bearer YOUR_API_KEY"},
        files={"file": f},
        data={"formats": '["markdown"]'},
    )

print(response.json()["scrape"]["result"]["data"]["markdown"])
FieldRequiredNotes
fileYesThe document. Max 25 MB.
formatsNoJSON array — markdown (default), html, rawHtml

A PDF with no extractable text layer — a scan, or a page of images — returns HTTP 422 Parsing failed rather than an empty document, so a failed extraction is never mistaken for an empty one. OCR is not applied.

Useful for

Research papers, SEC and regulatory filings, tax and government forms, investor decks, whitepapers, and any report published as a PDF — all converted to LLM-ready markdown in one call.

Anti-bot Scraping

Some sites block automated requests — either by rejecting datacenter IP ranges (Reddit, many company databases) or with a browser challenge like Cloudflare Turnstile (RocketReach, review sites). Set the proxy option to get past them. It runs a waterfall: the cheap path first, escalating only when a site actually blocks you — so you never overpay for pages that were easy.

{
  "url": "https://www.reddit.com/r/programming/",
  "type": "scrape",
  "options": { "proxy": "auto" }
}
proxyWhat it doesCredits (by tier reached)
basicFast datacenter fetch (default; omit proxy for this)1
autoDatacenter → residential IP → stealth browser, escalating only on a block1 / 2 / 5
stealthSkips straight to residential + a stealth browser, for sites you know are hard2 / 5

You are billed for the tier that succeeds, reported as result.data.metadata.proxyTier (datacenter, residential-fallback, or camoufox-stealth). An auto scrape of an unprotected page costs the same 1 credit as a basic scrape.

Blocked scrapes escalate automatically

You don't have to ask for this. If a normal scrape is blocked, it's retried through the waterfall automatically and billed by the tier that succeeds — pages that aren't blocked are untouched and still cost 1 credit. Setting proxy explicitly just skips straight to that tier. Turn the behaviour off per account in Profile → Scraping if you want blocked scrapes to fail for free instead (they are never charged when they fail).

What it does and doesn't beat

Anti-bot mode beats IP-reputation blocks (Reddit-class), fingerprint challenges like Cloudflare Turnstile (RocketReach-class), and Akamai Bot Manager. It does not unlock Google search results — use the search type for web search — or PerimeterX "Press & Hold" walls.

Credits

Each scrape costs 1 credit. If you include json format, it costs 5 credits (uses LLM extraction). Anti-bot scrapes (proxy: "auto" or "stealth") are billed 1–5 credits by the tier that succeeded — see above.

Getting Started

Turn any website into clean, LLM-ready data in seconds.

Screenshot

Capture viewport, full-page, responsive, scroll-sliced, and element-targeted screenshots of any URL.

On this page

Scrape a PageBasic UsagePDFs and DocumentsUploading a fileAnti-bot Scraping