Web Reader API
Extract clean, model-ready content from any page URL, with optional browser rendering.
Endpoint
POST https://serpforai.dev/api/v1/url
GET https://serpforai.dev/api/v1/url
Authentication is the same bearer header used by the SERP API:
Authorization: Bearer <userKey>
This single endpoint also powers File Extraction and Screenshot - which response you get
back depends only on which optional fields you set. Plain s gives you Reader; adding file: 1
switches to File Extraction; adding image: 1 or image: 2 switches to Screenshot.
Request parameters
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
s | string | yes | - | The page URL to extract. When file: 1, this is the document URL instead |
mode | number | no | 0 | Render mode: 0 = standard HTTP fetch, 1 = headless browser |
w | number | no | 5000 | Wait after page load, in milliseconds. Only used when mode: 1 |
d | number | no | 30000 | Request timeout, in milliseconds |
html | number | no | 0 | 1 = include the raw page HTML alongside the Markdown, 0 = omit |
proxy | number | no | 0 | Proxy tier: 0 none, 1 Shared Pool (+2 credits), 2 Datacenter (+5 credits), 3 Residential (+10 credits) |
captcha | number | no | 0 | 1 = attempt to solve simple CAPTCHAs automatically. No extra credits |
file | number | no | 0 | File Extraction: 1 = parse the document at s (PDF/DOCX/XLSX/…). No extra credits |
image | number | no | 0 | Screenshot: 1 = first screen only, 2 = full page. No extra credits |
Response
{
"code": 0,
"msg": "",
"data": {
"id": "...",
"markdown": "...",
"title": "...",
"description": "..."
}
}
| Field | Present when | Description |
|---|---|---|
id | always | Opaque identifier for this extraction |
markdown | always | The page content, converted to Markdown |
title | always | Page title |
description | always | Page description/summary |
html | html: 1 | Raw page HTML |
fileUrl | file: 1 | The original document URL, echoed back |
fileMarkdown | file: 1 | Parsed document content as Markdown |
imageUrl | image: 1|2 | Hosted PNG URL of the screenshot |
When to enable the browser
Set mode: 1 only for pages that render their main content with client-side JavaScript. Headless
rendering costs more time, so leaving mode: 0 (the default) keeps latency low for ordinary
article pages.
curl -X POST https://serpforai.dev/api/v1/url \
-H "Authorization: Bearer $SERPFORAI_KEY" \
-H "Content-Type: application/json" \
-d '{"s":"https://example.com/spa-article","mode":1,"w":2000,"d":15000}'
Proxies and CAPTCHAs
Most sites need nothing extra. Reach for proxy only when a target consistently blocks the
platform’s default IPs, starting with tier 1 (Shared Pool) before paying for 2 (Datacenter) or
3 (Residential). Set captcha: 1 if a page occasionally shows a simple CAPTCHA wall - it costs
no extra credits and is skipped automatically when there is nothing to solve.
File Extraction and Screenshot
# Parse a PDF into Markdown
curl -X POST https://serpforai.dev/api/v1/url \
-H "Authorization: Bearer $SERPFORAI_KEY" \
-H "Content-Type: application/json" \
-d '{"s":"https://example.com/whitepaper.pdf","file":1}'
# Full-page screenshot
curl -X POST https://serpforai.dev/api/v1/url \
-H "Authorization: Bearer $SERPFORAI_KEY" \
-H "Content-Type: application/json" \
-d '{"s":"https://example.com","image":2}'