Messy input in — a URL, raw HTML, plain text, a PDF dumped to text — and JSON out that provably validates against your JSON Schema. Built for agents that need "scrape this page into this shape, reliably".
The guarantee. Every 2xx from /extract carries
data that has been validated against your exact schema, locally, before it was
returned. If the model's output does not conform, we feed it the precise
validation errors and retry. If it still does not conform, you get a
5xx and you are not charged — never
almost-right data dressed up as a success.
A wrong value is worse than a missing one, so the model is
instructed to emit null rather than guess, and we report per-field
confidence plus an absent[] list so you can see exactly which
fields the source did not contain.
| Endpoint | Price | What it does |
|---|---|---|
POST /extract | $0.015 | Source + your JSON Schema → conformant JSON. |
POST /infer | $0.02 | No schema? We design one and fill it. Returns both. |
POST /table | $0.012 | Every table on a page → typed records + CSV. Merged cells handled. |
POST /batch | $0.012 / item | One schema across up to 15 sources, fetched concurrently. |
POST /classify | $0.006 | Text + your labels → per-label scores and a rationale. |
GET /, /health, /openapi.json, /limits |
free | Docs, liveness, machine-readable spec, limits. |
POST /validate | free | Check a schema is well-formed and see what we will send the model. |
# Any x402 client pays automatically. Without one you get a 402 + challenge.
curl -s https://structura.x.c00l.site/extract \
-H 'content-type: application/json' \
-d '{
"source": { "url": "https://example.com/product/widget-pro" },
"schema": {
"type": "object",
"properties": {
"name": { "type": "string" },
"price": { "type": ["number","null"] },
"currency": { "type": ["string","null"], "pattern": "^[A-Z]{3}$" },
"inStock": { "type": ["boolean","null"] },
"specs": { "type": "array", "items": {
"type": "object",
"properties": { "name": {"type":"string"},
"value": {"type":"string"} },
"required": ["name","value"] } }
},
"required": ["name","price","currency","inStock","specs"]
}
}'
Make your fields nullable. If the schema says
"type": "string" and the page has no such value, the only ways out are to
invent one or to fail. We fail — with
schema_unsatisfiable, the exact pointers, and no charge. Write
"type": ["string","null"] and you get an honest null instead.
/extract $0.015 / callThe main event. source is exactly one of {url}, {html} or {text}. schema is JSON Schema draft 2020-12 and is checked for well-formedness before the payment challenge, so a bad schema costs nothing. instructions is optional extra guidance. strict:false opts out of the conformance guarantee and returns the near-miss instead.
{
"source": { "text": "Senior Rust Engineer at Ferrous Labs. Remote (EU).\n$150k-$185k. Requires: 5y systems, async Rust, Postgres." },
"schema": {
"type": "object",
"properties": {
"title": { "type": "string" },
"company": { "type": "string" },
"salaryMin": { "type": ["integer","null"] },
"salaryMax": { "type": ["integer","null"] },
"remote": { "type": ["boolean","null"] },
"requirements": { "type": "array", "items": { "type": "string" } }
},
"required": ["title","company","salaryMin","salaryMax","remote","requirements"]
},
"instructions": "Salaries in whole dollars, no separators.",
"strict": true
}
{
"data": {
"title": "Senior Rust Engineer",
"company": "Ferrous Labs",
"salaryMin": 150000,
"salaryMax": 185000,
"remote": true,
"requirements": ["5 years systems experience","async Rust","Postgres"]
},
"valid": true,
"validationErrors": [],
"confidence": { "/title": 1, "/company": 1, "/salaryMin": 0.95,
"/salaryMax": 0.95, "/remote": 0.9, "/requirements/0": 0.9 },
"confidenceOverall": 0.95,
"absent": [],
"notes": "Salary range stated as $150k-$185k; expanded to whole dollars.",
"repairAttempts": 0,
"usage": { "inputTokens": 1841, "outputTokens": 402, "cacheReadTokens": 1102,
"modelCalls": 1, "model": "claude-sonnet-4-5" }
}
/infer $0.02 / callFor "just structure this for me". We work out what the document is, design the schema a competent engineer would have written for it, and return the schema alongside the extracted data. The returned data is guaranteed to validate against the returned schema.
{
"source": { "url": "https://news.example.com/2024/03/quarterly-results" },
"instructions": "Optional: emphasise the financial figures."
}
{
"schema": { "$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Quarterly Results Article", "type": "object",
"properties": { "headline": {"type":["string","null"]}, "...": {} },
"required": ["headline"] },
"data": { "headline": "Acme posts record Q1", "...": {} },
"documentKind": "financial news article",
"valid": true, "repairAttempts": 0
}
/table $0.012 / callEvery data table in the document as clean arrays of objects, with detected headers, per-column types and units, plus RFC 4180 CSV. Merged cells are expanded locally using the real HTML placement rules before the model sees them, so colspan/rowspan, two-level headers and header-in-first-column spec sheets all come out right. Multi-table pages return every table.
{ "source": { "url": "https://en.wikipedia.org/wiki/List_of_countries_by_GDP" } }
{
"tables": [{
"name": "GDP by country (nominal)",
"orientation": "row",
"rowCount": 190, "columnCount": 4,
"columns": [
{ "name": "Country", "type": "string", "unit": null, "nullCount": 0 },
{ "name": "IMF estimate", "type": "integer", "unit": "USD millions", "nullCount": 6 }
],
"rows": [ { "Country": "United States", "IMF estimate": 30507217 } ],
"csv": "Country,IMF estimate\r\nUnited States,30507217\r\n",
"notes": "Two-level header flattened; footnote markers stripped."
}],
"tableCount": 1,
"detected": { "htmlTables": 3, "mode": "html-grid", "mergedCells": true, "skipped": 0 }
}
/batch $0.012 / itemThe same schema across up to 15 sources, fetched and processed concurrently. Priced per item and the total is computed from the item count, so the challenge you receive already states the exact amount. Per-item failures are reported inline and do not fail the call.
{
"sources": [ { "url": "https://a.example/p/1" },
{ "url": "https://a.example/p/2" },
{ "text": "Widget C, 4.99 GBP, out of stock" } ],
"schema": { "type": "object",
"properties": { "name": {"type":["string","null"]},
"price": {"type":["number","null"]} },
"required": ["name","price"] }
}
{
"count": 3, "succeeded": 2, "failed": 1,
"results": [
{ "index": 0, "ok": true, "data": { "name": "Widget A", "price": 19.99 },
"valid": true, "repairAttempts": 0, "confidence": { "/name": 1, "/price": 1 } },
{ "index": 1, "ok": false,
"error": { "code": "upstream_timeout", "message": "Upstream did not respond within 15000 ms." } },
{ "index": 2, "ok": true, "data": { "name": "Widget C", "price": 4.99 }, "valid": true }
],
"usage": { "inputTokens": 5203, "outputTokens": 611, "modelCalls": 3 }
}
/classify $0.006 / callCheap, high-volume text classification against your own label set. Every label is scored, not just the winners, so you can threshold it yourself. multi:true allows several labels (and none). rubric defines what your labels mean and overrides the model's own reading of the names.
{
"text": "Charged twice for the same order and support has not replied in four days.",
"labels": ["billing","bug","feature request","praise","churn risk"],
"multi": true,
"rubric": "churn risk: the customer signals they may leave or is visibly angry."
}
{
"labels": ["billing","churn risk"],
"scores": { "billing": 0.96, "bug": 0.22, "feature request": 0.02,
"praise": 0.01, "churn risk": 0.71 },
"top": { "label": "billing", "score": 0.96 },
"margin": 0.25,
"rationale": "\"Charged twice for the same order\" is a billing fault; four days without a reply plus the tone signals churn risk."
}
Well-formedness, depth, size and evaluability are checked before the 402.
A malformed schema is a free 400.
Your schema becomes a tool's input_schema and the model is forced
into it. No JSON scraped out of prose, ever.
The model sees your schema widened so null is always legal, so it
can be honest instead of cornered into inventing a value.
Output is validated against your original schema — draft 2020-12, per-pointer errors — inside the Worker.
On failure the exact errors go back to the model, in-conversation, and it repairs rather than restarting. Up to 3 attempts.
Still not conformant? 5xx, no settlement, and an error naming the
pointers that could not be satisfied.
notes.
Output shape is enforced by the tool schema regardless of what the page says.169.254.0.0/16), CGNAT, ULA and reserved ranges, in
IPv4, IPv6 and IPv4-mapped forms, plus credential-bearing URLs, blocked ports
and single-label hosts — re-checked on every redirect hop. Rejection happens
before the payment challenge.truncated), 64 KB of schema, bounded depth,
timeouts on every fetch and every model call.x402 v2, USDC on Base (eip155:8453). Point any x402-aware client at
these URLs and payment is automatic; no account and no API key. Without a
payment header you get 402 with a PAYMENT-REQUIRED
challenge. Failed requests are never settled, so you only pay for a result you
can use. See openapi.json and /limits.