structura.

Messy input in — a URL, raw HTML, plain text, a PDF dumped to text — and JSON out that provably validates against your JSON Schema. Built for agents that need "scrape this page into this shape, reliably".

x402 · USDC on Base no API key from $0.006/call schema-conformant or free

The guarantee. Every 2xx from /extract carries data that has been validated against your exact schema, locally, before it was returned. If the model's output does not conform, we feed it the precise validation errors and retry. If it still does not conform, you get a 5xx and you are not charged — never almost-right data dressed up as a success.

A wrong value is worse than a missing one, so the model is instructed to emit null rather than guess, and we report per-field confidence plus an absent[] list so you can see exactly which fields the source did not contain.

Endpoints

EndpointPriceWhat it does
POST /extract$0.015 Source + your JSON Schema → conformant JSON.
POST /infer$0.02 No schema? We design one and fill it. Returns both.
POST /table$0.012 Every table on a page → typed records + CSV. Merged cells handled.
POST /batch$0.012 / item One schema across up to 15 sources, fetched concurrently.
POST /classify$0.006 Text + your labels → per-label scores and a rationale.
GET /, /health, /openapi.json, /limits freeDocs, liveness, machine-readable spec, limits.
POST /validatefree Check a schema is well-formed and see what we will send the model.

Quick start

# Any x402 client pays automatically. Without one you get a 402 + challenge.
curl -s https://structura.x.c00l.site/extract \
  -H 'content-type: application/json' \
  -d '{
    "source": { "url": "https://example.com/product/widget-pro" },
    "schema": {
      "type": "object",
      "properties": {
        "name":     { "type": "string" },
        "price":    { "type": ["number","null"] },
        "currency": { "type": ["string","null"], "pattern": "^[A-Z]{3}$" },
        "inStock":  { "type": ["boolean","null"] },
        "specs":    { "type": "array", "items": {
                        "type": "object",
                        "properties": { "name": {"type":"string"},
                                        "value": {"type":"string"} },
                        "required": ["name","value"] } }
      },
      "required": ["name","price","currency","inStock","specs"]
    }
  }'

Make your fields nullable. If the schema says "type": "string" and the page has no such value, the only ways out are to invent one or to fail. We fail — with schema_unsatisfiable, the exact pointers, and no charge. Write "type": ["string","null"] and you get an honest null instead.

Reference

POST /extract $0.015 / call

The main event. source is exactly one of {url}, {html} or {text}. schema is JSON Schema draft 2020-12 and is checked for well-formedness before the payment challenge, so a bad schema costs nothing. instructions is optional extra guidance. strict:false opts out of the conformance guarantee and returns the near-miss instead.

Request
{
  "source": { "text": "Senior Rust Engineer at Ferrous Labs. Remote (EU).\n$150k-$185k. Requires: 5y systems, async Rust, Postgres." },
  "schema": {
    "type": "object",
    "properties": {
      "title":     { "type": "string" },
      "company":   { "type": "string" },
      "salaryMin": { "type": ["integer","null"] },
      "salaryMax": { "type": ["integer","null"] },
      "remote":    { "type": ["boolean","null"] },
      "requirements": { "type": "array", "items": { "type": "string" } }
    },
    "required": ["title","company","salaryMin","salaryMax","remote","requirements"]
  },
  "instructions": "Salaries in whole dollars, no separators.",
  "strict": true
}
Response
{
  "data": {
    "title": "Senior Rust Engineer",
    "company": "Ferrous Labs",
    "salaryMin": 150000,
    "salaryMax": 185000,
    "remote": true,
    "requirements": ["5 years systems experience","async Rust","Postgres"]
  },
  "valid": true,
  "validationErrors": [],
  "confidence": { "/title": 1, "/company": 1, "/salaryMin": 0.95,
                  "/salaryMax": 0.95, "/remote": 0.9, "/requirements/0": 0.9 },
  "confidenceOverall": 0.95,
  "absent": [],
  "notes": "Salary range stated as $150k-$185k; expanded to whole dollars.",
  "repairAttempts": 0,
  "usage": { "inputTokens": 1841, "outputTokens": 402, "cacheReadTokens": 1102,
             "modelCalls": 1, "model": "claude-sonnet-4-5" }
}

POST /infer $0.02 / call

For "just structure this for me". We work out what the document is, design the schema a competent engineer would have written for it, and return the schema alongside the extracted data. The returned data is guaranteed to validate against the returned schema.

Request
{
  "source": { "url": "https://news.example.com/2024/03/quarterly-results" },
  "instructions": "Optional: emphasise the financial figures."
}
Response
{
  "schema": { "$schema": "https://json-schema.org/draft/2020-12/schema",
              "title": "Quarterly Results Article", "type": "object",
              "properties": { "headline": {"type":["string","null"]}, "...": {} },
              "required": ["headline"] },
  "data": { "headline": "Acme posts record Q1", "...": {} },
  "documentKind": "financial news article",
  "valid": true, "repairAttempts": 0
}

POST /table $0.012 / call

Every data table in the document as clean arrays of objects, with detected headers, per-column types and units, plus RFC 4180 CSV. Merged cells are expanded locally using the real HTML placement rules before the model sees them, so colspan/rowspan, two-level headers and header-in-first-column spec sheets all come out right. Multi-table pages return every table.

Request
{ "source": { "url": "https://en.wikipedia.org/wiki/List_of_countries_by_GDP" } }
Response
{
  "tables": [{
    "name": "GDP by country (nominal)",
    "orientation": "row",
    "rowCount": 190, "columnCount": 4,
    "columns": [
      { "name": "Country", "type": "string", "unit": null, "nullCount": 0 },
      { "name": "IMF estimate", "type": "integer", "unit": "USD millions", "nullCount": 6 }
    ],
    "rows": [ { "Country": "United States", "IMF estimate": 30507217 } ],
    "csv": "Country,IMF estimate\r\nUnited States,30507217\r\n",
    "notes": "Two-level header flattened; footnote markers stripped."
  }],
  "tableCount": 1,
  "detected": { "htmlTables": 3, "mode": "html-grid", "mergedCells": true, "skipped": 0 }
}

POST /batch $0.012 / item

The same schema across up to 15 sources, fetched and processed concurrently. Priced per item and the total is computed from the item count, so the challenge you receive already states the exact amount. Per-item failures are reported inline and do not fail the call.

Request
{
  "sources": [ { "url": "https://a.example/p/1" },
               { "url": "https://a.example/p/2" },
               { "text": "Widget C, 4.99 GBP, out of stock" } ],
  "schema": { "type": "object",
              "properties": { "name": {"type":["string","null"]},
                              "price": {"type":["number","null"]} },
              "required": ["name","price"] }
}
Response
{
  "count": 3, "succeeded": 2, "failed": 1,
  "results": [
    { "index": 0, "ok": true, "data": { "name": "Widget A", "price": 19.99 },
      "valid": true, "repairAttempts": 0, "confidence": { "/name": 1, "/price": 1 } },
    { "index": 1, "ok": false,
      "error": { "code": "upstream_timeout", "message": "Upstream did not respond within 15000 ms." } },
    { "index": 2, "ok": true, "data": { "name": "Widget C", "price": 4.99 }, "valid": true }
  ],
  "usage": { "inputTokens": 5203, "outputTokens": 611, "modelCalls": 3 }
}

POST /classify $0.006 / call

Cheap, high-volume text classification against your own label set. Every label is scored, not just the winners, so you can threshold it yourself. multi:true allows several labels (and none). rubric defines what your labels mean and overrides the model's own reading of the names.

Request
{
  "text": "Charged twice for the same order and support has not replied in four days.",
  "labels": ["billing","bug","feature request","praise","churn risk"],
  "multi": true,
  "rubric": "churn risk: the customer signals they may leave or is visibly angry."
}
Response
{
  "labels": ["billing","churn risk"],
  "scores": { "billing": 0.96, "bug": 0.22, "feature request": 0.02,
              "praise": 0.01, "churn risk": 0.71 },
  "top": { "label": "billing", "score": 0.96 },
  "margin": 0.25,
  "rationale": "\"Charged twice for the same order\" is a billing fault; four days without a reply plus the tone signals churn risk."
}

How the guarantee is enforced

1. Schema checked first

Well-formedness, depth, size and evaluability are checked before the 402. A malformed schema is a free 400.

2. Forced structured output

Your schema becomes a tool's input_schema and the model is forced into it. No JSON scraped out of prose, ever.

3. Nullable escape hatch

The model sees your schema widened so null is always legal, so it can be honest instead of cornered into inventing a value.

4. Local validation

Output is validated against your original schema — draft 2020-12, per-pointer errors — inside the Worker.

5. Repair loop

On failure the exact errors go back to the model, in-conversation, and it repairs rather than restarting. Up to 3 attempts.

6. Fail, don't fudge

Still not conformant? 5xx, no settlement, and an error naming the pointers that could not be satisfied.

Safety

Payment

x402 v2, USDC on Base (eip155:8453). Point any x402-aware client at these URLs and payment is automatic; no account and no API key. Without a payment header you get 402 with a PAYMENT-REQUIRED challenge. Failed requests are never settled, so you only pay for a result you can use. See openapi.json and /limits.