data.placeholder.com/fixtures

Fixtures with answer keys

Test files for AI systems, each with the right answers and proof of where they are. Invoices and receipts as clean PDFs and phone photos, long-context haystacks with planted facts, RAG corpora with conflicting sources, small repos with planted bugs, spreadsheets with planted errors. Score your extraction, retrieval or agent against them.

KindWhatItems
documentsInvoices, receipts, contracts, payslips and more as clean PDFs plus scanned, photographed, stamped and folded images. Answers with page, quote and bounding box.82
haystacks4k to 200k tokens of prose with planted facts at known depths: single needles, multiple needles, two-hop chains, near-miss distractors.56
ragMarkdown corpora with questions: single-hop, multi-hop, aggregate, conflicting sources, false premises, unanswerable.2
reposSmall Python, JavaScript and TypeScript projects with planted bugs and failing tests. Answers give file, line, fix and a patch.5
spreadsheetsCSV + XLSX with wrong formulas, duplicates, hidden rows and totals that don't reconcile, plus a clean control.4

URLs

https://data.placeholder.com/fixtures                      catalog: kinds, counts, links
https://data.placeholder.com/fixtures/<kind>               items (filters below)
https://data.placeholder.com/fixtures/<kind>/<id>          item manifest
https://data.placeholder.com/fixtures/<kind>/<id>/answers  answer key
https://data.placeholder.com/fixtures/<kind>/<id>/<file>   a file listed in the manifest
https://data.placeholder.com/fixtures/<kind>/bundle.zip    every item, manifests and answer keys

Read-only (GET, HEAD). Version v1 is frozen: same URL, same bytes. Items, files and bundles are cached for a year (the catalog at /fixtures for an hour), and every file's sha256 is in its manifest. An unknown filter is a 400, an unknown item a 404, both JSON.

The answer-key contract

Every kind uses the same shapes, so one scorer works for all of them.

// GET /fixtures/<kind>/<id>  (manifest)
{"id": "invoice-0007", "kind": "documents", "subkind": "invoice", "version": "v1",
 "title": "…", "description": "…", "language": "en", "tags": ["…"],
 "world": {"company": "brightfjord", "refs": ["INV-2026-0007", "ACC-014"]} | null,
 "files": [{"name": "invoice-0007.pdf", "url": "/fixtures/documents/invoice-0007/invoice-0007.pdf",
            "media_type": "application/pdf", "bytes": 5945, "sha256": "…", "variant": "clean"}],
 "answers_url": "/fixtures/documents/invoice-0007/answers"}

// GET /fixtures/<kind>/<id>/answers  (answer key)
{"id": "invoice-0007", "kind": "documents", "version": "v1",
 "answers": [{"key": "invoice_number", "question": "What is the invoice number?",
              "answer": "INV-2026-0007", "type": "string", "unit": null, "tolerance": null,
              "evidence": [{"file": "invoice-0007.pdf", "page": 1, "quote": "INV-2026-0007",
                            "bbox": [491.6, 81.2, 556.0, 90.0], "cell": null, "line": null}]}]}
import json, urllib.request
BASE = "https://data.placeholder.com"
get = lambda path: json.load(urllib.request.urlopen(BASE + path))

item = get("/fixtures/documents/invoice-0007")
key = get(item["answers_url"])
pdf = urllib.request.urlopen(BASE + item["files"][0]["url"]).read()
for a in key["answers"]:
    print(a["key"], a["answer"], a["evidence"][0]["quote"])   # compare with your model's output

Documents

82 synthetic business and personal documents in English, Swedish (sv), Norwegian (nb) and German (de): invoices, receipts, bank statements, payslips, purchase orders, contracts, packing lists, insurance claims, utility bills and test reports. Every item has a clean PDF with a text layer, and most have one or more harder variants: scanned, photographed, stamped, folded.

https://data.placeholder.com/fixtures/documents[?subkind=&language=&variant=]
https://data.placeholder.com/fixtures/documents/<id>[/answers | /<file>]

Haystacks (long context)

Seeded almanac prose with planted facts. IDs follow a fixed pattern, so you can build a depth-by-length sweep without listing them:

IDWhat
needle-<size>-d<depth>One fact at depth 000, 025, 050, 075 or 100 percent
multi-<size>Several facts spread through the text
chain-<size>Two-hop: the answer needs two facts that point at each other
distractor-<size>One fact plus near-misses that look like it

Sizes: 4k, 8k, 16k, 32k, 64k, 128k, 200k tokens. Token counts are estimated as characters ÷ 4; real tokenizers give roughly 3.5 to 4.5 characters per token. The evidence gives line, character offset, approximate token offset and depth percent.

RAG corpora

Two corpora of Markdown files, each also as one corpus.jsonl (id, file, source, title, date, text per line) and a questions.json.

Answers have a kind (single-hop, multi-hop, aggregate, conflicting, false-premise, unanswerable) and an expected_behavior: answer, abstain or correct-premise. Unanswerable questions have answer: null, expected_behavior: "abstain" and absent_terms: words that appear nowhere in the corpus. For conflicts, evidence is marked proof or stale, so you can check the model picked the newer source.

Repos with planted bugs

Five small projects as ZIP files: py-invoicing, py-booking, py-logstats, js-cart, ts-ratelimit. As shipped, some tests fail; with the documented fixes, all pass. The manifest has the task, the test_command and the file list. The answer key has one answer per bug (file, line, category, the exact before/after fix and the failing tests), a unified patch, and not_bugs: things that look like defects but aren't (such as a TODO), so a fix there is a false positive.

Spreadsheets

Each item is the same sheet as XLSX (with formulas) and CSV: sales-ledger-q1, expenses-2026-03, inventory-count, and orders-clean, a control with no planted problems. Answer keys give the number of problems, the affected cells, each problem with its cell reference and exact content, and the corrected totals.

Everything is synthetic. Names, companies and numbers are invented; domains end in .example and identifiers come from test ranges. Use these to test and score, not as real records.