Token Economics of LLM Extraction: A Cost Worksheet
Measure the real token cost of LLM extraction: tokens per page, per field, and per clean record, with a worksheet you can recompute on your own corpus.
On this page
Why Extraction Costs Differ from Chat: The Hidden Token Anatomy of a Prompt
Most cost estimates for LLM work assume a chat-shaped workload: short prompt, medium reply, done. Extraction is the opposite shape. The prompt is enormous — the entire page — and the reply is tiny. That inversion changes which levers matter.
Here is the anatomy of a typical extraction prompt for a store.example.com product page, with token counts measured with the same tokenizer the target model uses:
┌─────────────────────────────────────────────┬───────────────┐
│ Component │ Tokens │
├─────────────────────────────────────────────┼───────────────┤
│ System prompt ("You extract product data…") │ 180 │
│ Schema instructions (JSON Schema, 12 fields)│ 240 │
│ Few-shot example (1 filled record) │ 310 │
│ Page content (cleaned text) │ 1,900 │
│ User instruction + delimiters │ 45 │
├─────────────────────────────────────────────┼───────────────┤
│ TOTAL INPUT │ 2,675 │
└─────────────────────────────────────────────┴───────────────┘
Two things jump out. First, the page is 71% of the prompt. Everything you control at the prompt level — system text, schema, examples — is under 30%. Second, the fixed components get re-sent on every single request. That 575 tokens of scaffolding is pure overhead multiplied by your entire corpus.
Compare the three prompt shapes:
| Component | Chat prompt | Extraction prompt | Few-shot extraction prompt |
|---|---|---|---|
| System / instructions | ~40 tokens | ~180 tokens | ~180 tokens |
| Schema definition | 0 | ~240 tokens | ~240 tokens |
| Examples | 0 | 0 | ~310 tokens each |
| User content | ~50 tokens | ~1,900 tokens | ~1,900 tokens |
| Delimiters / wrappers | ~10 tokens | ~45 tokens | ~45 tokens |
| Total input | ~100 tokens | ~2,365 tokens | ~2,675 tokens |
| Typical output | ~300 tokens | ~120 tokens | ~120 tokens |
The chat prompt and the extraction prompt differ by a factor of 25 on input. And the output column is where extraction gets quietly favorable — you pay output prices (usually 3-5x input prices) on a small, structured reply rather than a long essay. This is why naive cost projections borrowed from chat benchmarks are wrong in both directions: people overestimate output cost and wildly underestimate input cost on large corpora.
Measuring Tokens per Page: A Reusable Counting Script
You cannot budget what you have not measured. Tokenizers differ between model families, so counting words or characters and dividing by four is a guess, not a measurement. Use the tokenizer that matches the model you will actually call.
This script fetches a page through a scraping API — once for raw HTML, once with boilerplate stripped — and reports token counts for both:
import requests
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # match your target model's tokenizer
API = "https://api.finedata.ai"
HEADERS = {"Authorization": "Bearer fd_your_api_key"}
def fetch(url, clean):
payload = {
"url": url,
"formats": ["text"] if clean else ["rawHtml"],
"only_main_content": clean,
}
r = requests.post(f"{API}/api/v1/scrape", json=payload, headers=HEADERS)
r.raise_for_status()
data = r.json()
key = "text" if clean else "rawHtml"
return data["data"][key]
def count_tokens(text):
return len(enc.encode(text))
url = "https://store.example.com/product/123"
raw = fetch(url, clean=False)
clean = fetch(url, clean=True)
print(f"raw HTML: {count_tokens(raw):>7,} tokens")
print(f"cleaned: {count_tokens(clean):>7,} tokens")
print(f"reduction: {1 - count_tokens(clean)/count_tokens(raw):.1%}")
Run this against a sample of your corpus, not one page. Page-to-page variance is the thing that kills budgets, because your cost model needs the distribution, not the mean. Here is a 10-page sample from a hypothetical example.com crawl:
| Page | Raw HTML tokens | Cleaned tokens | Reduction |
|---|---|---|---|
| /product/101 | 6,120 | 1,840 | 70% |
| /product/102 | 5,980 | 1,760 | 71% |
| /product/103 | 11,450 | 2,310 | 80% |
| /product/104 | 6,050 | 1,890 | 69% |
| /product/105 | 7,220 | 1,950 | 73% |
| /product/106 | 6,300 | 1,810 | 71% |
| /product/107 | 9,870 | 2,050 | 79% |
| /product/108 | 6,010 | 1,870 | 70% |
| /product/109 | 6,440 | 1,900 | 70% |
| /product/110 | 6,900 | 1,930 | 72% |
| Median | 6,470 | 1,895 | 71% |
Note page 103. It has 11,450 raw tokens — probably a review section or a comparison widget — and it drags the raw mean up 15% over the median. Cleaned, it is unremarkable. This is the argument for stripping boilerplate before extraction, and it is the single highest-leverage decision in the whole pipeline. If you want to go further and convert pages to markdown before they ever touch the model context, the approach in stop cleaning HTML by hand stacks nicely with token counting.
Tokens per Field: How Output Schema Width Drives Cost
Output tokens are the expensive ones. At a stated price of $15 per 1M output tokens, every field you extract is a recurring charge, and verbose field names cost more than terse ones. This is the part of extraction economics people skip entirely.
Here is a 12-field product schema next to a 4-field minimal one, each with a filled record:
// 12-field schema — ~240 tokens of schema instructions, ~120 tokens of output
{
"product_name": "Aurora Trail Runner",
"brand": "NorthPeak",
"price": 129.99,
"currency": "USD",
"availability": "in_stock",
"sku": "NPR-TR-042",
"color_options": ["slate", "sand", "moss"],
"size_range": "7-13",
"rating_average": 4.6,
"review_count": 312,
"shipping_days": 4,
"return_policy_days": 30
}
// 4-field schema — ~90 tokens of schema instructions, ~45 tokens of output
{
"name": "Aurora Trail Runner",
"price": 129.99,
"currency": "USD",
"availability": "in_stock"
}
The cost math per 1,000 records at $15 per 1M output tokens:
| Schema | Output tokens/record | Output cost per 1,000 records | Schema tokens in prompt (per request) |
|---|---|---|---|
| 12-field | ~120 | $1.80 | ~240 |
| 4-field | ~45 | $0.68 | ~90 |
| Delta | -62% | -$1.12 per 1,000 | -150 per request |
A dollar per thousand records sounds trivial until you are extracting two million records a month. Then it is $2,240 a month for fields you may not need. And the schema tokens compound it: the wide schema adds 150 input tokens to every request, which at $3 per 1M input tokens is another $0.45 per 1,000 records on top.
My rule: extract only what a downstream consumer reads. If nobody has queried return_policy_days in six months, it is not a field, it is a subscription fee you keep paying.
Retries, Validation Failures, and the Token Cost of Getting It Wrong
The sticker price assumes the model gets it right the first time. It will not. Schema violations, truncated JSON, hallucinated values that fail validation — each failure re-spends the full input tokens, because the retry sends the page again. This is the multiplier nobody puts in their spreadsheet.
Here is a bounded retry loop with the failure points annotated:
import json
from jsonschema import validate, ValidationError
MAX_RETRIES = 3
def extract_record(page_text, schema, client):
prompt = build_prompt(page_text, schema) # ~2,400 input tokens
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": prompt},
]
for attempt in range(MAX_RETRIES):
resp = client.chat.completions.create(
model=MODEL,
messages=messages,
response_format={"type": "json_object"},
)
# INPUT TOKENS RE-SPENT HERE on every loop iteration
# OUTPUT TOKENS RE-SPENT HERE too — the failed reply still bills
try:
record = json.loads(resp.choices[0].message.content)
validate(instance=record, schema=schema)
return record
except (json.JSONDecodeError, ValidationError) as e:
# TOKENS SPENT ON THE ERROR MESSAGE: small, but nonzero
messages.append({"role": "assistant",
"content": resp.choices[0].message.content})
messages.append({"role": "user",
"content": f"Your output failed validation: {e}. "
f"Return corrected JSON only."})
return None # page goes to the failure queue; all tokens already spent
Three token sinks hide in that loop: the re-sent page on every retry, the failed output itself (you pay output prices for invalid JSON), and the growing conversation as validation errors accumulate. Now the worked math, using our 2,365 input tokens and 120 output tokens per attempt at $3/$15 per 1M:
| Scenario | Attempts per accepted record | Input tokens | Output tokens | Effective cost per record |
|---|---|---|---|---|
| 0% failures | 1.00 | 2,365 | 120 | $0.0089 |
| 15% failures | 1.18 | 2,790 | 142 | $0.0105 |
| 30% failures | 1.43 | 3,382 | 172 | $0.0127 |
At a 30% failure rate, each clean record costs 43% more than the sticker price. That is not a rounding error. It is the difference between a pipeline that fits the budget and one that does not — and it is why measuring your real failure rate locally, before you model costs, matters so much. The method in audit field fill rates is the right companion here.
The Clean-Record Ratio: From Raw Pages to Usable Data
Every number so far feeds one metric. I call it the clean-record ratio, and it is the single best measure of extraction efficiency because it collapses input waste, output verbosity, retries, and acceptance into one figure:
CRR = (T_in × N_pages × R_retry + T_out × N_pages × R_retry)
─────────────────────────────────────────────────────
N_accepted
Where:
T_in = mean input tokens per attempt (page + prompt scaffolding)
T_out = mean output tokens per attempt
N_pages = pages submitted
R_retry = mean attempts per submitted page (1 / (1 - failure_rate))
— only valid when failures are independent and bounded
N_accepted = records that pass validation AND downstream quality checks
The denominator is the part people get wrong. A record that validates against the schema but has a hallucinated price is not a clean record. It is worse than a failure, because it silently corrupts your data while still costing you the tokens. Count only records you would actually ship.
Walkthrough on a 500-page example.com crawl with the 12-field schema:
- Mean input per attempt: 2,365 tokens (1,900 page + 465 scaffolding)
- Mean output per attempt: 120 tokens
- Failure rate: 15%, so R_retry = 1.18
- Submitted: 500 pages, attempts = 590
- Accepted after validation and spot-check: 460 records (92% acceptance)
Input spend: 2,365 × 590 = 1,395,350 tokens
Output spend: 120 × 590 = 70,800 tokens
CRR = 1,466,150 / 460 = 3,187 tokens per clean record
At $3/$15 per 1M, that is $0.0112 per clean record, or $5.15 for the crawl. The ratio is what you track over time. If it drifts from 3,187 to 4,500 with no schema change, something upstream broke — pages got heavier, a template changed, or the failure rate crept up. It works as a monitoring signal, not just a budget line.
The Cost Worksheet: A Fill-In Template for Your Corpus
Copy this, replace the placeholders, and recompute. Every derived cell has its formula stated so you can port it to a spreadsheet or a script:
| Row | Field | Value | Formula / source |
|---|---|---|---|
| A | Pages in corpus | 500 | your count |
| B | Mean input tokens per page (cleaned) | 1,900 | measured, tokenizer-matched |
| C | Prompt scaffolding tokens (system + schema + delimiters) | 465 | measured once |
| D | Mean input tokens per attempt | 2,365 | B + C |
| E | Mean output tokens per attempt | 120 | measured on sample |
| F | Failure rate (invalid or unparseable) | 0.15 | measured over N ≥ 200 attempts |
| G | Retry multiplier | 1.18 | 1 / (1 - F) |
| H | Acceptance rate (records you ship) | 0.92 | measured downstream |
| I | Accepted records | 460 | A × H |
| J | Price per 1M input tokens | $3.00 | your model’s pricing |
| K | Price per 1M output tokens | $15.00 | your model’s pricing |
| L | Total input tokens | 1,395,350 | D × A × G |
| M | Total output tokens | 70,800 | E × A × G |
| N | Input cost | $4.19 | L × J / 1e6 |
| O | Output cost | $1.06 | M × K / 1e6 |
| P | Total cost | $5.25 | N + O |
| Q | Cost per clean record | $0.0114 | P / I |
| R | Clean-record ratio (tokens) | 3,187 | (L + M) / I |
The filled row is the example.com walkthrough from the previous section, so you can verify your arithmetic against a known answer: 460 accepted records, $5.25 total, $0.0114 per record, 3,187 tokens per clean record. If your recomputation of row Q from rows P and I does not land on $0.0114, your formulas are off.
One honest caveat about row G: the 1/(1-F) formula assumes failures are independent. In practice failures cluster — a template change breaks 40 pages at once, and those pages fail all three retries. When failures cluster, your effective multiplier is worse than the formula predicts. Measure attempts per page empirically if you can; use the formula only as a first estimate.
Cost Levers Ranked: What Actually Moves the Number
Not all optimizations are equal. Here is my ranking by measured impact per unit of effort, using the worksheet numbers as the baseline:
| Rank | Lever | Token savings | Effort | Quality risk |
|---|---|---|---|---|
| 1 | Boilerplate stripping | ~70% of input | Low — one API flag or a readability pass | Low; occasionally strips real content |
| 2 | Schema narrowing | ~60% of output, ~150 input/request | Low — delete fields | None, if the field was unused |
| 3 | Prompt prefix caching | ~80-90% discount on scaffolding | Low — reorder prompt so fixed parts come first | None |
| 4 | Cheaper model tier | 50-90% on both prices | Medium — needs a quality eval first | Real; narrow schemas survive downgrades better |
| 5 | Few-shot example removal | ~310 input/request | Trivial | Medium; format adherence can degrade |
| 6 | Batched submission | No token savings; reduces overhead | Low | None |
Two opinions in here that people push back on. First, I rank boilerplate stripping above model downgrades, always. Downgrading the model to save money while still sending navigation bars and cookie banners into the context is paying a smart model to read garbage. Strip first, then decide if you still need the expensive model — most narrow-schema extraction tasks do not. Second, few-shot examples are overrated for schema-driven extraction. A tight JSON Schema plus one validation-retry loop usually matches the format adherence of few-shot prompting at 310 fewer input tokens per request.
Caching and batching in practice, using a batch submission against example.com pages with a narrow schema and a fixed, cache-friendly prompt prefix:
batch = {
"callback_url": "https://api.example.com/hooks/extraction-done",
"requests": [
{
"url": f"https://store.example.com/product/{pid}",
"formats": ["text"],
"only_main_content": True, # lever 1: strip boilerplate
"extract_schema": NARROW_SCHEMA, # lever 2: 4 fields, ~90 tokens
}
for pid in product_ids
],
}
requests.post(
"https://api.finedata.ai/api/v1/async/batch",
json=batch,
headers={"Authorization": "Bearer fd_your_api_key"},
)
The only_main_content flag handles lever 1 at the fetch layer, before tokens exist. The narrow schema caps both the schema-instruction tokens and the output tokens. On the model side, structure your prompt so the system text and schema sit at the front, identical across requests, with page content last — that is the layout prefix caching rewards. The scaffolding you re-sent 500 times in the worksheet drops to a discounted rate on 499 of them.
Wrap-Up
The unit economics of LLM extraction come down to four numbers you can measure this week: tokens per cleaned page, tokens per output record, the retry multiplier, and the acceptance rate. Multiply them out and you get the clean-record ratio — the one figure that tells you whether your pipeline costs a cent per record or a dime.
Measure before you optimize. The worksheet exists because every corpus is different; my 71% boilerplate reduction will not be yours, and your failure rate will not be my 15%. But the structure holds: strip the page, narrow the schema, cache the prefix, bound the retries, and count only records you would ship. Everything else is noise.
Related Articles
Forecast Scrape Spend: A Budget Model You Can Recompute
Turn success rates, retry multipliers, and rendering overhead into a scraping budget you can defend, with formulas and worked numbers you can recompute.
TechnicalWhat One Successful Page Actually Costs: A Unit Model
Build a per-page cost model for scraping: retries, rendering, proxy traffic, and CAPTCHA solving, with worked numbers you can recompute for your own budget.
TechnicalSuccess-Based vs Metered Scraping API Billing Models
Compare success-based and pay-per-request scraping API billing: how failed requests hit your budget, and how to model real cost per successful page.