What One AI Workflow Actually Costs: Unit Economics
Computed from published August 2026 rates: a support ticket conversation costs about 0.37 cents on Claude Haiku 4.5, a single-page invoice extraction costs 0.33 cents on Haiku 4.5 or 0.66 cents on Sonnet 5, and a 30,000-token contract review costs 8 cents on Sonnet 5 or 20 cents on Opus 5. Batch processing halves all of these. The useful budgeting unit is cost per completed task, not cost per million tokens.
Commercial Expertise
Need help with AI & Machine Learning?
Ortem deploys dedicated AI & ML Engineering squads in 72 hours.
Next Best Reads
Continue your research on AI & Machine Learning
These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.
AI & ML Solutions
Move from concept articles to real implementation planning for copilots, RAG, automation, and analytics.
Explore AI servicesAI Agent Development
See how Ortem builds autonomous workflows, tool-using agents, and human-in-the-loop systems.
View agent serviceAI Product Case Study
Study a production AI platform with architecture, launch scope, and operating model context.
Read case studyVendor pricing pages quote dollars per million tokens. No operations budget is denominated in tokens. The question a finance team actually asks is what it costs to process one invoice, resolve one ticket, or review one contract.
This converts current published rates into that unit.
Method
Every figure below is arithmetic on list prices published by Anthropic and OpenAI, read in August 2026 and linked in the sources. Token counts per task are stated assumptions, not measurements — substitute your own and the arithmetic holds.
Cost for any task is input tokens times the input rate, plus output tokens times the output rate. That is the whole model.
Current rates used
| Model | Input per 1M | Output per 1M |
|---|---|---|
| Claude Haiku 4.5 | $1.00 | $5.00 |
| Claude Sonnet 5 | $2.00 | $10.00 |
| Claude Opus 5 | $5.00 | $25.00 |
| gpt-4o-mini | $0.15 | $0.60 |
| gpt-4o | $2.50 | $10.00 |
Anthropic discounts batch processing 50 percent on both input and output, and prices prompt cache reads at 0.1x base input.
Cost per completed task
| Workflow | Assumed tokens | Model | Cost per task | Batch (50%) |
|---|---|---|---|---|
| Support ticket conversation | ~3,700 total | Haiku 4.5 | 0.37 cents | 0.19 cents |
| Invoice extraction, 1 page | 1,800 in / 300 out | Haiku 4.5 | 0.33 cents | 0.17 cents |
| Invoice extraction, 1 page | 1,800 in / 300 out | Sonnet 5 | 0.66 cents | 0.33 cents |
| Contract review, 30k-token doc | 30,000 in / 2,000 out | Sonnet 5 | 8.00 cents | 4.00 cents |
| Contract review, 30k-token doc | 30,000 in / 2,000 out | Opus 5 | 20.00 cents | 10.00 cents |
The support-ticket figure is Anthropic's own worked example: roughly 3,700 tokens per conversation on Haiku 4.5, stated as about $37 per 10,000 tickets. The others are computed from the rate table above using the stated token assumptions.
What the numbers actually say
Document length sets the order of magnitude. A contract review costs 24 times an invoice extraction on the same model. That gap is entirely input tokens. Before comparing vendors, ask whether the task genuinely needs the whole document in context — chunking, pre-filtering or routing only the relevant sections is a larger saving than any model switch.
Model choice is a 2–3x lever, not a 20x one. Within a workflow, moving between Haiku 4.5 and Sonnet 5 doubles cost; moving to Opus 5 triples it again. Real, but second-order against document length.
Batch halves everything that is not interactive. Invoice processing, contract review and overnight classification are all batchable. A workflow that never waits on a human should not be paying interactive rates.
At these unit costs, volume matters less than people expect. One hundred thousand invoices a month on Haiku 4.5 is $330, or $165 batched. The build and integration cost dominates the inference cost at almost any realistic volume.
Building your own unit cost
Count the tokens for one real example of the task — not an estimate, an actual run with the usage figures the API returns. Multiply by the published rates. Then multiply by monthly volume.
Two adjustments matter after that. If a stable system prompt or document preamble repeats across calls, prompt caching drops that portion to a tenth of input price after the first call. If the work is asynchronous, apply the batch discount.
What this model deliberately excludes: retries, failed extractions needing a second pass, human review time, and the engineering cost of building the pipeline. In most deployments the human review line is larger than the inference line, which is the strongest argument for measuring accuracy before optimising cost.
Working out whether an AI workflow pays for itself at your volumes, or building the pipeline that runs it? Ortem Technologies' AI and ML solutions practice works on both. See our LLM cost optimization guide for cutting an existing bill, or talk to our team →.
About Ortem Technologies
Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.
Get the Ortem Tech Digest
Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.
Sources & References
- 1.Claude API Pricing - Anthropic
- 2.OpenAI API Pricing - OpenAI
- 3.Customer Support Agent Guide - Anthropic
About the Author
Director – AI Product Strategy, Development, Sales & Business Development, Ortem Technologies
Praveen Jha is the Director of AI Product Strategy, Development, Sales & Business Development at Ortem Technologies. With deep expertise in technology consulting and enterprise sales, he helps businesses identify the right digital transformation strategies - from mobile and AI solutions to cloud-native platforms. He writes about technology adoption, business growth, and building software partnerships that deliver real ROI.
Frequently Asked Questions
- A single-page invoice extraction of roughly 1,800 input tokens and 300 output tokens costs about 0.33 cents on Claude Haiku 4.5 or 0.66 cents on Claude Sonnet 5, using published August 2026 rates. Batch processing halves both figures, to 0.17 and 0.33 cents respectively.
- Anthropic's published worked example puts a support conversation at roughly 3,700 tokens on Claude Haiku 4.5, costing about $37 per 10,000 tickets, or 0.37 cents per ticket. Batched, that falls to about 0.19 cents.
- A 30,000-token contract with a 2,000-token output costs about 8 cents on Claude Sonnet 5 or 20 cents on Claude Opus 5. Batch processing halves both, to 4 and 10 cents. Contract review costs roughly 24 times an invoice extraction on the same model, driven almost entirely by document length.
- Document length. Within the same model, a contract review costs 24x an invoice extraction, while moving between Haiku 4.5 and Sonnet 5 only doubles cost and Opus 5 roughly triples it again. Reducing how much of a document enters context saves more than any vendor switch.
Stay Ahead
Get engineering insights in your inbox
Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.
Ready to Start Your Project?
Let Ortem Technologies help you build innovative software solutions for your business.
You Might Also Like

AI App Development Cost in 2026: Real Numbers from Shipped Projects

AI Chatbot vs AI Agent: The Difference That Decides Your Budget in 2026

