At a 1 million chunk corpus, indexing costs about $10 one-time with text-embedding-3-small, vector storage runs $2.03 per month on Pinecone or $7.14 per month on Weaviate Flex, and 1,000 RAG queries cost $0.59 on gpt-4o-mini, $4.22 on Claude Haiku 4.5, $8.44 on Claude Sonnet 5 or $21.10 on Claude Opus 5. The generation model, not the vector database, is the dominant cost — a single day of moderate query traffic can exceed a month of storage.
Commercial Expertise
Need help with AI & Machine Learning?
Ortem deploys dedicated AI & ML Engineering squads in 72 hours.
Next Best Reads
Continue your research on AI & Machine Learning
These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.
AI & ML Solutions
Move from concept articles to real implementation planning for copilots, RAG, automation, and analytics.
Explore AI servicesAI Agent Development
See how Ortem builds autonomous workflows, tool-using agents, and human-in-the-loop systems.
View agent serviceAI Product Case Study
Study a production AI platform with architecture, launch scope, and operating model context.
Read case studyMost RAG cost discussions start with the vector database, because that is the component with a pricing page that looks like infrastructure. That turns out to be the wrong place to look.
This is a cost model built from published vendor rates, with every input stated so you can substitute your own.
Method
Corpus shape. Chunks of roughly 500 tokens, embedded with text-embedding-3-small at 1,536 dimensions, stored at 4 bytes per dimension.
Query shape. A 20-token question embedded, five chunks retrieved, and a generation call of roughly 2,720 input tokens (five chunks plus a 200-token system prompt plus the question) returning 300 output tokens.
Prices. Taken from the vendors' own pricing pages in August 2026 and listed in the sources below. Vector-database rates vary by cloud and region; the lower published figure is used throughout.
This is a model, not a measurement. Every number below is arithmetic on public list prices — reproduce it with your own chunk sizes and retrieval depth rather than adopting the totals.
Indexing and storage
| Corpus | Indexing (one-time) | Pinecone storage | Weaviate Flex storage |
|---|---|---|---|
| 10,000 chunks | $0.10 | $0.02/mo | $0.07/mo |
| 1,000,000 chunks | $10.00 | $2.03/mo | $7.14/mo |
| 10,000,000 chunks | $100.00 | $20.28/mo | $71.42/mo |
Indexing is a one-time charge at 500 tokens per chunk against text-embedding-3-small at $0.02 per million tokens. Pinecone storage is 6,144 bytes per vector at $0.33 per GB per month. Weaviate Flex is charged per million vector dimensions at $0.00465, before its $45 monthly minimum.
Note the minimums. On Weaviate Flex a 10,000-chunk corpus bills at the $45 floor, not $0.07 — at small scale you are paying for the plan, not the vectors.
Generation, per 1,000 queries
| Model | Cost per 1,000 queries | Per query |
|---|---|---|
| gpt-4o-mini | $0.59 | $0.0006 |
| Claude Haiku 4.5 | $4.22 | $0.0042 |
| Claude Sonnet 5 | $8.44 | $0.0084 |
| gpt-4o | $9.80 | $0.0098 |
| Claude Opus 5 | $21.10 | $0.0211 |
Query embeddings add $0.0004 per 1,000 queries — four hundredths of a cent, and safely ignorable. Retrieval reads on Pinecone add roughly $0.08 per 1,000 queries assuming five read units per query, which is the one figure here that depends heavily on your index shape.
The finding
At one million chunks, a month of vector storage costs $2.03 on Pinecone. One thousand queries on Sonnet 5 costs $8.44.
A production system serving 1,000 queries a day spends about $253 a month on generation against $2.03 on storage. The database is 0.8 percent of the bill.
This inverts how most teams optimise. Migrating vector databases to save 60 percent of storage saves about $1.20 a month at this scale. Moving the same workload from Sonnet 5 to Haiku 4.5 saves $127 a month, and moving to gpt-4o-mini saves $235.
Where the real savings are
Model choice dominates everything else. The spread between the cheapest and most expensive model in the table is 36x. No infrastructure decision in a RAG stack has that leverage.
Retrieval depth is a cost lever, not just a quality lever. Input tokens scale linearly with chunks retrieved. Dropping from five chunks to three cuts roughly 1,000 input tokens per query — about 20 percent off the Sonnet 5 line.
Prompt caching applies to the stable part. Anthropic prices cache reads at 0.1x base input. A 200-token system prompt is small, but if your RAG prompt carries a large fixed preamble, caching it removes almost all of that cost after the first call.
Batch pricing halves non-interactive work. Both Anthropic and OpenAI discount batch processing 50 percent. Any RAG workload that is not answering a user in real time — evaluation runs, backfills, bulk summarisation — should not be paying interactive rates.
Substituting your own numbers
The model is four multiplications. Indexing is chunks times tokens-per-chunk times the embedding rate. Storage is chunks times dimensions times the vendor rate. Query embedding is negligible. Generation is retrieved-tokens times the input rate plus output-tokens times the output rate.
If your chunks are 1,000 tokens rather than 500, indexing and per-query input both double. If you retrieve ten chunks rather than five, generation input roughly doubles. Those two choices move the total far more than any vendor selection.
Building a RAG system and want the cost model worked out against your actual corpus and query volume before you commit to a stack? Ortem Technologies' AI and ML solutions practice does this work. See our enterprise RAG implementation cost breakdown for the full build picture, or book a technical consultation →.
About Ortem Technologies
Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.
Get the Ortem Tech Digest
Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.
Sources & References
- 1.Pinecone Pricing - Pinecone
- 2.Weaviate Cloud Pricing - Weaviate
- 3.Claude API Pricing - Anthropic
- 4.OpenAI API Pricing - OpenAI
About the Author
Director – AI Product Strategy, Development, Sales & Business Development, Ortem Technologies
Praveen Jha is the Director of AI Product Strategy, Development, Sales & Business Development at Ortem Technologies. With deep expertise in technology consulting and enterprise sales, he helps businesses identify the right digital transformation strategies - from mobile and AI solutions to cloud-native platforms. He writes about technology adoption, business growth, and building software partnerships that deliver real ROI.
Frequently Asked Questions
- Using a 2,720-token input and 300-token output per query, 1,000 RAG queries cost $0.59 on gpt-4o-mini, $4.22 on Claude Haiku 4.5, $8.44 on Claude Sonnet 5, $9.80 on gpt-4o and $21.10 on Claude Opus 5, based on published August 2026 rates. Query embeddings add $0.0004 per 1,000 queries and retrieval reads add roughly $0.08.
- No. At a 1 million chunk corpus, Pinecone storage costs $2.03 per month while 1,000 queries on Claude Sonnet 5 cost $8.44. A system serving 1,000 queries a day spends about $253 per month on generation against $2.03 on storage, making the vector database roughly 0.8 percent of the bill. The generation model is the dominant cost.
- Indexing 1 million chunks of roughly 500 tokens each with text-embedding-3-small costs about $10 as a one-time charge, at $0.02 per million tokens. Storing those vectors at 1,536 dimensions costs $2.03 per month on Pinecone or $7.14 per month on Weaviate Flex, before Weaviate Flex's $45 monthly minimum.
- Change the generation model before touching infrastructure. The spread between the cheapest and most expensive model tested is 36x. Moving 1,000 daily queries from Claude Sonnet 5 to Haiku 4.5 saves about $127 a month, while switching vector databases to halve storage saves about $1. Reducing retrieval depth and enabling prompt caching and batch pricing are the next largest levers.
Stay Ahead
Get engineering insights in your inbox
Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.
Ready to Start Your Project?
Let Ortem Technologies help you build innovative software solutions for your business.
You Might Also Like
What One AI Workflow Actually Costs: Unit Economics

AI App Development Cost in 2026: Real Numbers from Shipped Projects

