Ortem Technologies
    AI & Machine Learning

    What RAG Actually Costs Per 1,000 Queries

    Praveen JhaAugust 28, 202610 min read
    What RAG Actually Costs Per 1,000 Queries
    Quick Answer

    At a 1 million chunk corpus, indexing costs about $10 one-time with text-embedding-3-small, vector storage runs $2.03 per month on Pinecone or $7.14 per month on Weaviate Flex, and 1,000 RAG queries cost $0.59 on gpt-4o-mini, $4.22 on Claude Haiku 4.5, $8.44 on Claude Sonnet 5 or $21.10 on Claude Opus 5. The generation model, not the vector database, is the dominant cost — a single day of moderate query traffic can exceed a month of storage.

    Commercial Expertise

    Need help with AI & Machine Learning?

    Ortem deploys dedicated AI & ML Engineering squads in 72 hours.

    Deploy Private AI

    Next Best Reads

    Continue your research on AI & Machine Learning

    These links are chosen to move readers from general education into service understanding, proof, and buying-context pages.

    RAG cost per thousand queries

    Most RAG cost discussions start with the vector database, because that is the component with a pricing page that looks like infrastructure. That turns out to be the wrong place to look.

    This is a cost model built from published vendor rates, with every input stated so you can substitute your own.


    Method

    Corpus shape. Chunks of roughly 500 tokens, embedded with text-embedding-3-small at 1,536 dimensions, stored at 4 bytes per dimension.

    Query shape. A 20-token question embedded, five chunks retrieved, and a generation call of roughly 2,720 input tokens (five chunks plus a 200-token system prompt plus the question) returning 300 output tokens.

    Prices. Taken from the vendors' own pricing pages in August 2026 and listed in the sources below. Vector-database rates vary by cloud and region; the lower published figure is used throughout.

    This is a model, not a measurement. Every number below is arithmetic on public list prices — reproduce it with your own chunk sizes and retrieval depth rather than adopting the totals.


    Indexing and storage

    CorpusIndexing (one-time)Pinecone storageWeaviate Flex storage
    10,000 chunks$0.10$0.02/mo$0.07/mo
    1,000,000 chunks$10.00$2.03/mo$7.14/mo
    10,000,000 chunks$100.00$20.28/mo$71.42/mo

    Indexing is a one-time charge at 500 tokens per chunk against text-embedding-3-small at $0.02 per million tokens. Pinecone storage is 6,144 bytes per vector at $0.33 per GB per month. Weaviate Flex is charged per million vector dimensions at $0.00465, before its $45 monthly minimum.

    Note the minimums. On Weaviate Flex a 10,000-chunk corpus bills at the $45 floor, not $0.07 — at small scale you are paying for the plan, not the vectors.


    Generation, per 1,000 queries

    ModelCost per 1,000 queriesPer query
    gpt-4o-mini$0.59$0.0006
    Claude Haiku 4.5$4.22$0.0042
    Claude Sonnet 5$8.44$0.0084
    gpt-4o$9.80$0.0098
    Claude Opus 5$21.10$0.0211

    Query embeddings add $0.0004 per 1,000 queries — four hundredths of a cent, and safely ignorable. Retrieval reads on Pinecone add roughly $0.08 per 1,000 queries assuming five read units per query, which is the one figure here that depends heavily on your index shape.


    The finding

    At one million chunks, a month of vector storage costs $2.03 on Pinecone. One thousand queries on Sonnet 5 costs $8.44.

    A production system serving 1,000 queries a day spends about $253 a month on generation against $2.03 on storage. The database is 0.8 percent of the bill.

    This inverts how most teams optimise. Migrating vector databases to save 60 percent of storage saves about $1.20 a month at this scale. Moving the same workload from Sonnet 5 to Haiku 4.5 saves $127 a month, and moving to gpt-4o-mini saves $235.


    Where the real savings are

    Model choice dominates everything else. The spread between the cheapest and most expensive model in the table is 36x. No infrastructure decision in a RAG stack has that leverage.

    Retrieval depth is a cost lever, not just a quality lever. Input tokens scale linearly with chunks retrieved. Dropping from five chunks to three cuts roughly 1,000 input tokens per query — about 20 percent off the Sonnet 5 line.

    Prompt caching applies to the stable part. Anthropic prices cache reads at 0.1x base input. A 200-token system prompt is small, but if your RAG prompt carries a large fixed preamble, caching it removes almost all of that cost after the first call.

    Batch pricing halves non-interactive work. Both Anthropic and OpenAI discount batch processing 50 percent. Any RAG workload that is not answering a user in real time — evaluation runs, backfills, bulk summarisation — should not be paying interactive rates.


    Substituting your own numbers

    The model is four multiplications. Indexing is chunks times tokens-per-chunk times the embedding rate. Storage is chunks times dimensions times the vendor rate. Query embedding is negligible. Generation is retrieved-tokens times the input rate plus output-tokens times the output rate.

    If your chunks are 1,000 tokens rather than 500, indexing and per-query input both double. If you retrieve ten chunks rather than five, generation input roughly doubles. Those two choices move the total far more than any vendor selection.


    Building a RAG system and want the cost model worked out against your actual corpus and query volume before you commit to a stack? Ortem Technologies' AI and ML solutions practice does this work. See our enterprise RAG implementation cost breakdown for the full build picture, or book a technical consultation →.

    About Ortem Technologies

    Ortem Technologies is a premier custom software, mobile app, and AI development company. We serve enterprise and startup clients across the USA, UK, Australia, Canada, and the Middle East. Our cross-industry expertise spans fintech, healthcare, and logistics, enabling us to deliver scalable, secure, and innovative digital solutions worldwide.

    📬

    Get the Ortem Tech Digest

    Monthly insights on AI, mobile, and software strategy - straight to your inbox. No spam, ever.

    RAG costvector database pricingPinecone pricingWeaviate pricingembedding costretrieval augmented generationAI unit economics

    Sources & References

    1. 1.Pinecone Pricing - Pinecone
    2. 2.Weaviate Cloud Pricing - Weaviate
    3. 3.Claude API Pricing - Anthropic
    4. 4.OpenAI API Pricing - OpenAI

    About the Author

    P
    Praveen Jha

    Director – AI Product Strategy, Development, Sales & Business Development, Ortem Technologies

    Praveen Jha is the Director of AI Product Strategy, Development, Sales & Business Development at Ortem Technologies. With deep expertise in technology consulting and enterprise sales, he helps businesses identify the right digital transformation strategies - from mobile and AI solutions to cloud-native platforms. He writes about technology adoption, business growth, and building software partnerships that deliver real ROI.

    Business DevelopmentTechnology ConsultingDigital Transformation
    LinkedIn

    Frequently Asked Questions

    Stay Ahead

    Get engineering insights in your inbox

    Practical guides on software development, AI, and cloud. No fluff — published when it's worth your time.

    Ready to Start Your Project?

    Let Ortem Technologies help you build innovative software solutions for your business.