RAG guide

Best AI models for RAG

This page ranks models that fit document-backed assistants and retrieval pipelines, with extra weight on tool support, structured output, and enough context to hold retrieved evidence.

Live signalTool support for retrieval and workflow integration
Live signalStructured output support for controlled downstream handling
Live signalEnough context to hold retrieved chunks and instructions
Live signalLive pricing for document-backed production traffic
Fast answer

Start with the live shortlist, then validate the route

Most RAG teams should compare tool support, context room, and structured output before raw benchmark scores, because those factors determine whether the model fits the actual retrieval stack.

Data source and freshness

Catalog-backed, not a static price sheet

TVP refreshes this page from the live OpenRouter model catalog. This render used 587 public model records and was synchronized Sep 11, 2026, 6:28 AM UTC. Pricing, availability, and context values can change.

Verify the exact route before sending production traffic, then use the linked model and provider pages as the source for current values. Read the TVP data methodology.

Shortlist

Top live candidates right now

x-ai

SpaceXAI: Grok 4.20

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Context2,000,000
Input$1.25
Output$2.50

2,000,000 token context, tool support, and structured output make it a better RAG candidate.

  • text
  • image
  • file
  • tools
  • structured
~deepseek

DeepSeek V4 Flash Latest

This model always redirects to the latest model in the DeepSeek V4 Flash family.

Context1,310,720
Input$0.04
Output$0.16

1,310,720 token context, tool support, and structured output make it a better RAG candidate.

  • text
  • tools
  • structured
deepseek

DeepSeek: DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

Context1,310,720
Input$0.065
Output$0.18

1,310,720 token context, tool support, and structured output make it a better RAG candidate.

  • text
  • tools
  • structured
~z-ai

Z.ai: GLM Flash Latest

This model always redirects to the latest model in the GLM Flash family.

Context1,310,720
Input$0.075
Output$0.25

1,310,720 token context, tool support, and structured output make it a better RAG candidate.

  • text
  • image
  • video
  • tools
  • structured
meta-llama

Meta: Llama 4 Scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

Context1,310,720
Input$0.1
Output$0.3

1,310,720 token context, tool support, and structured output make it a better RAG candidate.

  • text
  • image
  • tools
  • structured
FAQ

What buyers usually ask

What should RAG teams compare first?

Most RAG teams should compare tool support, context room, and structured output before raw benchmark scores, because those factors determine whether the model fits the actual retrieval stack.

Why does pricing matter in RAG?

RAG prompts are often large because they include retrieved context. That makes per-token cost a bigger operational factor than in short chat routes.

Next step

Use the guide, then validate the route in live TVP data.

TVP keeps the shortlist connected to the current catalog, provider coverage, and token pricing so buyers can move from research to routing without starting over.