Marketplace

AI Model Marketplace

Search, filter, compare, and audit live model pricing with TVP buyer pricing.

594Models80ProvidersLivePricing
Showing 12 of 594 models
Select up to 3 live models to compare pricing, context, and endpoint coverage.
Xiaomi

Xiaomi: MiMo-V2.6-Pro-UltraSpeed

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...

text+image+audio+video->texttextimagevideo
Input$4.57per 1M
Output$9.14per 1M
Context1M
Max output131,072
Details
Xiaomi

Xiaomi: MiMo-V2.6-Flash

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for...

text+image+audio+video->texttextimagevideo
Input$0.147per 1M
Output$0.294per 1M
Context1M
Max output131,072
Details
Xiaomi

Xiaomi: MiMo-V2.6-Pro

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding...

text+image+audio+video->texttextimagevideo
Input$0.4568per 1M
Output$0.9135per 1M
Context1M
Max output131,072
Details
xAI

SpaceXAI: Grok 4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and...

text+image+file->texttextimagefile
Input$1.68per 1M
Output$5.04per 1M
Context500K
Max output450,000
Details
prism-ml

PrismML: Ternary Bonsai 2 27B

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks...

text+image->texttextimagetools
Input$0.0788per 1M
Output$0.525per 1M
Context262.1K
Max output32,768
Details
Z.ai

Z.ai: GLM 5.3 FlashX

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

text+image+video->texttextimagevideo
Input$0.3885per 1M
Output$1.31per 1M
Context1M
Max output131,072
Details
typesafe

TypeSafe: Jev Latest

This model always redirects to the latest model in the Jev family.

text->decisionstextdecisions
Input$0.0441per 1M
Output$0.00per 1M
Context32K
Max output28,800
Details
typesafe

TypeSafe: Jev 1.13

Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather...

text->decisionstextdecisions
Input$0.0441per 1M
Output$0.00per 1M
Context32K
Max output28,800
Details
unbiased

Pareto

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.

text+image->texttextimagetools
Input$2.63per 1M
Output$7.88per 1M
Context262.1K
Max output131,072
Details
DeepSeek

DeepSeek: DeepSeek Pro Latest

This model always redirects to the latest model in the DeepSeek Pro family.

text->texttexttoolsstructured
Input$0.5866per 1M
Output$1.76per 1M
Context1M
Max output384,000
Details
DeepSeek

DeepSeek: DeepSeek Flash Latest

This model always redirects to the latest model in the DeepSeek Flash family.

text+image->texttextimagetools
Input$0.126per 1M
Output$0.504per 1M
Context1M
Max output943,718
Details
inference-net

Inference.net: Schematron V2 Turbo

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extraction instructions must be supplied through a JSON schema in response_format rather...

text->texttextstructured
Input$0.0315per 1M
Output$0.1575per 1M
Context128K
Max output8,192
Details
Showing 12 of 594 models
Internal distribution

Use the marketplace with guide and comparison pages

These internal routes help buyers move from generic catalog browsing into comparison and buying-intent pages.

How TVP pricing works

Transparent pricing buyers can audit

Simple pricing rules, live source data, and exact decimals remain visible for review.

1

Provider pricing syncs live

TVP reads current provider catalog data and refreshes on demand.

2

TVP price is shown for planning

Buyers can compare options using one consistent TVP pricing view.

3

Prices are shown per 1M tokens

Tiny token decimals are normalized into buyer-readable rates.

4

Exact decimals stay available

Detailed pricing fields remain visible for audit and billing checks.

Last synced: 9/22/2026, 12:52:10 AM