I

Inception: mercury-2.5:free

mercury-2.5:free
260K kontekst65.5K wyjścieRozumowanieNarzędziaRównoległe narzędziaCacheStrukturalneWiad. systemowa
Wydano Sep 8, 2026Zaktualizowano Sep 14, 2026

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs. Mercury 2 is 5x+ faster than leading speed-optimized LLMs like Claude 4.5 Haiku and GPT 5 Mini, at a fraction of the cost. Mercury 2 supports tunable reasoning levels, 128K context, native tool use, and schema-aligned JSON output. Built for coding workflows where latency compounds, real-time voice/search, and agent loops. OpenAI API compatible. Read more in the blog post.

Tryb chatTokenizer Other

Cennik

Cena wejścia
$0.00/ 1 mln tokenów
Cena wyjścia
$0.00/ 1 mln tokenów
Okno kontekstu 128K tokenówKompatybilne endpointy openaiDostawca Inception

Czas dostępności

Wydajność

Ładowanie danych wydajności...

Użycie i ranking

Ładowanie użycia...

Obsługiwane parametry

Wszyscy dostawcy = obsługiwany przez każdy upstream serwujący ten model. Niektórzy dostawcy = zależy od upstreamu obsługującego żądanie. Domyślnie = wartość wysyłana, gdy nic nie ustawisz.

ParametrDostawcyDomyślne
include_reasoningWszyscy dostawcy-
max_tokensWszyscy dostawcy-
reasoningWszyscy dostawcy-
reasoning_effortWszyscy dostawcy-
response_formatWszyscy dostawcy-
stopWszyscy dostawcy-
structured_outputsWszyscy dostawcy-
temperatureWszyscy dostawcy-
tool_choiceWszyscy dostawcy-
toolsWszyscy dostawcy-

Często zadawane pytania

Podobne modele