Compare flagship and efficient AI models for journalism — context windows, indicative token pricing, and newsroom workflow tags. Sources refresh from OpenRouter and Groq.
Could not refresh Groq prices from groq.com/pricing (Could not parse any pricing rows from Groq pricing page (markup may have changed).). Groq rows use bundled catalog fallback until the next sync.
›
Showing 21 of 53 in this view
Live
OpenAI: GPT-6 Astra
API / route id
gpt-6-astra
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon tasks that involve computer and browser use.
Modality: text+image+file->text. Inputs: file, image, text. Outputs: text
Best forDrafting & rewriting · Investigations & deep research · Structured extraction & fact packs · Multimodal & rich formats
Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Best forDrafting & rewriting · Investigations & deep research · Summaries & briefs · Multimodal & rich formats
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware.
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance. Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window.
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Best forMultimodal & rich formats · Long documents & transcripts · Structured extraction & fact packs · Summaries & briefs
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Best forMultimodal & rich formats · Long documents & transcripts · Structured extraction & fact packs · Summaries & briefs
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Modality: text+image+file->text. Inputs: text, image, file. Outputs: text
Best forEngineering & automation · Multimodal & rich formats · Long documents & transcripts
Whisper Large v3 is OpenAI's most advanced and capable speech recognition model, delivering state-of-the-art accuracy across a wide range of audio conditions and languages. This flagship model excels at handling challenging audio scenarios including background noise, accents, and technical terminology.
LightOnOCR-2 is an efficient end-to-end 1B-parameter vision-language model for converting documents (PDFs, scans, images) into clean, naturally ordered text without relying on brittle pipelines.
Best forStructured extraction & fact packs · Long documents & transcripts