Compare flagship and efficient AI models for journalism — context windows, indicative token pricing, and newsroom workflow tags. Sources refresh from OpenRouter and Groq.
Could not refresh Groq prices from groq.com/pricing (Could not parse any pricing rows from Groq pricing page (markup may have changed).). Groq rows use bundled catalog fallback until the next sync.
›
Showing 21 of 49 in this view
Live
Google: Gemini 3.7 Flash
API / route id
gemini-3.7-flash
Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance. Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware.
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context window.
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Best forMultimodal & rich formats · Long documents & transcripts · Structured extraction & fact packs · Summaries & briefs
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
Modality: text+image+file+audio+video->text. Inputs: text, image, video, file, audio. Outputs: text
Best forMultimodal & rich formats · Long documents & transcripts · Structured extraction & fact packs · Summaries & briefs
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Modality: text+image+file->text. Inputs: text, image, file. Outputs: text
Best forEngineering & automation · Multimodal & rich formats · Long documents & transcripts
Whisper Large v3 is OpenAI's most advanced and capable speech recognition model, delivering state-of-the-art accuracy across a wide range of audio conditions and languages. This flagship model excels at handling challenging audio scenarios including background noise, accents, and technical terminology.
LightOnOCR-2 is an efficient end-to-end 1B-parameter vision-language model for converting documents (PDFs, scans, images) into clean, naturally ordered text without relying on brittle pipelines.
Best forStructured extraction & fact packs · Long documents & transcripts
Claude Fable 5 is state-of-the-art on nearly all tested benchmarks of AI capability, showing exceptional performance in software engineering, knowledge work, vision, scientific research, and many other areas. The longer and more complex the task, the larger Fable 5’s lead over our other models. Modality: text+image+file->text. Inputs: text, image, file. Outputs: text
Best forStructured extraction & fact packs · Long documents & transcripts · Multimodal & rich formats · Investigations & deep research · Engineering & automation
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels. Modality: text+image+file->text. Inputs: text, image, file. Outputs: text
Mistral OCR 4 is a document-intelligence model that turns PDFs and other files into structured, richly annotated content instead of plain text. It extracts text with bounding boxes, block types, and confidence scores, enabling better RAG chunking, agentic workflows (like form filling or invoice processing), and robust ingestion/search pipelines across 170 languages.
Best forStructured extraction & fact packs · Tagging, routing & metadata · Long documents & transcripts · Investigations & deep research