Strong small GPT for coding subagents, quick tool use, and high-volume work. gpt-5.4-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Turn visual inputs into useful answers
Explore vision and multimodal models that interpret images and answer questions about visual content. Compare supported inputs, capabilities, and API prices for your application.
Featured models include gpt-5.4-mini, gpt-5.6-terra, and gpt-5.6-luna, ranked by live weekly usage when available.
Vision & Multimodal Models
Balanced GPT-5.6 model for capable, cost-efficient everyday work. gpt-5.6-terra by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Cost-efficient GPT-5.6 model for fast, high-volume workloads. gpt-5.6-luna by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
The dependable workhorse of the GPT line — strong general reasoning, reliable structured output and first-class tool calling, priced so you can put it on the hot path of a production app. GPT-5.6 Sol by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Strongest Claude Opus model for coding, agents, and professional work. claude-opus-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Default frontier GPT for coding, computer use, research, and knowledge work. gpt-5.5 by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Claude model for creative writing, analysis, and controlled agent workflows. claude-fable-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents. claude-opus-4-8 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Everyday Claude agent model for coding, planning, browsing, and general work. claude-sonnet-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
High-end Claude for difficult coding, planning, and slower expert reasoning. claude-opus-4-6 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Stronger Opus tier for advanced software work and high-stakes reasoning. claude-opus-4-7 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Claude workhorse for coding agents, careful analysis, and production cost control. claude-sonnet-4-6 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Fast Gemini workhorse for multimodal apps where latency and price matter. gemini-2.5-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents. gemini-2.5-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Google's proven reasoning model for coding, math, and multimodal analysis. gemini-2.5-pro by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs. gemini-3-flash-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Low-latency Gemini model for high-volume multimodal and agent workloads. gemini-3.1-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Reasoning-first Gemini preview for agentic coding and complex problem solving. gemini-3.1-pro-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Advanced Gemini model for complex reasoning, coding, and multimodal analysis. gemini-3.1-pro-preview-customtools by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.5-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.5-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.6-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
gemini-robotics-er-1.6-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Affordable GPT-4.1 lane for fast coding help and structured extraction. gpt-4.1-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Omni-era GPT for multimodal chat, practical coding, and general assistants. gpt-4o by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Small omni GPT for cheap multimodal assistance and production-scale traffic. gpt-4o-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Small GPT-5 for responsive agents, coding help, and everyday automation. gpt-5-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Agent-ready GPT for coding and computer-use workflows at a lower cost. gpt-5.4 by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation. gpt-5.4-nano by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.