Turn visual inputs into useful answers

Explore vision and multimodal models that interpret images and answer questions about visual content. Compare supported inputs, capabilities, and API prices for your application.

Featured models include gpt-5.4-mini, gpt-5.6-terra, and gpt-5.6-luna, ranked by live weekly usage when available.

Vision & Multimodal Models

1.
3,036,362,789,400 weekly usage

Strong small GPT for coding subagents, quick tool use, and high-volume work. gpt-5.4-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

400K context$0.6 / 1M tokens
View model details
1,488,445,691,500 weekly usage

Balanced GPT-5.6 model for capable, cost-efficient everyday work. gpt-5.6-terra by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.6 / 1M tokens
View model details
3.
1,224,744,420,700 weekly usage

Cost-efficient GPT-5.6 model for fast, high-volume workloads. gpt-5.6-luna by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.16 / 1M tokens
View model details
4.
1,060,552,118,600 weekly usage

The dependable workhorse of the GPT line — strong general reasoning, reliable structured output and first-class tool calling, priced so you can put it on the hot path of a production app. GPT-5.6 Sol by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$3.2 / 1M tokens
View model details
5.
claude-opus-5

Anthropic

967,117,925,000 weekly usage

Strongest Claude Opus model for coding, agents, and professional work. claude-opus-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$4 / 1M tokens
View model details
6.
gpt-5.5

OpenAI

684,912,587,600 weekly usage

Default frontier GPT for coding, computer use, research, and knowledge work. gpt-5.5 by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$4 / 1M tokens
View model details
7.
343,468,212,700 weekly usage

Claude model for creative writing, analysis, and controlled agent workflows. claude-fable-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$9 / 1M tokens
View model details
8.
335,043,223,300 weekly usage

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents. claude-opus-4-8 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$4 / 1M tokens
View model details
9.
308,805,481,000 weekly usage

Everyday Claude agent model for coding, planning, browsing, and general work. claude-sonnet-5 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.6 / 1M tokens
View model details
10.
240,832,691,600 weekly usage

High-end Claude for difficult coding, planning, and slower expert reasoning. claude-opus-4-6 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$4 / 1M tokens
View model details
11.

Stronger Opus tier for advanced software work and high-stakes reasoning. claude-opus-4-7 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$4 / 1M tokens
View model details
12.

Claude workhorse for coding agents, careful analysis, and production cost control. claude-sonnet-4-6 by Anthropic is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$2.4 / 1M tokens
View model details

Fast Gemini workhorse for multimodal apps where latency and price matter. gemini-2.5-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.072 / 1M tokens
View model details

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents. gemini-2.5-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.072 / 1M tokens
View model details

Google's proven reasoning model for coding, math, and multimodal analysis. gemini-2.5-pro by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.9 / 1M tokens
View model details

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs. gemini-3-flash-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.056 / 1M tokens
View model details

Low-latency Gemini model for high-volume multimodal and agent workloads. gemini-3.1-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.2 / 1M tokens
View model details

Reasoning-first Gemini preview for agentic coding and complex problem solving. gemini-3.1-pro-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.6 / 1M tokens
View model details

Advanced Gemini model for complex reasoning, coding, and multimodal analysis. gemini-3.1-pro-preview-customtools by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.6 / 1M tokens
View model details

Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.5-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.2 / 1M tokens
View model details

Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.5-flash-lite by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.24 / 1M tokens
View model details

Fast Gemini model balancing multimodal reasoning, tool use, and cost. gemini-3.6-flash by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$1.2 / 1M tokens
View model details

gemini-robotics-er-1.6-preview by Google is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

128K context$0.24 / 1M tokens
View model details
24.

Affordable GPT-4.1 lane for fast coding help and structured extraction. gpt-4.1-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$0.32 / 1M tokens
View model details
25.
gpt-4o

OpenAI

Omni-era GPT for multimodal chat, practical coding, and general assistants. gpt-4o by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

128K context$2 / 1M tokens
View model details
26.

Small omni GPT for cheap multimodal assistance and production-scale traffic. gpt-4o-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

128K context$0.12 / 1M tokens
View model details
27.

Small GPT-5 for responsive agents, coding help, and everyday automation. gpt-5-mini by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

400K context$0.2 / 1M tokens
View model details
28.
gpt-5.4

OpenAI

Agent-ready GPT for coding and computer-use workflows at a lower cost. gpt-5.4 by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

1.0M context$2 / 1M tokens
View model details
29.

Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation. gpt-5.4-nano by OpenAI is available through the Flatkey unified API and is included in our Vision & Multimodal Models collection. Review its current pricing, context, and availability before integrating it.

400K context$0.16 / 1M tokens
View model details