Choosing a vision model
Asha has a job called Vision: reading pictures. When another agent needs to understand an image — a screenshot, a design mock, a chart, a photo — the vision model is what looks at it. Your other jobs (talking, coding, writing) run on the models you chose for them.
What “can see” means
A model can only take this job if it accepts images as input. Text-only models cannot do it at all, however good they are at everything else. Asha reads that capability from each model’s own description, so the Vision list never offers you something that cannot see.
How to choose
- Free or paid? On a free plan only the free models will run; a paid model comes back with a credit error. Start free.
- What are you reading? Screenshots and UI mocks are easy — almost any vision model handles them. Dense documents, diagrams and handwriting reward a stronger model.
- Context is capacity, not quality. A larger context lets one request hold more (long screenshots, many pages). It says nothing about how well a model reads.
- Cost is per use, and small. Reading one screenshot costs a fraction of a cent on the cheap models below. You will notice the difference in quality long before you notice the bill.
Free — runs on a free plan today
| Model | Context | Price per 1M tokens (in / out) | Why it is here |
|---|---|---|---|
inclusionai/ling-3.0-flash-vl:free | 262k | Free | Good all-round reader — a sound default |
nex-agi/nex-n2.5-pro:free | 262k | Free | The stronger of the two Nex models |
nex-agi/nex-n2.5-mini:free | 262k | Free | Smaller, faster Nex |
google/gemma-4-31b-it:free | 262k | Free | Strong all-round vision |
google/gemma-4-26b-a4b-it:free | 262k | Free | Lighter Gemma |
qwen/qwen3.8-27b:free | 1M | Free | Very large context, slower |
dots-studio/dots-3-note-preview:free | 512k | Free | Reads documents and notes |
thinkingmachines/inkling:free | — | Free | General-purpose vision |
thinkingmachines/inkling-small:free | — | Free | Lighter Inkling |
Best value — needs a little credit
| Model | Context | Price per 1M tokens (in / out) | Why it is here |
|---|---|---|---|
qwen/qwen3-vl-32b-instruct | 131k | $0.10 / $0.42 | Best cheap dedicated vision model |
google/gemini-2.5-flash-lite | 1M | $0.10 / $0.40 | Cheapest reliable, huge context |
qwen/qwen3-vl-8b-instruct | 262k | $0.12 / $0.45 | Smallest paid vision model |
openai/gpt-4o-mini | 128k | $0.15 / $0.60 | Very consistent and well known |
google/gemini-2.5-flash | 1M | $0.30 / $2.50 | Best quality per dollar for screenshots |
anthropic/claude-haiku-4.5 | 200k | $1.00 / $5.00 | Careful reader, higher price |
Lowest cost of everything that can see
| Model | Context | Price per 1M tokens (in / out) | Why it is here |
|---|---|---|---|
openai/gpt-5-nano:batch | 400k | $0.02 / $0.20 | Cheapest of all |
qwen/qwen3.7-flash | 1M | $0.03 / $0.13 | Cheap, very large context |
google/gemma-3-4b-it | 131k | $0.05 / $0.10 | Tiny and cheap |
google/gemma-3-12b-it | 131k | $0.05 / $0.15 | Slightly larger Gemma |
google/gemini-2.5-flash-lite:batch | 1M | $0.05 / $0.20 | Batch pricing |
openai/gpt-4.1-nano:batch | 1M | $0.05 / $0.20 | Batch pricing |
openai/gpt-5-nano | 400k | $0.05 / $0.40 | Same model, pay-as-you-go |
inclusionai/ling-3.0-flash-vl | 131k | $0.06 / $0.18 | Paid twin of the free default |
How to set it in Asha
Open Settings → Models, choose the Vision row, pick your service if you have not already, then pick the model. The change takes effect immediately. Models marked Free run without credit.
Prices and context sizes are a snapshot taken 21 September 2026. Services change their catalogues constantly — Asha fetches the live list from the service itself, so your app will always show more than any page can. Prices are US dollars per one million tokens: input (what you send) / output (what it answers).