Choosing a vision model

Asha has a job called Vision: reading pictures. When another agent needs to understand an image — a screenshot, a design mock, a chart, a photo — the vision model is what looks at it. Your other jobs (talking, coding, writing) run on the models you chose for them.

What “can see” means

A model can only take this job if it accepts images as input. Text-only models cannot do it at all, however good they are at everything else. Asha reads that capability from each model’s own description, so the Vision list never offers you something that cannot see.

How to choose

  1. Free or paid? On a free plan only the free models will run; a paid model comes back with a credit error. Start free.
  2. What are you reading? Screenshots and UI mocks are easy — almost any vision model handles them. Dense documents, diagrams and handwriting reward a stronger model.
  3. Context is capacity, not quality. A larger context lets one request hold more (long screenshots, many pages). It says nothing about how well a model reads.
  4. Cost is per use, and small. Reading one screenshot costs a fraction of a cent on the cheap models below. You will notice the difference in quality long before you notice the bill.

Free — runs on a free plan today

ModelContextPrice per 1M tokens
(in / out)
Why it is here
inclusionai/ling-3.0-flash-vl:free262kFreeGood all-round reader — a sound default
nex-agi/nex-n2.5-pro:free262kFreeThe stronger of the two Nex models
nex-agi/nex-n2.5-mini:free262kFreeSmaller, faster Nex
google/gemma-4-31b-it:free262kFreeStrong all-round vision
google/gemma-4-26b-a4b-it:free262kFreeLighter Gemma
qwen/qwen3.8-27b:free1MFreeVery large context, slower
dots-studio/dots-3-note-preview:free512kFreeReads documents and notes
thinkingmachines/inkling:freeFreeGeneral-purpose vision
thinkingmachines/inkling-small:freeFreeLighter Inkling

Best value — needs a little credit

ModelContextPrice per 1M tokens
(in / out)
Why it is here
qwen/qwen3-vl-32b-instruct131k$0.10 / $0.42Best cheap dedicated vision model
google/gemini-2.5-flash-lite1M$0.10 / $0.40Cheapest reliable, huge context
qwen/qwen3-vl-8b-instruct262k$0.12 / $0.45Smallest paid vision model
openai/gpt-4o-mini128k$0.15 / $0.60Very consistent and well known
google/gemini-2.5-flash1M$0.30 / $2.50Best quality per dollar for screenshots
anthropic/claude-haiku-4.5200k$1.00 / $5.00Careful reader, higher price

Lowest cost of everything that can see

ModelContextPrice per 1M tokens
(in / out)
Why it is here
openai/gpt-5-nano:batch400k$0.02 / $0.20Cheapest of all
qwen/qwen3.7-flash1M$0.03 / $0.13Cheap, very large context
google/gemma-3-4b-it131k$0.05 / $0.10Tiny and cheap
google/gemma-3-12b-it131k$0.05 / $0.15Slightly larger Gemma
google/gemini-2.5-flash-lite:batch1M$0.05 / $0.20Batch pricing
openai/gpt-4.1-nano:batch1M$0.05 / $0.20Batch pricing
openai/gpt-5-nano400k$0.05 / $0.40Same model, pay-as-you-go
inclusionai/ling-3.0-flash-vl131k$0.06 / $0.18Paid twin of the free default

How to set it in Asha

Open Settings → Models, choose the Vision row, pick your service if you have not already, then pick the model. The change takes effect immediately. Models marked Free run without credit.

Prices and context sizes are a snapshot taken 21 September 2026. Services change their catalogues constantly — Asha fetches the live list from the service itself, so your app will always show more than any page can. Prices are US dollars per one million tokens: input (what you send) / output (what it answers).