SIMIT: Self-Improving Vision-Language
Models via Imagination at Test-Time

Mehmet Onurcan Kaya1,2 Desmond Elliott3,2 Dim P. Papadopoulos1,2

1Technical University of Denmark 2Pioneer Center for AI 3University of Copenhagen

TL;DR Before answering an unlabeled test question, a vision-language model imagines similar practice problems: it writes each question and its answer first, then creates an image that fits. It learns from them and then answers. No labels. No external models.

+7.20%
mean relative gain over BAGEL-7B on 17 benchmarks
+6.98%
training-free, with in-context learning only (SIMIT-ICL)
16 / 17
benchmarks improved by SIMIT-ICL, and one tie
0
labels, teacher models or external verifiers

The key idea

Don't learn from your own guesses.
Learn from problems you wrote yourself.

Test-time self-improvement usually learns from the model's own answers, such as majority votes or self-judgments. When the model is wrong, it teaches itself to stay wrong. SIMIT turns this around. The model writes a new question, decides the answer first, and only then creates an image that makes the answer true. Solving an unfamiliar question is hard. Writing a practice problem whose answer you already know is much easier.

A blurry photo of a long-haired brown cat on a person's lap next to a laptop keyboard.
Unlabeled test query ยท VizWiz-VQA

Q: What color is this cat?

BAGEL-7B, zero-shot
Unanswerable

Learning from its own guesses

Consensus-based self-training, e.g. TTRL

  1. 1
    Sample many answers to the same query
    UnanswerableUnanswerableBrownUnanswerableUnanswerableTanUnanswerable
  2. 2
    The majority vote becomes the pseudo-label
    “Unanswerable”
  3. 3
    Rewarding agreement with it reinforces the mistake

On VizWiz, “unanswerable” is TTRL's majority vote in 99.2% of training steps.

Sampled answers are illustrative.

Imagining practice problems

SIMIT: the same model, no labels

Q: Is the cat sitting or lying down?

A: Sitting

Imagined image: a fluffy brown cat sitting on a person's lap in front of a keyboard.
verified

Q: How many legs does the cat have?

A: Four

Imagined image: a fluffy brown cat sitting upright on a person's lap next to a keyboard.
verified

Answer fixed first, image generated to match it.

Learn from them in context, then answer: Brown

A model does not need to solve the test query correctly to write useful practice problems, because it controls the answers it writes.

How it works

One model plays every role

SIMIT runs on a unified multimodal model (BAGEL-7B) that can both understand and generate images. No teacher, retriever or verifier is used. The only other components are deterministic renderers, which draw whatever the model specifies.

  • Synthesizerwrites question, answer and image-description triplets
  • Routerdecides how each image should be made
  • Artistpaints natural images natively
  • Architectwrites specs that renderers draw exactly
  • Criticchecks that each image supports its answer
  • Solverlearns from the samples, then answers

Swipe the figure sideways to explore it.

SIMIT's data synthesis pipeline (Fig. 2 of the paper): adaptive budget allocation, triplet synthesis, visualization routing, image realization, verification and difficulty filtering.

Four ways to learn from imagined data

in-context+6.98%

SIMIT-ICL

Imagined samples become few-shot demonstrations. No weight access needed, so it also fits API-only models.

in-weight+1.69%

SIMIT-FT single-query

Fine-tune on one query's samples, answer it, then reset. Nothing carries over between queries.

in-weight+6.39%

SIMIT-FT test-set

Pool the samples from every test query and fine-tune once, so practice for one query helps the others.

both+7.20%

SIMIT-FT → ICL

Fine-tune on the pooled samples, then also show each query its own samples in context. Best overall.

Mean relative gain over zero-shot BAGEL-7B across 17 benchmarks.

Beyond natural images

Imagining charts, documents and diagrams

Unified models can paint photos, but they garble text, numbers and structure. For structured visuals, SIMIT gives the model a skill library. The model writes a compact specification, and a deterministic renderer draws it exactly. Drag the slider to compare.

BAGEL native generation with skill library

One pipeline, many kinds of images

Click a render to see the prompt that produced it.

Try it

Run SIMIT on your own images

Ask any question about any image in the online demo, or add SIMIT to your own code with one pip install.

Online demo

Upload an image, ask a question, and compare the base model's answer with SIMIT's self-improved one. Choose BAGEL-7B, Lance or Qwen3.8-27B. Free on Hugging Face ZeroGPU.

The SIMIT demo: for a photo of a woman with a conical hat, the base model answers Unanswerable, SIMIT answers Vietnam.

Python library

Imagine practice problems and answer with them in a few lines. Works with BAGEL, Lance and any Hugging Face vision-language model.

pip install simit
from simit import SIMIT

model = SIMIT.from_pretrained(
    "ByteDance-Seed/BAGEL-7B-MoT")
# imagine practice problems, then use them
demos = model.imagine(image, question)
answer = model.answer(image, question, demos)

Results

Beats test-time scaling and label-free RL

Compared with test-time scaling and label-free reinforcement learning on the same BAGEL-7B model, across 17 benchmarks in five domains. Extra compute does not explain the gap. SIMIT-ICL uses 25× the zero-shot runtime, while label-free RL methods that run longer reach at most +2.07%.

Mean relative gain over zero-shot BAGEL-7B

Each method at its best configuration. The right column shows end-to-end runtime relative to zero-shot. Hover a row for per-domain scores.

Show as table (Table 1 of the paper)

Gains per benchmark

Relative improvement over BAGEL-7B. Hover a bar for raw scores.

Show as table
+53.9%NoCaps CIDEr with SIMIT-ICL (0.755 → 1.162)
+44.0%OK-VQA accuracy with test-set SIMIT-FT (37.6 → 54.2)
+2.21% / +3.97%SIMIT-ICL on two other unified models, SenseNova-U1 and Lance

Why does it work?

It fixes mistakes instead of reinforcing them

Consensus rewards mostly sharpen answers the majority already gets right. Imagined supervision is created independently of the model's prediction on the test query, so it can correct errors.

Fixes more errors, breaks no more

Test queries are split by whether Self-Consistency already answers correctly. The chart shows how often each method moves them.

correct → wrong wrong → correct Δ (pp)

Cleaner supervision where the model is wrong

Supervision each method encourages on queries the base model gets wrong, as judged by a blind external evaluator.

correct none incorrect δ (pp)

SIMIT corrects 14.4% of initially wrong answers while breaking only 2.0% of correct ones. It is the only method whose supervision on hard queries is more often correct than incorrect. Its overall supervision-error rate is about 12%, against 33–38% for label-free RL. δ is the gap in supervision-error rate between queries the base model gets wrong and those it gets right. A large positive δ means errors pile up exactly where the model is already weak.

Explore

What did the model imagine?

22 real test queries from the paper (Fig. 4 and Appendix N). For each one, compare the practice problems SIMIT imagined with the human-labeled neighbors that retrieval baselines use, and see who answered correctly. All results are precomputed, so no model runs in your browser.

RICES and TTT-NN retrieve nearest neighbors from a labeled pool, which SIMIT never uses. For compact display these examples disable adaptive budgeting and difficulty filtering and use K = 2 samples. A ✓ follows the paper's per-example judgment.

Citation

BibTeX

@article{kaya2026simit,
  title   = {{SIMIT}: Self-Improving Vision-Language Models via Imagination at Test-Time},
  author  = {Kaya, Mehmet Onurcan and Elliott, Desmond and Papadopoulos, Dim P.},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}