Learning from its own guesses
Consensus-based self-training, e.g. TTRL
-
1
Sample many answers to the same queryUnanswerableUnanswerableBrownUnanswerableUnanswerableTanUnanswerable
-
2
The majority vote becomes the pseudo-label“Unanswerable”
-
3
Rewarding agreement with it reinforces the mistake
On VizWiz, “unanswerable” is TTRL's majority vote in 99.2% of training steps.
Sampled answers are illustrative.
Synthesizer
Router
Artist
Architect
Critic
Solver