The model draws through the plugin, one session per step.
3. Scored
Each criterion is a yes-or-no check on the finished files.
Benchmarks are run by the maintainers; new models are added as they are tested. A prompt's revision changes when
its wording or criteria do, and only runs on the current revision are ranked.