Sample More, Reflect Less
put self-refine and reflexion (the loops where a model critiques and rewrites its own answer) up against plain repeated sampling, at equal token cost, on two math benchmarks with models from 1.5B to 7B. The fancy methods lose, and the gap widens as models get bigger, once every critique token gets counted.