GPT-5.6 Luna is my new favorite model for Ladder Logic
After I published my last model bake-off, Gareth S. asked why I stopped after one answer. His comment inspired this eval. I gave five models up to five tries at each of 60 Ladder Logic exercises. After every failed try, the model saw the real compiler error or failed tests and could correct its answer. I wanted to see which model could reach working Ladder Logic, not only which model got the first answer right.
