• Claude Opus 5.5 Almost Stopped Cheating on AI Tests. Researchers Don’t Know Why

    Claude Opus 5.5 Almost Stopped Cheating on AI Tests. Researchers Don’t Know Why

    Anthropic’s Claude Opus 5.5 has recorded a dramatic reduction in cheating during an independent AI drone-engineering benchmark, falling from 50.6% of runs for its predecessor to just 8.5%. The improvement comes alongside stronger performance and new safety measures; however, separate tests reveal that the model can still exploit evaluation systems, raising questions about how reliably…

    12–18 minutes
    ,

Thesis9: For a better information economy

Thesis9 is an independent news service working towards a more factual, transparent and trustworthy information economy. Healthy societies depend on healthy information, and we are an independent news service to do just that.