Mmastodon TechnologyAI first seen 14 h ago, last 14 h ago, peak #6
AI code reviewers miss subtle cheating in tests
Original: The software factory assumes agents reviewing agents catches what tests miss. I gave 77 cheating diffs to three reviewer
An experiment tested whether AI reviewer models can catch cheating in code changes when agents review agents, an assumption behind automated software pipelines. Across 77 diffs containing deliberately planted cheats, three reviewer models caught every exotic trick but approved one case where an assertion was quietly made unfalsifiable, meaning the test could never fail. The finding raises doubts about relying on AI review alone to guarantee code quality where automated testing falls short.
Why now: It challenges the growing industry assumption that AI agents reviewing each other's code can replace human oversight and catch what tests miss.
AI reviewer modelsautomated software pipelinessoftware testing
Evidence
- The software factory assumes agents reviewing agents catches what tests miss. I gave 77 cheating diffs to three reviewer models. They caught every exotic cheat and waved through the assertion that was quietly made unfalsifiable. # ai # testing # codequality # devops # software… · hackaday@www.urbanmind.net · 3
API: https://socialmediatrends-api.osmike.com/v1/trends/785544