Can AI actually find IDORs in real code?
We tested top coding agents against real-world apps—and the results were mixed. The models discovered genuine vulnerabilities, but also generated large numbers of false positives and inconsistent findings. By dissecting results across multiple authorization complexity levels, we show where LLMs shine, where they fail, and why IDORs remain a uniquely hard class of bugs for AI to reason about.
Expect real examples, surprising failure modes, and practical lessons for anyone considering AI as a security testing assistant.