A contact who runs an engineering org at a growing startup told a story on LinkedIn recently that raises a shared experience among many engineering leaders. His team had mostly standardized on an AI development tool. It was their most-loved tool because of how productive the team felt using it. Five months later they abandoned it. You may guess which tool that was, but there were no migration or sunsetting plans. No formal policy change prompted (groan) the exodus. Engineers just wandered off and switched as they found something they liked even better. The team ended up split across three AI tools, which is a common concern many engineering and security leaders are debating.
Whether your teams are using Codex, Claude Code, Devin, Cursor, VS Code, and more likely a combination of many of these developer experience tools including IDEs and CLIs, adoption across an organization just, sorta, happens. Any IC working on production code will choose the tool that boosts their productivity, which is a decision commonly driven by development teams, not security leadership.
Codex is where a lot of that shuffling has been landing even for Semgrep’s engineering team lately. Ask an engineer why they moved over and model quality isn't the first thing they bring up. What comes up is that they stopped reviewing every pull request the way they had in the past.
Codex Gets a Leg Up with Autonomy
One engineer on our team switched about six months ago, around the GPT-5.5 release, and now does 90% of their work through Codex. They weren't a devotee of another tool before this, so it isn't a brand-loyalty abandonment. They described the shift in blunt terms: they don't babysit as much for certain types of tasks. They tell it to go do the thing, and it does the thing. Opening PRs, merging them, checking AWS usage, querying Datadog, or upgrading billing for a new model. Six months later, they've barely touched a terminal. It's all one integrated GUI.
That's the decision worth noting. The pitch for Codex wasn't "better code" output. It's "less of your attention," and for many engineers, that's the more valuable resource.
Ask around and the answers cluster. Codex got comfortable running unsupervised before Claude Code did for some. By several accounts Opus 4.8 produces output just as good, but it wasn't yet trusted to merge its own PR with nobody watching. Another engineer who moved over put it more simply: it's just the preferred agent now so earns trust.
Some of that trust is structural, not just vibes.
Codex reads AGENTS.md, an open, de facto standard for giving an agent project context, instead of insisting on its own proprietary format. If you're running more than one harness — and per the fragmentation argument above, a lot of teams are — that matters more than it sounds. You write the context once and every tool reads it whether Codex, Pi, or OpenCode.
Price Sensitivity with AI Harnesses
Cost plays in too. Several engineers described running Codex at a volume they wouldn't attempt with Claude Code, simply because a failed attempt costs little enough (both runtime and tokens) to just retry it. One framed it as the cost of a unit of intelligence decreasing over time at a fixed point on the open coding benchmarks: not that Codex is smarter, but that the same output costs less to get. For a wide swath of day-to-day work, the harness that asks the least of you now has a credible claim on the highest-leverage work too.
The conversations get less rosy once you ask about limits. Usage caps can be a source of friction. One engineer described burning through 94% of a five-hour limit. We find ourselves in a post-"tokenmaxxing" era and evaluate the tokenomics of any development workflow relying on AI reasoning more carefully.
Standardizing Development Workflows to Catch Mistakes
When we asked what engineers actually wanted more of, the answer wasn't more time to experiment. It was something more boring: a shared, blessed team configuration. Sixty percent wanted a common setup, sixty percent wanted common workflows. Nobody was asking to tinker as much as they were a year ago. They were asking the organization to converge on one setup, so they could stop reinventing everything themselves.
Here's the part that doesn't come up unprompted, but should. If an engineer has gone from reviewing every PR to reviewing none of them, something still has to be watching and acting as a policy gate. It just may not be a person anymore. More autonomy doesn't lower the need for a check between the agent and production. It raises it, because the one check that used to happen by default, a human reading the diff, has been voluntarily removed.
That's the argument we made for Cursor last year, when Semgrep Guardian started calling itself automatically after every file edit, deterministically gating the task, instead of trusting the agent to remember.
As of this week, Codex gets the same treatment. OpenAI shipped a real hooks framework — PreToolUse, PostToolUse, Stop, SessionStart, and a dozen other lifecycle events, each reviewed and trusted individually before Codex will run it. Semgrep Guardian registers on PostToolUse, the same event it already hooks into for Claude Code, so every file Codex writes gets scanned the moment it's written, not whenever the agent remembers to ask.
The three harnesses had a common problem and two of them landed on the same specification. Cursor's hooks afterFileEdit and stop cover the AI-code generation lifecycle as Claude Code and Codex which standardized on a much larger lifecycle, down to near-identical event names. This says something about how fast "hooks" went from a Cursor innovation to something more reliable than an MCP server alone.
The part that matters more for a security team: Codex's hooks can be managed the same way Claude Code's can. Push Semgrep Guardian through enterprise Mobile Device Management (MDM), your cloud config, or a requirements.toml then get back to shipping. An individual contributor won’t disable it from their own harness because it acts seamlessly in the background of Codex, the same way it does in Cursor and Claude Code.
This is a real answer to why devs are choosing Codex, with deterministic security gates freeing them to pursue the paved road toward go do the thing.
If you're running Codex, Semgrep Guardian now scans what it writes the moment it's written.
To get started, run:
$ codex plugin marketplace add https://github.com/semgrep/guardian
$ codex plugin add semgrep@semgrep-marketplace
Visit the documentation for additional instructions on enterprise wide deployment.