If you work with AI coding tools every day, you know the difference between an imperfect result and a system that has stopped being dependable. An imperfect result can be reviewed and corrected. A system that repeatedly wanders off task, disregards the working context, or makes a simple job take days creates a much larger problem: it moves the work back onto the developer.
That is the frustration behind this open letter to OpenAI. It is not written by someone who dislikes the product. It is written by someone who still likes it enough to expect more from it.
Dear OpenAI: this is a complaint from someone who wants you to succeed
I am a longtime user, a fan of ChatGPT, and a fan of Codex. In my experience, Codex has been one of the strongest harnesses for AI-assisted development. That is precisely why the recent experience has been so disappointing.
The concern is not that every response is perfect. No serious developer expects that. The concern is that routine work can feel less reliable than it used to. Tasks that once felt manageable can become prolonged sessions of explaining, redirecting, checking, and explaining again.
As the original complaint puts it: “I spend days trying to get simple task done now the system refuses to stay on task.” That is the kind of failure that matters more to a working developer than whether the system can produce an impressive demonstration.
The love and the frustration are connected. If the tool were irrelevant, a bad result would be easy to ignore. When the tool is deeply integrated into a developer’s workflow, declining reliability becomes a direct cost in time, attention, and money.
The real issue is not novelty. It is dependable execution.
AI products are often judged by what they can do in a demonstration: generate an animation, produce a polished interface, or complete a difficult-looking task in one pass. Those capabilities may be valuable. But for people using AI to build and maintain real systems, the foundation comes first.
Can the system understand the assignment? Can it preserve the constraints? Can it use the documentation without treating documentation as a substitute for the actual objective? Can it recognize when it has gone out of bounds? Can it complete a bounded task without requiring a developer to supervise every step?
These are not glamorous questions. They are operational questions. They determine whether AI reduces work or merely changes the shape of the work.
A tool that needs constant babysitting can still be impressive. It is simply not as useful as it could be. The developer remains responsible for the result, but also becomes responsible for keeping the system oriented, correcting avoidable detours, and recovering from decisions that were never requested.
For teams working at scale, that difference compounds. A few minutes of correction on one task can become hours across a workflow. A small reliability problem can become a review bottleneck. A model that appears faster in a demo can be slower in production if it requires repeated intervention.
Upgrades should make developers’ lives easier
The basic expectation of an upgrade is straightforward: the experience should improve in ways users can feel while doing ordinary work.
That does not mean every new release must be better at every task. It does mean the overall direction should be clear. If a new version introduces impressive features but becomes less consistent at following instructions, users are entitled to ask whether the trade-off is worth it.
The original user describes a perceived pattern in which performance seems to worsen before a new release. That perception may have multiple explanations, and it would require product telemetry or independent testing to establish what is happening. But the experience itself should not be dismissed. When users feel that reliability is changing underneath them, they need clear communication and stable ways to work.
Removing or changing access to a previously useful model can intensify that problem. Users may have built workflows around particular behaviors, response patterns, or levels of task adherence. When those options disappear, the cost is not only emotional attachment to an old tool. It can include rework, retraining, prompt changes, and new review procedures.
For AI coding users, model choice is not just a preference. It can be part of the operating environment. When the environment changes, the people building on top of it need to know what changed, why it changed, and how to adapt without guessing.

Documentation should support the work, not replace the work
Good documentation matters. A coding system should understand the relevant rules, repository context, and technical constraints. But documentation is not the same thing as the assignment.
A system can follow a written instruction literally while missing the purpose of the task. It can cite a rule while failing to make the requested change. It can optimize for compliance with a narrow interpretation instead of producing a useful result within the stated boundaries.
That is where task adherence becomes more important than surface-level obedience. The goal is not to remove safeguards or encourage reckless automation. The goal is to make the system capable of balancing instructions, context, constraints, and the intended outcome.
A practical standard for an AI coding system should include:
- Clear recognition of the requested outcome.
- Consistent preservation of constraints and project context.
- A useful response when instructions conflict or remain ambiguous.
- Early disclosure when the system cannot complete the task reliably.
- Less unnecessary repetition from the developer.
- A reviewable record of what was changed and why.
These principles also matter beyond coding. Any governed business automation system has to distinguish between following a rule and achieving the approved business objective. Automation that is technically compliant but operationally unhelpful still creates work for someone else.

The human cost of unreliable AI is easy to underestimate
Developers who understand code can often identify when an AI system has gone off course. They can inspect the changes, challenge the reasoning, roll back a bad decision, and try again.
That does not make the problem harmless. It means the developer has become the recovery system.
Less experienced users may not recognize that the system has misunderstood the task. They may accept a plausible but incorrect result, lose confidence in their own project, or spend more time trying to persuade the tool than learning what the tool actually changed. The more confidently a system presents an out-of-bounds answer, the harder that problem becomes to detect.
This is why reliability is not merely a convenience feature. It affects who can use the system safely and how much oversight the work requires.
The frustration is especially sharp for people trying to build durable foundations. They are not asking only for a spectacle. They are trying to create software, processes, and businesses that other people can depend on. Flashier output cannot compensate for a weak foundation.
What should OpenAI focus on next?
The request is not to stop innovating. It is to make reliability a visible product priority.
That could mean clearer performance reporting across ordinary tasks, better continuity for users whose workflows depend on a model’s behavior, and more transparent explanations when capabilities or model access change. It could also mean evaluating releases against the work people actually do repeatedly, not only against difficult benchmark questions or compelling demonstrations.
For AI coding, the most meaningful improvements may be unglamorous:
- Stay on the assigned task.
- Preserve context across the work.
- Ask a useful question when the request is genuinely unclear.
- Avoid inventing progress or implying that a task is complete when it is not.
- Make failures easier to inspect and recover from.
- Reduce the amount of supervision required for bounded, repeatable work.
Those improvements would not make AI less ambitious. They would make ambitious systems more usable.
So: am I unreasonable?
Maybe some of this experience is specific to particular workflows, model configurations, or release conditions. That is exactly why users need ways to compare notes and describe what they are seeing without being dismissed as irrational or nostalgic.
The central complaint is simple: an AI coding tool should not make a capable developer spend days babysitting a task that used to be straightforward. If that is happening, the answer should not be to tell the developer that the tool is powerful and therefore the problem is acceptable.
OpenAI has built products that many developers genuinely want to use. That creates a higher standard, not a lower one. Users who care enough to complain may be identifying the gap between what these systems can demonstrate and what they can reliably do in the middle of real work.
Are you having the same problems? Tell us about them. Reach out on social, share what you are seeing, and compare experiences with other people who use AI coding tools regularly. Maybe the problem is specific to certain workflows. Maybe it is broader. Either way, the conversation should be grounded in concrete examples, not dismissed before it begins.
