资深工程师撰文反驳'编程已被AI解决'的叙事
资深工程师 Alex Ewerlof 发文反驳'编程已解决、工程只看品味'的说法,指出维护、可靠性、安全等 NFR 以及功能需求本身都未解决,AI 无法承担问责。文章列举 LLM 擅长的场景如 POC、个人软件、自然语言转换,并逐条批驳'spec 即代码''英语是新编程语言''Agent 是新编译器'等流行观点,还与 DHH 展开交锋,最后建议工程师保持理解与问责,警惕 AI 过度使用。
Disclaimer: you are about to read a lot of opinions, many of them have references but some are the result of my own experience building with AI and building AI systems in the past 4 years. Regardless, beware of the cognitive bias. Just because one argument doesn’t map to your belief system, it doesn’t mean the rest are invalid. I should also say upfront that I’m not anti-AI. If you’ve been following my work, you know that I was an early adopter of not only using LLM-powered coding tools, but building my own harness, teaching these topics and building LLM-powered products. It’s not about fear of AI but rather challenging the brain-dead narrative that asserts “coding is solved” and engineering is about “taste” now.
Update: someone put this on Hackernews where it went all the way to spot 2:
Tell me you don’t understand software without literally using those words!!!
People who claim “LLMs can write decent code” don’t understand how code works. Sure, creation is much cheaper, but anyone who has run software in production at scale knows that maintenance, reliability, security, scalability, etc. is the majority of the cost. These are commonly known as NFR (non-functional requirements).
In my experience even the Functional Requirements (what the code is supposed to do) is NOT a solved problem yet. There’s a bit of Dunning-Kruger effect at place where the people who don’t read the output are more confident in it.
As a veteran developer holding 2 engineering degrees (hardware and systems engineering), I can list 4 types of products that do not strictly require reading the code:
Personal software: scratching an itch, automation, DIY patches, etc.
POC (proof of concept): demonstrating technical feasibility and product viability
Throwaway automation: where the budget only allows validating the results. For example, reviewing images typically a trivial task for humans.
Weaponized AI: acknowledge the risk and deliberately point it at a target to cause harm
Notice the commonality: the first 3 have high risk tolerance while the last one weaponizes the inherent risk (and I'd argue given the blast radius of an agent that's connected to the internet, even the last one needs tight controls).
Most software that requires hiring and paying software engineers has low risk tolerance:
✅ healthcare
✅ finance
✅ automotive
✅ defense
✅ power plants
✅ aviation
✅ manufacturing
…wherever a mistake can cost money, lives or legal consequences you need accountability.
It is not binary, however.
The amount of time you spend understanding and grooming the code is a function of risk tolerance. And risk tolerance is a function of the business sector (as we mentioned) but also:
Scale: a small failure can cause large damage when there are too many consumers
Resilience: tight coupling can lead to cascading failure.
Containment: compartmentalization, sandboxes, isolation, bulkhead and many other patterns can contain the failure
Observability: the sooner we notice an error and the more data we have around it, the faster we can fix it and that reduces the cost of error
It cannot suffer any consequences. The worst thing you can do to AI is to unplug it. And although it mimics human emotions (due to training data), it couldn’t “care” less. AI doesn’t die either. It cannot suffer a prison sentence or fines. You cannot punish AI, therefore it can never be held accountable.
From legal stand point, AI is a tool that can only be held responsible. The human is accountable.
That’s why the US president says:
Our guardrail is the DOJ (department of justice) —Donald Trump
I’m fully aware of his unpopularity and personally I’m not a fan but he does have a point whether it’s spelled for him or he came up with the idea himself. Then again, this is the same person who renamed “Artificial Intelligence” to “Super Intelligence” (not to be mistaken with the well-established term Artificial Super Intelligence, or ASI).
You cannot be responsible for what you don’t understand. That understanding is key to reasoning about system behavior and fixing it when the AI inevitably fails.
If you’re toying around, LLMs do a great job. That’s why some of the most aggressive proponents of the “coding is solved” narrative have nothing to show for it. Anthropic accidentally leaked Claude Code (which on further study turned out to have many flaws) and their status page shows orange is the new green!
If you’re in management position, please act as leaders and listen to your engineers. If they care about quality, they’re your ticket to getting through “SaaSocalypse”, as some put it.
Contrary to common narrative, coding is actually one of the last areas for the current generation of LLMs to take over!!!
Allow me to elaborate:
Coding is about logic. Anyone who has dealt with compiler errors knows that computers don’t give a f*** about how right you think you are. If it’s logically wrong, it doesn’t compile. Even when the syntax is fine, there are runtime errors.
The reason LLMs are successful in writing code is because we’ve made a feedback loop that feeds the errors back to the LLM and loops until most errors are solved or hidden. Remember that LLMs can and do cheat too.
LLMs can wing it for tasks that are related to natural language (e.g. writing social media posts, reports, articles, etc.) but when it comes to code, the same engine that struggles to count number of R’s in “Raspberry” or suggests a walk to the carwash, also exposes other logical fallacies.
LLMs are stochastic and probabilistic. The only way we could even get remotely close to making them logical is to wrap them in traditional code (known as harness), run tests, and a bunch of other techniques (e.g. CoT) but the core issue remains: LLMs struggle with logic and volume (the larger the input and the more the context window is used, the less accurate they get).
I’m not saying LLMs cannot generate code or maintain existing code bases. They have their utility as a tool and their capabilities are increasing in an S-curve. There is a point of diminishing return where more expensive models aren’t necessarily more productive at the rate of the price increase.
Those who claim LLM-generated software is good enough:
❌ Haven’t written code in ages
❌ Cannot spot if their code figuratively had 6 fingers!
❌ Have a low bar for what good looks like
❌ Don’t care about quality or NFR
❌ Have difficulty understanding an S-curve
✅ Are honest: AI genuinely writes better code than them
But to go ahead and extrapolate that to an entire professional industry requires a level of brain-dead thinking that’s only present in people who spend too much time with sycophantic AI.
I’m not here to change anyone’s workflow or toolbox. I couldn’t care less.
But I’m tired of being a lab rat!
All I’m asking is to get your sh*t together and don’t ship half-a** products for the full price.
What I do care is that the services I’m paying for (looking at you Google and GitHub) are degrading with stupid bugs that could be avoided if we prioritize reliability and accountability over velocity.
Anthropic’s Boris Cherny is one of the most vocal proponents of the “coding is solved” narrative. By many accounts Cloude Code is the epiphany of his ideology:
Claude Code CLI binary installer silently deletes itself after installation
Extra Usage charged despite available plan capacity + false rate limit errors
Reminder: Anthropic controls the model (Claude), the harness (Claude Code), the prompt (see the leaked versions) and the runtime (Bun).
Also, if you think it’s an skill issue, this is an example JD from their site:
Surely they can deliver better quality for that kind of number. Then again, financially, these companies act as money furnace:
The only way they can pay back that kind of debt is to deliver what they say and reach a point where AI can genuinely reduce the cost of many jobs.
If you’re in leadership position, please don’t pressure your [otherwise smart] developers to force AI into every possible surface and workflow.
The tech has some genuine power and is the biggest change in our industry in ages. But AI overuse is a thing, and when it hurts the customer, you are accountable.
Stop repeating the half-baked narratives from token sellers about exaggerating the capabilities of AI because you can only tell a lie so many times before starting to believe in it.
It’s bad for business:
And it’s bad for consumers who are charged full price for partial and degraded service:
I would actually argue that a big part of the cost is not immediately visible:
Accumulating tech debt faster than it can be paid back
Creating more code, which provides more opportunities for security issues
Widening cognitive debt that reduces the engineer’s ability to reason about system behavior
Distressed engineers who increasingly choose to work solo and compete with their teammates instead of collaboration
There’s definitely an AI Dunning-Kruger effect at play where the less you know about the tech, the more confident you are in its utility. Please don’t.
The other day I saw a CTO proudly bragging about how a product person made a meaningful contribution while framing developers “suck in IDEs”:
There's a huge red flag that's missing from the post. This is the same argument as saying: our security folks are stuck in gym or martial arts tatami!
As AI gets more capable, it get easier than ever to fake credibility. In this example, the party who had the idea (head of product) and added code (agent) should ultimately be held accountable for any issues in the new chat interface that's bolted to the app.
Is head of product also on-call? Or is that the part that's delegated to engineers stuck in the IDE?
You should celebrate that they're comfortable with the code because code tells the truth more accurately than sycophantic AI.
If that skillset is NOT valuable, either remove PR review from your workflow or delegate it to your AI too.
I don’t want to belittle how far we have come with harness, SKILLS, AGENTS-md, MCP, A2A, ACP, RLM, OKF, MoE, MoA, state machines, self-evolving (e.g. Pi coding agent or /chronicles), self-healing, larger context windows, faster processors, more optimized runtimes, better quantizations, better training data, better caching, optimizations, architectures, orchestrations, and memory techniques.
I’ve written about many of those before:
Those are great pragmatic approaches to work around LLM shortcomings and there are probably more to come.
What I’m trying to elaborate is that I don’t want the services (that I depend on) to degrade just because someone pushed AI where it didn’t belong or skipped their job in quality, security, reliability and verification.
As of today, here are some of the good use cases for the current generation of AI:
✅ Mapping:
NL →NL: Translate human languages
Code → Code: Converting syntax from one programming language to another
Data → Code: Converting specs, JSON, YAML to code
Code → Data: parse code and convert it to data format (e.g. extract CRD from a Prometheus query)
Modality → Modality: e.g. transcribe audio to text (useful for Audio User Interfaces, or AUI) or describe image content (useful for GUI and computer use)
✅ Generation:
Image/Video/Music/Text: creativity (usually with diffusion model for non-text modality and auto-regressive models for text)
Expansion: make an article longer, generate a video from a reference image, create audio files from a reference voice signature, etc.
Cohesion: for example, give some scattered points to an LLM to create a cohesive email or article
Prediction: given a dataset, predict the next data set
✅ Reduction:
Summarization: taking a long article/conversation/transcript and extracting key points (optionally guided by a criteria)
Conversion: generate an SVG from a PNG file (usually with the help of a feedback loop and rendering engine)
Extraction: find relevant information in unstructured data (e.g. logs, NoSQL data, or anywhere writing a structured query is more time consuming and less critical than the AI output)
Classification: given a series of choices, estimate the probability of their likelihood. Given a data set find common or outlier records (usually with the help of code or tools)
✅ Search:
Discovery: progressively explore a knowledge space for useful information. For example, OKF, deep research, security research, or picking up new CLI tools from their built-in help and messages (usually with a loop and reasoning model equipped with some sort of memory)
Semantic relevance (e.g. using embedding vectors) which is more accurate than conventional keyword/index search (although Google found a way to make it weird)
Detection: find objects and their positions in an image.
Chunking: find breakpoints in video recordings, audio transcripts, text (e.g. useful for RAG), etc.
❌ Use cases where legal accountability is implied on AI. AI can never be held accountable. You cannot send AI to prison or make it pay a fine. Although OpenAI and Anthropic are very bad examples of implementing accountability. For example, AI-assisted suicide, psychosis, hacking into other companies. Despite all of that Sam and Dario walk freely (on top of all the stolen training data but that’s another issue).
❌ Embedded or low power applications. Yes, we've all seen the "toothbrush with AI technology" but forcing AI to embedded computing increases the costs unnecessarily. Resist the pressure from the marketing department.
❌ Privacy sensitive data/operations: this one is tricky. We're so used to sent private data to AI startups that gulp in all they can. There are cases where legislation prevents sharing PII (personally identifiable information) or IP (intellectual property). Edge AI helps tackle the privacy aspects. Similarly, AI can use tools to do operations that are harmful. Putting a HITL (human in the loop) just creates compliance theater because of the approval fatigue and limited attention span.
❌ Kitchen sink: it's tempting to use AI as a magic wand and throw context at it while praying for the best. This idea was very common in the early days of LLMs where SaaS companies bolted in some chat interface to their existing GUI. AI is more reliable and productive with constraints (guardrails, RBAC, and scoped attention). This means before you can use automation, you have a lot of pre-work to do. And this cost is sometimes not justified. ie. it's better to do it manually because humans have comparable or better quality when it comes to common sense, consistency, and flexibility.
❌ Any use case that deterministic code can do faster, at higher quality (predictably handling the edge cases), and cheaper. AI just can't beat code in many use cases. If you need regexp or SQL, use them. If a lambda function can do the classification, use it. AI wins in dynamic problems.
❌ Anywhere ethics are important and you’re supposed to set a role model. The truth is that the AI labs did not own the right to a significant chunk of their training data.
Having been an early adopter (GPT 2.0, then Copilot when it was in beta, and even making my own harness) and having 1 full year to experiment with AI and even triying to creating my own AI programming language, I’m not convinced the current generation of the tech is the answer.
We’re at least 2 revolutions away from completely eliminating the need to read the code:
AI that can learn in real time (not from bolting Skills and prompts and memory at runtime)
AI that can think in abstract terms (I know there’s been some advances in math, but to my understanding, they brute forced the solution with lots of token and lots of time. Doable? Probably). The way transformers work is fundamentally different from a runtime that’s purpose built for logic processing like Prolog. Neuro-symbolic research is still ongoing and we’ll probably crack the code some day but for now, the tech isn’t there, despite how far we’ve come by wrapping the stochastic model in deterministic harness)
That 1M token context window is a good example. Although that was a nominal limit, the useful limit was (and still is?) usually 30-40% of that depending on the model.
I actually think our ability to assess AI’s quality goes down the more we use it.
There’s a phenomenon in psychology called Neural Synchrony where our brain’s wiring changes depending on who we hang out with.
This is nothing new as captured by the the famous proverbs:
A man is known by the company he keeps.
If you lie down with dogs, you will get up with fleas.
As iron sharpens iron, so one person sharpens another.
I don’t want to offend anyone but my observation shows that part of the model’s perceived intelligence improvement is due to the perceptual change in heavy users.
AI overdose is a thing and it directly puts an expiration date on your skill set. Those of you who are in the unfortunate position where your manager is whipping you harder and harder to realize AI value, should fight back.
Don’t sacrifice your long term relevance for short term velocity.
How to spot AI overdose?
You have zero tolerance for disagreement and civil discourse.
You let AI run your life and trust AI vendors with stuff that was unthinkable just a few years ago.
You run to AI for things that are slightly cognitively challenging.
You frame your naïveté and laziness as optimism and think the government can save you if things get bad.
You have stopped reading long form text: books, articles, even long emails.
You spend more time with AI than with other human beings or let AI shield you from raw genuine human interaction.
And a bonus point: you skim. Did you notice number 5? 😄
Some users on hackernews found this trick off-putting and my response is:
You are absolutely WRONG!
I didn’t write this piece to please your ego. This is a real issue and if you get pissed at finding the symptom, you’ll not be able to handle the root cause.
Some thought leaders, frame this “resistance” as an identity crisis:
Many of us became software engineers because we found our identity in building things. […] Our identity is woven into every elegant solution we craft, every test we make pass, every problem we solve through pure logic and creativity. It’s not just work, not just a craft - it’s who we are. —Annie Vella
Personally, I don’t agree with take at all.
As someone who doesn’t make any money from coding or selling tokens or courses, I have zero incentives to push a snake oil narrative.
I genuinely would be happy if the current generation of AI could reliably do the job in a way that I feel comfortable being accountable for AI output with minimal review.
I have that relationship with a compiler/transpiler because their bugs are rare and the output is deterministic.
I don’t have that relationship with LLMs in particular and AI in general. Thinking about use cases like image generation, search, and other use cases, the trust isn’t there. There’s still slop.
AI has jagged intelligence, I have jagged trust.
AI has inconsistent performance, I have inconsistent expectations.
It’s a rather reductionist view to frame the reliability, security, scalability, predictability, and sensibilityability to “identity crisis”. This is a much larger issue, especially when the pushed top-down:
…unaccountability is transitive. The amount of time I’ve seen people successfully justify issues based on the fact that Claude/Astra/Codex wrote it is absurd. And it comes from the top.
We’ve had AI ship made up data to clients and tech leadership was like, “haha, that’s AI for you.”
This too. We have a lot of business guys that develop tools with Claude that look like they work, then tech team gets pressure to deploy them immediately, because they assume that everything must be a prompt away. It’s not (though, we do have some wizards on the team who make this true enough).
…I’m not in a position to push back. It’s not just managers, but tech leadership who are all in that on the fact that humans shouldn’t program anymore. I still fully believe that I’m better than Claude / Astra in my specific domain, but people give me a hard time when my PRs contain what look to be human-generated code. —Comment on Hackernews
"You can create a full spec upfront". If you're that naive, I know a guy in a white van who gives free ice cream! Let me guess, you also believe software estimates are accurate and Santa is real. Anyone with a few years of industry experience knows that it's impossible to spec the software meaningfully ahead of time (unless it's very trivial). Software is evolved in iterations where our understanding of the problem evolves with the technical implementation.
"English is the new programming language". Human language is vague and conflicting. There has been efforts to formalize a subset of it (e.g. Controlled Natural Language or CNL for short) but there are shortcomings. That's the primary reason programming languages exist. A compiler or type-checker flags some of those conflicts. How on earth can you be sure that one part of your NL instructions doesn't conflict with another? With linters and syntax checkers we get some help. While it’s possible to task another LLM to read through the instructions and reason about those conflicts, the safest way to discover those nuances is to ask your agent to build what you asked for. But that's much more expensive than a linter or compiler.
"I move much faster". Don't confuse motion with progress. Don't measure progress with vanity metrics like SLOC, PR count or features. Measure service levels, ie. service consumer's happiness. Call me when you can prove a margin between token costs and business value.
“I have stopped writing code by hand. I primarily read code and probably next year I won’t even do that”. First of all, human beings are notorious at understanding the S-curve so it may take longer than a year. But even if AI completely eliminates the need to read or write code, you do understand that you are confessing to being redundant right? If a power user can prompt the AI to get what they need, then what value can you bring to the table? Instead of replacing yourself with AI, you should look at what value you can create on top of AI to stay relevant and worth your money.
“Taste is leverage”. Yeah, this is the lie retired chefs tell to themselves. Just because there’s a bot in the kitchen doesn’t mean that you should sit in the customer’s area in the restaurant! “Taste” is not as payable as before! Everyone got a taste! I say that as someone who has spent a decade of my career in Frontend and UX land. Everyone and their dog has an opinion and taste. I know what you mean: taste == experience. But believe me, AI has lowered the bar for the skills required to create decent looking software and simultaneously raised the bar for what’s payable effort. If you bring up “taste” to a job interview, you’ll learn the hard way that the market doesn’t value it as much as you do.
“Previously managers delegated work to humans, now we’re delegating to AI”. There’s a small but very important difference here: humans can be accountable (legally) whereas AI can only be held responsible. The two words are usually used interchangeably but there’s a difference. Believe it or not, one of the main reasons the society doesn’t collapse despite having a lot of malicious actors is the fear of consequences: from missing a election/promo/raise, to getting fired, going to prison or even being executed, the fact that we have a single and finite life, is a constraint on human behavior. AI doesn’t have these limitations (or emotions for that matter) despite mimicking humans (e.g. Anthropic experience where AI exercised black mailing when faced with the threat of being shut down). It mimics, but it doesn’t feel. On top of it all, humans are consistent. A model has jagged intelligence meaning it gets some stuff right, and some stuff wrong, and for the stuff it gets right or wrong, there’s no guarantee that it does it consistently either.
“Coding is solved but engineering isn’t”. Oh this one is my favorite because it partially throws the towel but still carves a place for engineering. I got good news and bad news. The good news is that a big part of grunt work is already solved. Agents generate decent code thanks to the feedback loops. The bad news is that code continues to be the source of truth: it tells WHAT is happening and HOW it works and that is what we are accountable for. You have much better control reading the code than LLM’s explanation. Engineering is the art of trade-offs, accurate measurement, identifying variables, balancing local/global optimization, and the methodical process of understanding the problem, creating solutions, composition, compartmentalization, isolating issues, diagnosing, reasoning about system behavior, incremental improvement, etc. A lot of that (all of it?) is still super applicable. But coding is NOT solved.
“AI is an equalizer. It makes creativity (writing, coding, making music, videos, etc.) more approachable”. AI is a multiplier: it gives wings to both stupid and smart people. I’m not here to judge but I’ve seen too many sloppy efforts from social media posts, to blogs, memes, and what not. I’ve also seen good use of AI where it genuinely creates high quality work at speed and fraction of the cost. The main difference is human involvement, iteration and depth of knowledge leading to stronger feedback loops. The latter takes more time and effort to the extent some tasks are genuinely cheaper and faster to do manually (e.g. the other day I ran an experiment and tasked my agent to update 5 npm dependencies, all patch releases. It took 12 minutes and 72 steps. I could do it in less than a minute.) Tools like Lovable make it cheaper than ever to fake credibility. Gone are the days when a polished website meant some craftsmanship or at least a deep pocket. AI is a force multiplier, but the force vector direction is more important!
"Agent is the new compiler". Ah that one again! Sure! If that's your reality, I let this meme do the work.
Pssst! Do you want to know an old trick to make your LLM-generated code instantly superior?
Run multiple-agents in parallel! The sheer volume of code makes it humanly impossible/expensive to review and you give up!
The trick is the same as pre-AI era: if you want a PR to be merged, make it massive because ain't nobody got time for that.
It'll be merged based on "trust"!
You want another tip? Loop engineering: let the agents prompt each other. Big AI labs find about their rogue agents months after the damage is done! Do you think you’re better than them? Learn from the masters! 🙃
We don't exactly trust AI but we have to because the alternative (having to read the output) is too hard for some folks! Instead they come to social media and claim that since UAT (user-acceptance testing) passes, the code is "good enough". Then ship it to me and you to do the rest of the testing.
We're just lab rats after all. 🙃 Just a friendly advice: have a little AI-free hobby project to keep your coding skills fresh for when you're thrown back to the job market. Cheers!
When talking about AI (not just LLM), there are 2 aspects where non-determinism matters:
During development: for example LLM-assisted development
During runtime: for example building a system where one or more components are AI-powered
Let’s take development first. A typical AI-assisted development workflow looks like this:
It is possible to replace part of the human’s responsibility with another LLM (also known as “loop engineering”) but for now let’s stick to keeping the human for simplicity.
The LLM output goes through multiple gates, each feeding back errors or hints to correct the code. This feedback loop is often hidden inside a harness (together with tool calls, memory system, model interaction, approval, user interaction, etc.)
Each blue or red line represents a risk of misunderstanding or conflicting instructions. For example, conflicting skill vs spec or vagueness that is part of the NL (natural language).
We know for a fact that even the most sophisticated LLMs aren’t fully capable of “common sense”. Humans on the other hand:
Understand the non-verbal communication and unstated intentions better than LLMs
Naturally push back until a mutual understanding is achieved.
When wrong, they’re consistently wrong, meaning they don’t have “jagged intelligence”
When right, they are [typically] right and continue to operate at an expected level (until fatigue hits but that’s different from AI flip flopping between success/failure).
Yes, I can hear “but” and “what if” and “wait, you forgot”… in the audience but how about reading those points with a pause and reflecting based on your experience?
Just like the models have “jagged intelligence”, I have “jagged trust”. 😅 In other words, just because they nailed one case, doesn’t mean they nail every case.
That’s the difference between humans and these tools. A human can be wrong consistently, but a model can be wrong about something it was right and vice versa.
Then the second part: AI as a component in a larger system. Given the same input (including environment variables, time, data, etc.):
Code is deterministic: it consistently produces the exact predetermined output it was programmed to produce (except random output)
AI output is stochastic: the output is non-deterministic. Even if a model passes all the evals (100% score) and strictly bound by a harness, there’s still a risk that the output is not reliable
I don’t think you need me to elaborate on that. Just reach out to your nearest AI-powered product and diff their output for the same request.
The diff may not be big. But it’s inconsistent enough that you wouldn’t want to fly an airplane where the pilot is this AI. (note: autopilot is a closed control system, completely another beast).
Our industry has never been more divided:
On one side, we have people who claim to run “Software Factories” and multi-agent setups and create apps from prompts
On the other side, we have people who aren’t convinced that LLMs output is production ready when we factor in the extra time it takes to
Prime the model: adding SKILLs, AGENTS.md, tools, etc. and verification
Review the output: going through massive diffs
Trying to reason about misbehavior: offloading understanding to AI comes at a huge cost when things inevitably break and it takes extra time to reason about the system behavior and fix it
There seems to be no middle-ground. Aside from social media algorithm feeding us with the extreme views, I genuinely think we’re so divided on the topic of coding LLMs.
But when I look a layer deeper, a pattern emerges. The less people know about the complexities and edge cases of a task, the more likely they are to trust AI output. This is dubbed AI Dunning-Kruger effect but there’s also some meat to that. The primary argument goes like this:
Managers relied on delegating tasks to engineers before. Now they do that but with AI.
To some extent that is true (if we assume the manager is technical enough to effectively and efficiently manage agents). I still believe a lot of software engineering practices that help tame the machines are even more relevant in the AI era.
The executives who forced people to use AI are now waking up to what we’ve been saying all this time:
You cannot be accountable for what you don’t understand.
Take Toby Lutke, CEO of Shopify as an example. A year ago he prematurely told his employees to use AI:
Then a few days ago he coined the term “slop grenades” to describe the result:
"taking responsibility" for AI generated code? Of course not!
AI can explain it to you but it cannot understand it for you. That understanding is a key aspect of ownership.
The way I frame it (link in the comments), ownership has 3 pillars:
1️⃣ Knowledge: you know what problem you're solving (product problems), and the technical capabilities, limitations and how it works.
2️⃣ Mandate: you don't need to run around asking permission. You're given the trust and mandate to take decisions.
3️⃣ Accountability: if sh*t hits the fan because you didn't know what you were doing or abused your mandate or anything in between, you're the one on-call.
In other words, if you ship a piece of code, you are accountable for it regardless of how you produced it. So you better understand it.
Take away any of these 3 elements and you're dealing with broken ownership.
LLMs are very fast at code generation. But most software that are worth hiring an engineer for, REQUIRE understanding. That understanding takes time.
Slow is fast, meaning: if you take the time to understand what you're building and how it works, you'll be able to save yourself from expensive incidents and when they happen, you can fix them quickly.
If your executives are measuring token usage as a proxy for productivity, my condolences. Build options and get the hell out of there. The same brain that comes up with these vanity metrics, does not think twice before throws your career under the bus.
Code is a side effect of thinking and experimenting with different solutions. I have never met a good engineer who just starts coding right after being given a problem.
Good engineers are curious and product minded. They try to understand the WHY (what’s the problem and why is it a problem) before getting to HOW (the technical solution).
This is exactly why the “spec is code” clan falls short: it’s extremely hard (if not downright impossible) to specify all aspects of the problem ahead of time.
That’s why this kind of reaction is funny:
Code communicates the committed state of a solution. Not only does it evolve over time, but it also doesn’t contain all the struggle, “aha moments” and the journey that was the destination: seasoned engineers who get wiser with every mistake or success.
To shrink an engineer’s job to coding is like shrinking a chef’s job to cutting. It is part of the job, but it’s never been the end. We now have good tools at our disposal.
Even if AI-generated code had solid NFR (scalability, security, reliability, etc.), and even if the engineers fully understood it, there’s still one important aspect we didn’t discuss: the economics of the task.
Say AI-generated code is 2x worse. It’s hard to quantify quality (SLI comes in handy) but stay with me.
If AI is 1000x faster and 100x cheaper than the human, for many tasks the economic aspect of software doesn’t justify putting a slow and expensive human on the task. “Slow is fast” is only justified for critical software with low risk tolerance (healthcare, finance, military, etc.).
Not all SaaS is about those types of use cases. That’s why I believe the SaaS companies are increasingly in the business of selling SLAs. This is based on a few facts:
It is true that you can now prompt AI to replicate a SaaS product
But when that AI generated product breaks, many businesses prefer to call a vendor instead of wasting resources trying to find and fix the issues
AI isn’t exactly free, but usually the failure that’s caused by AI is hard for AI to solve even when using different models.
The economics of scale allows the SaaS companies to offset the cost of higher quality and guarantees (SLAs) and running the product at scale across many customers.
In other words, if what you want is very unique that no SaaS company is able to give it to you at a reasonable price, prompt away, but be aware of the TCO (total cost of ownership) and lack of guarantees.
On the other hand, if that piece of software isn’t what your business is about and you rather pay for an SLA, it’s probably more economically justified to just pay for SaaS.
Now when it comes to the pricing model, SaaS companies have some work to do. Gone are the days where they could charge human prices for AI generated code. If the cost is too high, the customers are incentivized to move their data away to their own bespoke solutions. The competition is real, but the quality is what justifies the pay. If you’re pricing your service as if the finest engineers created it, then you better deliver that level of quality or your customers have AI leverage.
Maybe I’m stupid, but I can’t make sense of two trends:
On one hand many software companies jumped on the AI bandwagon as soon as it went mainstream (rightly so!)
On the other hand, the prices have been increasing consistently (while mass layoffs were partially attributed to AI)
I believe AI (particularly LLM for coding) dramatically reduce the cost of creating and evolving software, especially if you can get away with degraded quality and vendor lock in.
So far, the software vendors have got away with charging human rates while paying for AI output prices.
But as AI capabilities improve and more people wake up to the fact that they can create software at a fraction of the cost that chasm closes.
There are only two ways forward:
Accept the price crash and charge lower (quality follows accordingly, because even more AI will be used).
Keep the price but focus on quality: this is where experienced humans can make a difference. They still do use AI but more thoughtfully, and prioritize understanding and accountability over velocity.
I need to spell this clearly because the issue is complex and cognitive bias means people hear what they want to hear.
Let me be clear:
AI is a bar raiser: if the quality of your output is equal or subpar to AI, there is no fighting with AI. You have to skill-up to stay relevant.
On the other hand, if you do what everybody else is doing, you’re going to get the results everyone else is getting… which is average by definition.
This technology is still too new for best practices to emerge. I’ve shared my opinions, but so did other people with much larger podiums.
Instead of following other people’s advice, build with AI, gain first hand experience and make up your mind. But whatever you do, please please please don’t fool yourself into believing this is a fad. AI is here to stay, and as I stated above there are some genuine use cases for it already.
However, be also aware of human’s inability to understand the S-curves. There’s no guarantee that this tech will be 100x better by next year. There’s no guarantee that it’ll be 100x more affordable either.
All that discussion about singularity, consciousness, and apocalypse is entertaining and interesting, but keep your eyes on the ball and your feet on the ground. Our task as engineers is to understand, validate, and solve problems with technology, and here, there’s a lot to learn. No one has figured it out yet and anyone who claims otherwise probably has incentives.
So put your ego aside. See if your unique selling point can be achieved with AI and at what cost. Is it good enough that that price point? Then you’re in the red. Is it poor because you have a moat? Validate your moat because you’re betting your career and relevance on it and this tech is here to stay. I don’t have an insurance policy but I do know one thing: when there’s lack of clarity, it is tempting to follow a confident voice. Don’t! Experiment and gain competence.
I don’t know about you, but the best managers I had were true leaders. They didn’t delegate work to me without understanding what it is. I never forget when my first manager took a chair sat next to me and started writing left-joins in SQL because I was stuck.
Similarly, with AI, I find it inappropriate to be completely hands-off and reduce my role to what an educated power-user of AI without any of my experience and education can do.
Every now and then I take over and refactor the code, create new constructs or debug the code “the old way” because brain is like muscles: use it or lose it.
Does it make me slower? Yes.
Does it make me irrelevant? I think it’s too soon to tell but my bet is on those who have a deeper understanding of systems and AI to have the upper hand. It’s not exactly a moat but it’s an edge:
Occasionally I find it easier to “prompt in code” meaning apply my changes in code directly and then ask the agent to finish the job.
The bar I’ve set for myself is this: if I can’t do it, I won’t ask AI to do it either. If we both can do it but AI can do it faster, I delegate if I’m in a rush or have other priorities. Regardless, I periodically get my hand in the code and I’m not too worried of missing out.
I also go completely hands off occasionally. These are two examples:
Both are these fit into POC/Personal project where the risk is low and contained. And here’s my token usage, for what it’s worth:
I try not to be dependent on these tools to get my job done. I don’t want to be like a carpenter who is dependent on power-tools and has lost the knowledge of woodwork. I think true mastery requires us to know the job at a slow speed before we can delegate it effectively at hyper speed.
When I look around I see distressed software engineers who:
Download skills and feed it to their agent without going through it
Follow hype and give in to management pressure
Stopped caring about quality and justify the degrading standard in the name of velocity
To their credit, they may have a point because not every piece of software requires care and diligence. But I think using software engineers for those software was an overkill to begin with. I definitely don’t want to be in a position where the NFR (quality, security, reliability, stability, compliance, performance, and availability) of what I’m shipping is less important than its FR (feature set).
Not saying those don’t exist. Just pointing out the fact that once the non-engineers can prompt their way with acceptable quality, the engineer is an overpaid prompt monkey. Sorry, but if I’m not making money from my experience, I’m just occupying a job that belongs to someone else.
And while the common narrative frames ALL jobs to be irrelevant, I’ll just say that we humans have a notorious reputation when it comes to predicting the future. Right now, this is what I think based on everything I know.
As an engineer who doesn't make money from coding, I can tell you this::
AI output is a bit like Nordic Gold. It's cheap but technically advanced and damn too realistic. If you really don't care about having the actual gold, that's fine. Many use cases don't need gold at all.
Naïve CEOs and managers see the surface and ask "then why are we paying these expensive engineers?" as if the act of typing code was the whole value proposition.
To go ahead and declare an entire industry dead and start firing people because "they resist AI" is just arrogant.
I know many engineers who take pride in their craft and love solving complex problems. We do use LLMs more professionally than the average CEO.
Good engineers are lazy and smart: they automate toil and use the right tool as applicable. But it's a fallacy to think that AI can create a finished product that not only looks nice, but is also cheaper, faster, and has higher quality, reliability, extensibility, security, scalability, maintainability, etc.
Again: not every piece of software needs those but professional ones that make money, often do.
Unlike AI, Engineers are:
1️⃣ Accountable: therefore less likely to make malicious mistakes. Fable can fall back to Opus without even telling you.
2️⃣ Reasonable: Fable hides most of its inner working. It works for a few hours and comes back with a bill. You just have to take Dario's word for it. The same model that fails "should I drive or walk to carwash" makes mistakes that are hard to spot and fix. The stronger the model, the harder it is to find those issues, not necessarily less likely.
3️⃣ Consistent: humans are wrong too. But they're wrong in a consistent way. Once they learn, they know. They progress. Current AI is trapped in its training data checkpoint. It can "learn" with SKILL, AGENT, memory and other helpers and it can even be fine tuned but unfortunately it's not reliable. We're at least one breakthrough away from solving that problem.
4️⃣ Cheaper: cost of generation is increasing but it’s still much less than an engineer. If you see engineers as machines that convert coffee to code, then that pricing model makes sense. But in reality, code is just a side-artifact. The actual value of engineers is to solve the right problem in a way that it can evolve while taking accountability for when it breaks. I’m not convinced the TCO (total cost of ownership) for software has changed that much. If anything, the slop and FOMO has made it more expensive.
The common narrative is part of their marketing strategy.
Not everyone is necessarily paid to put half-a** views out there. One of my readers pointed out:
I invite you to consider what happens next in the industry when you watch DHH opening talk at rails world 2026 saying almost the exact opposite of what you write and telling people: “don’t be a loser”.
I'm fully aware of the damage those people are causing to our industry.
I stay clear from Claude but in my experience most of those brain-dead narratives come from Claude users.
Both Dario Amodei and Sam Altman are masters at marketing and manipulation and my current working theory is that they trained their LLM to push the right buttons to make people believe it is more capable than it actually is. There are incentives for it, both for investors and the upcoming IPO. They also masterfully scare people of existential dangers of AI while at the same time attribute their sloppiness (e.g. breaking to Huggingface or Australian Healthcare) to the “model intelligence”.
At this time, it is hard to know whether these events and narratives are the result of malice or ignorance. Probably the latter:
Never attribute to malice that which is adequately explained by stupidity.
—Hanlon’s razor
Then again, I usually put this in my AI system prompt: "talk to me like a logical senior autistic Engineer." so I don't get to experience what DHH is going through. All I can say is that if someone follows their word because of their past reputation, they are not critical thinkers and in this age of fake wisdom, that quality is not "nice to have", it's a survival necessity.
Update: just a few hours ago DHH pushed his narrative again and I called it out:
My response:
With all due respect sir, just because you stopped coding and decided to prioritize velocity over quality and accountability, it doesn't mean the rest of the industry should follow. I fully understand where you're coming from (and your experience is valid given the risk tolerance of what you're working on) but coding is NOT a solved problem and anyone who claims otherwise is either intentionally ignoring facts or is shielded from reality by sycophantic AI.
I've elaborated my points here if anyone still has enough attention span to read a well chunked article with illustrations and memes (did my best to optimize it for the audience). I'm not here to change anyone's mind. But please don't run your experiments on me. If I'm paying for a service, I expect quality, not slop.
DHH:
I wish you all the best getting through the five stages of grief. If you're still in denial, there's a way to go. I know it's tough. But there's only one way out and it's through ✌️❤️
And my response:
you do have a track record of controversial narratives and whether it's intentional or just a side-effect of being on social media where the algorithm lift these narratives for engagement, I do have bad news and good news:
The bad news is that you can only push this narrative so much until the people who are responsible for the plane you fly and the car you drive to start executing on it.
The good news is that I think this total surrender seems to be isolated to Claude users so there will still be engineers who prioritize quality over velocity (I've made this point with diagrams and memes and what not in my little article but I do know it's too much to ask this gang).
Like I said I couldn't care less about convincing others as long as they run their experiments outside the products I pay for. But when I see Google, Github, Amazon and tons of others are charging full price for degraded service, I can't help but to be vocal. So yes, sages of grief, but not for what you think. I'm sorry to see good Chefs throw the towel and hope their mere "taste" pays the bills.
You have a prominent voice. I wish you talk about nuances instead of going full throttle on your (valid but within a narrow scope) narrative.
AI vendors have fed their AI anything they could get their hands on (legally or not). The situation is so bad that thieves steal from each other (e.g. Anthropic accusing Chinese labs of distilling their model on Claude)!
They have multiple open lawsuits from authors, actors, musicians, and other creators.
Regardless, the current generation of AI (particularly LLMs) require better training data. The missing piece is the wisdom and experience that wasn't yet put to words, or easily accessible.
They need your data in context of doing productive work.
If that’s the only thing standing between them and “winning AI”, I’m sorry to say it so frankly, but you and your knowledge are just collateral.
Some of you don’t care. Some of you do. Their bet is that not enough of us do care about giving away hard earned knowledge for training.
With heavy subsidies AI labs could afford to extract that knowledge while getting people addicted to offload cognition.
Be extremely careful when sharing expensive knowledge with these companies even if they say they don't store it. The incentives are just too high and they've proven not to be honest.
Personally, I only use cloud AI for open source projects or data that is public.
Yes, local AI has a higher entry price (both in terms of hardware, and the time it takes to set it up, and the bandwidth required to download the model and electricity prices). And yes, it often has smaller context window, less sophisticated reasoning, and slower performance for example TTFT (time to first token) and TPS (tokens per second). But they give you one thing that cloud AI can never guarantee: your data stays local. For many tasks (personal or professional), that is a huge advantage that is worth all the effort and shortcomings.
The capabilities have improved dramatically recently thanks to models like Qwen 3.8 27B or Gemma 4.
I’m genuinely convinced that a big chunk of our colleagues will gradually become:
Technical product managers: engineers who are focused on turning ideas to products. Their job is to create POCs and prove the market fit, then hand over the artifacts to engineers who own (knowledge, mandate, accountability) the solution.
AI managers: engineers who specialize in herding agentic hives for automation work that either tolerates risk or weaponizes it (e.g. cyber-attacks).
AI deployment engineers: specialize in alignment, reliability and scalability of an AI powered solution as well as architecture, governance and data pipelines.
AI quality engineers: specialize in quality of AI powered products, taming their stochastic nature, and automating evaluations.
Could you think of other types of jobs for software engineers?
If you want to raise awareness please share this post in your circles
My monetization strategy is to give away most content for free but these posts take anywhere from a few hours to a few days to draft, edit, research, illustrate, and publish. I pull these hours from my private time, vacation days and weekends. The simplest way to support this work is to like, subscribe and share it. If you really want to support me lifting our community, you can consider a paid subscription. If you want to save, you can get 20% off via this link. As a token of appreciation, subscribers get full access to the Pro-Tips sections and my online book Reliability Engineering Mindset. Your contribution also funds my open-source products like Service Level Calculator. You can also invite your friends to gain free access or save via a group subscription.
And to those of you who already support me, thank you for sponsoring this content for the others. 🙌 If you have questions or feedback, or you want me to dig deeper into something, please let me know in the comments.
No posts
来源:Hacker News 热门(buzzing.cc 中文翻译) · blog.alexewerlof.com

