🔥 The Code Runs. The Logs Say Nothing.
📅 Tuesday, Oct 6, 2026
⏰ 8:00 PM
I once helped troubleshoot a problem in production and went looking for the application logs. There were none. No error explaining the failure, no info message saying the service had started, and no way to turn up logging while we investigated.
The story behind that silence was more complicated than a developer forgetting to add log statements. The team responsible for keeping the system running had built software that read the application’s logs to detect problems. At some point an included library changed its logging output. The operations team saw that as an unapproved change to something their tooling depended on. As I understood it, the engineers’ response was to stop emitting application logs in production.
When logs become an interface
That was an extreme outcome, and the disagreement underneath it was real. Once another system reads your logs, their format is an interface, and a change to a library’s logging output can break that interface while the product’s behavior stays exactly the same.
The fix was available. Agree on stable, structured events for the tooling, review format changes against the systems that consume them, and leave the diagnostic detail free to evolve underneath. OpenTelemetry’s log data model
is built for that split: EventName identifies the class of event, SeverityText and SeverityNumber carry the level, Attributes carry the detail that varies, and TraceId ties the record to the request it came from. Tooling can bind to an event name and the fields it needs without owning every string a developer writes.
Nobody took that path, and we ended up with a running application that could tell us almost nothing about itself. That is the one outcome OWASP’s logging cheat sheet rules out by name: “It should not be possible to completely deactivate application logging or logging of events that are necessary for compliance requirements.”
I think about that experience when I review AI-generated code.
The statement that gets cleaned up
I have watched AI add log statements to debug a problem, use them to get a test passing, and then remove them as part of its cleanup. Sometimes removal is exactly right. A temporary dump of a request body may be noisy or unsafe to keep. But sometimes the statement captured a decision, a retry, or a state transition that would be worth having the next time the problem showed up in production. The judgment call between those two cases is the entire job.
Production problems are rarely considerate enough to reproduce locally. If the useful statement has been deleted, raising the log level reveals nothing, and putting it back means a code change and another deployment, possibly while customers are already affected. Google’s SRE troubleshooting chapter is direct about what you want instead: “It’s really useful to have multiple verbosity levels available, along with a way to increase these levels on the fly,” without restarting the process.
What the research measured
Is this actually an AI problem? The research supports something narrower and more useful than a complaint.
Start with what models do well. A 2024 study in IEEE Transactions on Software Engineering , posted on arXiv under the better title “Can LLMs Log?”, built a benchmark of 6,849 logging statements pulled from GitHub and measured how well language models fill them in. The best model picked the correct log level 74.3% of the time, and every model in the study got the level right in at least 60% of cases. AI has a working concept of severity.
The content is where it comes apart. Log text topped out at a BLEU score of 0.249. And when the same code was mechanically transformed so the models had not seen it before, the degradation was not evenly spread: level accuracy fell by 1.4%, variable selection by 11.6%, and log text by 15%. The part that is a label held up. The part that carries the information did not.
A 2026 preprint on coding agents finds the same split at the system level. The researchers stripped the human-written observability out of 10 open-source and 8 industrial repositories and asked agents to put it back. Agents often picked the right places and then filled them with the wrong things. GPT-5.5 scored 0.551 on placement against 0.357 on diagnostic content; Claude Opus 4.8 scored 0.580 against 0.294.
Then they ran it. Two hundred microservice systems generated from specifications, deployed on Kubernetes, with 13 kinds of production fault injected across 1,615 failure instances. The systems produced logs. Explicit, fault-specific evidence showed up for 4.95% to 13.99% of the failures depending on the model. Between a quarter and a third of the generated systems never ran at all, and scoring only the ones that did lifts the best model to 20.62%. Either number tells the same story: the limitation was not missing logs, it was logs that could not say which failure had happened.
The uncomfortable part is what happened when the researchers asked for observability explicitly. The agents complied by volume. Statements per instance went from 2.1 to 4.9 and diagnostic tokens from 11.5 to 22.9, while quality went down, content F1 falling from 0.36 to 0.26 as precision dropped from 0.33 to 0.20. A packaged observability skill did better on fault signals, worth between one and nine percentage points, but moved the semantic scores by 0.003 to 0.015, which is the argument I made about Skills in The Jig showing up in someone else’s data.
Two caveats and then I will stop hedging. This is one experimental setup, and it cannot tell us how often this happens across production software generally. The agents were also generating whole systems from a specification, which is harder than the incremental work most of us actually hand them.
None of this is new. Developers have always logged too little, logged too much, or picked the wrong level. What AI changes is the ratio, because a great deal of working code can now be produced without anyone spending the corresponding time learning how it fails. When the task ends at a passing test, emitting a log line is easier than deciding what a future investigator will need to know.
Levels are a control surface
Good logging is a set of choices about content and about control, and the levels are where those two meet.
At INFO I want a small number of meaningful lifecycle and business events, including enough to know which version started and whether it became ready. At WARN and ERROR I want the outcome and the context needed to investigate it. At DEBUG and TRACE I want carefully chosen detail around decisions and state changes, off during normal operation and switchable when an incident calls for it. “Payment failed” is a label. An event that identifies the dependency, the operation, the failure category, the retry attempt, and the trace ID, without carrying payment data, is a diagnosis.
What the framework hands you
There is a stronger version of this argument that I do not think holds: that some languages and frameworks are for production and the rest are for throwaway proof-of-concept work. Instagram runs on Django. Python’s standard library logging module was modeled on log4j and has had hierarchical loggers and severity levels for more than twenty years. Serious money rides on TypeScript services. The language is not the thing that is missing.
What differs is the default, and defaults are what you get when nobody is paying attention.
Spring Boot is the clearest case in the other direction. Actuator ships a loggers endpoint
that reads and changes a logger’s level on a running process: POST a body of {"configuredLevel": "DEBUG"} and the level changes with no restart and no deployment. That is the SRE book’s requirement, already implemented. Worth being precise about the cost, because it is not zero: only /health is exposed over HTTP by default, so you opt in through management.endpoints.web.exposure.include and you put it behind authentication before it goes anywhere. Add SLF4J’s per-package levels and MDC for request-scoped context, and “turn up logging for this one package on this one pod” becomes a conventional request rather than a project. I made the adjacent argument about Micrometer and stable error events
years ago, for the same reason.
Express ships none of that. console.log is the path of least resistance, the logging library is a decision someone has to make, and the endpoint for changing a level at runtime is something you build. Flask and FastAPI land in between: the levels are there because Python gives them to you, the packaged surface for changing them on a running process is not.
Which is what makes this a generated-code problem rather than a taste problem. An agent reaches for the ecosystem’s default path, because that is what its training data is full of and what the scaffolding produces. Ask for a Spring Boot service and it inherits a facade, levels, and an actuator it may never have reasoned about. Ask for an Express service and it inherits console.log. Neither study I cited measured language choice, so this is an argument about defaults and not a finding. But inheriting a default is still a decision about what the service will be able to tell you, and it is worth making on purpose.
More is not better
Volume costs money and buries the signal, and the SRE chapter points out that turning on verbose logging can make a latency problem worse and confuse the result you were trying to read. Sensitive values turn a useful diagnostic record into a security problem, which is why OWASP’s list of what to keep out of logs includes session identifiers, access tokens, passwords, encryption keys, and personal data.
Logs are also one signal of three. Metrics tell you a service is failing, traces show where a request went, and logs explain the decision or state at a point along the way. OpenTelemetry treats all three as what instrumentation has to emit, and sets the bar this way: an application is properly instrumented when developers do not need to add more instrumentation to troubleshoot an issue.
That is the bar a deleted log statement fails.
The review question
So yes, logs still matter, and they matter more when code can be produced faster than a team can learn how it behaves. Where that gets settled is the runtime, which is the argument I keep coming back to .
The review question is no longer just whether the code works. It is what the people keeping it running will be able to see when it fails, and whether they can get the detail they need without deploying again.