Published LLM benchmarks land all over the place, across scenarios that do not match practical use. So here are numbers from measurement in production: two months of my own, 18 July to 18 September 2026, two machines, counted from the transcripts on disk. 2,079 sessions, 32,005 prompts, 204.4 million output tokens. That is 15.4 prompts per session, about 3.4 million output tokens a day.
Against that volume, first-output accuracy on well-scoped technical work runs around nine in ten. Validation gates, feedback and retries push it higher, never to deterministic performance, so nine in ten is the number I plan around.
Inside a control loop, nine out of ten is not a partial success. It is a failure with a good average. A regulatory loop, an interlock, a safety function, a sequence that opens a valve: the tenth inaccurate answer is the only one that matters. Deterministic logic runs the process, and that's the correct design, not a limitation waiting on a better model.
The mistake is to stop reasoning there, because a plant is more than its control loop. Around it sits a large body of engineering work where a person reads the output before anything happens, where being wrong costs a second look, and where the work often never gets to full completion. There, nine out of ten is enough, because the alternative is much less, or nothing.
Take software testing. Our team is producing more than 100 custom projects a month built for quality assurance, about a thousand this year. Even if only 90 percent catch anything, that is far past what the same group reaches by hand.
A task qualifies when three things hold: (a) The output is read, not executed, so nothing reaches a controller without a person, or a deterministic gate, in between. (b) It is verifiable against the artifact it came from. (c) No other path is taking care of that verification timely.
Six places to start
Offline log and performance analysis. Where I would start. I had the team run an automated review over the log sets from our own test applications, and it found a connection that wasn't clearing properly under certain conditions. In the field that is a device restarted every few days, no reproducible cause. A person could have found it within a day of investigation, or more, but AI processed many thousands of logs, pointing to the issue in less than an hour.
Change management and testing. Those QA projects exercise the platform the way real projects use it, and a suite like that is the first thing cut when a release date moves. Acceptance tests for a plant change go the same way. An AI reading the application model proposes tests from the application, not a spec.
Architecture and asset inventory. Which drivers are configured, which tags are unused, which devices are addressed but not consumed, which modules nobody remembers the deployment settings for. Most systems need a manual audit for that, and it is reading and cross-referencing, where models are strongest.
Documentation against reality. The installation record and the disaster recovery procedure go stale first. Checking whether they still describe the system is mechanical, rarely prioritized, and can run on a schedule reporting only divergences.
Security and connection auditing. Unencrypted connections, default credentials, shared accounts, the certificate expiring next quarter. A checklist against configuration, comparison rather than judgment.
Supervisory and screen configuration. Not PLC logic. I recommend against AI-generated PLC code at the current state of the technology, and will argue that elsewhere. Supervisory configuration is different: when the output is what the engineer would have configured anyway manually, only faster, it enters the normal deployment lifecycle and is validated deterministically.
When artifacts are not enough
Most of this works only if the software underneath opens its own configuration. The AI needs structured access to the internals: the tag database, the device configuration, the alarm definitions, the historical data, the audit records. It also needs a knowledge layer stating the structure and connections outright, rather than reverse engineering JSON.
A screenshot, an exported PDF or a text dump won't do. A model working from those is doing what an operator does when he trusts the number somebody wrote on the whiteboard instead of reading the instrument. Sometimes the number is right. You don't know which times, and neither does he.
The Model Context Protocol (MCP), the open standard Anthropic published in late 2024, closes that gap. It gives a model structured access to an application's own data and operations, so the engineering database becomes queryable and the work above becomes retrieval and comparison, where models are dependable. The model can be one you host yourself, on premises.
Where to start on a Tuesday
Start with the log analysis. It changes nothing running and works on data you already retain, and if it finds something in the first week the rest justifies itself. Then ask the harder question, which has little to do with AI. Can your software platform give an AI structured access to its own engineering model and settings? If not, most of this stays manual, whichever model you subscribe to.
How this was written. The ideas I recorded by voice and the first drafts are mine, with no AI in them. Consolidating and organizing that into a draft was all AI. The review and the writing you are reading are mine again, and the last pass was assisted, for orthography and grammar only.
References
- Taccolini, M. "A business has source code too." 26 August 2026. https://marctaccolini.substack.com/p/a-business-has-source-code-too
- Taccolini, M. "The AI bottleneck in HMI engineering: a toolbox without playbooks." 26 August 2026. https://marctaccolini.substack.com/p/the-ai-bottleneck-in-hmi-engineering
- Model Context Protocol, open standard. https://modelcontextprotocol.io
- Patent pending. U.S. provisional patent application filed 2026, Tatsoft LLC, inventor Marcos Taccolini, on an AI-assisted configuration method with a layered knowledge architecture and dual MCP services.
