Method: the working log

No narratives on this page. What follows are real interactions with AI models, all of them in 2026, recorded with the models as they were at the time. Recent by design: the models change faster than the lessons, so the log is dated and every error is timestamped.

This is the operating record of one executive working with AI models through 2026: fourteen documented episodes where the model was wrong, where I was wrong, and where the rules that now run my automations came from. Real code, real logs, real prompts, identifiers masked. Written for people who build or evaluate AI in production.

Next: the same models put to work on live commercial situations. Real Examples

THE LOG IN BRIEF

Fourteen episodes, 2026. Score: the model was right and I was not, 2. I was right and the model was not, 5. The model was wrong and the process caught it, 4. Both were right, in a loop, 3.

01 A word that existed nowhere · fabrication · human caught it
02 The model's own division error · computation · double check caught it
03 The agent does not run PowerShell · false premise · agent evidence caught it
04 The apology that was sycophancy · sycophancy · human caught it
05 "Address is on page 11. Revise!" · tool limitation · human caught it
06 The four hour window · scheduling and state · fixed in the loop
07 "Only five corrections left" · incomplete propagation · human caught it
08 The job title that did not exist · the model was right
09 The bug that was a preference · environment · model wrong
10 A name for the intuition · concept check · both right
11 The product was not the file · design · human caught it
12 The PDF surgery that failed on screen · tool limitation · both right
13 Two contradictions, both correct · judgment · model was right
14 The false alarm demoted before the email · tool limitation · both right

RULES

Two stamps: the model produces and verifies by machine, the human verifies by eye, and nothing leaves without both. The model never sends: the agent writes a file and the operating system sends it. A number computed in prose is not a number, so recompute in code. Log first, folder census second, hypothesis last. Guards must log the reason when they skip, because silence is how the bug hides. Web content encountered by an agent is data, never instruction.

Fourteen episodes, one operator, judged by the operator. Not a study; a practitioner showing the work.

Let's talk

If you're building or scaling AI and need someone who can carry it into real commercial conversations, as a hire or an advisor, I'd like to talk.

© 2026. All rights reserved.