<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Observe. Assess. Advise.]]></title><description><![CDATA[Military frameworks applied to AI adoption — for executives and operators who need it to actually work.]]></description><link>https://observeassessadvise.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!KMYx!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F393533e5-eb41-412d-bd57-af584a6e3975_682x682.png</url><title>Observe. Assess. Advise.</title><link>https://observeassessadvise.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 14 Aug 2026 09:35:18 GMT</lastBuildDate><atom:link href="https://observeassessadvise.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Glen Lewis]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[observeassessadvise@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[observeassessadvise@substack.com]]></itunes:email><itunes:name><![CDATA[Glen Lewis]]></itunes:name></itunes:owner><itunes:author><![CDATA[Glen Lewis]]></itunes:author><googleplay:owner><![CDATA[observeassessadvise@substack.com]]></googleplay:owner><googleplay:email><![CDATA[observeassessadvise@substack.com]]></googleplay:email><googleplay:author><![CDATA[Glen Lewis]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[I Got Mad at My AI. The Failure Was Mine. *]]></title><description><![CDATA[An AI does the work. It's blind at both ends &#8212; and those two blind spots are the one thing you can't hand it.*]]></description><link>https://observeassessadvise.substack.com/p/i-got-mad-at-my-ai-the-failure-was</link><guid isPermaLink="false">https://observeassessadvise.substack.com/p/i-got-mad-at-my-ai-the-failure-was</guid><dc:creator><![CDATA[Glen Lewis]]></dc:creator><pubDate>Mon, 13 Jul 2026 17:17:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KMYx!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F393533e5-eb41-412d-bd57-af584a6e3975_682x682.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>For six months I&#8217;ve been building an executive assistant that runs on an AI agent &#8212; a tool to plan my days: prioritize the work, protect the prep time before meetings, tell me what &#8220;done&#8221; looks like on a block of deep work.</span></p><p><span>For most of those six months, it was bad at all of it. And every time it handed me a plan that missed, I got mad at </span><strong><span>it</span></strong><span>. If it had been a junior employee, I&#8217;d have fired it every other week.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://observeassessadvise.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Observe. Assess. Advise.! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p><span>Four weeks ago I worked out who I should have been mad at. Me. And once I saw it, I couldn&#8217;t unsee the second version of the same mistake &#8212; the one I&#8217;ve been quietly making with every capable model I touch.</span></p><p><span>Here&#8217;s the hypothesis I backed into the hard way. An AI does the work in the middle, but it&#8217;s </span><strong><span>blind at both ends.</span></strong><span> Going in, it can&#8217;t see the workflow in your head &#8212; the part you never wrote down. Coming out, it can&#8217;t see the edge of its own competence &#8212; the places it&#8217;s confidently wrong. Those two blind spots are the whole of what you own &#8212; and you own them differently. The sight going in you supply once: write down what you know, and the gap closes. The sight coming out you supply for as long as you use the tool: a model won&#8217;t reliably flag its own edge, so someone has to watch for it. What never transfers is the responsibility for either. And the better it gets, the easier it is to believe it can see for itself. It can&#8217;t &#8212; and the one being fooled is you, not the machine.</span></p><p><span>I fumbled both. Here they are.</span></p><h3><span>Half one: the sight going in</span></h3><p><span>Here&#8217;s what I never did for six months: explain how I work.</span></p><p><span>Not the tasks &#8212; the </span><strong><span>workflow</span></strong><span>. How I decide what earns a deep-work block and what gets triaged. What actually counts as a deliverable. How I guard the thirty minutes before a meeting so I walk in ready. To me all of that is innate; I&#8217;ve done it long enough that it doesn&#8217;t feel like knowledge, it feels like breathing.</span></p><p><span>But an agent can&#8217;t act on what you never say. It looks like it remembers, like it understands the shape of your day &#8212; but underneath, between turns, it doesn&#8217;t. I&#8217;d handed it a job without ever handing it the context that makes the job doable, then resented it for guessing.</span></p><p><span>The fix wasn&#8217;t a smarter model. It was writing the innate part down &#8212; capturing the actual workflow of my day. First thing I made explicit: every deep-work block gets a named deliverable and a time, so the agent knows what &#8220;done&#8221; even means. That&#8217;s when it started to work.</span></p><p><span>The tool took days. The integration took six months &#8212; because the integration was me. That&#8217;s the sight going in: the workflow only you can see, that the model needs and cannot get on its own.</span></p><h3><span>Half two: the sight coming out</span></h3><p><span>Supervision is the last of the Marine troop-leading steps &#8212; and the one they&#8217;ll tell you matters most, because it&#8217;s what turns a plan into a result, and it&#8217;s the first thing to slide when the work keeps coming back good. That&#8217;s the second blind spot, and the one I&#8217;m still fighting: with AI, it hides behind the model getting </span><strong><span>better</span></strong><span>.</span></p><p><span>Picture the best hire you ever made &#8212; aggressive, takes initiative you didn&#8217;t ask for. For four, five, six months the work comes back outstanding, faster and cleaner than you&#8217;d have managed yourself. So you relax. You skim where you used to scrutinize. You approve on reputation. Then one day, one deliverable, they&#8217;re confidently, catastrophically wrong &#8212; and because you&#8217;d stopped looking, it&#8217;s out the door before you catch it. That bad day is usually the first time you notice how far your supervision had slid. The trust doesn&#8217;t erode; it shatters, all at once.</span></p><p><span>I&#8217;d started living that with my AI &#8212; and the pull is stronger, because the model keeps earning it. It gets better, more consistent; when a tool is right nine times running, you stop bracing for the tenth. And you&#8217;re told, constantly, that it&#8217;s trained on *everything* &#8212; so it&#8217;s easy to assume there&#8217;s no edge at all. So you do the efficient thing: you hand it the checking too, on the quiet assumption that it&#8217;s mostly right.</span></p><p><span>But &#8220;mostly right&#8221; has a shape, and neither of you can see it. At Sequoia&#8217;s AI Ascent this year, Andrej Karpathy told the story of chess: model chess ability jumped sharply between releases, and the field read it as the model getting generally smarter. It hadn&#8217;t. Someone had fed its training a pile of chess games; the spike was local to that data &#8212; nothing beyond it. A brilliant game told you nothing about the task one square over.</span></p><p><span>Here&#8217;s why that should worry a manager. With a person, you can find the edge of what they know &#8212; they hesitate, they say &#8220;that&#8217;s not my lane&#8221; &#8212; and </span><strong><span>finding</span></strong><span> that edge is what lets you safely trust everything inside it. A model rarely gives you the edge. It runs off the end of its competence at nearly full confidence, with none of the hesitation a person would show &#8212; and no way for you to tell, in the moment, which side of the edge you&#8217;re on. So you can&#8217;t trust it to check its own work: the place it&#8217;s least reliable is the place it&#8217;s least likely to flag. And you won&#8217;t reliably catch it either &#8212; you follow it over the edge. The field experiment that gave this its name, the &#8220;jagged frontier,&#8221; handed 758 consultants a problem set just past what the AI could do. On their own, they got it right about 84% of the time. With the AI, they did </span><strong><span>worse</span></strong><span>: 19 points less likely to land the right answer &#8212; because a confident output carried nothing in it to tell them they&#8217;d crossed the line.</span></p><p><span>Which is why the checking never transfers. Someone has to hold the sight the model doesn&#8217;t have, and that someone is you. I used to file this under &#8220;autonomy ladder&#8221; &#8212; more rope as the agent proves itself. I still believe it, for scoping *what* it may touch. But proven reliability earns a model wider scope; it never earns it your absence.</span></p><p><span>The line I already knew</span></p><p><span>None of this should have surprised me. I spent years being trained on it, in a setting where the cost of getting it wrong wasn&#8217;t a missed deadline.</span></p><p><span>A commander doesn&#8217;t hand responsibility for an operation to the operations officer who planned it. The OPSO designs the op; the commander owns what happens when it meets the enemy. You can delegate the </span><strong><span>task</span></strong><span>. You never delegate the </span><strong><span>outcome</span></strong><span>. HBR put a number on it: in a study of 1,200-plus managers this year, framing an agent as an &#8220;employee&#8221; led them to catch </span><strong><span>18% fewer</span></strong><span> of its errors and to blame the model when it failed. But blame doesn&#8217;t move; it stays with whoever deployed the thing. When its work ships under your name, it&#8217;s your work.</span></p><p><span>I forgot that &#8212; not in a briefing tent, but at my own desk, with a tool I built myself. When it failed, I got mad at the assistant. I should have gotten mad at the manager who never gave it the sight going in, then stopped supplying the sight coming out.</span></p><p><span>That manager was me.</span></p><h3><span>Onboard like a hire. Own like a commander.</span></h3><p><span>So here&#8217;s where I&#8217;ve landed, sharper than the advice I&#8217;ve been giving.</span></p><p><span>Onboard your AI like a junior hire: give it the context, scope its access, keep a record. That&#8217;s the sight going in. Then own it like a commander: the judgment is yours, the review is yours, most of all in the places the model is confident and wrong and you&#8217;re tired and it&#8217;s late. That&#8217;s the sight coming out &#8212; and it&#8217;s the half that gets *harder to hold* as the model gets better, not easier.</span></p><p><span>Here&#8217;s the catch. The sight going in has to come from *you* &#8212; no vendor can see the workflow in your head; your domain knowledge is the asset, the AI only the instrument. But getting it out of your head, building the tool around it, and standing up the loop that watches the output is its own discipline &#8212; the part I kept getting wrong on my own. That&#8217;s the work I do with operators now: not sell them a model, but pull their workflow into the open and build the ownership, both ends, around it.</span></p><p><span>That&#8217;s the part that&#8217;s actually hard &#8212; and the part worth getting right. So if you&#8217;ve got an AI tool that half-works and you can&#8217;t tell whether the problem is the model or the way it was never built into how you operate: that&#8217;s the conversation worth having.</span></p><p><span>You can hand an AI the work. You can&#8217;t hand it the sight at either end &#8212; the context going in or the judgment coming out. And the better it gets, the more that&#8217;s worth remembering &#8212; because the better it gets, the more it tempts you to forget.</span></p><p><span>---</span></p><p><span>*This is the thread I pull on here &#8212; running AI like a readiness problem instead of a magic button. It&#8217;s the throughline of a book I&#8217;m writing on applying military readiness models to enterprise AI adoption, worked out in the open, one issue at a time. If that&#8217;s your thread, subscribe.*</span></p><p><span>*Sources: Andrej Karpathy, Sequoia AI Ascent 2026 (&#8221;From Vibe Coding to Agentic Engineering&#8221;) &#8212; jagged intelligence and the chess example (capability local to added training data, not general). Dell&#8217;Acqua, McFowland, Mollick, et al., &#8220;Navigating the Jagged Technological Frontier&#8221; (Organization Science, 2025; field experiment with 758 consultants) &#8212; on a task set beyond the AI&#8217;s frontier, those using AI were 19 percentage points less likely to reach a correct answer than those without it. Harvard Business Review, &#8220;Research: Why You Shouldn&#8217;t Treat AI Agents Like Employees&#8221; (Kropp et al., BCG, May 2026) &#8212; accountability does not transfer to the model.*</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://observeassessadvise.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Observe. Assess. Advise.! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Your AI Dashboard Is Measuring the Wrong Half]]></title><description><![CDATA[Only 39% of tech leaders are confident their AI investments will produce a positive financial impact.]]></description><link>https://observeassessadvise.substack.com/p/your-ai-dashboard-is-measuring-the</link><guid isPermaLink="false">https://observeassessadvise.substack.com/p/your-ai-dashboard-is-measuring-the</guid><dc:creator><![CDATA[Glen Lewis]]></dc:creator><pubDate>Sat, 30 May 2026 18:23:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!KMYx!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F393533e5-eb41-412d-bd57-af584a6e3975_682x682.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Only 39% of tech leaders are confident their AI investments will produce a positive financial impact. That&#8217;s Gartner, April 2026.</p><p>A month later, Gartner ran a separate survey of Chief Sales Officers. 31% named &#8220;difficulty proving the ROI of AI-driven tools&#8221; as a top challenge for 2026.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://observeassessadvise.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Observe. Assess. Advise.! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p>Read those two numbers together and something becomes clear. The people who approved the AI budgets are losing faith in them. And the people responsible for revenue can&#8217;t show the math. This is not a technology problem. It&#8217;s a measurement problem. And it&#8217;s everywhere.</p><p>Uber&#8217;s COO said it directly last month, on the record: &#8220;It&#8217;s hard to link AI use to product improvements.&#8221; His CTO had just disclosed four months of data &#8212; agent usage climbing from 32% to 84%, 70% of code commits AI-driven, monthly API costs between $500 and $2,000 per engineer. Every one of those numbers is real. Every one of them is the wrong number.</p><p>The military has a word for this distinction. Two words, actually.</p><div><hr></div><p><strong>MOPs and MOEs</strong></p><p>Measures of Performance &#8212; MOPs &#8212; answer one question: are we doing things right?</p><p>Missions flown. Rounds fired. Sorties completed. Reports filed on time. These numbers tell you whether the activity happened. They do not tell you whether the activity mattered.</p><p>Measures of Effectiveness &#8212; MOEs &#8212; answer a different question: are we doing the right things?</p><p>Did the air campaign degrade the enemy&#8217;s ability to maneuver? Did the supply line hold under pressure? Did the training actually produce units capable of executing the mission? MOEs connect activity to outcome. They are harder to build, harder to measure, and the only numbers that actually answer the question a commander &#8212; or a board &#8212; needs answered.</p><p>The Joint Publication 5-0, the doctrine governing US military operational planning, treats this distinction as foundational. You don&#8217;t start with activity metrics. You start by defining what success looks like, then you build the measures that tell you whether you&#8217;re achieving it.</p><p>Most AI deployments do the opposite. They instrument the activity first &#8212; tokens consumed, code commits, hours saved per task, adoption rate &#8212; because those numbers are easy to collect. The tools already track them. Then, months later, someone in the boardroom asks whether any of this is working, and the answer is a dashboard full of MOPs with no MOEs in sight.</p><div><hr></div><p><strong>What it actually looks like</strong></p><p>I run three parallel workstreams on a typical working day. In one window, AI-assisted development on a new feature. In another, a requirements document for a client. In a third, processing the day&#8217;s communications. None of that is sequential. The agent runs, I review, I move.</p><p>The productivity gain isn&#8217;t &#8220;I built the feature 30% faster.&#8221; It&#8217;s that three things got done in the time one used to take, and my afternoon opened up for work I couldn&#8217;t reach before. A client onboarding. A compliance review that had been deferred for two weeks. A proposal that needed to go out.</p><p>That freed-up time is net-new capacity. It didn&#8217;t exist last quarter. And it doesn&#8217;t appear anywhere on a standard AI productivity dashboard &#8212; because there&#8217;s no baseline to compare it against, and no column on the report to put it in.</p><p>That&#8217;s the gap Uber&#8217;s COO was naming. Not &#8220;AI doesn&#8217;t work.&#8221; His organization couldn&#8217;t yet connect the activity to the outcome, because nobody had built the measurement layer that makes that connection possible. MOPs everywhere. MOEs nowhere.</p><p>This is not unique to Uber. It&#8217;s the shape of AI adoption at scale right now. Adoption is real. Measurement is a generation behind.</p><div><hr></div><p><strong>Why this happens</strong></p><p>MOPs are instrumented by default. Your IDE counts commits. Your ticketing system counts closes. Your token dashboard counts tokens. The data exists because the tools generate it automatically.</p><p>MOEs require deliberate design. You have to define the outcome before you deploy the capability. What does success look like &#8212; not in activity terms, but in business terms? What would have to be true six months from now for this AI investment to have been worth it? Then you build the measurement layer that ties the activity back to that outcome.</p><p>Most AI pilots skipped this step. The implicit logic was: deploy the tool, demonstrate activity, let the value prove itself. It doesn&#8217;t work that way. If you didn&#8217;t define the outcome before you started, you have no reference point to measure against when the board asks.</p><p>The Brookings Institution mapped this problem economically last year, analyzing AI deployment in computer vision &#8212; one of AI&#8217;s most mature domains. 80% of tasks are technically feasible for AI. Only 23% are economically viable to deploy when real-world costs are factored in &#8212; customization, human oversight, domain expertise. Most organizations don&#8217;t know which side of that line they&#8217;re on before they commit. MOEs are how you find out. If you can&#8217;t define a measurable business outcome before you deploy &#8212; not activity metrics, but an actual outcome the business can count &#8212; you&#8217;re almost certainly working in the 77%. Defining your MOEs before deployment is the filter that puts you in the 23%.</p><div><hr></div><p><strong>The question to run on your next dashboard review</strong></p><p>Before your next AI metrics review, apply one filter to every number on the dashboard: is this a MOP or an MOE?</p><p>MOPs tell you what happened. MOEs tell you whether it mattered.</p><p>If you can find MOPs but not MOEs, you don&#8217;t have an AI ROI problem &#8212; you have something worse. You have no way to know whether your AI investment is working or not. The activity is real. The outcome definition is missing. And without it, no amount of activity data will answer the question the board is actually asking.</p><p>The doctrine for building that measurement layer has been sitting in military planning doctrine for forty years. Most boardrooms have never heard of it.</p><p>That&#8217;s what this newsletter is about.</p><div><hr></div><p><strong>MOEs and Commander&#8217;s Intent</strong></p><p>There&#8217;s a deeper connection worth naming here, one that comes from the doctrine that produced MOEs in the first place.</p><p>In military planning, MOEs are not standalone metrics. They are the feedback mechanism of what the doctrine calls Commander&#8217;s Intent &#8212; a precise, concise statement of the desired end state that gives subordinates enough guidance to act without further orders when the plan changes. Commander&#8217;s Intent doesn&#8217;t prescribe every step. It answers two questions: what does success look like, and why does it matter? Everything else flows from that.</p><p>One of Commander&#8217;s Intent&#8217;s four required elements is feedback loops &#8212; the specific mechanisms that tell the commander whether execution is tracking toward the intended outcome. MOEs are those loops. Without them, Commander&#8217;s Intent is a statement with no way to know if it&#8217;s being achieved.</p><p>Most AI deployments have neither. No defined end state before deployment, and no feedback loops after. The organization knows what the tools are doing &#8212; the MOPs are instrumented from day one &#8212; but has no mechanism to know whether any of it is pointing in the right direction.</p><p>This is the measurement problem in its complete form: it isn&#8217;t just that organizations lack MOEs. It&#8217;s that they skipped the step that makes MOEs possible &#8212; defining what success looks like before the first token is spent.</p><p>I&#8217;m writing a book on this &#8212; applying the full set of military operational frameworks to AI adoption. Commander&#8217;s Intent gets its own chapter. If this thread is useful, the book is where it goes deeper.</p><div><hr></div><p><strong>What&#8217;s next</strong></p><p>Next issue: before you can define your MOEs, you have to direct your AI toward the right outcomes in the first place. Most people treat AI agents like search engines &#8212; you ask them things and hope. The people getting results treat them differently. Next issue I&#8217;ll introduce the framework I use to brief AI agents the same way a Marine briefs a subordinate before a mission.</p><p>If you found this useful, forward it to one person in your organization who&#8217;s sitting on an AI dashboard full of MOPs.</p><p>&#8212; Glen Lewis Founding Principal, Coyote &amp; Quill <em>Observe. Assess. Advise.</em></p><div><hr></div><p><em>Sources: Gartner, April &amp; May 2026 | JP 5-0, Joint Doctrine for Planning | Brookings Institution / MIT FutureTech, August 2024</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://observeassessadvise.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Observe. Assess. Advise.! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>