As AI takes over more of the work that once took teams of subject-matter experts, some companies are learning the hard way that their human in the loop is rather far outside the loop. In this week's feature, learn how your growing company can ensure your safeguards aren't asleep at the wheel.
| In this issue | ||||||
|
More than one in three firms in a September Richmond Fed survey said changing prices had become harder than a year earlier, and almost one-third already use inflation-linked or index-based pricing. An index clause settles part of each renewal increase at signing. If the customer asks for a cap, model it as a multi-year concession beside the upfront discount.
New Copilot seats to arrive with metered billing switched on
Microsoft is making pay-as-you-go billing the default on new Microsoft 365 Copilot Business licenses. The change starts Oct. 19 for seats bought directly from the company and Dec. 1 for seats bought through a reseller, a date Microsoft moved back from Nov. 2 (Microsoft).
The new default ceiling
Each new license comes with a spending limit of 4,000 Copilot Credits per user per month, and Microsoft bills for whatever gets used.
Microsoft's announcement gives no price per credit. CRN estimates that the default limit works out to roughly $40 per user per month. At that rate, 25 users could generate about $1,000 a month in usage charges. Copilot Cowork and Work IQ APIs are among the features that consume credits.
One new seat can change existing billing
For an existing Copilot customer, the new seat can reach well beyond the person who gets it.
Microsoft's notice to administrators says buying a new Copilot Business seat deploys usage-based billing across all Microsoft 365 Copilot subscriptions in the account. An existing spending policy stays in place. WindowsForum reports that added seats, upgrades and subscriptions moved from another reseller are covered as well.
Admins can set the limit to zero in Cost Management in the Microsoft 365 admin center. That shuts off metering, but it also blocks the features that require it.
Microsoft told resellers that early usage "can create a natural path to expanded consumption."
Founder read: The cost of your next Copilot seat can include new consumption from people who already had Copilot. When you model an added seat, measure the change in total account exposure. Keep the limit at zero until one documented workflow has been tested and priced.
Pay growth slowed to 3% as white-collar industries shed jobs
Employers added 29,000 jobs in September, while information, professional and business services, and financial activities each lost positions. Average hourly earnings rose 3% over the past 12 months, the slowest annual increase since May 2021 (SHRM). Founder read: A softer labor market gives you a chance to fix compensation debt that was harder to address when market pay was rising faster. Use some of the room to correct pay compression, stale leveling and people whose scope grew faster than their compensation before hiring accelerates again.
New Jersey says a 1099 by itself won't prove contractor status
A New Jersey rule that took effect Oct. 1 applies a three-part ABC test under the state's wage, sick leave and unemployment laws. The third part asks whether the worker is customarily engaged in an independently established business. A contractor agreement, separate registration and 1099 are each insufficient on their own (Insurance Journal). Founder read: Ask what business remains if you disappear. For a New Jersey contractor, look for other customers, marketing, its own employees, equipment and an operation that exists independently of your company. The paperwork matters less than whether an independent business exists.
Twenty-seven companies took a third of Q3 venture funding
Global venture funding reached $159 billion in the third quarter, up 53% from a year earlier. The 27 companies that raised at least $1 billion absorbed roughly one-third of it. Seed funding totaled $13 billion, and $2.6 billion of that went into rounds of $100 million or more (Crunchbase News). Founder read: Strip mega-rounds out of your fundraising comps. If you use broad seed statistics to set valuation, runway or hiring assumptions, rounds of $100 million or more are counted in those statistics and are a different financing market.
Blu Dot surpasses 2,000% ROAS with self-serve CTV ads
Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Here’s how:
After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.
The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.
“For CTV campaigns, Roku has been a top performer,” said Claire Folkestad, Paid Media Strategist, Blu Dot. “Comping to our other platforms, we have seen really strong ROAS… and highly efficient CPMs, lower than any other CTV partner we've worked with.”
Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.
The human in the loop has a maintenance problem
The last line of defense is being trained by the thing it defends against

A Sydney University researcher was reading a Deloitte report for the Australian government when he came across a book that did not exist.
The report credited the book to a law professor at his own university. Chris Rudge had never heard of the book and recognized the problem immediately.
"I instantaneously knew it was either hallucinated by AI or the world's best kept secret because I'd never heard of the book and it sounded preposterous," he told The Associated Press.
Deloitte later disclosed that the report had been produced with "a generative AI language system, Azure OpenAI." The firm revised it and agreed to refund the final installment of the US$290,000 contract. Deloitte said the updates did not change the report's substantive findings or recommendations.
The useful detail is how the mistake was found. Rudge did not need another AI system to tell him the citation looked wrong. He already knew enough about the subject to notice that something was off.
That ability is becoming a more valuable part of AI work.
When plausible is the problem
MIT research scientist Christian Catalini put the problem neatly in an Oct. 2 Harvard Business Review essay: "When execution is cheap, verification becomes more valuable."
His examples are deliberately mundane. An expert rejects an answer that looks perfectly plausible. A controller spots a revenue figure that includes deferred bookings. A security engineer stops code that passed its tests but would have given an AI agent write access to a production database.
These are harder failures than obvious nonsense. The output looks reasonable. Catching the problem requires somebody who knows what reasonable looks like.
Companies already rely on this arrangement. AI writes the code, drafts the customer response, prepares the analysis or produces the report. A person checks it before anything consequential happens.
That person inherits the accountability too.
A British Columbia tribunal made the point unusually concrete in 2024 when it ruled against Air Canada over incorrect information its chatbot gave a customer about a bereavement fare. The airline argued, among other things, that the chatbot was a separate legal entity responsible for its own actions. The tribunal did not find that persuasive. Air Canada was responsible for information on its website, including information produced by the bot.
Software developer Alex Ewerlöf makes the same argument about AI-generated code in less legal language. Whoever ships it, he writes, is accountable for it "regardless of how you produced it. So you better understand it."
Understanding it requires retaining some ability to do the work yourself.
When AI went away
A study of colonoscopies in Poland offers one of the stranger examples of what may happen when that ability changes.
Nineteen experienced endoscopists at four centers began routinely using AI assistance to help detect adenomas, the precancerous growths doctors look for during a colonoscopy. Each physician had already performed more than 2,000 procedures.
Researchers then examined colonoscopies those doctors performed without AI assistance.
Before AI had been introduced into their routine work, the physicians found adenomas in 28.4% of unassisted procedures. After they had become accustomed to AI, the detection rate during unassisted procedures was 22.4%.
The study was observational. The patients differed between the two periods, and the researchers could not establish that AI use caused the decline. Still, the result points toward a question that will matter well outside gastroenterology: what happens to unaided performance after assistance becomes normal?
Software researchers have found a related pattern during learning.
Anthropic ran a randomized experiment with 52 mostly junior engineers who were learning an unfamiliar Python library. One group could use AI. The other coded by hand.
On a quiz afterward, the AI-assisted engineers averaged 50%. The hand-coding group averaged 67%. The largest difference appeared on debugging questions.
For a company using AI to produce software, debugging is also one of the skills required to decide whether the AI produced something worth shipping.
The test
A larger experiment with nearly 1,000 high school math students in Turkey pushes the issue back another step.
Students using a standard ChatGPT-style chatbot performed much better while practicing. Their scores were 48% higher than the no-AI group's average.
Then the AI went away for the exam.
Those students scored 17% below the no-AI group's exam average, a difference of about 5 points out of 100. Researchers found that students often asked the chatbot for solutions and copied them. The paper also found evidence that the students did not realize how much the assistance had hurt their learning.
The researchers also tested a different version of the AI tutor. It was prompted to provide hints rather than simply hand over answers, and the system incorporated solutions and common mistakes prepared by two math teachers hired by the research team. Building the teachers' input was labor-intensive.
Students using that version performed about as well on the unaided exam as students who had no AI during practice. The difference was not statistically significant.
That result offers a more useful distinction than AI versus no AI. How the assistance is designed, and what the user is required to do while receiving it, may matter considerably.
Anthropic saw something similar in its engineering experiment. Participants who asked AI for explanations scored better than those who delegated the work more completely, although the researchers did not claim that asking for explanations caused the higher scores.
No clean check
Coding has one advantage over many kinds of business work: software can often be tested quickly.
Catalini argues that AI has moved fastest in areas where verification is relatively easy. Run the test. Check the output. Reject what fails.
Many important decisions have no equivalent test suite.
A financial forecast can be internally consistent and still rest on a bad assumption. A customer answer can sound polished while misstating policy. A market analysis can cite a book that does not exist. An AI-generated security recommendation can survive routine checks while introducing a risk nobody thought to test for.
In those situations, verification depends more heavily on judgment accumulated through experience. Catalini describes that judgment as the product of "years of friction" with customers, technologies, markets and regulators.
AI can remove some of that friction. That is part of its appeal.
Companies now have to decide which friction they can safely remove.
Ewerlöf says he still debugs code "the old way" periodically because "brain is like muscles: use it or lose it." He readily admits the cost: "Does it make me slower? Yes."
That cost becomes easier to justify when the person is responsible for approving work that cannot be checked mechanically.
For a founder, the useful audit is fairly concrete. Pick the AI-assisted work where a bad answer could cost you money, a customer, a security incident or a legal problem. Identify the person who gives the final approval. Then ask whether that person still does enough of the underlying work without AI to recognize a plausible mistake when one appears.
That leads to a harder operating decision: which jobs require the person approving AI's work to keep doing some of that work unaided? For those roles, make it part of the job. A security engineer might still review code without an AI summary. A finance lead might build or check a model from the source numbers. A customer support lead might periodically answer a difficult case without an AI draft.
Rudge caught Deloitte's invented book because he had never heard of the book and it sounded preposterous.
Before you call human review a safeguard, make sure the human can still do the work they're being asked to judge.
A pilot can hit every technical goal and still end in another sales meeting unless both sides agreed beforehand what success buys you. Crimmins gets specific about choosing metrics the pilot can move, securing both a business champion and a technical champion, and agreeing before kickoff that hitting the success criteria leads to a paid contract. The show notes have no chapters, so start at the beginning.
You can answer a surprising amount of the competitive-position question with closed opportunities you already have. Dave Kellogg asks four questions of them: how often a competitor appears, your win rate against it, how that changes by segment, and whether the competitor is appearing more often. Pull four quarters of closed deals and you have the beginnings of a competitive trend line without buying another research product. Start at "Then Track the Competitive Battle Deal by Deal." The sections above it can wait.
Monthly cash projections can make the quarter look survivable while one particular week is a mess. Float's free Excel template takes your bank balance, open invoices, approved bills and payroll dates and shows the week cash falls below the floor you set. There is no sign-up form. The example uses a UK company and pounds, so read "Adapting the example outside the UK" before entering tax payments. You can stop at "How Float fits."
SOC 2 Ready in 14 days. Three sessions from you.

Enterprise buyers will not put your product near their customer data without a SOC 2 report. Sprinto gets you audit ready in 14 days, across three working sessions. AI agents do the collecting, you approve. Your auditor signs off.
I run The Pricing Reset, a six-week engagement for founders of sales-led B2B SaaS companies whose product has outgrown the pricing they set a few years ago. I rebuild the packaging, price points and discount rules, then help put the new pricing live on new deals so the team can run it without every exception landing on the founder.
The work continues through a 90-day measurement window, with check-ins at days 30, 60 and 90 and a final readout against the company's own starting numbers.
If your pricing has not kept pace with the product, book a 20-minute call at redwoodridge.net/book, or reply with your current pricing page. I'll tell you whether The Pricing Reset applies and, if it doesn't, what I'd look at instead.
Thanks for reading. See you next Wednesday.
— Jason
P.S. Every experienced reviewer started as a junior. May's The Junior You Didn't Hire looks at what happens when you stop hiring them.



