Every new office technology creates a period when output rises before quality systems catch up. Email made more communication possible before firms learned how much of it was defensive noise. Slideware gave unfinished strategy a finished look. Dashboards made weak measures feel authoritative. Generative AI is doing the same thing to knowledge work.

The useful word for the failure mode is unpleasant: slop. It means output that looks like work, travels like work, and consumes review time like work, yet stays thin, generic, inaccurate, stale, unowned, or misaligned with the real decision. Factually correct slop can still fail as work. Plausible drafts often survive the first glance—and that is part of the problem.

The Society for Human Resource Management (SHRM) 2026 workplace research sharpens the edge. It reports that 41% of workers use AI in their work, and that just under half of those users identify their own output as "AI slop." The phrase matters less than the admission. Many workers know that some of what AI helps them produce sits below the standard the organization would want if anyone stopped to inspect it.

The risk is review-shaped work produced faster than managers can judge it.

The Slop Layer

Executives usually discuss AI quality in two categories: model capability and human review. Better models reduce some defects. Human review catches others. Both statements are partly true, and together they still fall short of a control system.

Slop appears between generation and decision. Fluent drafts can be strategically empty. Tidy customer summaries can miss the real objection. Market memos can cite a respectable source while drawing an unsupported conclusion. Performance reviews can turn a hard personnel signal into polished mush. Project plans can sound decisive while hiding dependencies, owners, and dates. Anyone who has reviewed AI-assisted work at speed knows the pattern: the document arrives looking complete, structure takes the first minute, substance hunting takes the next five, and reconstructing the judgment that should have existed before the draft takes fifteen more.

Why Adoption Metrics Mislead

Stanford Human-Centered Artificial Intelligence (HAI) institute's 2026 AI Index reports that organizational AI adoption continued to rise, while agent use remains early and productivity gains are strongest in structured, measurable work where outputs are easy to monitor. That caveat is the doctrine. Real gains arrive when the task has a clear unit of output, a visible quality bar, and fast feedback. Support tickets, code tasks, document operations, and marketing variants are easier to measure; strategy, negotiation, hiring judgment, regulatory interpretation, and board advice are not.

When executives celebrate adoption without measuring the review burden, they confuse volume with operating gain. The company may produce more drafts, summaries, briefs, tickets, and options—while labor shifts from creation to inspection, from junior staff to managers, and from visible hours to invisible judgment debt. Microsoft's 2026 Work Trend Index points to the same organizational gap from another angle: many workers are moving faster than the systems around them. Leaders need to rearchitect work, including the definition of finished work, not just distribute more tools.

The audit

An AI slop audit is a sample-based quality review of AI-assisted work where low-grade output becomes an operating cost. The aim is to find where AI improves throughput and where it creates a new layer of managerial cleanup.

Slop type Signal Test Stop rule
Accuracy Claims, numbers, names, dates, citations, product facts, policy references. Can the worker show the source trail without reconstructing it after review? Pause AI use for the workflow if unverified claims reach customers, partners, regulators, or executives.
Specificity Generic arguments, interchangeable recommendations, context-free summaries. Would the output still make sense if the company, customer, or market name were changed? Reject drafts that omit the decision, constraint, owner, or tradeoff.
Judgment Advice that sounds balanced but avoids a choice. Does the document say what should happen next and what evidence would change the recommendation? Do not escalate AI-assisted analysis until a human has made the call explicit.
Review Burden Manager time spent checking, rewriting, de-risking, or apologizing for output. Did AI reduce total cycle time, or only move labor into review? Stop scaling when reviewer load rises faster than accepted output.
Learning Workers become better at the work, or only faster at prompting. Can the employee explain the reasoning without the generated draft? Limit delegation if AI use weakens the skill the role is supposed to develop.

The Managerial Cost

Slop is expensive because it hides inside apparently cheap output. One AI-generated brief can take six minutes to produce and forty minutes to repair. Summaries can save an analyst time and cost a partner trust. Sales notes can improve reply speed and flatten the customer's actual concern. Policy drafts can look complete until legal finds three jurisdictions merged into one confident paragraph.

This cost is rarely booked. It shows up as meeting drift, rework, vague feedback, managerial fatigue, slower promotion readiness, and a quiet decline in the standard of internal writing. The firm still feels busy—and may even feel more modern—while the work grows harder to trust. The audit should count accepted output ahead of generated volume. Decision quality matters more than document length. Reviewer minutes matter more than user satisfaction alone. How fast employees feel is a weak primary metric. Defensible decisions with less total waste are the standard.

The Quality Contract

AI-assisted work needs a simple ownership contract. The worker may use AI and still owns the output. Drafts that make factual claims need a source trail. Workers should mark where AI was used for synthesis, wording, calculation, code, research, or decision support. Reviewers need to know whether they are checking prose, facts, judgment, compliance, or customer risk.

That contract separates a drafting engine from an unacknowledged junior employee with no training record. An organization that would bar an intern from sending unsupported analysis into a board pack should bar a model-assisted draft from the same path when only the grammar is better. The quality contract should stay light for low-risk work and strict for consequential work. Brainstorming notes need no forensic record; customer renewal recommendations do. Social captions tolerate ambiguity that an HR investigation summary cannot. Code comments are a lighter control surface than production changes. Control should follow consequence.

The Executive Move

Pick five workflows where AI-assisted output is now common—customer communication, internal analysis, recruiting, performance management, product documentation, sales enablement, legal intake, support operations, finance commentary, or software delivery. Pull a sample of recent AI-assisted artifacts. Skip tool-satisfaction scores and measure four things instead.

  1. What share of generated output was accepted without major rework?
  2. What defects appeared most often: accuracy, specificity, judgment, tone, compliance, or missing context?
  3. How many review minutes were needed per accepted artifact?
  4. What rule would have prevented the worst output from reaching review?

Then change the workflow. Add source trails where facts matter. Add human decision boxes where judgment matters. Use templates only when they improve specificity. Set manager standards before adding more seats. Reward employees who use AI to improve the work; skip rewards for pure machine-polished volume.

Strong AI organizations will be able to show, with evidence, the workflows where AI raised speed, raised quality, raised risk, or should stay off-limits for now. Generated volume alone is a weak success metric. That is the slop doctrine: audit the artifact, price the review burden, and make quality visible before volume becomes the strategy.

Source Notes