Goodhart's law -- when a measure becomes a target, it ceases to be a good measure -- was formulated about economic policy, but nowhere does it apply more cleanly than to workplace productivity scores. Once people know a specific number is being watched and tied to consequences, behavior reorganizes around moving that number, and the number stops reliably reflecting the thing it was designed to measure.

The specific ways scores get gamed

The gaming patterns are remarkably consistent across tools and organizations. Idle-time metrics get gamed with mouse jigglers -- physical or software tools that simulate activity during genuine breaks -- which are cheap, widely available, and functionally undetectable by most monitoring software. Application-category scoring gets gamed by keeping a 'productive' application open and in the foreground while doing something else in a different window or on a second monitor the software isn't tracking. Keystroke-based scoring gets gamed by typing filler text or repeatedly pressing a key that generates input events without producing meaningful output.

  • Mouse jigglers defeat idle-time detection almost universally
  • Keeping a 'productive' app in the foreground while working elsewhere defeats category scoring
  • Filler keystrokes defeat raw input-frequency scoring
  • Scheduling low-value tasks during monitored hours and real work during unmonitored personal time defeats almost any metric

The gaming isn't usually dishonesty -- it's a rational response

It's tempting to frame score-gaming as an integrity problem with the employee. In most documented cases it's better understood as a rational response to a measurement system the person correctly perceives as inaccurate: someone doing genuinely valuable work that the software scores poorly (deep thinking, phone-based client work, collaborative whiteboarding) has a reasonable incentive to protect their measured score from a metric they know doesn't reflect their actual contribution. The gaming behavior is evidence the metric is broken, not primarily evidence about the person's character.

When smart, otherwise honest employees start gaming a metric, that's usually the metric's fault, not theirs.

Why this matters even if you don't care about the ethics

Even setting aside the trust and fairness questions, gamed metrics are simply bad data. An organization making staffing, promotion, or performance decisions based on scores that a meaningful share of the workforce has learned to manipulate is making decisions on noise dressed up as signal, and the noise isn't randomly distributed -- it correlates with who figured out the gaming technique first, which tends to correlate with tenure and technical comfort rather than with actual performance.

What resists gaming better

Metrics tied to verifiable output -- tickets closed, code merged, deals moved to the next stage, deliverables submitted on time -- are harder to game because gaming them requires actually producing the fake output, which is a much higher bar than jiggling a mouse. These output-based metrics have their own weaknesses (they can undervalue collaboration, mentorship, and quality), but they are structurally more resistant to the specific gaming patterns that plague activity-based scores, because the thing being measured is closer to the thing that actually matters.

A specific example from a documented rollout

A customer support organization tied a portion of quarterly bonus calculation to an activity-based productivity score from its monitoring software. Within two months of the policy taking effect, mouse-jiggler devices -- inexpensive and widely available -- began appearing in expense reports and personal purchase histories at a noticeably higher rate among the team than before the policy, according to the organization's own later internal review. The team's average measured productivity score rose substantially during this period. Actual ticket-resolution volume, tracked separately and not directly tied to bonus calculation, stayed essentially flat. The organization had, in effect, paid for a higher number without any corresponding increase in the underlying work getting done, and only discovered the gap because it happened to also track ticket volume as an independent measure -- an organization relying on the activity score alone would have seen only a positive trend and no reason to question it. Readers comparing this approach with a commercial implementation can review see it here from Monitask.

Why punitive responses to gaming rarely fix the underlying problem

The instinctive organizational response to discovering score-gaming is often to tighten enforcement -- ban mouse jigglers explicitly, add stricter idle-time thresholds, increase monitoring depth to catch the workaround. This response treats the gaming as the problem rather than as a symptom, and it tends to produce an arms race: employees find the next workaround, the organization tightens further, and trust erodes further at each round, all while the underlying question -- does this metric actually reflect the work being done -- remains unaddressed. Organizations that instead treat sustained gaming as a signal to redesign the metric itself, moving toward output-based measurement where the work allows it, generally resolve the underlying tension rather than escalating it. For an independent reference, consult RescueTime blog.

The specific case of AI-assisted work complicating input-based metrics

The growing use of AI-assisted writing, coding, and analysis tools introduces a newer complication for input-based productivity metrics specifically: a person using an AI tool to draft a document in a few minutes, then spending most of the remaining time reviewing and refining the output, generates a very different keystroke and application-usage pattern than someone drafting the same document manually over a longer period, even though both may produce comparable final output quality. A raw activity-based score built around older assumptions of how long a given deliverable 'should' take, based on manual-drafting baselines, will systematically misread the AI-assisted workflow as unusually low-effort or suspiciously fast, when it may simply reflect a different, legitimate way of producing the same result.

Organizations still relying heavily on activity-based scoring should expect this specific measurement gap to widen as AI-assisted tools become more embedded in ordinary knowledge work, and it's one more reason -- alongside the gaming vulnerabilities discussed in the main article -- to weight output-based measures more heavily than activity-based ones going forward, since output measures remain meaningful regardless of which tools and workflow a person used to produce that output.

One further point worth making explicit to any team relying on activity-based scores: communicate clearly that gaming detection isn't the goal of a well-designed program -- fixing the underlying metric is. An organization known for escalating enforcement against gaming, rather than questioning why gaming became rational in the first place, tends to see the arms race described earlier accelerate rather than resolve.

Ultimately, a metric worth trusting is one an organization would be comfortable showing employees the exact mechanics of -- if a scoring formula would embarrass the organization to explain openly, that's usually a sign the formula needs to change before the gaming problem does.

Key takeaway: Treat any spike in score-gaming as a signal the metric is broken, not primarily a discipline problem -- and prefer output-based metrics over activity-based ones wherever the work allows it.