Quick summary

  • An Application Metrics dashboard showed an AI-suggested formula produced files up to 83% larger than predicted, while a bitrate formula stayed near 5% error.
  • AI-generated code still needs production telemetry and measurement tied directly to user-visible outcomes.
  • Instrument predicted-versus-actual ratios for generated logic, define error budgets, and validate against fixtures before broad rollout.

What happened

A plausible formula that was wrong

A video conversion tool needed to predict output size before users waited for encoding. The first estimator, suggested by an AI assistant, multiplied source size by the ratio of target height to source height. Resolution sounded relevant, but the formula ignored the core behavior of bitrate encoding: bits per second and duration are the primary drivers of approximate file size.

Sentry Application Metrics recorded the ratio between actual and estimated size. The dashboard showed an average ratio of 1.57 for the old formula and a worst observed value of 1.83, meaning output could be 83 percent larger than promised. Replacing it with a bitrate-based calculation moved the average to 0.97 and kept transcodes within roughly five percent in the sample.

Telemetry as a test for AI code

Numeric metrics are not sampled and can be grouped by estimator, source format, destination format, or conversion mode. That made the primary failure obvious and revealed two additional patterns: copy mode overestimated by about 16 percent and GIF estimates were consistently about nine percent high. Those observations translate directly into prioritized engineering work.

The lesson is not to avoid AI-generated code. It is to avoid treating generated output as proof of correctness. For quantitative logic, define invariants, maintain representative fixtures, record predictions alongside outcomes, and alert on error ratios. Production observability becomes the feedback loop that catches invalid assumptions before they erode user trust.

Source

Application Metrics caught my broken size estimator

Why developers should care

AI-generated code still needs production telemetry and measurement tied directly to user-visible outcomes.

  1. 1Instrument predicted-versus-actual ratios for generated logic, define error budgets, and validate against fixtures before broad rollout.