Mlls
A new product from the team at Sawmills
Get early access

Animated diagram: a field of healthy components; a few warm to a warning state, an investigation converges on them, and one verified issue is filed while the rest settle back to normal.

Live monitoring · 312 components Healthy Signal Investigating
Issues detected

AI that continuously watches over your production.

Watches everything · Investigates every signalAlerts only on verified issues · Identifies root causes and recommends fixes
Get early access How it works
Works with the stack you already run Datadog Prometheus Grafana Slack PagerDuty New Relic Splunk Coralogix +

Mills detects issues across your whole system, investigates so only real issues reach you, and hands you the fix, without adding noise.

Detect · Miss nothing real

Know of problems before your customers do

Mills watches every component across metrics, logs, events, and traces - without the noise or blind spots

Coverage map: every component and the detections watching it

The whole surface, watched

Every service, database, and dependency gets its own monitors, generated from a living map of your system.

Sensitive — without alerting you

Monitors are broad and sensitive on purpose — and none of it reaches your team. Only verified cases become an alert

Blind spots, hunted down

Mills audits its coverage against your topology, finds what isn't watched, and closes the gap itself.

Not just the pager-worthy

Slow regressions, stalled jobs, queues that keep growing. Nobody writes a threshold for those, so nobody sees them. Mills reports them at their real severity.

No thresholds to tune

Nobody sets the dial anymore

Mills writes its own monitors and recalibrates them as your system changes. No thresholds to pick, no monitors to maintain, no quarterly pass to turn the noise down and the coverage back up.

Nothing to pick

Every threshold is written & calibrated from the component and it's telemetry, not from a number someone guessed.

Nothing to re-tune

As traffic, topology, and release cadence change, the thresholds move with them.

Nothing to trade off

Sensitivity costs you nothing, because monitors never reaches your team unverified.

Investigate · Chase nothing false

Issues investigates before interrupting your team

Broad, sensitive monitors would bury a human in noise - so no human sees it. The agent does the triage, and your team is alerted only when an issue is verified and important.

Every signal investigated

Triggered monitors are grouped into cases — and every case is qualified based on observed facts.

A verifier gates every alert

Before a case escalates, Mills rechecks the evidence and rules out false positives.

Deliberately not alerting

Most cases are watched, correlated, and ruled out behind the scenes. That's a decision the agent makes

Investigation
Current explanation

A stale JWKS cache on one verifier pool best explains the regional concentration of verification failures.

Why these monitors belong together

4 monitors cover the same production service (token-service), triggered within four minutes, and use overlapping five-minute windows. Their error-budget, latency, 5xx-rate, and saturation readings moved at the same onset.

Still unknown

Customer impact is not confirmed from observability telemetry alone.

What we observed

Signature verification failures are clustered on one verifier pool and one cached key version.

Affected scope

3.1% of EU token-verification requests failed or retried during the observed window.

Evidence source

Datadog

Triggered monitors
token-service error budget burn
3.26% · trigger > 2% over 5m
token-service p95 latency
890.5 ms · trigger > 650 ms over 5m
token-service error rate
2.61% · trigger > 1.5% over 5m
token-service resource saturation
91.84% · trigger > 82% over 5m
Evidence timeline
Aug 2011:28 PM
token-service error budget burn triggered
Datadog read 3.26% against its trigger threshold > 2% over 5m.
11:29 PM+1m
token-service p95 latency triggered
Datadog read 890.5 ms against its trigger threshold > 650 ms over 5m.
11:30 PM+1m
token-service error rate triggered
Datadog read 2.61% against its trigger threshold > 1.5% over 5m.
11:31 PM+1m
token-service resource saturation triggered
Datadog read 91.84% against its trigger threshold > 82% over 5m.
An issue, investigated and verified before it pages you
Root cause analysis

Root cause, established

After a case escalates, the agent tests multiple hypotheses based on monitors and telemetry data, then adds the most likely root cause & evidence to the issue.

The investigation runs on whichever agent you point it at.

Mills itselfYour own agentAny vendor agent
Root cause analysis: root cause, confidence boundary, investigation activity, hypotheses
The root cause, with the activity that established it
Fix

From a verified issue to a fix.

Once the root cause is clear, Mills finds the faulty code, recommends a fix, and hands your coding agent a complete, evidence-backed brief or opens a pull request for review.

Fix contract: proposed change, code path, proposed tests, guardrails, draft pull request, and the patch diff
The fix contract, with the diff and the pull request
How it works

A good resolution requires perfecting every step

Five steps from a read-only connection to a fix in review.

Connect

Read-only access to the observability platform you already have. No code changes, no agents to deploy, no rip-and-replace.

Understand

Continuously builds a living map of your system, keeping it current as your environment changes.

Detect

Continuously watches metrics, logs, events, and traces - writing and calibrating its own monitors.

Investigate

Every monitor triggered is reviewed by an agent, not a human. only real issues reach you, with root cause attached.

Resolve

The recommended fix, with root cause attached goes to your coding agent as context, or Mills opens the pull request itself.

Your new normal.

Six things that change

Know before your customers do

Real issues surface early — including the ones no threshold would ever have caught.

No more blind spots

Coverage across your entire surface, audited proactively — instead of the fraction you had time to instrument.

A pager you trust again

The relevant alerts, and only the relevant alerts. Every page is worth reading — and acting on.

Nothing to maintain

No alert rules, no threshold tuning. Detections regenerate as your system changes.

Detect what you mean

Tell it what matters in plain language — "payments shouldn't queue too long" — and it's covered.

A closed loop

The verified issue goes to your developer or your SRE agent — or Mills runs the root-cause analysis itself.

Get early access

See Mills run against your own telemetry.

Hook it up to your existing telemetry — nothing to install, instrument, or configure. Within a day, Mills has mapped your system and started surfacing what your current alerting misses.

We'll reach out to you with next steps
Mlls Mills — The Production Intelligence Agent · from the team at Sawmills