epok

Logs, metrics and traces.
Investigate them together.

Send telemetry directly to Epok through supported open agents. Search logs, explore metrics and follow traces in one workspace. Included AI groups related alerts and drafts probable causes with source links, or reports insufficient evidence.

Google sign-in · ingest pauses after 14 days · no automatic charge

No autonomous production changes · synthetic demo data

The Epok incident workspace: a checkout saturation cascade with a reversible recovery recommendation and a probable-cause explanation with linked evidence, in a demo scenario demo scenario · synthetic data · no signup →
OPEN AGENTS

Documented open shippers send straight to Epok — there is no proprietary Epok agent for server-side telemetry.

ONE INCIDENT

Groups repeated alerts and related failures into one incident timeline.

CITED OR ABSTAINED

Every committed verdict links to the lines, spans and metrics behind it — or it abstains.

Send it once

Point the agents you already run at one destination.

Logs, metrics and traces land in one place, joined by trace and service.

as documented — install guides in the docs
Investigate without the AI

Search logs. Explore metrics. Follow traces. See services.

One workspace, before any AI is involved.

Epok log search: a query bar, a hits-over-time histogram, service, level and host facets, and a service-volume breakdown, on synthetic demo dataLogs — search, facets, patterns · synthetic dataEpok trace list with a selected trace opened as a span waterfall across browser, frontend, checkout, auth and payments services, on synthetic demo dataTraces — spans and waterfalls · synthetic dataEpok service map: request flow between eight services derived from trace spans, coloured by error rate, with a worst-first service table, on synthetic demo dataServices — request flow from spans · synthetic data

Metrics, live tail, dashboards and SLOs are in the same workspace. Write your own threshold or composite rules in the same search syntax. custom rules →

THE VERDICT THAT ADMITS WHEN IT DOESN'T KNOW

It commits when the evidence is there. When it isn't, it says so.

AI assistance is only useful if responders can distinguish evidence from fluency. Epok can abstain and show why the available signals do not support one cause. Use the trial to inspect both cases: committed verdicts and honest "not enough evidence" outcomes.

COMMITTED
Probable cause: checkout-service CPU saturation, sustained 97%. Cited to the metric + the two spans that corroborate it.
ABSTAINED
"Three candidates are close and none dominates — committing to one would be a guess." So it withholds a diagnosis and shows the three, ranked, with what each is missing.

IMMEDIATE — rule packs match supported evidence as it arrives. LEARNS — baselines need enough history for the service. WHEN WIRED — customer impact and trace-linked replay need a customer identifier and shared trace context. Each detector states what it needs. see the catalog →

proof →live incidentdetector cataloglimits

Recover · suggested, not applied

A next move, and the reason not to repeat it.

From a staged run on our own demo estate. Epok recommends; your responders execute.

Observed change: deploy payment v2.5.0suggested, not applied
Suggested: Roll back the deploy on payment

A change immediately preceded the errors; reverting it is the fastest reversible mitigation.

UNDO · EASYBLAST RADIUS · MEDIUMCONFIDENCE · LOW
⚠ This incident has re-fired 6× — if this was already attempted, escalate instead of repeating it.

$199/mo · 1 TB/month across logs, metrics and traces · 30-day retention · AI incident analysis included (daily allowance applies) · overage $0.20/GB · no host, query or cardinality charges.

plan limits →

See pricing →

Start with a representative service group.
Judge it on your next investigation.

Evaluating beside Datadog? Ship to both for 14 days — your paging stays authoritative. Plan your evaluation →