Observability & SRE

Metrics, logs, tracing, error budgets, and reliability engineering culture.

  • 19 Tracked terms
  • Last 30 days Feed window

What this topic collects on

An article joins this feed when it matches these terms. Each one is also a search of its own.

Latest in Observability & SRE


blogs.oracle.com > autonomous-ai-database > turn-any-sql-query-into-an-oci-monitoring-metric-from-inside-autonomous-ai-database

Turn Any SQL Query Into an OCI Monitoring Metric — From Inside Autonomous AI Database

2+ hour, 32+ min ago   (1158+ words) Information, tips, tricks and sample code for data management in an autonomous, cloud-native, data-driven world No agent. No function. No compute instance. Just a SELECT and a scheduler job. Autonomous AI Database gives you more than forty metrics in OCI…...


dev.to > martin_bammer_6838a4d3b65 > fastlogging-rs-high-performance-logging-for-many-different-programming-languages-772

fastlogging-rs: High-Performance Logging for many different Programming Languages

28+ min ago   (183+ words) Release, 0.9.0, is available with the following features: Compared to Python’s built-in logging module: When your application logs millions of messages, these speedups can turn minutes into seconds. fastlogging-rs is written in Rust and comes with thin wrappers for your favorite…...


gooddata.ai > blog > ai-observability-tools

AI Observability Tools: Do You Need a Separate One?

6+ hour, 22+ min ago   (17+ words) AI Observability Tools: Do You Need a Separate One, or Is It Already in Your Platform? GoodData...


dev.to > kharesam > the-three-events-that-turn-your-logs-into-an-slo-dashboard-1gie

The Three Events That Turn Your Logs Into an SLO Dashboard

1+ hour, 3+ min ago   (306+ words) That gap between "the infrastructure looks healthy" and "the system is doing its job" is where most observability budget quietly goes to waste. And closing it doesn't take a new tool. It takes logging the right three events per request,…...


dev.to > humzakt > four-stops-instead-of-thirty-rebuilding-the-dashboard-and-retiring-hand-rolled-ui-1a35

Four Stops Instead of Thirty: Rebuilding the Dashboard and Retiring Hand-Rolled UI

2+ hour, 34+ min ago   (1695+ words) An editor's real complaint about the main video-generation service's internal tool wasn't any single bug. It was the shape of using it: thirty confirmation stops on an attended run, a dashboard that buried the one thing you actually owed a…...


dev.to > npayyappilly > from-devops-to-sre-a-practitioners-roadmap-for-enterprise-transformation-252h

From DevOps to SRE: A Practitioner's Roadmap for Enterprise Transformation

4+ hour, 47+ min ago   (859+ words) This post maps the journey from DevOps foundations to SRE practice at the enterprise scale. It is not a theoretical model. It is a phased roadmap with specific exit gates for each phase, derived from the observable signals that distinguish…...


dev.to > tacio_souza > laboratorio-kubernetes-multi-no-no-home-lab-a-base-para-ckne-e-kubestronaut-156j

Laboratório Kubernetes Multi-Nó no Home Lab: A Base para CKNE e Kubestronaut

8+ hour ago   (206+ words) Subi meu laboratório de estudos para a jornada Kubernetes e não poderia estar mais animado com o que vem por aí. Na imagem, você vê a base de tudo: um cluster Kubernetes rodando em 3 VMs (1 Control Plane e 2 Workers) gerenciadas…...


dev.to > akarshan > the-coredns-black-hole-how-one-dead-dns-pod-broke-our-api-gateway-4hh3

The CoreDNS Black Hole: how one dead DNS pod broke our API gateway

8+ hour, 45+ min ago   (1087+ words) Our API gateway started handing this back to roughly a third of all requests: Restart the gateway and everything is fine. A day later, it's back. No deploy, no code change, no traffic spike, nothing in the backend logs — the…...


openobserve.ai > ai-observability

AI Observability | Monitor, Evaluate & Improve AI in Production

11+ hour, 3+ min ago   (741+ words) Most tools log what your AI did. AI observability on OpenObserve also tells you whether it was any good, and helps you make it better - by scoring quality on live traffic, running experiments before you ship, and turning human review into…...


dev.to > metalbear > how-mondaycom-runs-agent-evals-against-real-dependencies-webinar-recap-41ge

How monday.com Runs Agent Evals Against Real Dependencies: Webinar Recap

9+ hour, 37+ min ago   (529+ words) If you missed the session, here's the recap: what agent evals actually need to check, why mocks fall short compared to the agent calling real systems, and how monday.com built their agent evals infrastructure. One very important part in…...