Observability & SRE
Structured logging, metrics, tracing, health checks, and SLO-friendly instrumentation.
Generate Mock CPU and Memory Metrics in Python
Build a mock_host_metrics() generator that outputs realistic CPU and memory usage percentages for monitoring demos and tests.
import time
import random
def mock_host_metrics():
"""Generate mock CPU and memory metrics for a host."""
cpu_percent = round(random.uniform(10.0, 95.0), 1)
memory_percent = round(random.uniform(20.0, 90.0), 1)
memory_used_mb = round(random.uniform(512, 8192), 1)
return {
"timestamp": in…
How to Build a Consumer Lag Gauge in Python
Simulate Kafka consumer lag with a Python class that tracks lag over time and reports health and averages.
import time
import random
from collections import deque
class ConsumerLagGauge:
"""Mock consumer lag gauge measuring how far behind a consumer is."""
def __init__(self, producer_rate=10, consumer_rate=7, initial_lag=0):
self.producer_rate = producer_rate
self.consumer_rate = consumer_rate
…
How to Calculate Percentile Latency in Python
Generate mock latency samples with occasional spikes and compute 50th, 90th, 95th, and 99th percentile values in milliseconds.
import random
import statistics
def generate_latency_samples(n=1000):
"""Generate realistic mock latency data (ms) with occasional spikes."""
samples = []
for _ in range(n):
# Normal case: ~50ms with jitter
base = random.gauss(50, 5)
# 2% spike chance: slow downstream or GC pause
…
How to Calculate SLO Error Budget in Python
Simulate an SLO error budget by computing allowed downtime from a target availability percentage and mocking monthly incidents.
```python
import random
def calculate_error_budget(total_seconds: int, target_availability: float) -> float:
return (1.0 - target_availability) * total_seconds
def simulate_monthly_availability(seconds_in_month: int, budget_seconds: float) -> float:
# Mock: randomly consume a fraction of the error budget i…
How to Flush Metrics on Graceful Shutdown in Python
Register an atexit handler to automatically flush collected metrics when a Python process exits gracefully.
import atexit
import time
import random
class MetricsCollector:
def __init__(self):
self._metrics = []
atexit.register(self.flush)
def record(self, name, value):
self._metrics.append((name, value, time.time()))
def flush(self):
print(f"Flushing {len(self._metrics)} metri…
How to Mock HTTP Client Latency in Python
Simulate outbound HTTP request latency with configurable ranges to test timeouts, retries, and SLO monitoring without external services.
import time
import random
def mock_latency(host: str, min_ms: int = 100, max_ms: int = 500) -> dict:
"""Simulate an outbound HTTP request with mock latency."""
latency_ms = random.randint(min_ms, max_ms)
start = time.perf_counter()
time.sleep(latency_ms / 1000)
elapsed_ms = (time.perf_counter() - …
How to Process System Metrics (RSS, CPU) in Python
Simulate and aggregate RSS and CPU system metrics to compute averages and maximums for monitoring dashboards.
import random
import time
from collections import namedtuple
Metric = namedtuple("Metric", ["name", "value", "unit"])
def generate_metrics(num_metrics: int = 5) -> list:
"""Simulate a batch of system metrics."""
metrics = []
for i in range(num_metrics):
rss = random.randint(50, 500) # MB
…
How to Route Alerts by Severity in Python
Map alert severity levels to routing targets and simulate dispatching alerts to on-call pages, email, Slack, or logs.
def main():
# Severity levels with corresponding alert routing targets
routing_map = {
"critical": "call_page",
"high": "call_page",
"medium": "email_team",
"low": "slack_channel",
"info": "log_only"
}
# Simulated alerts with severity
alerts = [
{"na…
How to Simulate a Queue Depth Gauge in Python
Simulate a queue depth over time using a random enqueue/dequeue process, returning depth values that can be used for monitoring or testing dashboards.
import collections
import random
import time
def simulate_queue_depth(max_depth=10, steps=20):
queue = collections.deque()
depth_history = []
for _ in range(steps):
# Randomly enqueue or dequeue
if random.random() < 0.6 and len(queue) < max_depth:
queue.append("task")
…
How to mock SLI availability success ratio in Python
Simulate request outcomes with deterministic randomness and compute the SLI availability success ratio to check if a target is met.
import random
from collections import defaultdict
def mock_availability(num_requests=1000, target_ratio=0.995):
"""
Simulate request outcomes and compute the SLI availability success ratio.
Args:
num_requests: Total number of requests to simulate
target_ratio: Target availability rati…
Browse by section
Each section groups closely related Python snippets.
Observability & SRE — Python code examples
What you will find here
This page collects observability & sre snippets — short, copy-ready Python you can paste into our free online IDE and run without installing anything. Each sample includes a plain-English explanation and the full source code.
Samples vs tutorials and challenges
Samples are quick reference — one concept per page. For step-by-step teaching, use our Python tutorials. To test yourself, try quizzes or coding challenges. Clean up style with the Python formatter.