Skip to main content
CloudCareerLabs Logo CloudCareerLabs
Observability Systems

Monitoring & Observability Prep.

Master PromQL metrics aggregation, Grafana dashboards, Alertmanager configurations, distributed tracing propagation, and structured log indexing.

SM
Written by Sachin Mehta • Founder & Principal Cloud Architect Principal cloud infrastructure specialist and systems architect. Former SRE.

Interactive Flashcards

# 1
Unreviewed

What is the difference between Monitoring and Observability?

# 2
Unreviewed

What are the "Three Pillars of Observability"?

# 3
Unreviewed

What is Prometheus, and how does it collect metrics?

# 4
Unreviewed

How do you monitor an Amazon EKS cluster?

# 5
Unreviewed

What is Grafana, and what is its role in the observability stack?

# 6
Unreviewed

What is PromQL and how is it used?

# 7
Unreviewed

What is the role of an Exporter in Prometheus monitoring?

# 8
Unreviewed

What is Structured Logging, and why is it preferred in production?

# 9
Unreviewed

What is Distributed Tracing, and what problem does it solve?

# 10
Unreviewed

How do you configure alerting thresholds to avoid alert fatigue?

# 11
Unreviewed

How do monitoring tools collect metrics in terms of push vs pull models?

# 12
Unreviewed

What observability stack is standard for Kubernetes environments?

Scenario Challenges

Select a scenario below to test your troubleshooting workflow.

Topic: Diagnosing Pod OOM Leak

An application pod in your cluster is crashing repeatedly due to OOMKilled errors. Isolate the issue and apply resource constraints.

Analyze the Grafana dashboard and run PromQL query topk(5, container_memory_working_set_bytes) to isolate the leaking pod. Click to select
Run kubectl get pod <pod> -o yaml to check if CPU/Memory resources.limits are configured. Click to select
Modify the Deployment manifest, adding explicit memory requests and limits to prevent host kernel out-of-memory triggers. Click to select
Selected Sequence

No steps selected yet. Click options above in sequence.