Written by Sachin Mehta • Founder & Principal Cloud ArchitectPrincipal cloud infrastructure specialist and systems architect. Former SRE.
Interactive Flashcards
# 1
Unreviewed
What is the difference between Monitoring and Observability?
Answer Guide
Monitoring tells you *when* a system is broken (using predefined metrics and thresholds, e.g. "CPU is > 90%"). Observability allows you to understand *why* a system is broken by analyzing internal states, outputs, and telemetry data (logs, metrics, traces) to debug novel problems.
Key Concepts Checklist
Evaluate difficulty:
# 2
Unreviewed
What are the "Three Pillars of Observability"?
Answer Guide
1. Metrics: Numeric values measured over time (e.g., CPU load, request count). 2. Logs: Timestamped text records of discrete events (e.g., error trace, auth success). 3. Traces: End-to-end paths of requests through a distributed system.
Key Concepts Checklist
Evaluate difficulty:
# 3
Unreviewed
What is Prometheus, and how does it collect metrics?
Answer Guide
Prometheus is a pull-based open-source monitoring and alerting system. It collects metrics by scraping HTTP endpoints (typically "/metrics") exposed by targets (like applications or exporters) at configured intervals, storing them as time-series data.
Key Concepts Checklist
Evaluate difficulty:
# 4
Unreviewed
How do you monitor an Amazon EKS cluster?
Answer Guide
Deploy Prometheus agents (or Prometheus Operator) to scrape container metrics, Node Exporter for node metrics, and Kube-state-metrics for Kubernetes API object metrics. Use Grafana to visualize the metrics, and integrate CloudWatch Container Insights for native AWS logging/metrics.
Key Concepts Checklist
Evaluate difficulty:
# 5
Unreviewed
What is Grafana, and what is its role in the observability stack?
Answer Guide
Grafana is an open-source analytics and visualization web application. It connects to data sources (like Prometheus, Elasticsearch, CloudWatch, Loki) and allows you to create interactive, real-time dashboards with charts and graphs.
Key Concepts Checklist
Evaluate difficulty:
# 6
Unreviewed
What is PromQL and how is it used?
Answer Guide
PromQL (Prometheus Query Language) is the query language used for Prometheus time-series data. It allows you to select, aggregate, and filter metrics in real-time (e.g. calculating the rate of HTTP requests: "sum(rate(http_requests_total[5m]))").
Key Concepts Checklist
Evaluate difficulty:
# 7
Unreviewed
What is the role of an Exporter in Prometheus monitoring?
Answer Guide
An Exporter is a helper agent that translates metrics from third-party systems (that don't natively output Prometheus format) into Prometheus-compatible metrics. Examples include Node Exporter (OS metrics), Blackbox Exporter (network endpoints), and pg_exporter (PostgreSQL).
Key Concepts Checklist
Evaluate difficulty:
# 8
Unreviewed
What is Structured Logging, and why is it preferred in production?
Answer Guide
Structured Logging writes logs in a standardized machine-readable format (typically JSON) instead of unstructured plain text. This allows log aggregation systems (like Elasticsearch, Loki, Splunk) to parse and index log fields automatically, enabling fast filtering and querying.
Key Concepts Checklist
Evaluate difficulty:
# 9
Unreviewed
What is Distributed Tracing, and what problem does it solve?
Answer Guide
Distributed Tracing tracks the flow of a single request across multiple microservices. By injecting a unique correlation ID into HTTP headers, it allows developers to visualize latency bottlenecks and locate the exact microservice where a request fails.
Key Concepts Checklist
Evaluate difficulty:
# 10
Unreviewed
How do you configure alerting thresholds to avoid alert fatigue?
Answer Guide
1. Alert only on user-facing symptoms (like high latency or high error rate) rather than cause-based warnings (like CPU spikes). 2. Define clear severity levels (e.g., Page vs Ticket). 3. Use grouping, throttling, and silencing rules in Alertmanager.
Key Concepts Checklist
Evaluate difficulty:
# 11
Unreviewed
How do monitoring tools collect metrics in terms of push vs pull models?
Answer Guide
In the Pull model (e.g., Prometheus), the server requests metrics from targets via HTTP scrape calls. In the Push model (e.g., Datadog, InfluxDB), the targets or local agents stream metrics outbound to a central daemon or API gateway. Pull requires active service discovery, while Push handles dynamic clients easily.
Key Concepts Checklist
Evaluate difficulty:
# 12
Unreviewed
What observability stack is standard for Kubernetes environments?
Answer Guide
The standard cloud-native stack is PLG (Prometheus for metrics, Loki for logs, Grafana for visualization) or EFK (Elasticsearch, Fluentd, Kibana). Often, OpenTelemetry is used for distributed tracing alongside collectors like Tempo or Jaeger.
Key Concepts Checklist
Evaluate difficulty:
Scenario Challenges
Select a scenario below to test your troubleshooting workflow.
Topic: Diagnosing Pod OOM Leak
An application pod in your cluster is crashing repeatedly due to OOMKilled errors. Isolate the issue and apply resource constraints.
Analyze the Grafana dashboard and run PromQL query topk(5, container_memory_working_set_bytes) to isolate the leaking pod.Click to select
Run kubectl get pod <pod> -o yaml to check if CPU/Memory resources.limits are configured.Click to select
Modify the Deployment manifest, adding explicit memory requests and limits to prevent host kernel out-of-memory triggers.Click to select
Selected Sequence
No steps selected yet. Click options above in sequence.
Topic: Blackbox Uptime Probing
Configure synthetic endpoint checking to verify if the external API gateway is returning HTTP 2xx statuses.
Deploy the Prometheus Blackbox Exporter in the Kubernetes cluster namespace.Click to select
Add the target endpoint domain to the scrape config in prometheus.yml pointing to the blackbox service.Click to select
Query the probe_success metric in Grafana to build an uptime alert dashboard.Click to select
Selected Sequence
No steps selected yet. Click options above in sequence.
Topic: Remediating Alert Fatigue
Restructure your Prometheus alert rules to stop pager notifications during transient 2-second CPU usage spikes.
Locate the CPU high utilization alert rule in the Prometheus configuration files.Click to select
Change the duration threshold from "for: 1m" to "for: 15m" to ignore transient short-lived spikes.Click to select
Configure Alertmanager grouping rules to combine similar alerts into a single Slack message notification.Click to select
Selected Sequence
No steps selected yet. Click options above in sequence.
Topic: Distributed Tracing Setup
Trace a slow request bottleneck traversing from your frontend Node service down to a Postgres database.
Add OpenTelemetry SDK instrumentation to the application source code files.Click to select
Ensure HTTP headers include W3C Trace Context propagation variables across outbound calls.Click to select
Forward spans to Jaeger/Tempo collector services and query trace visualizations for bottleneck segments.Click to select
Selected Sequence
No steps selected yet. Click options above in sequence.
Topic: Loki JSON Log Aggregation
Structure your application logs and ingest them in Grafana Loki with proper indexing variables.
Change the application logging framework output type to structured JSON format.Click to select
Configure Promtail DaemonSet to parse the JSON logs on EKS nodes.Click to select
Define static labels (like app, environment, and level) inside the Promtail pipeline stages config.Click to select
Selected Sequence
No steps selected yet. Click options above in sequence.
We value your privacy
We use cookies to analyze site traffic, personalize content, and support our free educational platforms. By clicking "Accept All", you consent to our use of cookies.