Performance Monitoring in Cassandra
Keep an eye on your database health! Performance monitoring helps you track how fast and stable Cassandra is running, so you can fix problems early and maintain smooth operations.
đ The Story: Rachel's Silent Performance Degradation
Rachel's Cassandra cluster was "fine" - no alerts, no obvious errors. But users complained about slowness. She checked - P99 latency had crept from 10ms to 250ms over 3 months. Nobody noticed because there was no monitoring. By the time she investigated, 847 pending compactions, 15GB heap at 98%, and 200+ SSTables per table. A slow-motion disaster.
đą The Silent Degradation
How It Started (Month 1):
The Slow Creep (Month 2):
The Crisis (Month 3):
Root Causes Nobody Caught:
- Compaction Falling Behind: Write rate > compaction rate
- Heap Leak: Slow memory leak over 3 months
- SSTable Explosion: STCS with no size limit
- GC Pauses Growing: From 100ms â 15 seconds
- No Alerts: Everything just slowly got worse
- No Baseline: Didn't know what "good" looked like
The Painful Recovery:
- đ° Downtime: 6 hours maintenance window
- âī¸ Manual Compactions: Forced on all tables
- đ Heap Dump Analysis: Found memory leak in driver
- đ Revenue Impact: $50K lost from slow performance
- đ Team Stress: Weekend emergency work
â The Monitoring Solution
What Rachel Implemented:
1. Comprehensive Metrics Collection
2. Proactive Alerting
3. Baseline Tracking
The Results:
- ⥠Early Detection: Caught issues in minutes, not months
- đ P99 Latency: Stable at 8-12ms (instead of 250ms)
- đ Alert Fired: Week 2 - caught heap leak immediately
- đ° Zero Downtime: No more emergency maintenance
- đ Users Happy: Consistent performance
- đ¯ Proactive: Fix issues before users notice
Rachel learned: What you don't monitor will eventually fail! đ
đ Monitoring Fundamentals
What to monitor and why it matters!
đ¯ The Four Pillars of Monitoring
- Latency: How fast are requests? (P50, P99, P999)
- Throughput: How many ops/sec? (reads, writes, total)
- Errors: What's failing? (timeouts, unavailable, errors)
- Saturation: How full are resources? (CPU, disk, memory)
Remember: If you can't measure it, you can't improve it!
Critical Metrics by Category
Performance
- Read latency: P99 < 50ms
- Write latency: P99 < 20ms
- Throughput: ops/sec
- Timeouts: Should be 0%
- Unavailable errors: Track rate
Resources
- CPU usage: < 70%
- Heap usage: 60-80%
- Disk I/O: < 80%
- Network: < 70%
- Disk space: < 70% full
Cassandra Internals
- Pending compactions: < 20
- SSTable count: < 100/table
- GC pause time: < 200ms
- Dropped messages: Track rate
- Hints: Should be minimal
Cluster Health
- Node status: All UP
- Gossip state: Consistent
- Schema agreement: All nodes
- Streaming: Track repairs
- Ownership: Balanced tokens
Monitoring Levels
Level 1: Basic (Minimum)
What: Essential survival metrics
- Node up/down status
- CPU, memory, disk usage
- Basic throughput
- P99 latency
Good for: Small deployments, getting started
Level 2: Production (Recommended)
What: Comprehensive monitoring
- All Level 1 metrics
- Compaction metrics
- GC statistics
- SSTable counts
- Dropped messages
- Cache hit rates
Good for: Production clusters, most use cases
Level 3: Advanced (Enterprise)
What: Deep visibility
- All Level 2 metrics
- Per-table statistics
- Distributed tracing
- Anomaly detection
- Capacity forecasting
- SLA tracking
Good for: Large scale, critical workloads
đ§ nodetool Monitoring Commands
Essential commands for quick health checks!
Quick Health Check
Detailed Metrics Commands
Table Statistics
Resource Usage
Performance Metrics
Monitoring Script
đĄ JMX Metrics Export
Exporting metrics to monitoring systems!
Prometheus + Grafana Setup
đ¯ Best Practice Stack
Prometheus scrapes metrics from JMX exporter â stores time-series data â Grafana visualizes with dashboards. This is the industry-standard setup for Cassandra monitoring.
Step 1: Install JMX Exporter
Step 2: Configure Cassandra
Step 3: Configure Prometheus
Step 4: Create Grafana Dashboards
Key Metrics to Export
| Metric | JMX Path | Alert Threshold |
|---|---|---|
| Read Latency P99 | org.apache.cassandra.metrics:type=ClientRequest,scope=Read,name=Latency | > 50ms |
| Write Latency P99 | org.apache.cassandra.metrics:type=ClientRequest,scope=Write,name=Latency | > 20ms |
| Pending Compactions | org.apache.cassandra.metrics:type=Compaction,name=PendingTasks | > 20 |
| Heap Usage | java.lang:type=Memory | > 85% |
| Dropped Messages | org.apache.cassandra.metrics:type=DroppedMessage | > 0 |
đ¨ Alerting Rules
Setting up proactive alerts!
Prometheus Alert Rules
Alert Priority Levels
| Severity | Response Time | Examples |
|---|---|---|
| Critical đ´ | Immediate (< 15 min) | Node down, dropped messages, OOM |
| Warning â ī¸ | Within hours | High latency, compaction backlog |
| Info âšī¸ | Track trends | Capacity planning, growth rates |
đ Dashboard Essentials
What to display on your monitoring dashboards!
The Perfect Cassandra Dashboard
Performance Panel
- Read latency (P50, P99, P999)
- Write latency (P50, P99, P999)
- Operations per second
- Timeout rate
- Error rate
âąī¸ 1-minute intervals
Cassandra Health Panel
- Pending compactions
- SSTable count per table
- GC pause times
- Dropped messages
- Hints stored
đ 5-minute intervals
Resources Panel
- CPU usage per node
- Heap usage (%)
- Disk I/O (read/write)
- Network throughput
- Disk space remaining
đ 1-minute intervals
Cluster Overview Panel
- Node status (UP/DOWN)
- Schema agreement
- Ownership distribution
- Active repairs
- Streaming status
đ 5-minute intervals
Dashboard Design Tips
1. Most Important Metrics on Top
What you check first should be at the top!
- Row 1: P99 latency, throughput, error rate
- Row 2: Node status, pending compactions, heap
- Row 3: Detailed metrics, per-table stats
2. Use Color Coding
Visual cues for quick assessment!
- đĸ Green: Healthy (< warning threshold)
- đĄ Yellow: Warning (needs attention)
- đ´ Red: Critical (immediate action)
3. Historical Context
Show trends, not just current values!
- Last 6 hours for operational view
- Last 7 days for trend analysis
- Overlay events (deployments, repairs)
The "NOC Screen" Test
If you display your dashboard on a wall monitor, can you tell cluster health from 10 feet away?
- â Big numbers for P99 latency
- â Status indicators (green/yellow/red)
- â Trend graphs (up/down/stable)
- â Tiny text, complex tables
đŧ Interview Questions & Expert Answers
Master performance monitoring for your interview!
Answer: (1) P99 read/write latency - user experience, (2) Pending compactions - long-term stability, (3) Heap usage - OOM risk, (4) Dropped messages - overload indicator, (5) Node status - availability. These five catch 90% of problems before users notice.
The Critical Five:
1. P99 Latency (Most Important!)
2. Pending Compactions
3. Heap Usage
4. Dropped Messages
5. Node Status
Why These Five?
- Comprehensive: Cover performance, stability, capacity
- Early Warning: Catch issues before users notice
- Actionable: Clear remediation steps
- Essential: 90% of problems show up here
Key Takeaway: Monitor user experience (latency), long-term health (compactions), resource limits (heap), overload signals (drops), and availability (node status). These five metrics catch most problems early!
Answer: Use tiered alerting with leading indicators: (1) Warning alerts at 60-80% of critical thresholds, (2) Trend alerts for slow degradation, (3) Rate-of-change alerts for sudden spikes, (4) Anomaly detection for unusual patterns. Alert on metrics that predict problems (compactions, heap) before they affect users (latency).
Tiered Alerting Strategy:
Level 1: Leading Indicators (Predict Problems)
Level 2: Trend Alerts (Catch Slow Degradation)
Level 3: Rate-of-Change Alerts (Sudden Problems)
Level 4: Anomaly Detection (Unusual Patterns)
Complete Alert Configuration:
Alert Fatigue Prevention:
- â Tune thresholds: Adjust based on normal patterns
- â Require duration: "for: 10m" prevents flapping
- â Group related alerts: Don't spam with 100 alerts
- â Actionable only: Every alert needs clear action
- â Don't alert on: Metrics you can't act on
Key Takeaway: Proactive monitoring uses leading indicators (compactions, heap) and trends (slow degradation) to catch problems hours or days before users complain. Alert on predictive metrics, not just symptoms!
Answer: Monitoring answers "is it broken?" with predefined metrics (CPU, latency). Observability answers "why is it broken?" with arbitrary queries into system behavior. For Cassandra: monitoring catches known problems (high GC), observability debugs unknown problems (why is this specific query slow?). Need both.
Monitoring (Traditional Approach):
Observability (Modern Approach):
Cassandra-Specific Examples:
| Problem | Monitoring Says | Observability Reveals |
|---|---|---|
| Slow reads | P99 = 250ms â ī¸ | Trace: Reading from 237 SSTables! |
| High CPU | CPU = 95% â ī¸ | Logs: Compaction stuck on huge partition |
| Timeouts | Timeout rate = 15% â ī¸ | Trace: CL=ALL on cross-DC query! |
| Disk full | Disk = 95% â ī¸ | Query: Table X has 500GB uncompacted! |
Implementing Observability:
Why You Need Both:
- Monitoring: Catches known problems automatically
- Observability: Debugs unknown/new problems
- Together: Monitoring alerts, observability investigates
Key Takeaway: Monitoring tells you THAT there's a problem. Observability tells you WHY. For Cassandra: use monitoring for alerts, observability for debugging!
Answer: (1) Take heap dump with jmap, (2) Analyze with Eclipse MAT or jhat to find large objects, (3) Check nodetool tablestats for huge partitions, (4) Check row cache size, (5) Review memtable settings. Common causes: row cache too large, giant partitions, memtables oversized, or actual memory leak in driver/app code.
Diagnostic Process:
Step 1: Take Heap Dump
Step 2: Analyze Heap Dump
Step 3: Check Cassandra Statistics
Step 4: Common Causes & Fixes
| Cause | Diagnosis | Fix |
|---|---|---|
| Row Cache Too Large | Heap dump shows cache > 50% | Reduce row_cache_size_in_mb |
| Giant Partitions | tablestats shows > 1GB partitions | Fix data model (add bucketing) |
| Memtables Oversized | memtable_heap_space > 25% heap | Reduce memtable size |
| Driver Leak | Heap dump shows driver objects | Update driver, fix app code |
| Too Many SSTables | Bloom filters + indexes huge | Force compaction |
Step 5: Apply Fixes
Key Takeaway: 95% heap usually means: row cache too large, giant partitions, or memtables oversized. Heap dump analysis reveals the culprit. Fix by tuning caches, fixing data model, or reducing memtable size!
Answer: 5-minute health check: (1) nodetool status - all nodes UN?, (2) nodetool tpstats - any pending/blocked?, (3) nodetool compactionstats - < 20 pending?, (4) nodetool gcstats - < 200ms pauses?, (5) nodetool netstats - zero drops?. These five commands catch 90% of problems.
The 5-Minute Health Check:
Command 1: nodetool status
Command 2: nodetool tpstats
Command 3: nodetool compactionstats
Command 4: nodetool gcstats
Command 5: nodetool netstats
Bonus Commands (If Time Permits):
Decision Tree:
Key Takeaway: Five nodetool commands in 5 minutes: status (availability), tpstats (bottlenecks), compactionstats (long-term health), gcstats (JVM health), netstats (overload). These catch 90% of problems!
đ Chapter Summary: Monitoring Mastery
You now understand production-grade Cassandra monitoring!
The Essential Five Metrics:
- ⥠P99 Latency: User experience (< 50ms)
- đ Pending Compactions: Long-term stability (< 20)
- đž Heap Usage: OOM risk (60-80%)
- đ Dropped Messages: Overload signal (= 0)
- â Node Status: Availability (all UN)
5-Minute Health Check:
Monitoring Stack:
Prometheus (metrics) + Grafana (visualization) + Alerting (proactive)
Remember Rachel: What you don't monitor will eventually fail! đ
đą Responsive Ad đą