Performance Tuning

Performance Monitoring in Cassandra

Keep an eye on your database health! Performance monitoring helps you track how fast and stable Cassandra is running, so you can fix problems early and maintain smooth operations.

📖 The Story: Rachel's Silent Performance Degradation

Rachel's Cassandra cluster was "fine" - no alerts, no obvious errors. But users complained about slowness. She checked - P99 latency had crept from 10ms to 250ms over 3 months. Nobody noticed because there was no monitoring. By the time she investigated, 847 pending compactions, 15GB heap at 98%, and 200+ SSTables per table. A slow-motion disaster.

😱 The Silent Degradation

How It Started (Month 1):

-- Initial state (nobody monitoring): P99 latency: 10ms ← Good! Pending compactions: 5 ← Normal Heap usage: 60% ← Healthy SSTable count: 45 ← Fine Everything looks fine... or does it?

The Slow Creep (Month 2):

-- What Rachel didn't see: P99 latency: 35ms ← Getting worse! Pending compactions: 45 ← Growing! Heap usage: 75% ← Creeping up! SSTable count: 95 ← Doubling! // But no alerts! // No monitoring! // Nobody noticed!

The Crisis (Month 3):

-- When users started complaining: P99 latency: 250ms ← DISASTER! Pending compactions: 847 ← BACKLOG! Heap usage: 98% ← CRITICAL! SSTable count: 237 ← OUT OF CONTROL! GC pause max: 15,000ms ← 15 SECONDS! Users: "Why is everything so slow?" Rachel: "Let me check... OH NO!"

Root Causes Nobody Caught:

  1. Compaction Falling Behind: Write rate > compaction rate
  2. Heap Leak: Slow memory leak over 3 months
  3. SSTable Explosion: STCS with no size limit
  4. GC Pauses Growing: From 100ms → 15 seconds
  5. No Alerts: Everything just slowly got worse
  6. No Baseline: Didn't know what "good" looked like

The Painful Recovery:

  • 💰 Downtime: 6 hours maintenance window
  • âš™ī¸ Manual Compactions: Forced on all tables
  • 🔄 Heap Dump Analysis: Found memory leak in driver
  • 📉 Revenue Impact: $50K lost from slow performance
  • 😓 Team Stress: Weekend emergency work

✅ The Monitoring Solution

What Rachel Implemented:

1. Comprehensive Metrics Collection

-- Prometheus + Grafana setup - JMX metrics export (every 15 seconds) - System metrics (CPU, disk, network) - Application metrics (latency, throughput) - Custom dashboards for each node

2. Proactive Alerting

-- Alert rules (catch problems early!): - P99 latency > 50ms (warn), > 100ms (critical) - Pending compactions > 20 (warn), > 50 (critical) - Heap usage > 80% (warn), > 90% (critical) - GC pause > 1s (critical) - SSTable count > 100 per table (warn)

3. Baseline Tracking

-- Establish "normal" for each metric - Historical graphs (30 days) - Trend analysis (detect slow degradation) - Anomaly detection (catch spikes)

The Results:

  • ⚡ Early Detection: Caught issues in minutes, not months
  • 📊 P99 Latency: Stable at 8-12ms (instead of 250ms)
  • 🔔 Alert Fired: Week 2 - caught heap leak immediately
  • 💰 Zero Downtime: No more emergency maintenance
  • 😊 Users Happy: Consistent performance
  • đŸŽ¯ Proactive: Fix issues before users notice

Rachel learned: What you don't monitor will eventually fail! 📈

📊 Monitoring Fundamentals

What to monitor and why it matters!

đŸŽ¯ The Four Pillars of Monitoring

  1. Latency: How fast are requests? (P50, P99, P999)
  2. Throughput: How many ops/sec? (reads, writes, total)
  3. Errors: What's failing? (timeouts, unavailable, errors)
  4. Saturation: How full are resources? (CPU, disk, memory)

Remember: If you can't measure it, you can't improve it!

Critical Metrics by Category

⚡

Performance

  • Read latency: P99 < 50ms
  • Write latency: P99 < 20ms
  • Throughput: ops/sec
  • Timeouts: Should be 0%
  • Unavailable errors: Track rate
💾

Resources

  • CPU usage: < 70%
  • Heap usage: 60-80%
  • Disk I/O: < 80%
  • Network: < 70%
  • Disk space: < 70% full
🔄

Cassandra Internals

  • Pending compactions: < 20
  • SSTable count: < 100/table
  • GC pause time: < 200ms
  • Dropped messages: Track rate
  • Hints: Should be minimal
📈

Cluster Health

  • Node status: All UP
  • Gossip state: Consistent
  • Schema agreement: All nodes
  • Streaming: Track repairs
  • Ownership: Balanced tokens

Monitoring Levels

Level 1: Basic (Minimum)

What: Essential survival metrics

  • Node up/down status
  • CPU, memory, disk usage
  • Basic throughput
  • P99 latency

Good for: Small deployments, getting started

Level 2: Production (Recommended)

What: Comprehensive monitoring

  • All Level 1 metrics
  • Compaction metrics
  • GC statistics
  • SSTable counts
  • Dropped messages
  • Cache hit rates

Good for: Production clusters, most use cases

Level 3: Advanced (Enterprise)

What: Deep visibility

  • All Level 2 metrics
  • Per-table statistics
  • Distributed tracing
  • Anomaly detection
  • Capacity forecasting
  • SLA tracking

Good for: Large scale, critical workloads

🔧 nodetool Monitoring Commands

Essential commands for quick health checks!

Quick Health Check

#################################### # 5-MINUTE HEALTH CHECK #################################### -- 1. Cluster status nodetool status /* Datacenter: datacenter1 Status=Up/Down, State=Normal/Leaving/Joining UN 192.168.1.10 156.48 GB 32 100.0% abc123... UN 192.168.1.11 158.32 GB 32 100.0% def456... UN 192.168.1.12 155.91 GB 32 100.0% ghi789... ✅ All nodes UN (Up Normal) ✅ Balanced ownership (~33% each) ✅ Similar data size */ -- 2. Compaction status nodetool compactionstats /* pending tasks: 12 ← Should be < 20 SSTable count: 87 ← Should be < 100 per table âš ī¸ If pending > 50 → falling behind! */ -- 3. GC statistics nodetool gcstats /* Interval Max GC Total GC Stdev GC# 600 sec 157 ms 487 ms 23 ms 12 ✅ Max < 200ms ✅ Total < 500ms per interval */ -- 4. Thread pools nodetool tpstats /* Pool Name Active Pending Blocked ReadStage 8 0 0 ← Should be 0 MutationStage 4 0 0 ← Should be 0 âš ī¸ Pending > 100 → bottleneck! */ -- 5. Dropped messages (CRITICAL!) nodetool netstats | grep -i dropped /* Dropped: 0 ← MUST be 0! ❌ Any drops = serious problem! */

Detailed Metrics Commands

Table Statistics

-- Per-table metrics: nodetool tablestats keyspace.tablename /* Key metrics to watch: - Space used: Is it growing as expected? - SSTable count: Should be < 100 - Bloom filter size: Grows with data - Read latency: P99 should be < 10ms - Write latency: P99 should be < 5ms */ -- Quick check all tables: nodetool tablestats | grep -E "Table:|SSTable count:"

Resource Usage

-- Memory usage: nodetool info | grep -E "Heap|Off" /* Heap Memory (MB): 11245 / 16384 (69%) Off Heap Memory (MB): 87 */ -- Cache statistics: nodetool info | grep Cache /* Row Cache: size 2048 MB, capacity 4096 MB, hit rate 0.89 Key Cache: size 256 MB, capacity 512 MB, hit rate 0.95 ✅ High hit rates (> 80%) */ -- Data size per node: nodetool status | awk '{print $6}'

Performance Metrics

-- Latency histograms: nodetool proxyhistograms /* proxy histograms Percentile Read Latency Write Latency 50% 1.23 ms 0.87 ms 75% 2.45 ms 1.23 ms 95% 8.76 ms 3.45 ms 98% 15.2 ms 5.67 ms 99% 28.4 ms 8.92 ms ← Watch this! Min 0.42 ms 0.31 ms Max 247 ms 156 ms ← Watch for spikes! */ -- Timeout statistics: nodetool tablestats | grep -i timeout

Monitoring Script

-- health_check.sh (run every 5 minutes): #!/bin/bash LOGFILE="/var/log/cassandra/health.log" ALERT_FILE="/var/log/cassandra/alerts.log" echo "=== $(date) ===" >> $LOGFILE # Check compaction PENDING=$(nodetool compactionstats | grep "pending tasks" | awk '{print $3}') if [ $PENDING -gt 50 ]; then echo "ALERT: Pending compactions: $PENDING" >> $ALERT_FILE fi # Check heap HEAP=$(nodetool info | grep "Heap Memory" | awk -F'[(/]' '{print $2*100/$3}') if (( $(echo "$HEAP > 90" | bc -l) )); then echo "ALERT: Heap usage: $HEAP%" >> $ALERT_FILE fi # Check GC MAX_GC=$(nodetool gcstats | grep sec | awk '{print $2}' | sed 's/ms//') if (( $(echo "$MAX_GC > 1000" | bc -l) )); then echo "ALERT: Max GC pause: ${MAX_GC}ms" >> $ALERT_FILE fi # Check dropped messages DROPPED=$(nodetool netstats | grep Dropped | awk '{sum+=$2} END {print sum}') if [ $DROPPED -gt 0 ]; then echo "ALERT: Dropped messages: $DROPPED" >> $ALERT_FILE fi nodetool status >> $LOGFILE nodetool tpstats >> $LOGFILE

📡 JMX Metrics Export

Exporting metrics to monitoring systems!

Prometheus + Grafana Setup

đŸŽ¯ Best Practice Stack

Prometheus scrapes metrics from JMX exporter → stores time-series data → Grafana visualizes with dashboards. This is the industry-standard setup for Cassandra monitoring.

Step 1: Install JMX Exporter

-- Download JMX exporter: wget https://repo1.maven.org/maven2/io/prometheus/jmx/jmx_prometheus_javaagent/0.18.0/jmx_prometheus_javaagent-0.18.0.jar \ -O /opt/cassandra/jmx_exporter.jar -- Create config file (jmx_exporter.yaml): --- lowercaseOutputName: true lowercaseOutputLabelNames: true whitelistObjectNames: - org.apache.cassandra.metrics:* - java.lang:type=GarbageCollector,* - java.lang:type=Memory - java.lang:type=OperatingSystem

Step 2: Configure Cassandra

-- Add to jvm.options: -javaagent:/opt/cassandra/jmx_exporter.jar=7070:/opt/cassandra/jmx_exporter.yaml -- Restart Cassandra: sudo systemctl restart cassandra -- Verify metrics exposed: curl http://localhost:7070/metrics /* # HELP cassandra_table_readlatency_99thpercentile # TYPE cassandra_table_readlatency_99thpercentile gauge cassandra_table_readlatency_99thpercentile 8.7 ... */

Step 3: Configure Prometheus

-- prometheus.yml: global: scrape_interval: 15s evaluation_interval: 15s scrape_configs: - job_name: 'cassandra' static_configs: - targets: - 'node1:7070' - 'node2:7070' - 'node3:7070' labels: cluster: 'production'

Step 4: Create Grafana Dashboards

-- Import pre-built dashboard: // 1. Go to grafana.com/dashboards // 2. Search "Cassandra" // 3. Import dashboard ID: 11298 (popular one) -- Or create custom panels for: - Read/Write latency (P50, P99, P999) - Throughput (ops/sec) - Pending compactions - Heap usage - GC pause times - SSTable counts

Key Metrics to Export

Metric JMX Path Alert Threshold
Read Latency P99 org.apache.cassandra.metrics:type=ClientRequest,scope=Read,name=Latency > 50ms
Write Latency P99 org.apache.cassandra.metrics:type=ClientRequest,scope=Write,name=Latency > 20ms
Pending Compactions org.apache.cassandra.metrics:type=Compaction,name=PendingTasks > 20
Heap Usage java.lang:type=Memory > 85%
Dropped Messages org.apache.cassandra.metrics:type=DroppedMessage > 0

🚨 Alerting Rules

Setting up proactive alerts!

Prometheus Alert Rules

#################################### # cassandra_alerts.yml #################################### groups: - name: cassandra_alerts interval: 30s rules: # ========== PERFORMANCE ALERTS ========== - alert: HighReadLatency expr: cassandra_table_readlatency_99thpercentile > 50 for: 5m labels: severity: warning annotations: summary: "High read latency on {{ $labels.instance }}" description: "P99 read latency is {{ $value }}ms (> 50ms)" - alert: HighWriteLatency expr: cassandra_table_writelatency_99thpercentile > 20 for: 5m labels: severity: warning annotations: summary: "High write latency on {{ $labels.instance }}" description: "P99 write latency is {{ $value }}ms (> 20ms)" # ========== COMPACTION ALERTS ========== - alert: CompactionBacklog expr: cassandra_compaction_pendingtasks > 20 for: 10m labels: severity: warning annotations: summary: "Compaction backlog on {{ $labels.instance }}" description: "{{ $value }} pending compaction tasks" - alert: CompactionCritical expr: cassandra_compaction_pendingtasks > 100 for: 5m labels: severity: critical annotations: summary: "CRITICAL: Compaction backlog on {{ $labels.instance }}" description: "{{ $value }} pending tasks - compaction falling behind!" # ========== MEMORY ALERTS ========== - alert: HighHeapUsage expr: (jvm_memory_used_bytes{area="heap"} / jvm_memory_max_bytes{area="heap"}) > 0.85 for: 10m labels: severity: warning annotations: summary: "High heap usage on {{ $labels.instance }}" description: "Heap usage at {{ $value | humanizePercentage }}" - alert: CriticalHeapUsage expr: (jvm_memory_used_bytes{area="heap"} / jvm_memory_max_bytes{area="heap"}) > 0.95 for: 5m labels: severity: critical annotations: summary: "CRITICAL: Heap usage on {{ $labels.instance }}" description: "Heap at {{ $value | humanizePercentage }} - OOM imminent!" # ========== GC ALERTS ========== - alert: LongGCPauses expr: cassandra_storage_gc_pausetime_max > 1000 for: 5m labels: severity: critical annotations: summary: "Long GC pauses on {{ $labels.instance }}" description: "Max GC pause: {{ $value }}ms (> 1s)" # ========== MESSAGE DROPS ========== - alert: DroppedMessages expr: rate(cassandra_droppedmessage_dropped_count[5m]) > 0 for: 2m labels: severity: critical annotations: summary: "Dropped messages on {{ $labels.instance }}" description: "Dropping {{ $value }} messages/sec" # ========== NODE STATUS ========== - alert: NodeDown expr: up{job="cassandra"} == 0 for: 2m labels: severity: critical annotations: summary: "Cassandra node DOWN: {{ $labels.instance }}" description: "Node has been down for 2 minutes" # ========== DISK SPACE ========== - alert: LowDiskSpace expr: (node_filesystem_avail_bytes / node_filesystem_size_bytes) < 0.15 for: 10m labels: severity: warning annotations: summary: "Low disk space on {{ $labels.instance }}" description: "Only {{ $value | humanizePercentage }} space remaining" - alert: CriticalDiskSpace expr: (node_filesystem_avail_bytes / node_filesystem_size_bytes) < 0.10 for: 5m labels: severity: critical annotations: summary: "CRITICAL: Disk space on {{ $labels.instance }}" description: "Only {{ $value | humanizePercentage }} remaining!"

Alert Priority Levels

Severity Response Time Examples
Critical 🔴 Immediate (< 15 min) Node down, dropped messages, OOM
Warning âš ī¸ Within hours High latency, compaction backlog
Info â„šī¸ Track trends Capacity planning, growth rates

📊 Dashboard Essentials

What to display on your monitoring dashboards!

The Perfect Cassandra Dashboard

⚡

Performance Panel

  • Read latency (P50, P99, P999)
  • Write latency (P50, P99, P999)
  • Operations per second
  • Timeout rate
  • Error rate

âąī¸ 1-minute intervals

🔄

Cassandra Health Panel

  • Pending compactions
  • SSTable count per table
  • GC pause times
  • Dropped messages
  • Hints stored

📊 5-minute intervals

💾

Resources Panel

  • CPU usage per node
  • Heap usage (%)
  • Disk I/O (read/write)
  • Network throughput
  • Disk space remaining

📈 1-minute intervals

📡

Cluster Overview Panel

  • Node status (UP/DOWN)
  • Schema agreement
  • Ownership distribution
  • Active repairs
  • Streaming status

🔍 5-minute intervals

Dashboard Design Tips

1. Most Important Metrics on Top

What you check first should be at the top!

  • Row 1: P99 latency, throughput, error rate
  • Row 2: Node status, pending compactions, heap
  • Row 3: Detailed metrics, per-table stats

2. Use Color Coding

Visual cues for quick assessment!

  • đŸŸĸ Green: Healthy (< warning threshold)
  • 🟡 Yellow: Warning (needs attention)
  • 🔴 Red: Critical (immediate action)

3. Historical Context

Show trends, not just current values!

  • Last 6 hours for operational view
  • Last 7 days for trend analysis
  • Overlay events (deployments, repairs)

The "NOC Screen" Test

If you display your dashboard on a wall monitor, can you tell cluster health from 10 feet away?

  • ✅ Big numbers for P99 latency
  • ✅ Status indicators (green/yellow/red)
  • ✅ Trend graphs (up/down/stable)
  • ❌ Tiny text, complex tables

đŸ’ŧ Interview Questions & Expert Answers

Master performance monitoring for your interview!

1 What are the top 5 metrics you would monitor for a production Cassandra cluster? â–ŧ

Answer: (1) P99 read/write latency - user experience, (2) Pending compactions - long-term stability, (3) Heap usage - OOM risk, (4) Dropped messages - overload indicator, (5) Node status - availability. These five catch 90% of problems before users notice.

The Critical Five:

1. P99 Latency (Most Important!)

-- Why P99, not average: Average latency: 8ms ← Looks great! P99 latency: 250ms ← Users suffering! -- What it tells you: // - User experience (1 in 100 requests) // - GC pause impact // - Compaction interference // - Overload signals -- Alert thresholds: P99 read > 50ms: Warning P99 read > 100ms: Critical P99 write > 20ms: Warning P99 write > 50ms: Critical

2. Pending Compactions

-- Why it matters: Pending: 847 tasks ← Disaster brewing! SSTables: 237 ← Read amplification! -- Leading indicator: // Day 1: 5 pending (normal) // Week 1: 20 pending (warning!) // Week 2: 50 pending (critical!) // Week 4: 500 pending (too late!) -- Catches problems early!

3. Heap Usage

-- OOM predictor: Heap: 98% ← OOM imminent! GC frequency: 1/sec ← Thrashing! -- Sweet spot: 60-80%: Healthy ✅ 80-90%: Warning âš ī¸ > 90%: Critical ❌ -- Prevents OOM crashes!

4. Dropped Messages

-- Cluster overload indicator: Dropped: 1247 messages/sec ← Overloaded! -- What causes drops: // - Thread pool saturation // - Too much load // - Slow nodes // - Network issues -- Must be ZERO! Any drops = serious problem!

5. Node Status

-- Availability: nodetool status UN: 3 nodes ← All UP ✅ DN: 0 nodes ← None DOWN ✅ -- What to watch: // - Node down (immediate alert!) // - Gossip issues // - Schema disagreement // - Ownership imbalance

Why These Five?

  • Comprehensive: Cover performance, stability, capacity
  • Early Warning: Catch issues before users notice
  • Actionable: Clear remediation steps
  • Essential: 90% of problems show up here

Key Takeaway: Monitor user experience (latency), long-term health (compactions), resource limits (heap), overload signals (drops), and availability (node status). These five metrics catch most problems early!

2 How would you set up proactive alerting to catch performance degradation before users complain? â–ŧ

Answer: Use tiered alerting with leading indicators: (1) Warning alerts at 60-80% of critical thresholds, (2) Trend alerts for slow degradation, (3) Rate-of-change alerts for sudden spikes, (4) Anomaly detection for unusual patterns. Alert on metrics that predict problems (compactions, heap) before they affect users (latency).

Tiered Alerting Strategy:

Level 1: Leading Indicators (Predict Problems)

-- Alert BEFORE users affected: # Compaction backlog (leading indicator): 5 pending: Normal 20 pending: âš ī¸ WARNING (investigate) 50 pending: ❌ CRITICAL (act now!) # Heap creeping up: 60-75%: Normal 75-85%: âš ī¸ WARNING (monitor closely) > 85%: ❌ CRITICAL (add capacity/tune) These warn hours/days before users notice!

Level 2: Trend Alerts (Catch Slow Degradation)

-- Detect gradual worsening: # P99 latency trending up: if P99_today > P99_last_week * 1.5: alert("Latency degrading: was 10ms, now 15ms") # SSTable count growing: if SSTables_today > SSTables_last_month * 2: alert("SSTables doubled: compaction not keeping up") Catches Rachel's 3-month degradation!

Level 3: Rate-of-Change Alerts (Sudden Problems)

-- Catch sudden spikes: # Latency spike: if rate(P99_latency[5m]) > 50%: alert("Latency spiked 50% in 5 minutes!") # Error rate surge: if errors_per_sec > baseline * 10: alert("Error rate 10x normal!") Catches deployment issues, outages

Level 4: Anomaly Detection (Unusual Patterns)

-- Machine learning for anomalies: # Traffic pattern unusual: if current_ops > mean + 3 * stddev: alert("Traffic 3΃ above normal") # GC pattern changed: if gc_frequency_today != gc_pattern_last_30d: alert("GC behavior changed") Catches weird, unexpected issues

Complete Alert Configuration:

-- Prometheus alert rules: # Level 1: Leading indicators - alert: CompactionBacklogWarning expr: cassandra_compaction_pendingtasks > 20 for: 10m severity: warning # Level 2: Trend - alert: LatencyTrending expr: | cassandra_latency_p99 > cassandra_latency_p99 offset 7d * 1.5 for: 1h severity: warning # Level 3: Rate of change - alert: LatencySpike expr: | (cassandra_latency_p99 - cassandra_latency_p99 offset 5m) / cassandra_latency_p99 offset 5m > 0.5 for: 2m severity: critical # Level 4: Anomaly - alert: TrafficAnomaly expr: | abs(cassandra_ops_per_sec - avg_over_time(cassandra_ops_per_sec[7d])) > 3 * stddev_over_time(cassandra_ops_per_sec[7d]) for: 10m severity: warning

Alert Fatigue Prevention:

  • ✅ Tune thresholds: Adjust based on normal patterns
  • ✅ Require duration: "for: 10m" prevents flapping
  • ✅ Group related alerts: Don't spam with 100 alerts
  • ✅ Actionable only: Every alert needs clear action
  • ❌ Don't alert on: Metrics you can't act on

Key Takeaway: Proactive monitoring uses leading indicators (compactions, heap) and trends (slow degradation) to catch problems hours or days before users complain. Alert on predictive metrics, not just symptoms!

3 What's the difference between monitoring and observability? Why does it matter for Cassandra? â–ŧ

Answer: Monitoring answers "is it broken?" with predefined metrics (CPU, latency). Observability answers "why is it broken?" with arbitrary queries into system behavior. For Cassandra: monitoring catches known problems (high GC), observability debugs unknown problems (why is this specific query slow?). Need both.

Monitoring (Traditional Approach):

-- Known metrics, predefined dashboards: # What you track: - CPU usage - Memory usage - P99 latency - Throughput - Error rate # Questions it answers: "Is latency high?" → Yes, P99 = 250ms "Is heap full?" → Yes, 95% # What it CAN'T answer: "WHY is latency high?" → 🤷 "WHICH queries are slow?" → 🤷 "WHAT changed?" → 🤷

Observability (Modern Approach):

-- Flexible exploration, arbitrary queries: # Tools: - Distributed tracing (see request flow) - Structured logging (query by any field) - Metrics with high cardinality # Questions it answers: "WHY is latency high?" → Trace shows: 200ms in compaction "WHICH queries are slow?" → SELECT * FROM huge_table (no LIMIT!) "WHAT changed at 3pm?" → Deployment introduced N+1 query pattern

Cassandra-Specific Examples:

Problem Monitoring Says Observability Reveals
Slow reads P99 = 250ms âš ī¸ Trace: Reading from 237 SSTables!
High CPU CPU = 95% âš ī¸ Logs: Compaction stuck on huge partition
Timeouts Timeout rate = 15% âš ī¸ Trace: CL=ALL on cross-DC query!
Disk full Disk = 95% âš ī¸ Query: Table X has 500GB uncompacted!

Implementing Observability:

-- 1. Distributed Tracing (Jaeger/Zipkin): # See request flow across nodes Request → Node1 (coordinator) → Node2 (replica) → Node3 (replica) 2ms 150ms (GC pause!) 3ms Now you know: Node2's GC caused slowness! -- 2. Structured Logging: log.info({ "query": query, "table": "users", "latency_ms": 250, "sstables_read": 237, ← Aha! "consistency_level": "QUORUM" }) Query: "Show me all queries with sstables_read > 100" -- 3. High-Cardinality Metrics: cassandra_latency{ table="users", operation="read", consistency="QUORUM" } = 250ms Query: "Which tables have high latency?"

Why You Need Both:

  • Monitoring: Catches known problems automatically
  • Observability: Debugs unknown/new problems
  • Together: Monitoring alerts, observability investigates

Key Takeaway: Monitoring tells you THAT there's a problem. Observability tells you WHY. For Cassandra: use monitoring for alerts, observability for debugging!

4 A node shows 95% heap usage but no obvious memory leak. How would you diagnose this? â–ŧ

Answer: (1) Take heap dump with jmap, (2) Analyze with Eclipse MAT or jhat to find large objects, (3) Check nodetool tablestats for huge partitions, (4) Check row cache size, (5) Review memtable settings. Common causes: row cache too large, giant partitions, memtables oversized, or actual memory leak in driver/app code.

Diagnostic Process:

Step 1: Take Heap Dump

-- Capture heap snapshot: jmap -dump:format=b,file=/tmp/heap.bin -- Or wait for automatic dump (jvm.options): -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/cassandra/heapdump.hprof -- Check size: ls -lh /tmp/heap.bin // 14.8 GB ← Close to 16GB heap limit!

Step 2: Analyze Heap Dump

-- Use Eclipse MAT (Memory Analyzer Tool): // 1. Open heap dump // 2. Run "Leak Suspects Report" // 3. Look at "Dominator Tree" -- Common findings: # Finding 1: Huge row cache org.apache.cassandra.cache.RowCacheKey: 8.5 GB (57%) ↑ Row cache is 8.5GB! # Finding 2: Giant partition ByteBuffer[]: 3.2 GB └─ users:user_12345: 3.2 GB ← One partition! # Finding 3: Memtable bloat org.apache.cassandra.db.Memtable: 6.8 GB ↑ Memtables too large!

Step 3: Check Cassandra Statistics

-- Check row cache size: nodetool info | grep "Row Cache" /* Row Cache: size 8532 MB, capacity 8192 MB ↑ ↑ Using 104%! Configured max Problem: Row cache oversized! */ -- Check for huge partitions: nodetool tablestats keyspace.table | grep partition /* Compacted partition maximum bytes: 3,421,883,152 ↑ 3.2 GB partition! */ -- Check memtable settings: grep memtable /etc/cassandra/cassandra.yaml /* memtable_heap_space_in_mb: 8192 ← Way too high! */

Step 4: Common Causes & Fixes

Cause Diagnosis Fix
Row Cache Too Large Heap dump shows cache > 50% Reduce row_cache_size_in_mb
Giant Partitions tablestats shows > 1GB partitions Fix data model (add bucketing)
Memtables Oversized memtable_heap_space > 25% heap Reduce memtable size
Driver Leak Heap dump shows driver objects Update driver, fix app code
Too Many SSTables Bloom filters + indexes huge Force compaction

Step 5: Apply Fixes

-- Fix 1: Reduce row cache ALTER TABLE keyspace.table WITH caching = { 'keys': 'ALL', 'rows_per_partition': 'NONE' ← Disable row cache }; -- Fix 2: Reduce memtable size (cassandra.yaml) memtable_heap_space_in_mb: 2048 # 25% of 8GB heap -- Fix 3: Fix giant partition (data model change) // Add bucket column to split large partition CREATE TABLE users_v2 ( user_id uuid, bucket int, ← New! created_at timestamp, PRIMARY KEY ((user_id, bucket), created_at) ); -- Restart Cassandra and monitor sudo systemctl restart cassandra

Key Takeaway: 95% heap usually means: row cache too large, giant partitions, or memtables oversized. Heap dump analysis reveals the culprit. Fix by tuning caches, fixing data model, or reducing memtable size!

5 How would you use nodetool commands to perform a quick health check on a production cluster? â–ŧ

Answer: 5-minute health check: (1) nodetool status - all nodes UN?, (2) nodetool tpstats - any pending/blocked?, (3) nodetool compactionstats - < 20 pending?, (4) nodetool gcstats - < 200ms pauses?, (5) nodetool netstats - zero drops?. These five commands catch 90% of problems.

The 5-Minute Health Check:

Command 1: nodetool status

nodetool status /* Datacenter: datacenter1 Status=Up/Down, State=Normal/Leaving/Joining UN 192.168.1.10 156.48 GB 32 33.3% abc123... UN 192.168.1.11 158.32 GB 32 33.3% def456... UN 192.168.1.12 155.91 GB 32 33.4% ghi789... */ -- What to check: ✅ All nodes: UN (Up, Normal) ✅ Ownership: Balanced (~33% each) ✅ Load: Similar across nodes ❌ Red flags: - DN (Down, Normal) = node down! - Ownership: 60%, 20%, 20% = imbalanced! - Load: 200GB, 50GB, 50GB = hotspot!

Command 2: nodetool tpstats

nodetool tpstats /* Pool Name Active Pending Completed Blocked ReadStage 8 0 45M 0 MutationStage 4 0 127M 0 CompactionEx 2 5 15K 0 */ -- What to check: ✅ Pending: 0 (or very low) ✅ Blocked: 0 (MUST be zero!) ❌ Red flags: - Pending > 100 = bottleneck! - Blocked > 0 = thread starvation! - Growing pending = can't keep up!

Command 3: nodetool compactionstats

nodetool compactionstats /* pending tasks: 12 SSTable count: 87 ... */ -- What to check: ✅ Pending: < 20 ✅ SSTable count: < 100 per table ❌ Red flags: - Pending > 50 = falling behind! - SSTables > 200 = read amplification! - Growing pending = disaster brewing!

Command 4: nodetool gcstats

nodetool gcstats /* Interval Max GC Total GC Stdev GC# 600 sec 157 ms 487 ms 23 ms 12 */ -- What to check: ✅ Max GC: < 200ms ✅ Total: < 500ms per interval ✅ GC#: < 20 per interval ❌ Red flags: - Max > 1000ms = long pauses! - Total > 1000ms = GC pressure! - GC# > 50 = thrashing!

Command 5: nodetool netstats

nodetool netstats | grep -i dropped /* Dropped: 0 ← MUST be zero! */ -- What to check: ✅ Dropped: 0 (CRITICAL!) ❌ Red flags: - Any drops = overload! - Growing drops = serious problem!

Bonus Commands (If Time Permits):

-- Memory usage: nodetool info | grep -E "Heap|Cache" -- Latency check: nodetool proxyhistograms -- Per-table health: nodetool tablestats keyspace.table

Decision Tree:

-- All checks pass: ✅ Cluster healthy! -- status shows DN: ❌ CRITICAL: Node down, investigate immediately! -- tpstats shows pending/blocked: âš ī¸ WARNING: Bottleneck, check load/hardware -- compactionstats shows backlog: âš ī¸ WARNING: Add compaction throughput/nodes -- gcstats shows long pauses: âš ī¸ WARNING: Tune JVM or reduce heap -- netstats shows drops: ❌ CRITICAL: Cluster overload, add capacity!

Key Takeaway: Five nodetool commands in 5 minutes: status (availability), tpstats (bottlenecks), compactionstats (long-term health), gcstats (JVM health), netstats (overload). These catch 90% of problems!

🎓 Chapter Summary: Monitoring Mastery

You now understand production-grade Cassandra monitoring!

The Essential Five Metrics:

  • ⚡ P99 Latency: User experience (< 50ms)
  • 🔄 Pending Compactions: Long-term stability (< 20)
  • 💾 Heap Usage: OOM risk (60-80%)
  • 📉 Dropped Messages: Overload signal (= 0)
  • ✅ Node Status: Availability (all UN)

5-Minute Health Check:

nodetool status # All nodes UN? nodetool tpstats # Pending = 0? nodetool compactionstats # < 20 pending? nodetool gcstats # < 200ms pauses? nodetool netstats # Drops = 0?

Monitoring Stack:

Prometheus (metrics) + Grafana (visualization) + Alerting (proactive)

Remember Rachel: What you don't monitor will eventually fail! 📊

Advertisement

📱 Responsive Ad 📱