Cluster Management Tool

Master Nodetool

Your Swiss Army knife for Cassandra cluster administration!

๐Ÿ”ง What is Nodetool?

Nodetool = Your cluster management Swiss Army knife!

It's THE command-line tool for managing and monitoring Cassandra clusters. Think of it as:

  • ๐Ÿ”ง Administrator's toolkit for cluster operations
  • ๐Ÿ“Š Monitoring dashboard in your terminal
  • ๐Ÿ” Debugging tool for investigating issues
  • ๐Ÿ’พ Backup manager for snapshots
  • ๐Ÿฅ Health checker for cluster status

โœจ What You Can Do

  • โœ… Check cluster health: Node status, ring topology
  • โœ… Maintenance tasks: Repair, compact, flush, cleanup
  • โœ… Monitor performance: Stats, metrics, thread pools
  • โœ… Manage nodes: Add, remove, decommission, rebuild
  • โœ… Backup/restore: Snapshots, incremental backups
  • โœ… Tune settings: Cache, compaction, streaming

๐ŸŽฏ When to Use Nodetool

  • ๐Ÿ“Š Daily: Check cluster status
  • ๐Ÿ”ง Weekly: Run repairs, check stats
  • ๐Ÿšจ Troubleshooting: Investigate issues, view logs
  • โž• Scaling: Add/remove nodes
  • ๐Ÿ’พ Backups: Create snapshots
  • โšก Performance: Analyze bottlenecks

๐ŸŽ“ Master nodetool = Master Cassandra operations!

Running Nodetool

# Linux/Mac (if in PATH) nodetool status # Full path /opt/cassandra/bin/nodetool status # Docker docker exec cassandra-node1 nodetool status # Remote node nodetool -h 192.168.1.100 status # With JMX authentication nodetool -u jmxuser -pw jmxpassword status

โญ Essential Commands - Use These Daily

The commands you'll use most often!

๐Ÿฅ nodetool status

THE most important command! Shows cluster health.

# Check cluster status nodetool status # Output explanation: Datacenter: datacenter1 Status=Up/Down |/ State=Normal/Leaving/Joining/Moving -- Address Load Tokens Owns Host ID UN 172.18.0.2 69 KiB 16 33.3% abc123... UN 172.18.0.3 65 KiB 16 33.3% def456... DN 172.18.0.4 73 KiB 16 33.4% ghi789... # Status codes: # UN = Up and Normal โœ… # DN = Down and Normal โŒ # UL = Up and Leaving # UJ = Up and Joining # Specific keyspace nodetool status my_keyspace

โ„น๏ธ nodetool info

Detailed node information

# Get node details nodetool info # Shows: # - ID and generation # - Gossip active: true # - Load: 69.91 KiB # - Heap Memory: 256/1024 MB # - Off Heap Memory # - Data Center / Rack # - Exceptions # - Key Cache / Row Cache

๐Ÿ” nodetool describecluster

Cluster overview

# Get cluster information nodetool describecluster # Shows: # - Cluster Name # - Snitch # - DynamicEndPointSnitch # - Partitioner # - Schema versions (should be 1!)

๐Ÿ’ nodetool ring

Token ring distribution

# View token ring nodetool ring # Shows which node owns which token ranges # Useful for debugging data distribution # For specific keyspace nodetool ring my_keyspace

๐Ÿ“‹ nodetool version

Check Cassandra version

# Get version nodetool version # Output: ReleaseVersion: 4.1.3

๐Ÿ”ง Maintenance Commands

Keep your cluster healthy!

๐Ÿ”„

nodetool repair

CRITICAL: Ensures data consistency across replicas

# Full repair (all keyspaces) nodetool repair # Specific keyspace nodetool repair my_keyspace # Specific table nodetool repair my_keyspace users # Full repair (recommended weekly) nodetool repair -full # Primary range only (faster) nodetool repair -pr # Parallel repair (use with caution) nodetool repair -j 4 # 4 threads

โฑ๏ธ Time: Can take minutes to hours depending on data size

๐Ÿ“… Schedule: Weekly for critical data, monthly for non-critical

๐Ÿ—œ๏ธ

nodetool compact

Merge SSTables to reclaim space and improve performance

# Compact all keyspaces nodetool compact # Specific keyspace nodetool compact my_keyspace # Specific table nodetool compact my_keyspace users # Split output (create multiple SSTables) nodetool compact -s my_keyspace users

โš ๏ธ Warning: Use sparingly! Compaction happens automatically

๐Ÿ’ก When to use: After bulk deletes, when disk is low

๐Ÿงน

nodetool cleanup

Remove data this node no longer owns (after adding nodes)

# Cleanup all keyspaces nodetool cleanup # Specific keyspace nodetool cleanup my_keyspace # Specific table nodetool cleanup my_keyspace users

๐ŸŽฏ Purpose: Reclaim disk space after cluster expansion

๐Ÿ“… When: After adding new nodes to cluster

๐Ÿ’พ

nodetool flush

Write memtables to disk immediately

# Flush all keyspaces nodetool flush # Specific keyspace nodetool flush my_keyspace # Specific table nodetool flush my_keyspace users

๐Ÿ’ก When to use: Before backups, before shutdown, testing

โšก Fast: Usually completes in seconds

Maintenance Best Practices

  • โœ… Repair: Run weekly on production clusters
  • โš ๏ธ Compact: Only when necessary (not routinely)
  • โœ… Cleanup: Always after adding nodes
  • โœ… Flush: Before backups and shutdowns
  • โš ๏ธ I/O Impact: These operations use disk heavily
  • ๐Ÿ’ก Scheduling: Run during low-traffic periods

๐Ÿ“Š Monitoring Commands

Monitor performance and health!

๐Ÿ“Š nodetool tablestats

Table-level statistics

# All keyspaces nodetool tablestats # Specific keyspace nodetool tablestats my_keyspace # Specific table (with human-readable sizes) nodetool tablestats -H my_keyspace.users # Shows: # - Space used # - SSTable count # - Read/Write count # - Read/Write latency # - Pending compactions

๐Ÿงต nodetool tpstats

Thread pool statistics

# Thread pool stats nodetool tpstats # Shows for each pool: # - Active threads # - Pending tasks # - Completed tasks # - Blocked threads # Watch for: # - High pending counts (backpressure) # - Many blocked threads (contention)

๐Ÿ“ˆ nodetool proxyhistograms

Latency histograms

# Latency distribution nodetool proxyhistograms # Shows percentiles: # - 50%, 75%, 95%, 98%, 99%, 99.9% # For reads, writes, range queries # Example output: Percentile Read Latency Write Latency 50% 1.2 ms 0.8 ms 95% 5.4 ms 3.2 ms 99% 12.1 ms 8.7 ms

๐Ÿ—„๏ธ nodetool cfstats

Column family (table) statistics - Detailed

# All tables (verbose) nodetool cfstats # Specific keyspace nodetool cfstats my_keyspace # More detailed than tablestats # Includes bloom filter stats, compression, etc

๐Ÿ“‰ nodetool compactionstats

Active compactions

# Check compaction status nodetool compactionstats # Shows: # - Active compactions # - Progress percentage # - Type (major, minor, cleanup, etc) # - SSTables being compacted # Use with -H for human-readable sizes nodetool compactionstats -H

๐Ÿ’พ nodetool getlogginglevels

View current log levels

# Show all log levels nodetool getlogginglevels # Set log level dynamically nodetool setlogginglevel org.apache.cassandra DEBUG # Reset to default nodetool setlogginglevel org.apache.cassandra INFO

๐ŸŒ Cluster Operations

Scale and manage your cluster!

โž–

nodetool decommission

CRITICAL: Gracefully remove a node

# Decommission THIS node # Run on the node you want to remove nodetool decommission # Process: # 1. Streams data to other nodes # 2. Updates token ownership # 3. Leaves cluster # Time: 5 minutes to hours (depends on data size)

NEVER Skip Decommission!

Always decommission before removing a node!

  • โœ… Correct: decommission โ†’ wait โ†’ stop node
  • โŒ Wrong: Just stop/kill node
  • โš ๏ธ Skipping = permanent data loss!
๐Ÿ—‘๏ธ

nodetool removenode

Force remove a dead node (emergency only!)

# First, get the host ID nodetool status # Remove dead node (run from live node) nodetool removenode # Check removal status nodetool removenode status # Force removal (if stuck) nodetool removenode force

โš ๏ธ Use only when:

  • Node is permanently dead
  • Can't run decommission on it
  • Hardware failure, etc
๐Ÿ”„

nodetool rebuild

Stream data from another datacenter

# Rebuild from another DC nodetool rebuild -- datacenter1 # Rebuild specific keyspace nodetool rebuild -- datacenter1 -ks my_keyspace # Use when: # - Adding new datacenter # - Replacing failed node # - Restoring from another DC
๐Ÿ“ก

nodetool gossipinfo

View gossip state

# Show gossip information nodetool gossipinfo # Shows node states, heartbeat, etc # Useful for debugging cluster communication

๐Ÿ’พ Backup Commands

Protect your data!

๐Ÿ“ธ

nodetool snapshot

Create point-in-time backup

# Snapshot all keyspaces nodetool snapshot -t backup-20250106 # Specific keyspace nodetool snapshot -t backup-20250106 my_keyspace # Specific table nodetool snapshot -t backup-20250106 my_keyspace.users # List snapshots nodetool listsnapshots # Snapshots stored in: # /var/lib/cassandra/data/keyspace/table/snapshots/
๐Ÿ—‘๏ธ

nodetool clearsnapshot

Delete old snapshots

# Clear all snapshots nodetool clearsnapshot # Clear specific snapshot nodetool clearsnapshot -t backup-20250106 # Clear snapshots for keyspace nodetool clearsnapshot my_keyspace # Automate cleanup (keep last 7 days) find /var/lib/cassandra/data -name snapshots \ -type d -mtime +7 -exec rm -rf {} \;

Backup Strategy

Recommended approach:

  1. Daily snapshots: Automated with cron
  2. Copy to S3/Object Storage: Off-server backup
  3. Test restores: Verify backups work
  4. Retention: Keep 7-30 days
  5. Monitor disk: Snapshots use space!

๐Ÿ” Troubleshooting Commands

Debug issues quickly!

๐Ÿ” Quick Health Check

# 1. Check all nodes are UP nodetool status # 2. Check heap usage nodetool info | grep "Heap Memory" # 3. Check pending compactions nodetool compactionstats # 4. Check thread pools nodetool tpstats # 5. Check schema agreement nodetool describecluster

๐ŸŒ Slow Queries

# Check latencies nodetool proxyhistograms # Check if compaction is lagging nodetool compactionstats # Check pending tasks nodetool tpstats # Enable query tracing in cqlsh # TRACING ON;

๐Ÿ’พ High Disk Usage

# Check space per table nodetool tablestats -H # List snapshots (often the culprit) nodetool listsnapshots # Clear old snapshots nodetool clearsnapshot # Force major compaction (last resort) nodetool compact

๐Ÿง  High Memory Usage

# Check heap usage nodetool info # If > 75%: Need to increase heap or reduce load # Check cache sizes nodetool info | grep "Cache" # Trigger garbage collection (testing only) nodetool garbagecollect

๐Ÿ”Œ Node Not Joining Cluster

# Check gossip state nodetool gossipinfo # Check if node sees others nodetool describecluster # Check cluster name matches grep "cluster_name" /etc/cassandra/cassandra.yaml # Check seeds are correct grep "seeds" /etc/cassandra/cassandra.yaml

๐Ÿ’ก Best Practices

Use nodetool like a pro!

โœ…

DO

  • Run status daily
  • Schedule weekly repairs
  • Take daily snapshots
  • Monitor tpstats
  • Always decommission nodes
  • Use -H for human sizes
  • Check logs after commands
โŒ

DON'T

  • Run compact routinely
  • Skip decommissioning
  • Ignore high pending counts
  • Forget to clear snapshots
  • Run repairs too frequently
  • Force commands without checking
  • Ignore schema disagreements

Maintenance Schedule

Recommended routine:

  • Daily: Check nodetool status
  • Daily: Take snapshots (automated)
  • Weekly: Run nodetool repair
  • Weekly: Check tablestats, tpstats
  • Monthly: Clean up old snapshots
  • After scaling: Run cleanup
  • Before upgrades: Take snapshot, run flush

๐ŸŽ‰ You're a Nodetool Expert!

Congratulations! You now know how to use nodetool effectively!

๐ŸŽ“ What You Learned:

  • โญ Essential commands: status, info, describecluster, ring
  • ๐Ÿ”ง Maintenance: repair, compact, cleanup, flush
  • ๐Ÿ“Š Monitoring: tablestats, tpstats, proxyhistograms
  • ๐ŸŒ Cluster ops: decommission, removenode, rebuild
  • ๐Ÿ’พ Backups: snapshot, clearsnapshot, listsnapshots
  • ๐Ÿ” Troubleshooting: Debug common issues
  • ๐Ÿ’ก Best practices: Maintenance schedules, DO/DON'T

๐Ÿ’ก Essential Commands Cheat Sheet:

# Daily checks nodetool status # Cluster health nodetool info # Node details # Weekly maintenance nodetool repair -full # Data consistency nodetool tablestats -H # Check stats # Monitoring nodetool tpstats # Thread pools nodetool proxyhistograms # Latencies # Backups nodetool snapshot -t backup-$(date +%Y%m%d) nodetool listsnapshots nodetool clearsnapshot # Emergency nodetool decommission # Remove node nodetool removenode # Force remove

๐Ÿ”ง Nodetool is your best friend for Cassandra operations!

Advertisement

Responsive Ad