System Requirements
Everything you need to run Cassandra in development and production environments!
๐ The $50K Hardware Mistake
Meet Tom, a startup CTO who just got $500K in funding. He decided to run Cassandra for their new social media app...
โ What Tom Did Wrong
- ๐ฅ Bought cheap servers: 4GB RAM, 2 cores, spinning HDDs
- ๐ฅ Ran Java 8: Didn't check Java version requirements
- ๐ฅ Windows servers: Thought "OS doesn't matter"
- ๐ฅ Single SSD: No RAID, no redundancy
- ๐ฅ 1Gbps network: Shared with other services
Result:
- โฑ๏ธ Queries taking 30+ seconds
- ๐ฅ Constant out-of-memory crashes
- ๐ฅ Data loss from disk failures
- ๐ก Investors furious at slow demo
- ๐ธ Had to spend $50K replacing everything!
โ The RIGHT Way
After reading system requirements:
- โ 16GB RAM minimum per node
- โ 8+ CPU cores for production
- โ Linux servers (Ubuntu 20.04 LTS)
- โ NVMe SSDs in RAID 10
- โ 10Gbps network dedicated to Cassandra
- โ Java 11 (OpenJDK recommended)
Result: Queries in milliseconds, investors happy, system stable! ๐
Let's make sure YOU don't repeat Tom's mistakes!
โก Quick Summary
TL;DR for the busy developer!
Development
- RAM: 4GB minimum
- CPU: 2 cores
- Disk: 20GB any type
- OS: Any (Mac/Linux/Win)
- Java: 11 or 17
- Network: Standard LAN
Production
- RAM: 16-32GB (64GB ideal)
- CPU: 8-16 cores
- Disk: 1-4TB NVMe SSD
- OS: Linux (Ubuntu/RHEL)
- Java: 11 LTS (recommended)
- Network: 10Gbps dedicated
High Performance
- RAM: 64-128GB+
- CPU: 16-32+ cores
- Disk: 4TB+ NVMe RAID 10
- OS: Tuned Linux kernel
- Java: 11 + custom JVM flags
- Network: 25-100Gbps
๐ฅ๏ธ Hardware Requirements
Let's break down each component in detail!
๐พ RAM (Memory)
Why RAM Matters
Cassandra uses RAM for:
- MemTables: Recent writes before flushing to disk
- Row Cache: Hot data cached in memory
- Key Cache: Partition key locations
- JVM Heap: Java object storage
- OS Page Cache: File system caching
| Environment | Minimum RAM | Recommended RAM | Why? |
|---|---|---|---|
| Development | 4GB | 8GB | Basic testing, single node |
| Staging | 8GB | 16GB | Realistic workload testing |
| Production (Light) | 16GB | 32GB | < 1TB data per node |
| Production (Heavy) | 32GB | 64-128GB | > 1TB data per node |
Common RAM Mistakes
- โ Too little RAM: Constant swapping, slow queries
- โ Huge JVM heap (> 16GB): Long garbage collection pauses
- โ No room for OS cache: Disk reads instead of memory
- โ Rule of thumb: JVM heap = 1/4 of RAM, max 16GB
๐ข CPU (Processor)
What Cassandra Does with CPU
CPU intensive operations:
- Compaction: Merging SSTables (background process)
- Read/Write threads: Handling client requests
- Gossip: Cluster state communication
- Compression: Data compression on writes
- Repairs: Data consistency checks
| Environment | CPU Cores | Notes |
|---|---|---|
| Development | 2 cores | Minimum for single node |
| Production (Small) | 4-8 cores | Light to moderate traffic |
| Production (Medium) | 8-16 cores | Standard production workload |
| Production (Large) | 16-32+ cores | High throughput, many tables |
CPU Best Practices
- โ More cores > Higher frequency: Cassandra is highly concurrent
- โ Avoid hyperthreading issues: Count physical cores
- โ Leave headroom: Don't run CPU at 100%
- โ Monitor compaction: Can be CPU intensive
๐ฟ Disk Requirements
Disk is CRITICAL for Cassandra
This is the #1 performance bottleneck!
- โ HDDs: 100-200 IOPS = SLOW
- โ SSDs: 10,000+ IOPS = FAST
- โ NVMe SSDs: 100,000+ IOPS = BLAZING
| Disk Type | IOPS | Latency | Use Case |
|---|---|---|---|
| HDD (7200 RPM) | ~100 | 10-20ms | โ NOT for Cassandra |
| SATA SSD | ~10,000 | 0.1-1ms | โ Development/Small prod |
| NVMe SSD | ~100,000 | 0.01-0.1ms | โ โ Production (recommended) |
| NVMe RAID 10 | ~400,000+ | <0.01ms | ๐ High-performance prod |
Storage Layout Best Practices
- โ Separate disk for commit log: Isolate sequential writes
- โ Separate disk for data: Random read/writes
- โ RAID 10 for production: Performance + redundancy
- โ Don't use RAID 5/6: Write penalty hurts Cassandra
- โ XFS or ext4 filesystem: Best tested with Cassandra
๐ง Operating System Requirements
Cassandra runs on multiple OS, but Linux is STRONGLY recommended for production!
Linux (Recommended)
Best for production!
- Ubuntu 20.04/22.04 LTS
- RHEL 8/9
- CentOS 8
- Debian 11+
Why Linux?
- Best performance tuning
- Most production deployments
- Better community support
macOS (Development Only)
OK for local dev
- macOS 11+ (Big Sur)
- Works with Docker
- Good for testing
Limitations:
- โ NOT for production
- โ Performance issues
- โ Missing Linux features
Windows (Not Recommended)
Avoid if possible
- Windows Server 2016+
- Experimental support
- Limited tooling
Why avoid?
- Poor performance
- Few deployments
- Limited documentation
Linux Kernel Requirements
- Minimum: Linux kernel 3.10+
- Recommended: Linux kernel 4.15+ or 5.x
- File descriptor limit: At least 100,000 (ulimit -n)
- Max map count: vm.max_map_count >= 1048575
โ Java/JVM Requirements
Cassandra is a Java application - choosing the right JVM is crucial!
Supported Java Versions
| Cassandra Version | Java Versions | Recommended |
|---|---|---|
| Cassandra 4.x | Java 8, 11 | โ Java 11 |
| Cassandra 5.x | Java 11, 17 | โ Java 11 or 17 |
OpenJDK (Recommended)
- Free and open source
- Well tested
- Community support
- Azul Zulu
- Amazon Corretto
Oracle JDK
- Works fine
- Commercial license
- $ Cost in production
- Not necessary
JVM Heap Size Rules
Critical for performance!
- โ Heap = 1/4 of total RAM
- โ NEVER exceed 16GB heap (GC pauses!)
- โ Young gen = 1/4 of heap
- โ Don't use tiny heaps (< 2GB)
๐ Network Requirements
Cassandra is a distributed system - network performance is CRITICAL!
Network Bandwidth Requirements
| Environment | Minimum | Recommended | Why? |
|---|---|---|---|
| Development | 100 Mbps | 1 Gbps | Local testing |
| Production (Small) | 1 Gbps | 10 Gbps | Replication traffic |
| Production (Large) | 10 Gbps | 25-100 Gbps | High throughput + repairs |
Network Latency Matters!
- โ Same datacenter: < 1ms latency ideal
- โ ๏ธ Cross datacenter: < 100ms acceptable
- โ > 200ms latency: Will cause problems!
- ๐ก Tip: Use dedicated network for Cassandra if possible
Required Ports
| Port | Protocol | Purpose | Direction |
|---|---|---|---|
| 7000 | TCP | Inter-node communication (cluster) | Node โ Node |
| 7001 | TCP | Inter-node SSL | Node โ Node (encrypted) |
| 7199 | TCP | JMX monitoring | Admin โ Node |
| 9042 | TCP | CQL native transport (clients) | App โ Node |
| 9160 | TCP | Thrift (deprecated, legacy) | App โ Node |
๐ญ Production Environment Best Practices
Real-world production requirements from companies running Cassandra at scale!
๐ Netflix Production Setup
Netflix's Cassandra Nodes
- CPU: 16-32 cores (AWS c5.4xlarge or larger)
- RAM: 32-64GB
- Disk: 2TB NVMe SSD (i3 instances)
- Network: 10 Gbps
- OS: Ubuntu 20.04 LTS
- Java: OpenJDK 11
- Cluster Size: 100s of nodes per region
๐ฑ Instagram Production Setup
Instagram's Cassandra Nodes
- CPU: 24 cores minimum
- RAM: 64-128GB
- Disk: 4TB NVMe SSD RAID 10
- Network: 25 Gbps bonded
- OS: CentOS 7 (custom kernel)
- Java: OpenJDK 11 with G1GC
- Data per node: ~3TB
๐ณ Apple Production Setup
Apple's Cassandra Nodes
- CPU: 32+ cores
- RAM: 128GB+
- Disk: 8TB+ NVMe RAID 10
- Network: 40-100 Gbps
- OS: RHEL 8
- Cluster: Largest Cassandra deployment (75,000+ nodes!)
โ Production Checklist
Before Going to Production
Infrastructure:
- โ 16GB+ RAM per node (32GB recommended)
- โ 8+ CPU cores per node
- โ NVMe SSDs (NOT HDDs!)
- โ 10 Gbps+ network
- โ Redundant network paths
- โ Separate disk for commit log
Software:
- โ Linux OS (Ubuntu/RHEL)
- โ Java 11 OpenJDK
- โ Latest stable Cassandra version
- โ Tuned JVM settings
- โ Proper ulimit settings
Cluster:
- โ At least 3 nodes (5+ recommended)
- โ Replication Factor = 3
- โ NetworkTopologyStrategy
- โ Time synchronization (NTP)
- โ Monitoring setup (Prometheus/Grafana)
- โ Backup strategy in place
โ๏ธ Cloud-Specific Recommendations
Optimized instance types for AWS, Azure, and GCP!
AWS Recommendations
Recommended Instance Types:
- Dev/Test: m5.large (2 vCPU, 8GB)
- Small Prod: i3.2xlarge (8 vCPU, 61GB, 1.9TB NVMe)
- Medium Prod: i3.4xlarge (16 vCPU, 122GB, 3.8TB NVMe)
- Large Prod: i3.8xlarge (32 vCPU, 244GB, 7.6TB NVMe)
- High Perf: i4i.8xlarge (32 vCPU, 256GB, 7.5TB NVMe)
Azure Recommendations
Recommended VM Sizes:
- Dev/Test: Standard_D2s_v3 (2 vCPU, 8GB)
- Small Prod: Standard_L8s_v2 (8 vCPU, 64GB, 1.92TB NVMe)
- Medium Prod: Standard_L16s_v2 (16 vCPU, 128GB, 3.84TB NVMe)
- Large Prod: Standard_L32s_v2 (32 vCPU, 256GB, 7.68TB NVMe)
GCP Recommendations
Recommended Machine Types:
- Dev/Test: n2-standard-2 (2 vCPU, 8GB)
- Small Prod: n2-standard-8 + Local SSD (8 vCPU, 32GB)
- Medium Prod: n2-standard-16 + Local SSD (16 vCPU, 64GB)
- Large Prod: n2-standard-32 + Local SSD (32 vCPU, 128GB)
Cloud Best Practices
- โ Use local NVMe SSDs: Don't use network-attached storage (EBS/Azure Disk)
- โ Spread across availability zones: Use rack awareness
- โ Use placement groups: Low-latency networking (AWS)
- โ Enable enhanced networking: SR-IOV for better throughput
- โ Avoid burstable instances: (t2/t3) - not suitable for Cassandra
๐ผ Interview Questions & Expert Answers
Ace your Cassandra interview with these system requirements questions!
Answer:
SSDs are critical because Cassandra performs heavy random I/O operations. HDDs provide only ~100-200 IOPS, while SSDs provide 10,000+ IOPS.
Why Cassandra needs high IOPS:
- Compaction: Reads multiple SSTables, writes new ones
- Read path: May need to check multiple SSTables
- Repair: Heavy I/O during data consistency checks
- Bloom filters: Random disk access patterns
Impact of using HDDs:
- ๐ฅ Read latency: 10-50ms (vs 0.1-1ms on SSD)
- ๐ฅ Compaction takes 10x longer
- ๐ฅ Cluster can't keep up with writes
- ๐ฅ Repairs timeout and fail
Recommendation: Always use NVMe SSDs in production. HDDs are only acceptable for cold archival data that's rarely accessed.
Answer:
Large JVM heaps (> 16GB) cause long garbage collection (GC) pauses that can freeze your Cassandra node.
The Problem with Large Heaps:
- GC pause time: Increases with heap size
- 16GB heap: GC pauses ~200-500ms (acceptable)
- 32GB heap: GC pauses 1-3 seconds (BAD!)
- 64GB heap: GC pauses 5-10+ seconds (DISASTER!)
What happens during long GC pauses:
- ๐ฅ Node appears dead to cluster (gossip timeout)
- ๐ฅ Other nodes mark it as DOWN
- ๐ฅ Hints start accumulating
- ๐ฅ Client queries timeout
- ๐ฅ When GC finishes, flood of messages overwhelms node
The Solution:
Cap heap at 16GB. Cassandra uses off-heap memory for many operations (compression buffers, bloom filters, etc.), so the rest of your RAM is still used effectively.
Answer:
Cassandra generates massive inter-node network traffic for replication, repairs, and streaming operations. 1 Gbps becomes a bottleneck quickly.
Network Traffic Sources:
- Writes (RF=3): Each write sends data to 3 nodes
- Repairs: Compares and streams gigabytes of data
- Bootstrap: New node streams entire dataset
- Compaction: Streaming across nodes
- Read repair: Synchronizing inconsistent replicas
Example Scenario:
Problems with 1 Gbps:
- โฑ๏ธ Slow cluster expansion (adding nodes)
- โฑ๏ธ Slow repairs (can take days!)
- ๐ฅ Write throughput limited by network
- ๐ฅ Network saturation during repairs
Recommendation: 10 Gbps minimum for production. 25-100 Gbps for large deployments.
Answer:
| Component | Development | Production | Why Different? |
|---|---|---|---|
| RAM | 4-8GB | 16-64GB+ | Dev: small dataset. Prod: GB-TB data + caching |
| CPU | 2 cores | 8-32 cores | Dev: light load. Prod: concurrent users + compaction |
| Disk | Any (even HDD) | NVMe SSD only | Dev: latency OK. Prod: IOPS critical |
| Network | 100 Mbps | 10 Gbps+ | Dev: single node. Prod: multi-node replication |
| Nodes | 1 | 3-100+ | Dev: testing. Prod: HA + scalability |
Key Differences Explained:
- Data Volume: Dev has MB/GB, prod has TB/PB
- Traffic: Dev has 1-10 req/sec, prod has 10K-1M req/sec
- Uptime: Dev can crash, prod needs 99.99% uptime
- Performance: Dev accepts seconds, prod needs milliseconds
Common Mistake: Running dev-sized hardware in production! Always test with production-equivalent specs in staging.
Answer:
Follow this step-by-step capacity planning process:
Step 1: Calculate Storage Requirements
Step 2: Calculate Throughput Requirements
Step 3: Validate Latency Requirements
- Target p99: < 10ms?
- Need: NVMe SSDs + sufficient RAM for caching
- RAM rule: Allocate 1GB RAM per 100GB data
Step 4: Add Headroom
Final Recommendation:
12 nodes, each with:
- 32GB RAM
- 16 CPU cores
- 2TB NVMe SSD
- 10 Gbps network
๐ Chapter Summary: System Requirements Mastery
You now know exactly what hardware and software Cassandra needs!
Production Minimums (Never Go Below):
- ๐พ 16GB RAM per node (32GB+ recommended)
- ๐ข 8 CPU cores per node (16+ recommended)
- ๐ฟ NVMe SSD storage (HDDs = disaster)
- ๐ 10 Gbps network (dedicated to Cassandra)
- ๐ง Linux OS (Ubuntu 20.04/22.04 or RHEL 8/9)
- โ Java 11 OpenJDK (LTS version)
Critical Rules:
- โ JVM Heap โค 16GB (avoid GC pauses)
- โ Heap = 1/4 of RAM (leave room for OS cache)
- โ Separate disks for commit log and data
- โ Test in staging with production-equivalent hardware
๐ Next: Let's install Cassandra with these requirements!
Responsive Ad