When to Use Cassandra
Complete decision guide with real scenarios, flowcharts, anti-patterns, and practical checklists. Make the right database choice!
📖 Uber's Crisis: When PostgreSQL Couldn't Scale
In 2015, Uber faced a critical infrastructure crisis. Their PostgreSQL database was buckling under 1 million trip requests per second across 300+ cities...
🚗 The Breaking Point
- Trip Data: Real-time location updates every 4 seconds from millions of drivers
- Write Storm: 50K+ writes per second overwhelming single PostgreSQL master
- Surge Pricing Calculation: Needed to query last 10 minutes of trip data instantly
- Global Expansion: Adding new cities caused cascading database failures
✅ Why Cassandra Was Perfect
- Write-Optimized: Handles 300K+ writes/sec per node (trip updates, GPS coordinates)
- Time-Series Data: Natural fit for time-stamped location data
- Linear Scalability: Add nodes = add capacity (no complicated sharding)
- Multi-Datacenter: Replicate across cities with low latency
- Always Available: Zero downtime even during datacenter failures
📊 Results
50x improvement in write throughput
99.99% uptime (from 99.5% with PostgreSQL)
$2M+ savings annually in infrastructure costs
This is a textbook example of when Cassandra is the RIGHT choice!
Let's learn how to identify if YOUR project needs Cassandra...
🎯 The Ultimate Decision Flowchart
Follow this flowchart to determine if Cassandra is right for your project in under 2 minutes!
✅ Perfect Fit: When Cassandra Shines
These are the scenarios where Cassandra is the BEST choice - not just good, but truly exceptional.
Time-Series Data
Perfect when you have:
- Timestamped events (logs, metrics, sensor readings)
- Append-only workloads
- Query patterns: "last 24 hours", "this month"
- TTL requirements (auto-delete old data)
Real Examples:
- IoT Sensors: Temperature, pressure, GPS coordinates every second
- Application Logs: 1M+ log entries per minute
- Financial Ticks: Stock prices, trades, market data
- Website Analytics: Page views, clicks, user journeys
Why Cassandra Wins:
- Write optimization perfect for continuous data streams
- Time-based partition keys natural fit
- Built-in TTL automatically removes old data
- Compaction strategies optimized for time-series
⭐ Perfect Match Score: 10/10
Extreme Write Workloads
Perfect when you need:
- 50K+ writes per second per node
- Millions of concurrent writers
- Write latency under 5ms
- No write bottlenecks
Real Examples:
- User Activity: Likes, views, follows, notifications
- Gaming Events: Player actions, scores, achievements
- Ad Impressions: Billions of ad views tracked
- Message Queues: High-throughput messaging systems
Why Cassandra Wins:
- Writes are sequential (no random disk seeks)
- No single master bottleneck
- Every node handles writes equally
- Linear scalability: 2x nodes = 2x writes/sec
⭐ Perfect Match Score: 10/10
Multi-Datacenter Applications
Perfect when you have:
- Users across multiple continents
- Need local latency everywhere
- Regulatory data residency requirements
- Disaster recovery across regions
Real Examples:
- Social Media: Users in US, Europe, Asia need fast access
- Streaming Services: Netflix serving 150M users globally
- Mobile Games: Players worldwide, low latency critical
- CDN Analytics: Edge locations reporting metrics
Why Cassandra Wins:
- Built-in multi-DC replication
- Configurable per datacenter replication factors
- LOCAL_QUORUM queries (low latency)
- Automatic conflict resolution
⭐ Perfect Match Score: 10/10
Zero-Downtime Requirements
Perfect when you need:
- 99.99%+ uptime SLA (4 minutes/month)
- No maintenance windows allowed
- Survive datacenter failures
- Continuous availability during upgrades
Real Examples:
- Payment Systems: Can't go down, ever
- Trading Platforms: Markets never sleep
- Emergency Services: 911 systems, healthcare
- Live Sports: Real-time scores, betting platforms
Why Cassandra Wins:
- No single point of failure
- Rolling upgrades with zero downtime
- Automatic failover (no election delays)
- Proven: Apple's 99.9999% uptime
⭐ Perfect Match Score: 10/10
Unpredictable Growth
Perfect when you have:
- Rapid user growth (10x in 6 months)
- Viral potential (could explode overnight)
- Need to scale without redesign
- Want predictable performance at scale
Real Examples:
- Startups: Instagram scaled 1M→2B users same architecture
- Seasonal Traffic: Black Friday, holiday spikes
- Viral Apps: TikTok challenges, trending content
- New Product Launches: Don't know demand yet
Why Cassandra Wins:
- True linear scalability proven
- Add nodes with no downtime
- No resharding required
- Performance predictable at any scale
⭐ Perfect Match Score: 9/10
Predictable Access Patterns
Perfect when you have:
- Well-defined query patterns
- Queries by primary key or partition
- No ad-hoc analytics needed
- Read/write patterns stable
Real Examples:
- User Profiles: Always query by user_id
- Shopping Carts: Query by cart_id or user_id
- Session Storage: Query by session_id
- Product Catalogs: Query by product_id or category
Why Cassandra Wins:
- Optimized for known queries
- Denormalization pre-joins data
- Single-partition reads extremely fast
- Can design perfect data model
⭐ Perfect Match Score: 9/10
Perfect Fit Summary
Use Cassandra when you have 2+ of these characteristics:
- Time-series or append-only data
- Extremely high write volume (50K+ writes/sec)
- Multi-datacenter deployment needed
- 99.99%+ uptime requirement
- Need to scale to 100+ nodes
- Known, predictable query patterns
Real Companies: Netflix, Apple, Instagram, Uber, Discord, eBay all fit these criteria!
👍 Good Fit: When Cassandra Works Well
These scenarios work well with Cassandra, though alternatives might also be viable. Cassandra is a strong choice but not the only choice.
Mobile App Backends
Good fit because:
- Mobile apps generate high write volume
- User sessions, preferences, offline sync
- Global user base benefits from multi-DC
- Eventual consistency often acceptable
Alternatives to Consider:
- MongoDB: If you need flexible schema, ACID transactions
- DynamoDB: If you're AWS-only, want managed service
👍 Good Match Score: 7/10
Messaging Platforms
Good fit because:
- High message volume
- Time-ordered message history
- Needs to scale to millions of users
- Availability critical for user experience
Alternatives to Consider:
- MongoDB: If you need rich queries on messages
- Specialized: RabbitMQ, Kafka for queue semantics
👍 Good Match Score: 7/10
E-commerce Product Catalogs
Good fit because:
- Read-heavy with predictable queries
- Search by category, brand, price range
- Can denormalize for fast lookups
- High availability for shopping experience
Alternatives to Consider:
- MongoDB: Better for flexible product attributes
- Elasticsearch: Better for full-text search
- Hybrid: Cassandra + Elasticsearch often best
👍 Good Match Score: 6/10
Content Feeds & Timelines
Good fit because:
- Chronological ordering natural
- High write volume (new posts)
- Denormalized feeds per user
- TTL for old content
Alternatives to Consider:
- Redis: For small, hot feeds (recent posts)
- MongoDB: If you need complex filtering
- Hybrid: Redis cache + Cassandra storage
👍 Good Match Score: 8/10
Good Fit Guidance
For "good fit" scenarios, ask yourself:
- Do I have the team expertise to operate Cassandra?
- Is the operational complexity worth it for my scale?
- Could MongoDB or a managed service be simpler?
- Am I starting small but expect massive growth?
Rule of Thumb: If you're unsure, start with MongoDB. You can always migrate to Cassandra later when you hit scale issues!
❌ Bad Fit: When NOT to Use Cassandra
These are the scenarios where Cassandra will cause pain. Save yourself the headache - use something else!
Financial Transactions
BAD FIT because:
- Requires ACID transactions
- Strong consistency mandatory
- Complex multi-table updates
- Regulatory compliance needs
The Problem:
Bank transfer: Debit from Account A, Credit to Account B must be atomic. Cassandra's lightweight transactions (LWT) are too slow and limited.
What to Use Instead:
- PostgreSQL: Full ACID, proven in banking
- MySQL: Battle-tested for transactions
- MongoDB 4.0+: ACID transactions across documents
❌ DON'T USE CASSANDRA! Use PostgreSQL/MySQL
Ad-Hoc Analytics & BI
BAD FIT because:
- No JOINs across tables
- Limited aggregation support
- Can't do complex GROUP BY
- Analytics tools expect SQL
The Problem:
"Show me total sales by region, product category, and month" requires complex JOINs and GROUP BY - Cassandra can't do this efficiently.
What to Use Instead:
- MySQL/PostgreSQL: Full SQL support
- BigQuery: Purpose-built for analytics
- Snowflake: Data warehouse
- Hybrid: Cassandra → Spark → Analytics DB
❌ DON'T USE CASSANDRA! Use data warehouse
Small Datasets (<10GB)
BAD FIT because:
- Cassandra overhead not worth it
- Need minimum 3 nodes (over-provisioned)
- Operational complexity too high
- PostgreSQL would fit in RAM
The Problem:
Running 3 Cassandra nodes for 5GB of data is like using a Ferrari to go to the grocery store. Overkill and expensive.
What to Use Instead:
- PostgreSQL: Single instance handles 10GB easily
- SQLite: Embedded, perfect for small apps
- Redis: If data fits in memory
❌ DON'T USE CASSANDRA! Complete overkill
Unknown Query Patterns
BAD FIT because:
- Must design tables for queries upfront
- Can't adapt to changing requirements
- Denormalization locks you in
- Exploration/prototyping difficult
The Problem:
New startup doesn't know how users will query data. Cassandra requires you to know this before designing schema. Can't pivot easily.
What to Use Instead:
- MongoDB: Flexible schema, easy to change
- PostgreSQL: Can add indexes later
- Start Simple: Use SQL until patterns emerge
❌ DON'T USE CASSANDRA! Too inflexible
Highly Relational Data
BAD FIT because:
- No foreign keys
- No referential integrity
- Must denormalize everything
- Data duplication nightmare
The Problem:
ERP system with customers, orders, products, invoices, shipments - all interconnected. Normalizing in Cassandra is a mess.
What to Use Instead:
- PostgreSQL: Designed for relational data
- MySQL: JOINs work perfectly
- Graph DB: Neo4j if relationships are complex
❌ DON'T USE CASSANDRA! Use relational DB
Team Lacks Expertise
BAD FIT because:
- Steep learning curve
- Distributed systems knowledge required
- Complex operational requirements
- Debugging is hard
The Problem:
Team knows SQL. Cassandra requires understanding: partition keys, consistency levels, compaction, repairs, tombstones, etc. 6+ months ramp-up.
What to Use Instead:
- MongoDB Atlas: Managed, easier to learn
- AWS RDS: Managed PostgreSQL/MySQL
- Hire Experts: Or use managed Cassandra (DataStax Astra)
❌ Use managed service or simpler database
Red Flags - DON'T Use Cassandra If:
- You need ACID transactions across multiple records
- You need to do complex JOINs regularly
- Your dataset is under 100GB
- You don't know your query patterns yet
- Your team doesn't have distributed systems expertise
- You need ad-hoc analytics and reporting
- Strong consistency is mandatory for all operations
Save yourself pain: Use PostgreSQL, MySQL, or MongoDB instead!
📋 The Ultimate Decision Checklist
Use this comprehensive checklist to make your final decision. Score your project honestly!
✅ Cassandra Strengths - Check What Applies to You
❌ Cassandra Weaknesses - Check What Applies to You
Scoring Your Decision
Count your checkmarks:
- 7+ Green Checks (Strengths): Cassandra is likely an excellent fit!
- 5-6 Green Checks: Cassandra is a good option, evaluate alternatives too
- 3-4 Green Checks: Consider MongoDB or MySQL instead
- <3 Green Checks: Cassandra is overkill, use simpler database
RED FLAGS:
- ANY Red Check (Weaknesses): Seriously reconsider Cassandra
- 3+ Red Checks: DO NOT use Cassandra - it will cause pain
💡 Rule of Thumb: If you're unsure, start with PostgreSQL or MongoDB. You can migrate to Cassandra later when scale demands it!
🌍 Real-World Decision Examples
Let's see how real companies made the Cassandra decision - both when they chose it and when they didn't.
✅ Discord: Perfect Cassandra Use Case
Challenge: Store billions of messages for 150M+ users with instant retrieval
Why They Chose Cassandra:
- Write Volume: 120K+ messages per second during peak
- Time-Series: Messages naturally ordered by timestamp
- Query Pattern: Always "get messages for channel X"
- Availability: Chat must always work, no downtime acceptable
- Scale: Growing 50% year over year
Data Model:
CREATE TABLE messages_by_channel ( channel_id BIGINT, message_id TIMEUUID, user_id BIGINT, content TEXT, created_at TIMESTAMP, PRIMARY KEY (channel_id, message_id) ) WITH CLUSTERING ORDER BY (message_id DESC); -- Query: Get last 50 messages from channel SELECT * FROM messages_by_channel WHERE channel_id = 12345 LIMIT 50;
Results:
- Handles 1 trillion+ messages
- Sub-10ms read latency at scale
- 99.99% uptime
- Saves millions on infrastructure
⭐ Textbook Perfect Cassandra Use Case!
✅ Uber: Migrated FROM PostgreSQL TO Cassandra
Challenge: PostgreSQL couldn't handle real-time location updates from millions of drivers
Why PostgreSQL Failed:
- Single master bottleneck (50K writes/sec limit)
- Vertical scaling too expensive
- Downtime during scaling operations
- Cross-region replication lag
Why Cassandra Succeeded:
- Write Scale: 300K+ writes/sec per node
- Geographic: Multi-DC for each city
- Time-Series: Trip data with timestamps
- Zero Downtime: Add capacity without stopping service
Migration Results:
- 50x improvement in write throughput
- 99.5% → 99.99% uptime improvement
- $2M+ annual infrastructure savings
- Enabled expansion to 300+ cities
✅ Right database for the right scale!
❌ Digg: When Cassandra Was the WRONG Choice
What Happened: In 2010, Digg rewrote their entire platform on Cassandra. It was a disaster.
Why They Chose Cassandra (Wrongly):
- Hype around "web scale" NoSQL
- Wanted to be cutting-edge
- Thought it would solve all problems
- Didn't understand their actual needs
Why It Failed:
- Complex Queries: Needed JOINs for friend relationships, votes, comments
- Unknown Patterns: Users wanted features requiring queries they hadn't designed for
- Small Scale: Traffic didn't justify Cassandra complexity
- Team Expertise: Spent 6 months learning instead of building features
- Data Model: Denormalization caused massive data duplication and sync issues
The Damage:
- Site crashes on launch day
- Users fled to Reddit
- 6-month delay in features
- Eventually had to rewrite AGAIN back to MySQL
- Company never recovered (sold for $500K after being valued at $160M)
⚠️ Lesson: Don't choose technology for hype. Choose for YOUR requirements!
🎯 Spotify: Smart Hybrid Approach
Strategy: Use the right database for each use case
Where They Use Cassandra:
- User Playlists: High write volume, simple queries by user_id
- Listening History: Time-series data, billions of events
- Recommendations: Precomputed, query by user_id
Where They DON'T Use Cassandra:
- Payments: PostgreSQL (ACID required)
- Subscriptions: PostgreSQL (transactions)
- Music Metadata: BigTable (different access patterns)
- Analytics: BigQuery (complex queries)
💡 Best Practice: Polyglot Persistence - Use multiple databases for different needs!
Key Learnings from Real Companies
Successful Cassandra Adoptions:
- Discord, Uber, Netflix, Instagram - All had massive scale needs
- All had clear, known query patterns
- All needed write-heavy, time-series capabilities
- All could tolerate eventual consistency
Failed Cassandra Adoptions:
- Digg, early startups - Chose for hype, not requirements
- Needed complex queries Cassandra couldn't handle
- Team lacked expertise, underestimated complexity
- Scale didn't justify operational overhead
💼 Top 10 Interview Questions - When to Use Cassandra
Master these decision-making questions to demonstrate strategic thinking in interviews!
Answer:
Choose Cassandra over MongoDB when:
1. Write Volume is Extreme (50K+ writes/sec):
- Cassandra: 300K+ writes/sec per node (optimized for writes)
- MongoDB: 150K writes/sec (good but not extreme)
- Example: IoT sensors sending data every second from millions of devices
2. Always-On Availability is Critical:
- Cassandra: No single point of failure, peer-to-peer
- MongoDB: Primary node is single point during failover (10-30 sec)
- Example: Payment processing, 911 systems, live sports scores
3. Time-Series Data is Primary Use Case:
- Cassandra: Optimized for time-series with TTL support
- MongoDB: Can do time-series but not specialized
- Example: Application logs, metrics, user activity streams
4. Linear Scalability to 100+ Nodes:
- Cassandra: Proven linear scaling to 1000+ nodes
- MongoDB: Good scaling but requires careful sharding
- Example: Instagram scaled from 1M to 2B users same architecture
Choose MongoDB over Cassandra when:
- Need flexible schema that changes frequently
- Need complex queries, aggregations, or lookups (MongoDB's aggregation pipeline)
- Need ACID transactions (MongoDB 4.0+ has full ACID)
- Team is smaller or lacks distributed systems expertise
- Dataset is moderate (<10TB) and won't explode in size
Decision Rule: If you need 2+ of: extreme writes, always-on, time-series, 100+ nodes → Cassandra. Otherwise → MongoDB.
Answer:
CRITICAL Red Flags - Don't Use Cassandra:
1. Need ACID Transactions:
- Cassandra only has lightweight transactions (LWT) which are slow
- Can't do multi-row atomic updates
- Example: Bank transfer (debit + credit must be atomic) → Use PostgreSQL
2. Need Complex JOINs:
- Cassandra has NO JOIN support
- Must denormalize everything (data duplication)
- Example: E-commerce with orders + products + customers → Use MySQL
3. Unknown or Changing Query Patterns:
- Cassandra requires designing tables for specific queries upfront
- Changing patterns = redesign entire schema
- Example: New startup exploring features → Use MongoDB
4. Ad-Hoc Analytics Required:
- No GROUP BY across partitions
- Limited aggregation support
- Example: Business intelligence, reporting dashboards → Use data warehouse
5. Small Dataset (<100GB):
- Cassandra overhead not worth it
- Need minimum 3 nodes (over-provisioned)
- Example: Small SaaS app with 5GB data → Use PostgreSQL
6. Team Lacks Distributed Systems Expertise:
- Steep learning curve (6+ months ramp-up)
- Complex operational requirements
- Example: 3-person startup → Use managed MongoDB Atlas
7. Strong Consistency Mandatory:
- Cassandra default is eventual consistency
- Tuning to strong consistency hurts performance
- Example: Inventory management (can't oversell) → Use MySQL
Real Example: Digg chose Cassandra in 2010, hit all these red flags, and failed spectacularly. Company sold for $500K after $160M valuation.
Answer:
Evaluate Linear Scalability Need with These Questions:
1. What's Your Growth Trajectory?
- Need Cassandra: 10x growth in 6-12 months, viral potential, unpredictable spikes
- Don't Need: Steady 20% year-over-year growth
- Example: Instagram grew 1M→100M→2B users, needed Cassandra's scaling
2. What's Your Target Scale?
- Need Cassandra: Plan to handle 100M+ users, petabytes of data
- Don't Need: Staying under 10M users, terabytes of data
- Rule: If single MySQL instance can handle it → don't need Cassandra
3. Do You Need Predictable Performance at Any Scale?
- Cassandra Benefit: Add nodes = proportional capacity increase
- 3 nodes = 100K ops/sec → 6 nodes = 200K ops/sec (linear!)
- MySQL/MongoDB: Eventually hit diminishing returns with sharding
4. Can You Afford Resharding Downtime?
- Cassandra: Add capacity with zero downtime
- MySQL: Manual sharding requires downtime, engineering effort
- Example: Uber couldn't take downtime to reshard → migrated to Cassandra
5. What's the Cost of Over-Provisioning?
- With MySQL: Must provision for peak (expensive)
- With Cassandra: Add nodes only when needed (cost-effective at scale)
Decision Framework:
Current State: 1M users, 100GB data, 10K writes/sec ↓ Projected State (12 months): 50M users, 10TB data, 500K writes/sec ↓ Can single MySQL instance handle projected state? NO Can you afford 2-week downtime for resharding? NO Need predictable scaling without redesign? YES ↓ → Use Cassandra (linear scaling required)
Counter-Example: If you're a B2B SaaS with 1000 customers, growing to 5000 → MySQL is fine!
Answer:
Scenario: Smart Home IoT Platform
Requirements:
- 10 million smart homes globally
- Each home has 20 sensors (temperature, humidity, motion, etc.)
- Sensors report every 10 seconds
- Total: 200M sensors × 6 readings/minute = 1.2 billion writes/minute
- Need to query: "Last 24 hours of data for sensor X"
- Auto-delete data older than 90 days
Why Cassandra is Perfect:
1. Write Optimization:
- Cassandra handles 300K+ writes/sec per node
- Sequential writes (append-only) extremely fast
- MySQL would need hundreds of shards to handle this volume
2. Time-Based Partition Key:
CREATE TABLE sensor_readings ( sensor_id UUID, bucket DATE, -- Partition key (day) reading_time TIMESTAMP, -- Clustering key temperature FLOAT, humidity FLOAT, PRIMARY KEY ((sensor_id, bucket), reading_time) ) WITH CLUSTERING ORDER BY (reading_time DESC); -- Query last 24 hours (one partition) SELECT * FROM sensor_readings WHERE sensor_id = ? AND bucket = '2025-01-15' AND reading_time > '2025-01-15 00:00:00';
3. TTL (Time To Live):
- Automatically delete data after 90 days
- No batch jobs needed, happens during compaction
- Saves storage costs automatically
INSERT INTO sensor_readings (...) USING TTL 7776000; -- 90 days in seconds
4. Multi-Datacenter for Global Deployment:
- Sensors in US write to US datacenter (low latency)
- Sensors in Europe write to EU datacenter
- Data replicated for disaster recovery
5. Linear Scalability:
- 10M homes today → Add 10 nodes
- 50M homes tomorrow → Add 40 more nodes (linear!)
- No redesign, no downtime
Why Alternatives Fail:
- MySQL: Can't handle 1B+ writes/minute, sharding nightmare
- MongoDB: Good but not optimized for pure time-series, no built-in TTL
- InfluxDB: Purpose-built for time-series but doesn't scale like Cassandra
Real Companies Using This Pattern:
- Apple iCloud: Device telemetry from 1B+ devices
- Netflix: Viewing events, billions per day
- Uber: Driver location updates every 4 seconds
Answer:
Eventual Consistency is ACCEPTABLE when:
1. Social Media Interactions:
- Likes/Views Counters: Post shows 99 vs 100 likes briefly → no big deal
- Followers Count: Off by a few for seconds → acceptable
- Comments: Might not see your own comment for 1-2 seconds → tolerable
- Example: Instagram, Twitter use Cassandra with eventual consistency
2. Shopping Carts:
- User adds item to cart across two datacenters
- Might briefly see duplicate items
- Resolved at checkout (merge conflicts)
- Example: Amazon chose AP (availability) over consistency for carts
3. User Preferences/Settings:
- Update profile picture, brief delay in replication
- Theme preference takes 1-2 seconds to apply everywhere
- Not critical for user experience
4. Analytics/Metrics:
- Dashboard shows 98.5% vs 98.7% accuracy → close enough
- Real-time monitoring with eventual consistency acceptable
- Aggregates converge quickly
5. Content Delivery:
- Blog post published, might not appear everywhere instantly
- CDN cache invalidation takes time
- News feed updates eventually
Strong Consistency is MANDATORY when:
1. Financial Transactions:
- Bank Transfers: Account balance must be exact, always
- Payments: Can't charge customer twice
- Trading: Stock price must be accurate
- Why: Money + wrong data = lawsuits
- Use: PostgreSQL, MySQL with ACID
2. Inventory Management:
- E-commerce: Can't oversell last item in stock
- Hotel bookings: Can't double-book room
- Flight reservations: Can't oversell seats
- Why: Overselling = angry customers, refunds
- Use: MySQL with row-level locking
3. Authentication/Security:
- User changes password → must take effect immediately everywhere
- Revoke access token → can't allow stale access
- 2FA codes → must be consumed atomically
- Why: Security vulnerabilities
4. Regulatory/Compliance:
- Healthcare: Patient records must be accurate
- Legal: Audit trails must be consistent
- GDPR: Delete requests must be honored immediately
- Why: Legal requirements
5. Collaborative Editing:
- Google Docs: Everyone sees same version
- Code collaboration: Merge conflicts must be resolved
- Why: User confusion, data loss
Decision Framework:
Ask these questions: 1. Could wrong/stale data cause financial loss? YES → Strong consistency required 2. Could wrong data cause legal/compliance issues? YES → Strong consistency required 3. Is user expecting immediate reflection of change? - Changed password? YES → Strong - Liked post? NO → Eventual OK 4. Can conflicts be merged/resolved? - Shopping cart? YES → Eventual OK - Bank balance? NO → Strong required 5. What's cost of being wrong? - Social media counter? Low → Eventual - Payment amount? High → Strong
Real Example - Spotify:
- Eventual (Cassandra): Playlists, listening history, recommendations
- Strong (PostgreSQL): Payments, subscriptions, billing
Answer:
Team Readiness Assessment - 5 Key Areas:
1. Technical Expertise Required:
- Distributed Systems Knowledge:
- Understand CAP theorem, eventual consistency, quorum
- Know how to design for partition tolerance
- Debug network partitions, split-brain scenarios
- Data Modeling Skills:
- Query-driven design (one query = one table)
- Denormalization strategies
- Partition key selection, clustering columns
- CQL Proficiency:
- Not just SQL with different syntax!
- Understand write path, read path, compaction
2. Operational Capabilities:
- Can your team handle:
- Running nodetool repairs weekly
- Monitoring 50+ metrics per node
- Tuning compaction strategies
- Managing tombstones and TTL
- Rolling upgrades across clusters
- Capacity planning for multi-DC
3. Team Size & Dedication:
- Minimum Team: 2-3 dedicated DBAs or SREs for production
- Startup (3-5 people): Too small → Use managed service
- Mid-size (20+ engineers): Can dedicate 1-2 people → Feasible
- Large (100+ engineers): Full platform team → Perfect
4. Time to Productivity:
- SQL Background: 6-9 months to become proficient
- NoSQL Experience: 3-6 months ramp-up
- Distributed Systems Expert: 1-3 months
- Question: Can you afford this learning curve?
5. Budget for Training:
- DataStax Academy courses: 2-3 months per person
- External consultants: $200-300/hour
- Conferences, certifications: $5-10K per person
- Opportunity cost: Features delayed during learning
Readiness Scorecard - Check All That Apply:
□ At least 1 team member has distributed systems experience □ Team comfortable with Linux, networking, JVM tuning □ Can dedicate 1-2 people full-time to database operations □ Have 6+ months for learning curve before going to production □ Budget for training, consultants, or managed service □ Have monitoring/alerting infrastructure in place □ Comfortable with command-line tools (nodetool, cqlsh) □ Team can debug network issues, packet loss, timeouts □ On-call rotation can handle 24/7 database incidents Score: 7-9 checks: Team is ready ✅ 4-6 checks: Use managed Cassandra (DataStax Astra) 0-3 checks: Use MongoDB or MySQL instead
Alternative: Managed Services
- DataStax Astra: Fully managed Cassandra
- Handles operations, repairs, upgrades
- Team only needs to know data modeling
- Reduces readiness requirements by 70%
- AWS Keyspaces: Cassandra-compatible managed service
- Serverless, auto-scaling
- Minimal operational knowledge needed
Real Example - Startup Mistake:
A 5-person startup chose self-hosted Cassandra. Spent 8 months learning instead of building features. Competitors captured market. Eventually migrated to MongoDB Atlas. Cost: $500K in lost opportunity.
Recommendation:
- Small team? → Use managed Cassandra or choose simpler database
- Large team with budget? → Self-hosted Cassandra feasible
- Medium team? → Start managed, move to self-hosted at scale
Answer:
Discovery Questions to Ask Stakeholders:
1. Business & Scale Questions:
- "How many users do you have today and expect in 12 months?"
- Looking for: 10x growth or 10M+ users
- If steady growth <50% → simpler database fine
- "What's the cost of downtime per hour?"
- $100K+/hour → Need Cassandra's availability
- $5K/hour → MongoDB with replica sets OK
- "Are you planning global expansion?"
- Multiple continents → Multi-DC advantage clear
- Single region → Not a deciding factor
2. Data Pattern Questions:
- "Describe your top 5 most frequent queries"
- All query by ID? → Good fit
- Need JOINs, GROUP BY? → Bad fit
- "What's your read:write ratio?"
- Write-heavy (80:20) → Cassandra strength
- Read-heavy (20:80) → MongoDB might be better
- "Is your data timestamped?"
- Yes, time-series → Perfect for Cassandra
- No, relational → Consider alternatives
- "Do you need to delete old data automatically?"
- Yes → Cassandra TTL is perfect
- No → Not a deciding factor
3. Consistency Requirements:
- "If a user updates data, must they see it immediately?"
- Financial data, passwords → Strong consistency needed
- Social feeds, preferences → Eventual OK
- "Can you tolerate brief inconsistencies (1-2 seconds)?"
- Yes → Cassandra default works
- No → Need ACID, use PostgreSQL
- "Do you need multi-record transactions?"
- Yes (bank transfers) → Cassandra wrong choice
- No → Cassandra viable
4. Technical Constraints:
- "What's your team's database expertise?"
- SQL experts → Learning curve expensive
- NoSQL experience → Easier adoption
- "Do you have DevOps/SRE team?"
- Yes → Can handle Cassandra ops
- No → Use managed service or simpler DB
- "What's your budget for database infrastructure?"
- $50K+/year → Can run Cassandra cluster
- <$10K/year → Single PostgreSQL instance
5. Future Proofing:
- "How will your data model evolve?"
- Stable, known patterns → Cassandra good
- Rapidly changing → MongoDB more flexible
- "Will you need analytics/BI on this data?"
- Yes → Need data warehouse separately
- Cassandra alone can't do complex analytics
Example Conversation:
Engineer: "How many users do you expect in 12 months?" Stakeholder: "We're at 100K now, could hit 5M if viral" Engineer: "What's downtime cost?" Stakeholder: "$50K per hour - we're e-commerce" Engineer: "Describe your main query pattern" Stakeholder: "Users view their order history by user_id" Engineer: "Do you need to join orders with products?" Stakeholder: "Yes, frequently for analytics reports" Engineer: "Can reports be slightly delayed?" Stakeholder: "No, real-time dashboards required" → Decision: Use MongoDB for primary DB + Cassandra for logs MongoDB handles joins, Cassandra for write-heavy activity logs
Red Flag Responses:
- "We don't know our query patterns yet" → Don't use Cassandra
- "We need to analyze data in many different ways" → Don't use Cassandra
- "Team only knows SQL" → High learning curve cost
- "Budget is tight" → Managed service or simpler DB
Answer:
Migration Justification Framework:
When to ADVOCATE for Migration to Cassandra:
1. Build Business Case with Numbers:
- Current Pain Points:
- "MySQL master at 95% CPU during peak hours"
- "Sharding will require 3-month engineering effort"
- "Downtime for scaling costs $200K in lost revenue"
- "Query latency increased 10x in 6 months"
- Quantify Benefits:
- "Cassandra eliminates single point of failure → 99.5% to 99.99% uptime"
- "Save $100K/year in downtime costs"
- "Linear scaling vs exponential engineering effort"
- "Support 10x growth without architecture redesign"
2. Present Risk Analysis:
STAY WITH MYSQL: Risks: - Next scaling event in 6 months → 2 weeks downtime - Manual sharding complexity → 6-month engineering effort - Single point of failure → potential outages Costs: - Engineering: $300K (sharding implementation) - Downtime: $200K (lost revenue) - Opportunity cost: Delayed features MIGRATE TO CASSANDRA: Risks: - Learning curve → 3-6 months - Migration complexity → potential bugs - Team expertise gap → need training Costs: - Migration: $150K (one-time) - Training: $50K - Infrastructure: +$20K/year Benefits: - Zero downtime scaling - Linear capacity growth - 99.99% uptime SLA - Future-proof for 10x growth
3. Propose Phased Migration:
- Phase 1 (3 months): Migrate activity logs, analytics to Cassandra
- Low risk, high learning
- Team gains expertise
- Immediate write performance improvement
- Phase 2 (6 months): Migrate user sessions, preferences
- Medium risk
- Validate dual-write pattern
- Phase 3 (9 months): Migrate core application data
- Only if Phases 1-2 successful
- Keep MySQL for transactions
When to ADVOCATE AGAINST Migration:
1. MySQL Is Actually Fine:
- Say this if:
- "We're at 60% capacity, can scale vertically for 2+ years"
- "Read replicas solve our current scaling needs"
- "Dataset is 100GB, will stay under 1TB"
- Recommendation: "Upgrade MySQL instance, add read replicas. Save $200K migration cost."
2. Need ACID Transactions:
- "Our core business logic requires multi-row atomic updates"
- "E-commerce: inventory + orders + payments must be transactional"
- Recommendation: "Keep MySQL for transactions, use Cassandra only for logs/analytics"
3. Team Not Ready:
- "Team is 5 people, all know only SQL"
- "6-month learning curve will delay product roadmap"
- Recommendation: "Use managed PostgreSQL (RDS) with multi-region. Simpler + similar benefits."
Presentation to Management:
Slide 1: Current State - MySQL at 90% capacity - Next scaling event in 4 months - Projected to need resharding in 6 months Slide 2: Option A - Stay with MySQL - Cost: $300K engineering + $200K downtime - Timeline: 6 months - Risk: Still single point of failure Slide 3: Option B - Migrate to Cassandra - Cost: $200K one-time migration - Timeline: 9 months (phased) - Benefit: Future-proof for 10x growth Slide 4: Recommendation - Hybrid approach: Cassandra for logs/activity - Keep MySQL for transactions - Best of both worlds - 70% cost savings vs full MySQL sharding
Real Example - Uber's Pitch:
Uber's engineering team showed management:
- PostgreSQL hit limits at 1M rides/day
- Cassandra could handle 10M rides/day on same budget
- 99.99% uptime vs 99.5% = $10M/year revenue difference
- Management approved based on ROI, not technology preference
Key Point: Frame in business terms (cost, revenue, risk), not technical terms (peer-to-peer, eventual consistency). Management cares about outcomes!
Answer:
Common Misconceptions - Debunked:
Misconception 1: "Cassandra is always faster than MySQL"
- Reality: Cassandra is faster for writes, but MySQL can be faster for reads
- Truth:
- Writes: Cassandra 300K/sec > MySQL 50K/sec ✓
- Reads (indexed): MySQL 200K/sec ≈ Cassandra 200K/sec
- Complex queries: MySQL much faster (JOINs, aggregations)
- Lesson: Choose based on workload, not general speed claims
Misconception 2: "If you're doing NoSQL, use Cassandra"
- Reality: MongoDB, DynamoDB, Redis are also NoSQL and often better choices
- Truth:
- Cassandra is ONE type of NoSQL (wide-column)
- MongoDB (document) better for flexible schemas
- Redis (key-value) better for caching
- DynamoDB better for AWS-native apps
- Lesson: NoSQL ≠ Cassandra. Evaluate each NoSQL database independently
Misconception 3: "Cassandra scales infinitely"
- Reality: Linear scaling, but has practical limits
- Truth:
- Proven to 1000+ nodes (Netflix, Apple)
- Beyond that: coordination overhead increases
- Network becomes bottleneck at extreme scale
- Operational complexity grows with cluster size
- Lesson: "Scales really well" ≠ "infinite scaling"
Misconception 4: "Eventual consistency means data loss"
- Reality: Eventual consistency ≠ lost data, just brief delay
- Truth:
- Data is replicated and durable
- "Eventual" typically means milliseconds to seconds
- All replicas eventually converge to same value
- No data loss, just temporary inconsistency
- Example: Post gets 100 likes, you see 99 for 1 second → eventually shows 100
- Lesson: Eventual consistency is a delay, not data loss
Misconception 5: "Big data = Need Cassandra"
- Reality: Size alone doesn't determine database choice
- Truth:
- 1TB relational data with complex queries → PostgreSQL fine
- 100GB time-series with high writes → Cassandra good
- 10TB analytics → Data warehouse (Snowflake, BigQuery)
- Lesson: Access patterns matter more than size
Misconception 6: "Cassandra is easier to scale than SQL"
- Reality: Easier to scale, but not easier to operate
- Truth:
- Scaling: Just add nodes (easier than MySQL sharding) ✓
- Operations: More complex (repairs, tombstones, tuning) ✗
- Day-to-day: MySQL simpler for most teams
- Lesson: Easy to scale ≠ easy to run
Misconception 7: "You can just migrate anytime"
- Reality: Migration is expensive and risky
- Truth:
- MySQL → Cassandra requires complete data model redesign
- 6-12 month effort for large systems
- Risk of downtime, data inconsistency
- Cost: $100K-$1M+ depending on size
- Lesson: Choose right database upfront, migration is painful
Misconception 8: "Cassandra replaces caching layer"
- Reality: Cassandra is fast, but not a cache
- Truth:
- Cassandra read latency: 5-10ms
- Redis latency: <1ms
- Still need caching for hot data
- Architecture: Redis (cache) + Cassandra (persistent storage) common pattern
Misconception 9: "Use Cassandra because Netflix/Uber uses it"
- Reality: Their scale ≠ your scale
- Truth:
- Netflix: 150M users, 1 trillion requests/day
- Your startup: 10K users, 1M requests/day
- Their problems ≠ your problems
- Lesson: Don't cargo cult. Evaluate YOUR requirements
- Example: Digg tried to copy Facebook's use of Cassandra, failed miserably
Misconception 10: "Cassandra is 'web scale'"
- Reality: Meaningless buzzword
- Truth:
- Facebook runs on MySQL (very "web scale"!)
- Twitter uses MySQL + Manhattan
- Scale depends on architecture, not database choice alone
- Lesson: Ignore marketing buzzwords, focus on technical requirements
Key Takeaway: Cassandra is a powerful tool for SPECIFIC use cases (high writes, time-series, always-on). It's not a universal solution. Evaluate based on YOUR requirements, not hype or what big tech companies do!
Answer:
Complete Decision Matrix:
═══════════════════════════════════════════════════════════════════════════════
DATABASE DECISION MATRIX
═══════════════════════════════════════════════════════════════════════════════
FACTOR | CASSANDRA | MONGODB | MySQL
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
WRITE PERFORMANCE | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐
READ PERFORMANCE | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
HORIZONTAL SCALING | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐
VERTICAL SCALING | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
AVAILABILITY (UPTIME) | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐
CONSISTENCY | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
COMPLEX QUERIES | ⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
JOINS | ⭐ | ⭐⭐ | ⭐⭐⭐⭐⭐
AGGREGATIONS | ⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
TRANSACTIONS (ACID) | ⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
SCHEMA FLEXIBILITY | ⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
LEARNING CURVE | ⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
OPERATIONAL COMPLEXITY | ⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
COMMUNITY & SUPPORT | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐
━━━━━━━━━━━━━━━━━━━━━━━━━━|━━━━━━━━━━━|━━━━━━━━━━|━━━━━━━━━━━━━━━
MULTI-DATACENTER | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐
TIME-SERIES DATA | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐
═══════════════════════════════════════════════════════════════════════════════
Use Case Scoring System:
USE CASE CHECKLIST - Score each requirement (1-5):
□ Write Volume (1=low, 5=extreme)
□ Read Volume (1=low, 5=extreme)
□ Scalability Need (1=1-10K users, 5=10M+ users)
□ Availability Requirement (1=99%, 5=99.999%)
□ Consistency Requirement (1=eventual OK, 5=strong required)
□ Query Complexity (1=simple lookups, 5=complex JOINs)
□ Schema Stability (1=changes daily, 5=fixed)
□ Team Expertise (1=SQL only, 5=distributed systems)
SCORING:
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
IF (Write Volume >= 4 AND Scalability >= 4):
→ CASSANDRA (60% probability best choice)
IF (Query Complexity >= 4 OR Consistency >= 4):
→ MySQL (70% probability best choice)
IF (Schema Stability <= 2 AND Team Expertise <= 3):
→ MongoDB (65% probability best choice)
IF (Availability >= 5 AND Write Volume >= 4):
→ CASSANDRA (80% probability best choice)
IF (Write Volume <= 2 AND Scalability <= 2):
→ MySQL (75% probability best choice)
DEFAULT (no clear winner):
→ MongoDB (best all-around balance)
Real-World Decision Tree:
START: │ ├─ Need ACID transactions? │ ├─ YES → MySQL or PostgreSQL ✓ │ └─ NO → Continue │ ├─ Write volume > 50K/sec? │ ├─ YES → Cassandra (strong fit) ✓ │ └─ NO → Continue │ ├─ Need 99.99%+ uptime? │ ├─ YES → Cassandra (strong fit) ✓ │ └─ NO → Continue │ ├─ Time-series or IoT data? │ ├─ YES → Cassandra (good fit) ✓ │ └─ NO → Continue │ ├─ Need complex queries/JOINs? │ ├─ YES → MySQL (strong fit) ✓ │ └─ NO → Continue │ ├─ Schema changes frequently? │ ├─ YES → MongoDB (strong fit) ✓ │ └─ NO → Continue │ ├─ Team only knows SQL? │ ├─ YES → MySQL (practical choice) ✓ │ └─ NO → Continue │ └─ DEFAULT → MongoDB (best balance) ✓
Example Application:
SCENARIO: Social Media App for Gamers Requirements: ✓ Write Volume: 4/5 (user posts, game scores, achievements) ✓ Read Volume: 4/5 (feeds, leaderboards) ✓ Scalability: 5/5 (targeting 10M+ users) ✓ Availability: 5/5 (gamers online 24/7) ✓ Consistency: 2/5 (eventual OK for likes, scores) ✓ Query Complexity: 2/5 (mostly by user_id, game_id) ✓ Schema Stability: 3/5 (adding game features monthly) ✓ Team Expertise: 3/5 (mix of SQL and NoSQL) MATRIX EVALUATION: ───────────────────────────────────────────────────────────── Database | Total Score | Recommendation ───────────────────────────────────────────────────────────── Cassandra | 35/40 | ⭐⭐⭐⭐⭐ BEST FIT MongoDB | 30/40 | ⭐⭐⭐⭐ Good Alternative MySQL | 20/40 | ⭐⭐ Not Recommended ───────────────────────────────────────────────────────────── DECISION: Use Cassandra for: - User activity logs - Game scores and achievements - Leaderboards (time-series) - User feeds Keep PostgreSQL for: - User authentication - Payment processing - Admin tools This hybrid approach gets best of both worlds!
Interview Tip: Always present the decision matrix with real numbers and use cases. Don't just say "Cassandra is good for writes" - quantify it! "Cassandra handles 300K writes/sec vs MySQL's 50K, making it 6x better for write-heavy workloads like IoT sensor data."
Responsive Ad