Interview Preparation

100 Interview Questions

Master Cassandra interviews with comprehensive Q&A!

Basics & Fundamentals
15 questions • Foundation concepts
Q1 What is Apache Cassandra? Easy
Answer: Apache Cassandra is a highly scalable, distributed NoSQL database designed to handle large amounts of data across many commodity servers without any single point of failure. It provides high availability with no single point of failure and linear scalability.
  • Type: Wide-column store (NoSQL)
  • Architecture: Peer-to-peer, masterless
  • CAP: AP (Availability + Partition tolerance)
  • Use cases: Time-series data, IoT, messaging, user profiles
Q2 What is the CAP theorem and where does Cassandra fit? Easy
Answer: CAP theorem states that a distributed system can only guarantee two of three properties: Consistency, Availability, and Partition tolerance. Cassandra is an AP system - it prioritizes Availability and Partition tolerance over strong consistency. However, Cassandra allows tunable consistency through consistency levels, so you can achieve strong consistency when needed by using QUORUM or ALL.
Q3 What is a keyspace in Cassandra? Easy
Answer: A keyspace is the outermost container for data in Cassandra, similar to a database in relational systems. It defines replication strategy and replication factor for the tables it contains.
  • Contains one or more tables
  • Defines replication strategy (SimpleStrategy or NetworkTopologyStrategy)
  • Defines replication factor (how many copies of data)
Example: CREATE KEYSPACE ecommerce WITH REPLICATION = {'class': 'NetworkTopologyStrategy', 'dc1': 3};
Q4 Explain the difference between RDBMS and Cassandra Medium
Answer: Key differences:
  • Data Model: RDBMS uses tables with fixed schema; Cassandra uses flexible wide-column model
  • Joins: RDBMS supports joins; Cassandra does not (denormalization required)
  • Transactions: RDBMS has ACID transactions; Cassandra has eventual consistency
  • Scaling: RDBMS scales vertically; Cassandra scales horizontally
  • Architecture: RDBMS has master-slave; Cassandra is peer-to-peer
  • Query-first: Cassandra requires designing tables for specific queries
Q5 What is CQL? Easy
Answer: CQL (Cassandra Query Language) is the query language for Cassandra, similar to SQL but designed for Cassandra's data model. It provides a familiar SQL-like syntax for creating tables, inserting data, and querying, but doesn't support joins, subqueries, or complex transactions.
Q6 What is a column family in Cassandra? Easy
Answer: A column family (now called a "table" in CQL) is a container for rows. It's similar to a table in relational databases but with a flexible schema. Each row can have different columns, and columns are stored sorted by clustering key.
Q7 What are the advantages of Cassandra? Medium
Answer: Key advantages:
  • High availability: No single point of failure, always-on architecture
  • Linear scalability: Add nodes to increase capacity without downtime
  • Fast writes: Optimized for write-heavy workloads
  • Tunable consistency: Balance between consistency and performance
  • Distributed: Data distributed across multiple nodes/datacenters
  • Fault tolerant: Automatic replication and recovery
  • Flexible schema: Easy to add columns without migrations
Q8 What are the limitations of Cassandra? Medium
Answer: Key limitations:
  • No joins: Must denormalize data and duplicate information
  • No transactions: No ACID transactions across rows (only per-row atomicity)
  • Query limitations: Can only query by partition key efficiently
  • No aggregations: Limited support for SUM, COUNT, AVG in queries
  • Eventual consistency: By default, not immediately consistent
  • Disk space: Requires more space due to replication and denormalization
  • Learning curve: Different data modeling approach than RDBMS
Q9 What is eventual consistency? Easy
Answer: Eventual consistency means that after a write, all replicas will eventually become consistent given enough time, but reads immediately after writes may return stale data. Cassandra uses this model by default for better availability and performance, but you can achieve strong consistency using QUORUM or ALL consistency levels.
Q10 What companies use Cassandra in production? Easy
Answer: Major companies using Cassandra:
  • Netflix: Streaming metadata, user profiles (largest deployment)
  • Apple: 160,000+ nodes for iCloud and services
  • Uber: Trip data, user data (100M+ ops/sec)
  • Instagram: User feeds, stories
  • Discord: Messaging (handles trillions of messages)
  • Spotify: User playlists, recommendations
Q11 What is denormalization and why is it important in Cassandra? Medium
Answer: Denormalization is storing duplicate data across multiple tables to support different query patterns. It's important in Cassandra because:
  • No joins are supported, so related data must be stored together
  • Each query pattern needs its own table
  • Duplicating data trades disk space for query performance
  • Design principle: "One table per query"
Example: Store user data in both users table (by user_id) and users_by_email table (by email) for different lookup patterns.
Q12 What is the difference between NoSQL and RDBMS transactions? Medium
Answer: RDBMS provides full ACID transactions (Atomicity, Consistency, Isolation, Durability) across multiple rows and tables. Cassandra provides:
  • Atomicity: Only at the row level (all columns in a row updated atomically)
  • Isolation: Lightweight transactions (LWT) using Paxos for compare-and-set
  • No multi-row transactions: Can use batches, but they're not transactional
  • BASE model: Basically Available, Soft state, Eventual consistency
Q13 What is a wide-column store? Easy
Answer: A wide-column store is a NoSQL database that stores data in columns rather than rows. Each row can have different columns, and columns are grouped into column families. Cassandra is a wide-column store where:
  • Rows can have thousands of columns
  • Columns are stored sorted by clustering key
  • Efficient for reading subsets of columns
  • Good for time-series and analytical workloads
Q14 When should you use Cassandra? Easy
Answer: Use Cassandra when you need:
  • High availability: 24/7 uptime, no single point of failure
  • Linear scalability: Handle growing data by adding nodes
  • Fast writes: Write-heavy workloads (IoT, logging, events)
  • Time-series data: Sensor data, metrics, logs
  • Geographic distribution: Multi-datacenter deployment
  • Large datasets: Petabytes of data
Q15 When should you NOT use Cassandra? Medium
Answer: Avoid Cassandra when you need:
  • ACID transactions: Complex multi-row transactions
  • Joins: Complex queries across multiple tables
  • Aggregations: Heavy analytical queries (SUM, AVG, GROUP BY)
  • Small dataset: Less than 100GB (overkill for small data)
  • Unpredictable queries: Ad-hoc queries without known patterns
  • Strong consistency always: Require immediate consistency for all reads
Better alternatives: PostgreSQL, MySQL (RDBMS), MongoDB (flexible queries), Redis (small dataset, caching)
Architecture & Internals
15 questions • System design & internals
Q16 Explain Cassandra's architecture (peer-to-peer) Medium
Answer: Cassandra uses a peer-to-peer (P2P) distributed architecture where all nodes are equal:
  • No master node: No single point of failure, any node can serve requests
  • Gossip protocol: Nodes communicate to share cluster state
  • Consistent hashing: Data distributed across nodes using hash ring
  • Virtual nodes (vnodes): Each node owns multiple token ranges for better distribution
  • Replication: Data automatically copied to multiple nodes
  • Snitch: Determines network topology (datacenters, racks)
Q17 What is the gossip protocol? Medium
Answer: Gossip protocol is a peer-to-peer communication protocol used for failure detection and cluster state sharing:
  • Nodes exchange information every second with 1-3 random nodes
  • Shares: node status (UP/DOWN), load, token ranges, schema versions
  • Failure detection: If node doesn't gossip for ~10 seconds, marked as DOWN
  • Eventually, all nodes have consistent view of cluster state
  • Efficient: O(log N) message complexity
Q18 Explain the write path in Cassandra Hard
Answer: Write path flow:
  1. Commit log: Write is first appended to commit log on disk (durability)
  2. Memtable: Write is then added to memtable (in-memory structure)
  3. Acknowledgment: Client gets success response
  4. Flush to SSTable: When memtable full, flushed to disk as immutable SSTable
  5. Compaction: SSTables merged and old data removed periodically
Key points: Sequential writes (fast), immutable SSTables, periodic compaction
Q19 Explain the read path in Cassandra Hard
Answer: Read path flow:
  1. Check memtable: Look for data in memory first
  2. Row cache: Check row cache (if enabled)
  3. Bloom filter: Check if data might be in each SSTable
  4. Key cache: Find position in SSTable
  5. Read SSTable: Read data from disk
  6. Merge: Combine data from memtable + SSTables, apply timestamps
  7. Return: Return latest data to client
Key points: Multiple SSTables read, timestamps determine latest value
Q20 What is a commit log? Medium
Answer: Commit log is an append-only log file on disk that records every write before it's applied to memtable. Purpose:
  • Durability: Ensures no data loss on node crash
  • Recovery: Replays writes on restart if memtable not flushed
  • Sequential writes: Very fast (append-only)
  • Cleared: After memtable flushed to SSTable
Config: commitlog_sync can be 'periodic' or 'batch'
Q21 What is a memtable? Medium
Answer: Memtable is an in-memory data structure that stores writes before they're flushed to disk:
  • Per table: Each table has its own memtable
  • Sorted: Keeps data sorted by partition key + clustering key
  • Flushed: When full (size threshold), written to SSTable
  • Fast writes: All writes go to memory first (very fast)
  • Size limits: Configurable via memtable_heap_space_in_mb
Q22 What is an SSTable? Medium
Answer: SSTable (Sorted String Table) is an immutable file on disk that stores data:
  • Immutable: Never modified after creation (only compacted)
  • Sorted: Data sorted by partition key + clustering key
  • Components: Data file, Index file, Bloom filter, Statistics
  • Created: When memtable is flushed
  • Compaction: Multiple SSTables merged to remove old data
Q23 What is a bloom filter and why is it used? Medium
Answer: Bloom filter is a probabilistic data structure that checks if data might exist in an SSTable:
  • Purpose: Avoid unnecessary disk reads
  • How it works: Checks if partition key might be in SSTable
  • False positives: May say "yes" when data doesn't exist (but never "no" when it does)
  • Memory efficient: Very small (a few MB per SSTable)
  • Read optimization: Prevents reading SSTables that definitely don't have the data
Q24 What is compaction and why is it important? Hard
Answer: Compaction merges multiple SSTables into fewer, larger SSTables:
  • Removes deleted data: Tombstones discarded after gc_grace_seconds
  • Merges updates: Keeps only latest version of data
  • Improves reads: Fewer SSTables to scan
  • Reclaims space: Deletes old versions and tombstones
Strategies: SizeTieredCompactionStrategy (STCS), LeveledCompactionStrategy (LCS), TimeWindowCompactionStrategy (TWCS)
Q25 What are vnodes (virtual nodes)? Medium
Answer: Virtual nodes (vnodes) allow each physical node to own multiple token ranges on the hash ring:
  • Default: 256 vnodes per node
  • Better distribution: Data more evenly distributed across cluster
  • Faster bootstrapping: New node gets data from many nodes in parallel
  • Easier operations: No manual token assignment needed
  • Load balancing: Automatic with no hot spots
Config: num_tokens: 256 in cassandra.yaml
Q26 Explain consistent hashing in Cassandra Hard
Answer: Consistent hashing distributes data across nodes using a hash ring:
  1. Hash ring: Token space (2^64) arranged in a ring
  2. Node tokens: Each node assigned token positions on ring
  3. Data placement: Partition key hashed, data goes to next node clockwise
  4. Replication: Data copied to N nodes (RF) clockwise
  5. Adding nodes: Only affects immediate neighbors (minimal data movement)
Benefits: Scalable, minimal reorganization when adding/removing nodes
Q27 What is a snitch in Cassandra? Medium
Answer: A snitch determines network topology (datacenters and racks) for routing and replica placement:
  • SimpleSnitch: Single datacenter (treats all nodes as in same rack)
  • GossipingPropertyFileSnitch: Multi-DC, reads dc/rack from cassandra-rackdc.properties
  • PropertyFileSnitch: Legacy, complex config
  • Ec2Snitch/Ec2MultiRegionSnitch: AWS-specific
  • GoogleCloudSnitch: GCP-specific
Purpose: Ensures replicas placed in different racks/DCs for fault tolerance
Q28 What is hinted handoff? Medium
Answer: Hinted handoff is a mechanism where a coordinator node temporarily stores writes for a down replica:
  • Scenario: Write goes to coordinator, but replica node is down
  • Hint: Coordinator stores the write as a "hint"
  • Delivery: When replica comes back online, coordinator delivers hints
  • Max storage: Hints stored for 3 hours by default
  • Purpose: Improves write availability and eventual consistency
Q29 What is read repair? Medium
Answer: Read repair synchronizes replicas when inconsistent data is detected during reads:
  • How: Coordinator compares timestamps from all replicas
  • Detects: If replicas have different data versions
  • Fixes: Sends latest version to out-of-date replicas
  • Automatic: Happens in background (configurable with read_repair_chance)
  • Purpose: Maintains eventual consistency passively
Q30 What is anti-entropy repair (nodetool repair)? Hard
Answer: Anti-entropy repair is a full synchronization process across replicas:
  • Command: nodetool repair
  • How: Uses Merkle trees to efficiently compare data across replicas
  • Finds: All inconsistencies and missing data
  • Fixes: Streams missing/outdated data between replicas
  • When: Run regularly (weekly/monthly) or after node downtime
  • Important: Must run within gc_grace_seconds to prevent zombie data
Data Modeling & Keys
20 questions • Schema design & keys
Q31 What is a partition key? Hard
Answer: The partition key determines which node stores the data and which partition the data belongs to:
  • Hash: Partition key is hashed to determine token/node
  • Data distribution: All rows with same partition key stored together
  • Query requirement: Must include partition key in WHERE clause
  • Syntax: PRIMARY KEY (partition_key) or PRIMARY KEY ((compound, partition), clustering)
  • Example: PRIMARY KEY (user_id) - all data for user_id on same node
Q32 What is a clustering key? Hard
Answer: Clustering key determines sort order of rows within a partition:
  • Sorting: Rows stored sorted by clustering key
  • Within partition: Only applies to rows with same partition key
  • Multiple columns: Can have compound clustering key
  • Order: Can specify ASC or DESC with WITH CLUSTERING ORDER BY
  • Example: PRIMARY KEY (user_id, timestamp) - sorts by timestamp within each user
Q33 What is the difference between partition key and clustering key? Medium
Answer:
  • Partition key: Determines which node stores data (horizontal distribution)
  • Clustering key: Determines sort order within a partition (vertical organization)
  • Query: Partition key required in WHERE; clustering key optional but must be in order
  • Example: PRIMARY KEY ((user_id), timestamp, event_type)
    • Partition key: user_id (distribution)
    • Clustering keys: timestamp, event_type (sorting)
Q34 What is a composite partition key? Medium
Answer: A composite partition key uses multiple columns to determine data distribution:
  • Syntax: PRIMARY KEY ((col1, col2), clustering) - double parentheses!
  • Purpose: Better data distribution (avoid hot partitions)
  • Hash: All columns in composite key hashed together
  • Query: Must specify ALL partition key columns in WHERE
  • Example: PRIMARY KEY ((sensor_id, date), hour) - distributes by sensor+date combination
Q35 How do you design tables in Cassandra (query-first design)? Hard
Answer: Cassandra requires query-first design (opposite of RDBMS):
  1. List queries: Define all queries your application needs
  2. One table per query: Create separate table for each query pattern
  3. Choose partition key: Column(s) used in WHERE clause for distribution
  4. Choose clustering keys: Columns for sorting/filtering within partition
  5. Denormalize: Duplicate data across tables to support different queries
Example: User system needs "get user by id" AND "get user by email" → create 2 tables
Q36 What is a wide row/partition? Medium
Answer: A wide row/partition contains many rows with the same partition key:
  • Structure: One partition key with thousands/millions of clustering key combinations
  • Use case: Time-series data, user activities, sensor readings
  • Example: PRIMARY KEY (user_id, timestamp) - one user, many timestamps
  • Limit: Keep partitions under 100MB (best practice)
  • Pattern: Common in Cassandra (one-to-many relationships)
Q37 How do you handle one-to-many relationships? Medium
Answer: Use wide rows with clustering keys:
  • Pattern: Partition key = "one", Clustering key = "many"
  • Example: User has many orders
    • PRIMARY KEY (user_id, order_id)
    • All orders for a user stored in one partition
  • Collections: Can also use set/list/map for small relationships (< 100 items)
  • Avoid: Don't create separate "join" tables like in RDBMS
Q38 How do you handle many-to-many relationships? Medium
Answer: Create two tables (one for each query direction):
  • Example: Users and Groups (many-to-many)
    • Table 1: users_by_group - PRIMARY KEY (group_id, user_id)
    • Table 2: groups_by_user - PRIMARY KEY (user_id, group_id)
  • Queries supported:
    • "Get all users in group X" → Query Table 1
    • "Get all groups for user Y" → Query Table 2
  • Trade-off: Duplicate data but fast queries
Q39 What is partition size and why does it matter? Hard
Answer: Partition size is the total size of all rows with the same partition key:
  • Limit: Keep partitions under 100MB (recommendation)
  • Problems if too large:
    • Slow queries (must read entire partition)
    • Memory pressure (loaded into memory for queries)
    • Compaction issues (huge compactions)
    • Repair problems (streaming large partitions)
  • Solution: Use time bucketing (e.g., partition by sensor_id + date, not just sensor_id)
Q40 What is time bucketing? Medium
Answer: Time bucketing limits partition size by adding time periods to partition key:
  • Problem: Time-series data grows unbounded in one partition
  • Solution: Add time bucket (day, hour, month) to partition key
  • Example: PRIMARY KEY ((sensor_id, date), timestamp)
    • Without bucket: All sensor data in one partition (grows forever)
    • With bucket: One partition per sensor per day (bounded size)
  • Bucket size: Choose based on write rate (high rate = smaller buckets)
Q41 What are collection types in Cassandra? Medium
Answer: Cassandra supports three collection types:
  • Set: set<text> - Unordered, unique values
  • List: list<int> - Ordered, can have duplicates
  • Map: map<text, int> - Key-value pairs
Limitations: Max 64KB per collection, don't store huge lists. Use for small data like tags, preferences, metadata.
Q42 What is a User Defined Type (UDT)? Medium
Answer: UDT is a custom data type that groups related fields together:
  • Purpose: Store structured data in a single column
  • Example:
    CREATE TYPE address (
      street text,
      city text,
      zip int
    );
    
    CREATE TABLE users (
      id uuid PRIMARY KEY,
      name text,
      home_address frozen<address>
    );
  • Frozen: UDTs must be frozen (treated as blob, can't update individual fields)
Q43 What is a counter column? Medium
Answer: Counter is a special column type for distributed counting:
  • Type: counter
  • Operations: Can only increment/decrement, not set directly
  • Use case: Page views, likes, follower counts
  • Table restriction: Counter tables can only have counter columns (besides primary key)
  • Example:
    CREATE TABLE page_views (
      page_id text PRIMARY KEY,
      view_count counter
    );
    UPDATE page_views SET view_count = view_count + 1 WHERE page_id = 'home';
Q44 What is a secondary index and when should you use it? Hard
Answer: Secondary index allows querying by non-primary key columns:
  • Syntax: CREATE INDEX ON table (column)
  • How: Creates local index on each node
  • Query: Must still specify partition key OR use ALLOW FILTERING
  • When to use: Low cardinality columns (few unique values), small datasets
  • When NOT to use: High cardinality, large datasets, production (better to create proper table)
  • Performance: Can be slow (queries all nodes)
Better alternative: Create a dedicated table for the query pattern
Q45 What is a materialized view? Hard
Answer: A materialized view is an automatically maintained table with different primary key from base table:
  • Purpose: Support different query patterns without manual denormalization
  • Auto-sync: Cassandra automatically updates view when base table changes
  • Syntax:
    CREATE MATERIALIZED VIEW users_by_email AS
      SELECT * FROM users
      WHERE email IS NOT NULL
      PRIMARY KEY (email, user_id);
  • Limitations: All base table PK columns must be in view PK
  • Performance: Adds write overhead (writes to base + view)
Q46 What is ALLOW FILTERING and should you use it? Medium
Answer: ALLOW FILTERING allows querying non-indexed columns but requires full table scan:
  • What it does: Scans all partitions, filters in memory
  • Example: SELECT * FROM users WHERE age = 25 ALLOW FILTERING;
  • Performance: VERY slow, reads entire table
  • When to use: Development/testing only, small tables, one-off queries
  • Production: Never use in production (sign of poor data model)
  • Solution: Create proper table with age in primary key
Q47 How do you model time-series data? Medium
Answer: Use time bucketing with timestamp clustering:
  • Pattern: PRIMARY KEY ((entity_id, time_bucket), timestamp)
  • Example: PRIMARY KEY ((sensor_id, date), timestamp)
  • Compaction: Use TimeWindowCompactionStrategy (TWCS)
  • TTL: Set automatic expiration for old data
  • Clustering order: DESC for newest-first queries
This keeps partitions bounded and enables efficient queries for time ranges.
Q48 What are common data modeling anti-patterns? Hard
Answer: Common mistakes to avoid:
  • Unbounded partitions: No time bucketing for time-series
  • Using ALLOW FILTERING: In production
  • Secondary indexes everywhere: Instead of proper tables
  • Normalizing data: Trying to avoid duplication
  • Hot partitions: Poor partition key choice (e.g., all data to one partition)
  • Large collections: Storing huge lists/sets in one column
  • Querying without partition key: Requires full cluster scan
Q49 How do you update data in Cassandra? Medium
Answer: Updates in Cassandra work differently than RDBMS:
  • Upsert: INSERT and UPDATE are the same (both upsert)
  • No read-before-write: Directly write new value with timestamp
  • Syntax: UPDATE table SET col = value WHERE partition_key = ?
  • Timestamp: Latest timestamp wins (last write wins)
  • Partial updates: Can update single column without reading entire row
  • Collections: Can add/remove items without replacing entire collection
Q50 How do you delete data in Cassandra? Medium
Answer: Deletes create tombstones (markers) rather than immediately removing data:
  • Tombstone: Delete writes a marker with timestamp
  • Why: Distributed system needs to track deletes across replicas
  • Cleanup: Tombstones removed after gc_grace_seconds (default 10 days)
  • Compaction: Actually removes data during compaction
  • TTL alternative: Use TTL for auto-expiration (better than DELETE)
  • Performance: Too many tombstones can slow reads
CQL & Queries
15 questions • Query language & syntax
Q51 What is the difference between INSERT and UPDATE? Medium
Answer: In Cassandra, INSERT and UPDATE are functionally identical (both are upserts):
  • Both: Write data with timestamp, last write wins
  • INSERT: Typically used when creating new row
  • UPDATE: Typically used when modifying existing data
  • No difference: Both overwrite existing data if primary key matches
  • Condition: Can add IF NOT EXISTS to INSERT for conditional writes
Q52 What is a batch statement? Easy
Answer: Batch statement groups multiple INSERT/UPDATE/DELETE into single operation:
  • Syntax: BEGIN BATCH ... APPLY BATCH
  • Atomicity: All succeed or all fail
  • Performance: Sends multiple writes in one network roundtrip
  • Logged: Default is logged batch (uses distributed log)
  • Unlogged: UNLOGGED batch faster but only for same partition
  • Limit: Keep batches small (50-100 statements max)
Consistency & Replication
10 questions • Consistency levels & replication
Q66 Explain replication factor (RF) Hard
Answer: Replication Factor determines how many copies of data exist:
  • Definition: Number of nodes that store each piece of data
  • Example: RF=3 means 3 copies across 3 different nodes
  • Per keyspace: Set during keyspace creation
  • Per DC: Can set different RF for each datacenter
  • Common: RF=3 is production standard (balance availability/cost)
  • Trade-off: Higher RF = more availability but more storage cost
Q67 What are consistency levels in Cassandra? Hard
Answer: Consistency levels determine how many replicas must respond for success:
  • ONE: 1 replica responds (fastest, least consistent)
  • TWO: 2 replicas respond
  • THREE: 3 replicas respond
  • QUORUM: Majority of replicas (RF/2 + 1)
  • LOCAL_QUORUM: Majority in local datacenter
  • EACH_QUORUM: Quorum in each datacenter
  • ALL: All replicas (slowest, strongest consistency)
Formula for strong consistency: R + W > RF (e.g., QUORUM read + QUORUM write)
Operations & Performance
15 questions • Operations, tuning & monitoring
Q76 What is nodetool and its common commands? Medium
Answer: nodetool is command-line utility for managing Cassandra:
  • status: Show cluster status, node health
  • repair: Sync data across replicas
  • compaction: Trigger manual compaction
  • cleanup: Remove data after decommission
  • flush: Force memtable flush to disk
  • tpstats: Thread pool statistics
  • tablestats: Per-table statistics
  • describecluster: Cluster information
Advanced Topics
10 questions • LWT, security, best practices
Q91 What are Lightweight Transactions (LWT)? Hard
Answer: LWT provides linearizable consistency using Paxos consensus:
  • Syntax: IF NOT EXISTS, IF condition
  • Use case: Ensure uniqueness, compare-and-set operations
  • Example: INSERT INTO users (id, email) VALUES (?, ?) IF NOT EXISTS
  • Performance: 4-10x slower than regular writes (requires consensus)
  • Consensus: Uses Paxos algorithm across replicas
  • When to use: Only when you absolutely need strong consistency
Q100 What are Cassandra best practices for production? Hard
Answer: Key production best practices:
  • RF=3: Use RF=3 per datacenter minimum
  • NetworkTopologyStrategy: Always use (even single DC)
  • vnodes: Enable with num_tokens=256
  • Consistency: LOCAL_QUORUM for multi-DC
  • Monitoring: Track latency, disk usage, GC pauses
  • Backups: Regular snapshots + repair
  • Repair: Run weekly/monthly within gc_grace_seconds
  • Capacity: Keep disk <70%, partition sizes <100MB
  • G1GC: Use for heaps >6GB
  • Prepared statements: Always prepare queries
Advertisement

Responsive Ad