Advanced Topics

Graph Data

Model relationships and connections in Cassandra - social networks, recommendations, knowledge graphs!

🕸️ What is Graph Data?

The Social Network Problem 👥

You're building LinkedIn. Users can:

  • 🔗 Connect with other users (friends/follows)
  • 👀 See their network (1st, 2nd, 3rd degree connections)
  • 💡 Get friend recommendations ("People you may know")
  • 🔍 Find shortest path between two people
  • 📊 Discover mutual friends

Challenge: How do you model relationships and traverse networks efficiently in Cassandra?

Graph Data Explained

Graph data = Data focused on relationships between entities (nodes/vertices connected by edges).

Key Concepts:
  • 🔵 Nodes (Vertices): Entities (users, products, places)
  • ➡️ Edges: Relationships (follows, likes, bought)
  • 🏷️ Properties: Data on nodes/edges (name, timestamp, weight)
  • 🎯 Traversal: Following edges to explore connections

Graph Use Cases

👥

Social Networks

  • Friend connections
  • Follower/following
  • Mutual friends
  • Social influence
  • Community detection
💡

Recommendations

  • Product suggestions
  • Friend recommendations
  • Similar items
  • Collaborative filtering
  • Content discovery
🧠

Knowledge Graphs

  • Entity relationships
  • Semantic search
  • Question answering
  • Inference engines
  • Ontologies
🛣️

Network Analysis

  • Shortest paths
  • Route optimization
  • Supply chain
  • Dependency graphs
  • Infrastructure

⚡ Cassandra vs Dedicated Graph Databases

Important: Cassandra is NOT a Graph Database

Cassandra can model and query graph data, but it's not optimized for deep graph traversals like Neo4j or Amazon Neptune. It excels at shallow traversals (1-2 hops) at massive scale.

When to Use Each

✅ Use Cassandra For

  • Shallow traversals (1-2 hops)
  • Massive scale (billions of edges)
  • High writes (social activity)
  • Simple patterns (followers, friends)
  • Time-series graphs (activity feeds)
  • Known start nodes (user A's friends)

Example: Facebook news feed, Twitter followers, LinkedIn 1st connections

✅ Use Graph DB For

  • Deep traversals (4+ hops)
  • Complex queries (pattern matching)
  • Path algorithms (shortest path)
  • Graph analytics (PageRank, centrality)
  • Unknown patterns (discover relationships)
  • Ad-hoc traversals (exploratory)

Example: Fraud detection, knowledge graphs, recommendation engines

Comparison Table

Feature Cassandra Neo4j
Traversal Depth 1-2 hops ✅ Unlimited ✅
Scale (edges) Billions ✅ Millions
Write throughput Very high ✅ Moderate
Query language CQL Cypher ✅
Path finding Application-side Built-in ✅
Best for Social feeds Fraud detection

🎨 Modeling Graphs in Cassandra

Pattern 1: Adjacency Lists (Most Common)

Model: Store each user's connections

-- Store who Alice follows CREATE TABLE user_follows ( user_id uuid, ← Alice's ID followed_at timestamp, ← When she followed followed_id uuid, ← Who she follows followed_username text, PRIMARY KEY (user_id, followed_at) ) WITH CLUSTERING ORDER BY (followed_at DESC); -- Query: Who does Alice follow? SELECT followed_id, followed_username FROM user_follows WHERE user_id = 'alice-uuid'; -- Result: Bob, Carol, David (instant lookup!)

Pattern 2: Bidirectional Edges (Mutual Relationships)

-- For friend relationships (mutual), store BOTH directions CREATE TABLE friendships ( user_id uuid, friend_since timestamp, friend_id uuid, friend_username text, PRIMARY KEY (user_id, friend_since) ); -- When Alice and Bob become friends, write TWO rows: INSERT INTO friendships (user_id, friend_since, friend_id, friend_username) VALUES ('alice-uuid', now(), 'bob-uuid', 'Bob'); INSERT INTO friendships (user_id, friend_since, friend_id, friend_username) VALUES ('bob-uuid', now(), 'alice-uuid', 'Alice'); -- Now both queries work: -- "Alice's friends" AND "Bob's friends"

Pattern 3: Reverse Index (Who Follows Me?)

-- Forward: Who I follow CREATE TABLE user_following ( user_id uuid, followed_at timestamp, followed_id uuid, PRIMARY KEY (user_id, followed_at) ); -- Reverse: Who follows me CREATE TABLE user_followers ( user_id uuid, ← Person being followed follower_since timestamp, follower_id uuid, ← Person following follower_username text, PRIMARY KEY (user_id, follower_since) ); -- When Alice follows Bob, write to BOTH tables: -- 1. user_following: Alice → Bob -- 2. user_followers: Bob ← Alice

Pattern 4: Edge Properties (Weighted Relationships)

-- Store metadata about the relationship CREATE TABLE user_connections ( user_id uuid, connected_at timestamp, connected_id uuid, connection_type text, ← 'friend', 'family', 'colleague' strength int, ← Interaction frequency (1-100) mutual_friends int, ← Number of mutual connections last_interaction timestamp, PRIMARY KEY (user_id, connected_at) ); -- Use for: Friend recommendations, feed ranking

Pattern 5: Multi-Hop Pre-Computation

-- Pre-compute 2nd degree connections CREATE TABLE second_degree_connections ( user_id uuid, connection_id uuid, mutual_friend_count int, common_interests set, PRIMARY KEY (user_id, connection_id) ); -- Computed offline (batch job): -- For Alice, find all friends-of-friends -- Store: Alice → [Potential friends with mutual connection count] -- Query: "People you may know" SELECT * FROM second_degree_connections WHERE user_id = 'alice-uuid' ORDER BY mutual_friend_count DESC LIMIT 10;

📊 Common Query Patterns

Query 1: Direct Connections (1-Hop)

-- Get Alice's followers SELECT * FROM user_followers WHERE user_id = 'alice-uuid'; -- Get who Alice follows SELECT * FROM user_following WHERE user_id = 'alice-uuid'; -- ✅ O(1) partition lookup - instant!

Query 2: Mutual Friends (Application-Side Join)

# Python - Find mutual friends between Alice and Bob def get_mutual_friends(alice_id, bob_id): # 1. Get Alice's friends alice_friends = session.execute(""" SELECT friend_id FROM friendships WHERE user_id = %s """, (alice_id,)) alice_set = set(row.friend_id for row in alice_friends) # 2. Get Bob's friends bob_friends = session.execute(""" SELECT friend_id FROM friendships WHERE user_id = %s """, (bob_id,)) bob_set = set(row.friend_id for row in bob_friends) # 3. Intersection (application-side) mutual = alice_set & bob_set return list(mutual) # ✅ Two O(1) queries + in-memory intersection

Query 3: 2-Hop Traversal (Friends-of-Friends)

# Find Alice's friends-of-friends (2nd degree) def get_friends_of_friends(user_id): # 1. Get direct friends (1st degree) first_degree = session.execute(""" SELECT friend_id FROM friendships WHERE user_id = %s """, (user_id,)) friends_of_friends = set() # 2. For each friend, get THEIR friends (2nd degree) for friend in first_degree: second_degree = session.execute(""" SELECT friend_id FROM friendships WHERE user_id = %s """, (friend.friend_id,)) for row in second_degree: # Exclude self and direct friends if row.friend_id != user_id and row.friend_id not in first_degree: friends_of_friends.add(row.friend_id) return list(friends_of_friends) # ⚠️ N+1 queries (one per friend) # ✅ OK if N < 500, slow if N > 5000

Query 4: Activity Feed (Time-Based Graph)

-- Get recent activity from people Alice follows CREATE TABLE user_feed ( user_id uuid, ← Alice's ID post_time timestamp, ← When posted post_id uuid, author_id uuid, ← Who posted content text, PRIMARY KEY (user_id, post_time) ) WITH CLUSTERING ORDER BY (post_time DESC); -- Write pattern: When Bob posts, copy to ALL his followers' feeds -- This is "fan-out on write" -- Read pattern: Alice's feed (instant!) SELECT * FROM user_feed WHERE user_id = 'alice-uuid' LIMIT 50; -- ✅ Single partition read - sub-10ms!

Query 5: Recommendation Score

-- Pre-computed friend recommendations CREATE TABLE friend_recommendations ( user_id uuid, recommendation_score double, ← Computed: mutual friends + interests candidate_id uuid, candidate_username text, mutual_friends int, common_interests set, PRIMARY KEY (user_id, recommendation_score) ) WITH CLUSTERING ORDER BY (recommendation_score DESC); -- Query: Top 10 recommendations for Alice SELECT * FROM friend_recommendations WHERE user_id = 'alice-uuid' LIMIT 10; -- ✅ Pre-computed by batch job (Spark)

🎯 Real-World Use Cases

Use Case 1: Social Network (Twitter-Style)

Twitter's Follow Model

-- Following (who I follow) CREATE TABLE following ( user_id uuid, followed_at timestamp, followed_id uuid, followed_username text, followed_bio text, PRIMARY KEY (user_id, followed_at) ) WITH CLUSTERING ORDER BY (followed_at DESC); -- Followers (who follows me) CREATE TABLE followers ( user_id uuid, follower_since timestamp, follower_id uuid, follower_username text, PRIMARY KEY (user_id, follower_since) ) WITH CLUSTERING ORDER BY (follower_since DESC); -- Stats (for profile) CREATE TABLE user_stats ( user_id uuid PRIMARY KEY, follower_count counter, following_count counter ); -- When Alice follows Bob: -- 1. INSERT into following (Alice → Bob) -- 2. INSERT into followers (Bob ← Alice) -- 3. UPDATE counters

Scale: Handles billions of follow relationships (Twitter has 500M+ users)

Use Case 2: Product Recommendations (Amazon-Style)

-- User viewing history → Similar products CREATE TABLE product_similarity ( product_id uuid, similarity_score double, similar_product_id uuid, similar_product_name text, co_purchase_count int, ← Bought together count PRIMARY KEY (product_id, similarity_score) ) WITH CLUSTERING ORDER BY (similarity_score DESC); -- Computed by: Collaborative filtering on purchase history -- "Users who bought X also bought Y" -- Query: When viewing product A, show similar products SELECT * FROM product_similarity WHERE product_id = 'product-A-uuid' LIMIT 5;

Use Case 3: Knowledge Graph (Wikipedia-Style)

-- Entity relationships CREATE TABLE entity_relationships ( entity_id uuid, ← "Albert Einstein" relationship_type text, ← "worked_at", "born_in" related_entity_id uuid, ← "Princeton University" related_entity_name text, relationship_strength double, PRIMARY KEY ((entity_id, relationship_type), related_entity_id) ); -- Query: Where did Einstein work? SELECT * FROM entity_relationships WHERE entity_id = 'einstein-uuid' AND relationship_type = 'worked_at';

✅ Best Practices for Graph Data in Cassandra

✅ DO These

  • Model for 1-2 hop queries
  • Denormalize heavily
  • Store bidirectional edges
  • Pre-compute deep traversals
  • Use counters for stats
  • Fan-out writes for feeds
  • Limit graph traversal depth
  • Batch compute recommendations

❌ DON'T Do These

  • Try deep traversals (4+ hops)
  • Real-time path finding
  • Ad-hoc graph queries
  • Recursive traversals
  • Complex graph algorithms
  • Global graph analytics
  • Pattern matching queries
  • Unknown start nodes

When Cassandra Works Well

  • ✅ Social feeds: Twitter timeline, Facebook news feed
  • ✅ Direct connections: LinkedIn 1st degree, follower lists
  • ✅ Pre-computed recs: "People you may know" (batch computed)
  • ✅ Activity graphs: User interaction history
  • ✅ Massive scale: Billions of edges, millions of writes/sec

When to Use Graph DB Instead

  • ⚠️ Fraud detection: Multi-hop pattern detection
  • ⚠️ Recommendation engines: Complex collaborative filtering
  • ⚠️ Shortest path: Route finding, network optimization
  • ⚠️ Graph analytics: PageRank, community detection
  • ⚠️ Knowledge graphs: Semantic queries, inference

🎯 Graph Data Summary

You now understand graph modeling in Cassandra!

📚 Key Takeaways:

  • 🕸️ Cassandra excels at shallow (1-2 hop) graph queries
  • 📊 Use adjacency lists (store connections per user)
  • 🔄 Store bidirectional edges for mutual relationships
  • ⚡ Pre-compute deep traversals (batch jobs)
  • 📝 Fan-out writes for activity feeds
  • 🎯 Model for specific query patterns
  • 📈 Scales to billions of edges
  • ⚠️ Use Neo4j for deep traversals (4+ hops)

Cassandra + Graph = Fast social networks at massive scale! 🕸️🚀

Advertisement

Responsive Ad