Geospatial Data
Master location-based queries with Cassandra using GeoHash, quadkeys, and spatial indexing for ride-sharing, delivery, and location services!
🌍 What is Geospatial Data?
The Ride-Sharing Problem 🚗
You're building a ride-sharing app like Uber. A user opens the app in San Francisco:
- 📍 User location: 37.7749° N, 122.4194° W (lat/lon)
- 🚕 Need to find: All available drivers within 5km
- ⚡ Response time: < 100ms
- 🔄 Updates: Every 10 seconds as drivers move
- 📊 Scale: 10,000+ drivers in San Francisco alone
Challenge: How do you efficiently query "find all points within radius" in Cassandra?
Problem: ❌ Can't query by distance natively
Solution: ✅ Use GeoHash or QuadKeys to partition the world!
Geospatial Data Explained
Geospatial data = Location data (latitude/longitude) that requires proximity-based queries like "find nearby", "within radius", "along route".
Common Geospatial Queries:
- 🔍 Proximity Search: "Find restaurants within 2km"
- 📍 Nearest Neighbor: "Find closest 10 stores"
- 🗺️ Bounding Box: "Show all pins on this map view"
- 🛣️ Route Query: "Find gas stations along route"
- 🏢 Containment: "Is location inside polygon?"
Geospatial Use Cases
Ride-Sharing
Uber, Lyft, Grab
- Find nearby drivers
- Real-time location tracking
- Route optimization
- Surge pricing zones
Food Delivery
DoorDash, UberEats
- Restaurant proximity
- Delivery zones
- Driver dispatch
- ETAs calculation
Social Networks
Instagram, Snapchat
- Nearby friends
- Location-based posts
- Geotagged stories
- Event discovery
Travel & Hospitality
Airbnb, Booking.com
- Property search
- Map-based browsing
- Nearby attractions
- Availability zones
❓ The Geospatial Challenge in Cassandra
Why Geospatial is Hard in Cassandra
Cassandra has NO built-in geospatial functions. No ST_Distance, no spatial indexes, no geometry types.
What Doesn't Work
The Core Problem
Why Coordinate Queries Fail
- 🔑 Partition Key Required: Must include partition key in WHERE clause
- 📊 No Range Queries: Can't do latitude BETWEEN without partition key
- 🌍 2D Problem: Lat/lon are 2 dimensions, Cassandra thinks 1D (row order)
- ⚡ Performance: Scanning all rows to check distance = SLOW!
The Solution: Spatial Partitioning
Convert 2D coordinates (lat, lon) into 1D partition keys that group nearby locations together!
This is what GeoHash and QuadKeys do - they divide the Earth into a grid and give each grid cell a unique string ID.
🔢 GeoHash - The Most Popular Approach
What is GeoHash?
GeoHash encodes latitude/longitude into a short string like 9q8yy that represents a geographic area.
How GeoHash Works
Visual Example:
San Francisco (37.7749, -122.4194) → GeoHash: 9q8yy9mf1h5v
Character Precision:
9(1 char) = ±2,500 km (whole region)9q(2 chars) = ±630 km9q8(3 chars) = ±78 km9q8y(4 chars) = ±20 km9q8yy(5 chars) = ±2.4 km ← Good for city searches!9q8yy9(6 chars) = ±610 m9q8yy9m(7 chars) = ±76 m ← Street-level precision9q8yy9mf(8 chars) = ±19 m
Key Insight: Locations with the same GeoHash prefix are geographically close!
9q8yyxyz and 9q8yyabc are ~100m apart.
Implementing GeoHash in Cassandra
Python Implementation
GeoHash Precision Guide
GeoHash Edge Cases
- ⚠️ Border Problem: Points 1m apart can have different geohashes if on grid border
- ⚠️ Solution: Always check neighboring geohashes (8 neighbors)
- ⚠️ Polar Regions: Grid cells become smaller near poles
- ⚠️ Date Line: Special handling for longitude ±180°
📐 QuadKeys - Microsoft's Alternative
What are QuadKeys?
QuadKeys are Microsoft's geospatial indexing system used in Bing Maps. Similar to GeoHash but uses quadtree division.
QuadKey vs GeoHash
GeoHash
- Encoding: Base32 (0-9, a-z)
- Example: 9q8yy9mf
- Grid: Alternating lat/lon bits
- Pro: Shorter strings
- Con: Complex edge cases
QuadKey
- Encoding: Base4 (0, 1, 2, 3)
- Example: 02301012133
- Grid: Recursive quadtree
- Pro: Map tile aligned
- Con: Longer strings
Python QuadKey Implementation
🎨 Data Modeling for Geospatial Queries
Pattern 1: Single GeoHash Table (Simple)
Pattern 2: Multi-Level GeoHash (Advanced)
Pattern 3: Hybrid Bucketing (Production)
🔍 Common Query Patterns
Query 1: Find Nearby (Proximity Search)
Query 2: Bounding Box Search
Query 3: K-Nearest Neighbors
🎯 Real-World Use Cases
Use Case 1: Ride-Sharing Driver Matching
Uber/Lyft Architecture
Result: Sub-100ms driver search even with 100,000+ active drivers!
Use Case 2: Food Delivery Zone Assignment
Use Case 3: Social "Nearby Friends"
✅ Geospatial Best Practices
✅ DO These
- Use GeoHash 5-6 chars for cities
- Check neighboring geohashes
- Calculate actual distance (Haversine)
- Use TTL for moving objects
- Add city/region bucket
- Index status fields (available)
- Monitor hot partitions
- Cache frequent queries
❌ DON'T Do These
- Store only lat/lon (no geohash)
- Use too high precision (8+)
- Forget edge cases (borders)
- Skip distance calculation
- Create global hot spots
- Update location too frequently
- Assume geohash = distance
- Ignore privacy concerns
Precision Selection Guide
- ✅ Ride-sharing: 5-6 chars (city to neighborhood)
- ✅ Restaurant search: 6-7 chars (street level)
- ✅ Social nearby: 5-6 chars (privacy balance)
- ✅ Delivery zones: 4-5 chars (regional)
- ✅ Asset tracking: 7-8 chars (building level)
🎯 Geospatial Summary
You now understand geospatial queries in Cassandra!
📚 Key Takeaways:
- 🌍 Cassandra has NO native geospatial - need GeoHash/QuadKeys
- 🔢 GeoHash converts lat/lon → string partition key
- 📏 5 chars = ~2.4km (perfect for city searches)
- 🎯 Always check 9 cells (geohash + 8 neighbors)
- 📐 Calculate actual distance with Haversine formula
- 🏙️ Add city bucket to prevent hot partitions
- ⚡ Real use: Uber, DoorDash, Instagram, Airbnb
GeoHash + Cassandra = Fast proximity queries at any scale! 🌍🚀
Responsive Ad