User-Defined Types (UDT)
Create custom complex data structures! Bundle related fields together like building LEGO blocks for your data model.
📖 The Story: Mike's Address Nightmare
Mike is building a user management system. Every user has an address with street, city, state, and zip code. He tried storing it THREE different ways...
❌ Attempt 1: Separate Columns (Messy!)
Problems:
- 🤯 16 columns for just 2 addresses! What if we add shipping address? 24 columns!
- 📝 Updating address requires 4 UPDATE statements
- 🐛 Easy to miss a field (forgot to update zip code!)
- ❌ Can't validate complete address (partial data corruption)
⚠️ Attempt 2: JSON String (Hacky!)
Problems:
- 🔍 Can't query by city (it's inside JSON string!)
- ❌ No type validation (stored "12345" for city name!)
- 🐌 Must parse JSON every read (slow!)
- 💥 JSON parsing errors at runtime
✅ The RIGHT Way: User-Defined Types!
Benefits:
- ✅ Clean schema: 3 columns instead of 12!
- ✅ Reusable: Define address ONCE, use everywhere
- ✅ Type-safe: Cassandra validates each field
- ✅ Atomic updates: Update entire address in one operation
- ✅ Easy to add: Need billing address? Just add one column!
Mike's code is now clean, maintainable, and production-ready! 🎉
🎭 What are User-Defined Types?
UDTs let you create custom complex data structures by combining multiple primitive types into a single reusable unit.
Simple Definition
User-Defined Type (UDT): A custom data structure that groups related fields together, like a mini-table or struct.
Think of UDT as:
- 📦 LEGO Blocks: Build complex structures from simple pieces
- 🧩 Puzzle Pieces: Snap related fields together
- 🏗️ Building Blocks: Create reusable components
- 📚 Templates: Define structure once, use many times
When to Use UDT
Perfect For
- Grouped Data: Address, phone number, geolocation
- Repeating Structures: Multiple addresses per user
- Complex Objects: Payment info, ratings, coordinates
- Atomic Updates: Update all fields together
- Cleaner Schema: Reduce column count
Avoid For
- Frequently Queried Fields: Can't index UDT fields directly
- Partial Updates: Must replace entire UDT
- Large Data: Keep UDTs small (< 1KB typical)
- High Mutation Rate: Whole UDT must be rewritten
- Query Filters: Can't WHERE on UDT subfields
📝 UDT Syntax & Usage
Master the complete UDT lifecycle: CREATE, USE, ALTER, DROP.
Creating a UDT
Using UDT in Tables
Why FROZEN?
FROZEN means the UDT is treated as a single BLOB - you cannot update individual fields.
What FROZEN Does:
- Entire UDT stored as one unit (blob)
- Cannot update single field (must replace whole UDT)
- Better performance (single read/write)
- Simpler consistency model
Inserting Data with UDT
Querying UDT Data
🏗️ Real-World UDT Examples
Production-ready UDT patterns from real applications!
Example 1: E-Commerce Product Catalog
Example 2: Social Media User Profile
Example 3: IoT Sensor Data
Example 4: Payment System
🔗 Nested User-Defined Types
UDTs can contain other UDTs! Build complex hierarchical structures.
Nested UDT Best Practices
- Limit Nesting Depth: Max 2-3 levels deep (readability!)
- Create Base Types First: Build from bottom-up
- Always FROZEN: All nested UDTs must be FROZEN
- Document Structure: Complex hierarchies need good docs
- Consider Performance: Deep nesting = larger blobs
🖥️ Interactive UDT Console
Practice UDT commands in our safe simulator!
Try the examples or create your own UDT...
Available Examples:
• Example 1: Create address type
• Example 2: Use UDT in table
• Example 3: Nested UDT
⭐ UDT Best Practices & Common Mistakes
Production-proven strategies and pitfalls to avoid!
DO's
- Keep UDTs Small: < 1KB typical, max 1MB
- Group Related Data: Address, contact info, coordinates
- Always Use FROZEN: Required for UDT columns
- Reuse Types: Define once, use everywhere
- Name Clearly: address_type vs addr (be descriptive!)
- Document Schema: Especially for nested UDTs
- Version Types: Consider address_v2 for breaking changes
DON'Ts
- Don't Query by Subfields: Can't WHERE on UDT.field
- Don't Partial Update: Must replace entire UDT
- Don't Overuse: Not every column needs UDT
- Don't Make Huge: Large UDTs = slow reads/writes
- Don't Nest Too Deep: Max 2-3 levels
- Don't Skip FROZEN: Non-frozen UDTs fail
- Don't Index UDTs: Can't create secondary index
Pro Tips
- Denormalize for Queries: If you need to filter by city, add city column
- Use Collections: LIST<FROZEN<udt>> for multiple
- Null Individual Fields: Missing fields become null
- Monitor Size: Track UDT size in production
- Plan for Evolution: How will you handle schema changes?
- Test Performance: Large UDTs impact throughput
- Consider Alternatives: Sometimes separate tables better
⚠️ Common Mistake: Trying to Query by UDT Subfield
The Problem: Developer tries to find all users in NYC...
The Fix:
💼 Interview Questions & Expert Answers
Ace your Cassandra interview with these UDT questions!
Answer:
UDTs and Collections serve different purposes:
User-Defined Type (UDT):
- Purpose: Group different data types together
- Structure: Named fields with specific types
- Example: address {street: TEXT, city: TEXT, zip: TEXT}
- Use Case: Composite data that belongs together
Collection (LIST/SET/MAP):
- Purpose: Store multiple values of same type
- Structure: Homogeneous elements
- Example: LIST<TEXT> → ['tag1', 'tag2', 'tag3']
- Use Case: Multiple similar items
Combined: You can have LIST<FROZEN<address>> - multiple addresses!
Answer:
FROZEN treats the entire UDT as a single immutable blob, which simplifies Cassandra's internals.
Why FROZEN is Required:
- Consistency: Entire UDT is atomic - all fields update together
- Performance: Single read/write operation (not field-by-field)
- Indexing: UDT stored as one blob, easier to handle internally
- Comparison: Can compare entire UDT for equality
Trade-off:
Cannot update individual fields - must replace entire UDT:
Answer: NO - This is a major UDT limitation!
You CANNOT do this:
Why Not:
- UDT is stored as single blob (FROZEN)
- Cassandra doesn't index individual UDT fields
- Would require deserializing every row (super slow!)
Workarounds:
- Denormalize: Add city as separate column
- Separate Table: Create users_by_city table
- Application Filter: Fetch all, filter in app (small datasets only)
Answer: You can ADD fields, but cannot REMOVE or RENAME!
Adding a Field (Allowed):
Cannot Remove Fields:
Cannot Rename Fields:
Workaround for Breaking Changes:
- Create new UDT (address_v2)
- Create new table using address_v2
- Migrate data from old table to new
- Drop old table and UDT
Answer: Existing data remains valid - new field is NULL for old rows!
Scenario:
What Happens:
- ✅ Old data still works perfectly
- ✅ Queries on old data return NULL for country
- ✅ New inserts can include country field
- ✅ You can UPDATE old rows to add country
Best Practice: Plan for evolution - add optional fields, never remove!
🎓 Chapter Summary: UDT Mastery
Congratulations! You now understand User-Defined Types at a production level!
Key Concepts Mastered:
- UDT Definition: Custom data structures grouping related fields
- FROZEN Requirement: UDTs must be FROZEN (atomic updates)
- Nested UDTs: UDTs can contain other UDTs (max 2-3 levels)
- Collections: LIST/SET/MAP can contain UDTs
- Limitations: Cannot query by subfields, cannot partial update
The Golden Rule:
UDT = Grouping Related Data Together
Keep UDTs small, reusable, and always FROZEN!
When to Use UDT:
- ✅ Address: street, city, state, zip together
- ✅ Contact Info: phone, email, social links
- ✅ Coordinates: latitude, longitude, altitude
- ✅ Metadata: width, height, format, size
- ❌ Don't use for: Fields you need to query/filter by
🚀 You're now equipped to design clean, maintainable schemas with UDT!
Responsive Ad