FlockDB was an open-source, distributed, fault-tolerant graph database developed by Twitter for managing wide but shallow network graphs. It was primarily used to store relationships between Twitter users, such as followings and favorites.
Overview
FlockDB was designed for rapid set operations on very large adjacency lists rather than for traditional graph traversal. It differs from other graph databases of its era, such as Neo4j, in that it was not optimized for multi-hop graph queries. Instead, it was optimized for fast reads, writes, and pageable set arithmetic operations — a usage pattern comparable to that of Redis sets.
Architecture and Design
FlockDB stored graph data as sets of directed edges between nodes identified by 64-bit integers. Its query model focused on adjacency-list operations such as adding, removing, and counting edges, as well as set arithmetic across multiple edge sets. The database intentionally avoided multi-hop traversal features in favor of high performance in real-time, large-scale environments.
The distributed datastore was queried through Twitter's Gizzard framework, a Scala-based sharding and replication framework that was released shortly before FlockDB itself. This combination allowed FlockDB to achieve fault tolerance and horizontal scalability.
Release History
FlockDB was released by Twitter in April 2010. The project's source code was subsequently published to GitHub. The final release was version 1.8.5, dated February 23, 2012.
Technical Details
- Original authors: Nick Kallen, Robey Pointer, John Kalucki, and Ed Ceaser
- Developer: Twitter
- Written in: Scala, Java, Ruby
- License: Apache License 2.0
Current Status
Twitter no longer maintains FlockDB. The project is archived on GitHub under the "twitter-archive" organization, and the company has stated it is not responding to issues or pull requests regarding the project.
Significance
FlockDB is notable as one of the early examples of a purpose-built, distributed graph database deployed at massive internet scale. Its design decisions — prioritizing adjacency-list performance and set operations over general-purpose graph traversal — reflected the specific workloads of large social platforms. The project influenced subsequent approaches to storing social graph data and remains referenced in discussions of graph database design trade-offs.
See Also
- Gizzard (Scala framework)