---
title: "Announcing blob-stream: a Kafka alternative for no fuss, low cost high volume streaming"
slug: "blob-stream-kafka-alternative"
blurb: "Today I’m thrilled to announce the [open source release of blob-stream](https://github.com/bitdriftlabs/blob-stream), a [Kafka](https://kafka.apache.org/)-like streaming system for high-throughput workloads that prioritizes total cost of ownership over ultra-low latency. I suspect that a large number of cloud workloads using Kafka today can switch to blob-stream as an alternative. The why is simple: to vastly reduce costs, and possibly more importantly, vastly reduce operational burden.\n\nThis is a long post, so I won’t be offended if you skip to the sections that most interest you. I’m going to cover what blob-stream is, why I built blob-stream, and possibly most interesting, *how* I built it. (teaser: yes a lot of AI, but also a lot of good old fashioned systems engineering.)"
metaDescription: ""
cover:
  url: "/assets/posts/blob-stream-kafka-alternative/feature-blob-stream-blog-desktop@1x.webp"
  alt: "Announcing blob-stream: a Kafka alternative for no fuss, low cost high volume streaming"
socialThumbnail:
  url: "/assets/posts/blob-stream-kafka-alternative/feature-blob-stream-blog-desktop@1x.webp"
  alt: "Announcing blob-stream: a Kafka alternative for no fuss, low cost high volume streaming"
author:
  - "matt"
publishedDate: "2026-09-28T19:11:40.688Z"
modifiedDate: "2026-09-28T19:11:40.688Z"

---

# What is blob-stream?

## Similarities to Kafka

From a usage perspective, blob-stream is a streaming system that looks a lot like Kafka:

1. **There are topics and partitions.** Producers hash records via a user provided key onto a partition for locality.
2. **Producers send batches of partition owned records to brokers** that own those partitions and make the records durable before acknowledging success to the producer.
3. **Consumers organize themselves** into groups, divide the partitions for a topic amongst themselves, and then process partition records. Typical Kafka-like semantics are supported such as committing, seeking, assignment and revocation, and so on.
4. **The broker is a standalone binary.** As of this writing, there is a Rust producer library and consumer library. These libraries are similar to the role of [rdkafka](https://github.com/confluentinc/librdkafka) in the Kafka ecosystem. Producer applications use the producer side of the library to send records to brokers. Consumer applications use the consumer side of the library to receive records from assigned partitions.

## blob-stream goals

Kafka is the gold standard for high volume durable streaming applications for good reason. It is open source, extremely widely used, and extremely stable. Replacing it for similar use cases would have to come along with some very good reasons.

To this end, blob-stream has the following goals:

1. **Zero cross-Availability Zone (AZ) network traffic**. At volume, even with rack-aware consumers, cross-AZ networking costs dominate in a Kafka cloud deployment.

2. **Stateless brokers with zero local storage**. It is well known that managing Kafka broker storage is operationally complicated. Kafka brokers are stateful. Partitions have to be carefully balanced across brokers and rebalancing is a task that very few want to take on. This is so painful that at [bitdrift](https://bitdrift.io/), historically we didn’t even bother. We just made new clusters, and our software knew how to gracefully cutover from one to the other. Afterwards we would delete the first cluster.

3. **Brokers can rotate and autoscale at will without downtime**. Related to (2), traditional Kafka deployments require a complicated dance and significant over-provisioning to deal with regular maintenance such as OS/security patches and scaling up and down for peak events.

4. **No independent control plane or cluster manager beyond stock Kubernetes**. Traditional Kafka deployments and the more modern diskless replacements still require an independent control plane to manage the brokers. This creates additional operational burden and more moving pieces.

5. **Significantly lower cost than even the best modern diskless Kafka solutions**. Removing cross-AZ networking is only part of the picture. A highly efficient implementation means less compute costs across producers, brokers, and consumers. Further, not being beholden to the Kafka API and backwards compatibility with existing Kafka libraries allows for significant efficiency gains across the entire system which I will describe in more detail below.

6. **Excellent observability via metrics, OTLP traces, and admin HTTP state snapshots**. This is not unique to blob-stream, but I strongly believe that all systems software should be easy to observe and blob-stream follows through.

## blob-stream architecture at a high level

<Image alt="blob-stream architecture overview" asset="/assets/posts/blob-stream-kafka-alternative/diagram@1x.webp" width={500} />

This section is a very abbreviated overview of the blob-stream architecture that covers the highlights. For in-depth treatment please see the [design overview in the repo](https://github.com/bitdriftlabs/blob-stream/tree/main/docs/design).

In order to achieve the above goals, and in order to take advantage of primitives now commonly found in modern infrastructures, blob-stream diverges from Kafka in some significant ways.

### Clock synchronization

First and foremost, **blob-stream requires an accurate clock** shared between brokers and consumers. Exploiting clock synchronization in distributed systems to allow linearization without locking started with Google Spanner and is now found in many derivative systems. Blob-stream was designed and implemented against the [AWS Time Sync Service](https://aws.amazon.com/blogs/aws/keeping-time-with-amazon-time-sync-service/) which uses GPS/atomic clocks to provide extremely accurate system time readings to machines running within AWS datacenters, similar to Google TrueTime and other related solutions. How blob-stream utilizes accurate clocks to reduce locking and overall costs will be described more below.

### Partitioning model

Like Kafka, blob-stream topics have some number of partitions. Unlike Kafka, blob-stream partitions are availability-zone (AZ) local. Each AZ is assigned a *writer ID*. So for example a blob-stream deployment in 3 AZs would have the AZs assigned writer IDs 0 \- 2\. Every *logical* partition in each AZ becomes a *virtual* partition globally. For example, logical partition 1 becomes 3 virtual partitions globally across 3 AZs. Consumers consume from *all* virtual partitions, and the virtual partitions for a single logical partition may or may not wind up assigned to the same consumer. Workloads that require a single record hash to yield a global order in a single partition cannot currently be satisfied by blob-stream, [though this is likely solvable in the future](https://github.com/bitdriftlabs/blob-stream/blob/main/docs/faq.md#how-do-partitions-relate-to-virtual-partitions).

### Broker discovery and partition ownership

In production, producers discover AZ local brokers via Kubernetes service discovery. Consistent hashing is used to assign all topic logical partitions to discrete brokers. This means that all producers have a consistent view of which brokers should receive which partition records. Blob-stream relies on Kubernetes to do the heavy lifting of broker scheduling, health checking, replacement, discovery, and so on.

Brokers use the same consistent hashing of themselves within an AZ to understand which topic logical partitions they own. They use a NoSQL database (in this case DynamoDB) to acquire leases for those partitions that allow them to reserve offset sequence ranges, write durable record blocks, and write metadata that describes those durable record blocks. During periods of broker rotations and service discovery convergence, if a broker does not own a lease for partition records it receives, it responds with a specific “lease not held” response. Producers then back off and wait for discovery convergence before retrying. In practice with graceful draining and active discovery watches, the system converges within seconds.

### Offset allocation

Like Kafka, logical partition record offsets are monotonically increasing and non-overlapping. Unlike Kafka, logical partition offsets **can** have gaps. Brokers allocate large blocks of offsets against a partition lease. They then allocate segment offsets from this allocation directly from memory. Segments are larger units of contiguous record offsets that are durably written in a block. Offset allocation in this manner is the first major cost optimization. Per-segment offset allocation from pre-allocated memory blocks avoids transactional offset allocation on every segment write. Gaps follow when brokers restart, either normally or abnormally. For safety reasons an allocated block is never reused even if it was never fully sub-allocated to written segments. When brokers do not restart at steady state, offset allocations *are* gapless. This property makes the system easier to verify which I will cover in more detail below.

### Writing segments to blob storage

As is probably obvious from the name “blob-stream,” brokers write segments directly from memory to blob storage (in this case S3). Like all other diskless Kafka solutions, the major tradeoff when using blob storage is latency vs. API costs. For this type of system, storage is mostly “free” compared to the cost of writing blobs in API costs. Writing blobs too often will destroy any overall cost savings. Blob-stream doesn’t offer much new in this area; it provides all the normal knobs around max flush delay, max segment size, and so on. Blob-stream does implement an adaptive flush delay feature which allows the flush delay to float between a max and min, attempting to keep blob segments below the split point. Splitting a segment in half doubles S3 PUT costs. In these cases it makes more sense to attempt to lower overall system latency if extra PUT costs are going to be paid regardless.

### The metadata store: primary and range keys

A large part of a streaming system like Kafka is being able to determine what segments (and contained offsets) exist for a topic partition so they can be delivered to consumers. Blob-stream relies on DynamoDB as the metadata store. Blob-stream makes use of DynamoDB as a streaming metadata store in a fairly unique way that avoids any transactions in the common segment write path and allows for trivial discovery of available segments.

1. The primary key of the segment table is a combination of the topic name and a 5 minute window in unix epoch seconds. For example telemetry\#900.
2. The range key of the segment table is a type of [Snowflake ID](https://en.wikipedia.org/wiki/Snowflake_ID). Snowflake IDs have the property that using only 64-bits they are both unique within a machine ID domain and most importantly lexicographically ordered by time. They accomplish this by combining a time component, a machine identifier (typically the private IP address of the allocator), and an incrementing sequence used to allocate within a single time slice (blob-stream uses 10ms time slices).
3. The payload of each segment row describes the path to the blob that contains segment data, as well as metadata about every contained topic virtual partition including which range bytes within the blob contain it.

Using a primary and range key as described above allows *all* brokers globally to write into the segment table without any coordination or transactions. However, this is why blob-stream **requires** accurate clocks. Consumer scanning uses both the primary key window as well as the Snowflake range key to efficiently perform both recovery and fast frontier scanning at very low cost. Without a shared accurate clock between brokers and consumers this type of coordination would be impossible.

### Consumer coordination

While producer/broker coordination are AZ local and occur via consistent hashing based on Kubernetes service discovery, as described above, consumer coordination is global. Consumers must discover all consumers across all AZs and appropriately assign all virtual partitions amongst the population. At a high level the way this works is as follows:

1. **Consumers register themselves in a DynamoDB membership table.** They periodically heartbeat this table to verify continued presence.
2. **Consumers self elect a planner through DynamoDB.** If a planner dies another consumer will self-elect in its place.
3. **The planner performs a sticky assignment** for all consumers against all topic virtual partitions and writes the plan back to DynamoDB.
4. **All consumers read the active plan, and then acquire partition leases** for their assigned partitions. The consumer library handles assignment and revocation callbacks during this process just like Kafka.
5. **Consumers then begin segment scanning** using source checkpoint and cursor information stored in the partition lease record. A source checkpoint is a topic window and snowflake lower bound which avoids overscanning. The cursor is the actual application committed offset which further filters segment data to records that need to be delivered.

### Consumer scanning

The consumer scanning algorithm is the most complicated part of blob-stream by far and this abbreviated introduction does not do it full justice. I would again encourage you to [read the full design document](https://github.com/bitdriftlabs/blob-stream/tree/main/docs/design) for more information and various fully worked reader examples which will make things much more clear.

## A summary of how blob-stream achieves low cost

After the dense architecture summary in the previous section, let me summarize the mechanisms that blob-stream uses to optimize both total cost and operational cost.

1. **Producers discover and write to brokers in the local AZ only.** There is zero cross-AZ traffic in the production path. Topic logical partition assignment is handled solely via consistent hashing on top of Kubernetes service discovery and DynamoDB leases.
2. **Brokers allocate offset blocks periodically and write segments using Snowflake IDs** as the source of deployment uniqueness. None of this requires transactions in the hot path.
3. **The use of Snowflake IDs as the range key allows efficient scanning of available segments** based on seek/recovery time as well as last seen segment data on the fast frontier. All of this does not require any transactions.
4. **Consumers query for available segments and also fetch blobs via the brokers in their local AZ.** This uses the same consistent hashing based on Kubernetes service discovery to assign metadata queries and blob fetches to the same broker, greatly reducing the amount of DynamoDDB RRUs and S3 GET API calls required to service records. The default configuration coalesces metadata requests for 250ms to further reduce DynamoDB queries. Fully controlling the producer and consumer libraries (and not being forced to support Kafka libraries) makes the caching process much simpler and more efficient: the libraries go directly to the broker that is likely to yield a cache hit or successful coalesced operation and avoid any broker-to-broker cache sharing.
5. **Blobs contain data for all topics packed together,** up to the max segment size limit. This means that in the write path costs scale with throughput, not the number of topics and partitions.

## blob-stream reliability, performance and cost in the real world

bitdrift uses blob-stream as its primary durable streaming system across many different clusters at a relatively large scale. As of this writing, in our largest cluster:

1. There are 4 topics.
2. There are 100 logical partitions across 3 AZs (300 virtual partitions) across all topics.
3. At peak \~6 million records per second are processed, with a total uncompressed data throughput of \~11Gbps.
4. This is handled by 6 brokers total, 2 per AZ, with each broker allocated 2 8g-family Graviton vCPUs and 4GiB of RAM. These brokers run in our normal Kubernetes clusters amongst all other workloads and are not allocated special Kubelet nodes.
5. The brokers are configured with a max flush delay of 500ms and a max segment size of 16MiB. Depending on the moment, blob-stream’s variable flush delay feature adjusts down towards 400ms to keep pace with max segment size.
6. Total end to end record latency is \~2s.

Approximate total cost:

1. \~\$55 per day in S3 storage (3 day retention, ZSTD compressed), S3 API, and DynamoDB API at list prices.
2. \~\$12 per day in 8g EC2 family VCPU and RAM costs at list prices.
3. Rounding up, let’s call it **\$70 per day, or about \$2,100 per month.** This is obviously not including compute resources used within the producers and consumers as it’s hard to tease apart blob-stream processing vs. application processing.

I am quite certain that this is about as cheap as it gets for this type of system, and it does not get any easier to operate.

With all that said, and before the distributed systems correctness folks jump all over me, I want to be very honest that this deployment does *not* configure metadata to lease write fencing, so is susceptible to stalled writer wakeup loss. At bitdrift we accept this loss case for our workload given the realities of how unlikely this is in our real world deployment. [See this FAQ entry for more detail](https://github.com/bitdriftlabs/blob-stream/blob/main/docs/faq.md#what-data-loss-conditions-are-tolerated).

# Why did I build a Kafka alternative?

After reading the previous section, you would be justified in thinking: this is insane. Why did you build a Kafka replacement when you surely have something better to do building a startup? Or more specifically, this looks a lot like [WarpStream](https://www.warpstream.com/), why not just use that?

Before going further, let me give credit where credit is due. WarpStream is a great solution and I have a lot of respect for what the team there has built. The fine folks at WarpStream showed this type of system is possible and I am following in their footsteps.

So with that said, and beyond the fact that I have made a career out of doing ([and succeeding at](https://github.com/envoyproxy/envoy)) things people think are insane, here is why I built this, why we don’t use WarpStream, and why we now rely on blob-stream at bitdrift.

We have used standard AWS MSK deployments since the beginning of bitdrift. It’s an extremely reliable piece of software (based on open source Kafka) and has served us well. However, it’s both expensive and hard/impossible to autoscale in response to large load spikes. At bitdrift we serve some of the largest companies in the world and those companies have large load spikes, [such as 121 million concurrent users for the Cricket World Cup](https://aws.amazon.com/blogs/architecture/how-bitdrift-scaled-to-121-million-concurrent-grpc-connections-on-amazon-cloudfront-for-live-telemetry-sporting-events/). This has caused both operational burden for us and extreme risk: a traditional Kafka deployment is not something we can scale easily so it has to be significantly overprovisioned in anticipation of any large amount of load.

What I said above is true about WarpStream: I have a lot of respect for what they built. But I also have an extreme aversion to my business critically depending on a closed source piece of software. That has two implications:

1. I cannot control the pricing over time, and
2. More importantly, I cannot see the code running in my clusters if something goes haywire. For this reason alone I have kicked the can down the road on replacing MSK for a long time.

You are probably now asking: but you must use other proprietary software at bitdrift, right? And indeed we do\! The best example of this is our reliance on [ClickHouse Cloud](https://clickhouse.com/cloud) for substantial portions of our service. I am a huge fan of ClickHouse Cloud and I think they have built an amazing product which is why we use it (and, shameless plug, we have also recently [partnered with them](https://blog.bitdrift.io/post/clickhouse-integration)). However, I can still look at the source code of ClickHouse, because most of what they run in the cloud version is free and open. And we have relied on seeing the source code to help us debug things and to implement some core features of our system that are a bit, shall we say, off the ClickHouse beaten path (more on that some other time\!). And if push comes to shove I know that we could run ClickHouse ourselves. So to me, using ClickHouse Cloud is not a risk.

The other reasons are probably going to be obvious at this point:

1. Modern AI tooling has changed the equation on what can be built in a given period of time, especially in the right hands.
2. It is my nature to ask: is there a better way to do this? What could we build if we did not have to implement the Kafka API and could fully control both producers and consumers? I have a lot of curiosity about what is possible when taking a fresh look at a problem.

The honest reality is that I have always wanted to build a system like this and have never had the opportunity. I’ve worked on systems my entire career: operating systems, hypervisors, networking, cloud infrastructure, etc. But I’ve never worked for real on durable storage systems. And this is a big regret for me. I find them very interesting.

Pre-AI I never would have been able to justify spending enough time on this in my day job to make this system work, and I don’t think I would have had the energy to get the system into a working state in my “spare time.” But we no longer live in the pre-AI world, so blob-stream started this past January as an AI assisted side project on a weekend where I asked: what could I build without constraints?

# How did I build blob-stream?

I haven’t written anything publicly about how my engineering practice has evolved in the new AI-assisted world we live in. I’ve thought about doing it many times, but every time it seems like I would have very little to offer beyond a “me too” that I’ve barely written a line of code by hand in a year. Part of the inertia issue is also the reality that I know that anything I write is going to be out of date in a matter of months. So with that said, take this section for what it is, a historical anecdote relevant to the time it was written.

Before I provide a broad overview of my creation process, I want to point out that the [blob-stream commit history](https://github.com/bitdriftlabs/blob-stream/commits/main/) has **not** been squashed. I am providing the full history from the very first commit. When I released Envoy I was petrified of what the world would think and it seemed standard practice to squash code before releasing. 10+ years later I’m older, more confident in my standing as an engineer, and I frankly don’t care if you all see all my mistakes along the way. Perhaps something can be learned from the history. My only regret is that I did not think to somehow capture the model and chat history that produced each PR and commit. I’m quite sure there are startups working on this now and in the future the model back and forth will be part of the historical record.

## The planning process

Last January, the state of the art models readily available to me were Gemini 3.0 Pro, Opus 4.5, and GPT 5.2. I began by writing down a broad overview of what I wanted to accomplish. I then used a combination of all three models to begin exploring more details of how the system might work in aggregate. From the start, I knew the broad strokes of what I wanted:

1. No control plane (DynamoDB/S3 only)
2. Stateless brokers
3. A producer/consumer library that I controlled

Kafka compatibility was never a goal, though the system had to look enough like Kafka that I could swap it in at the application layer without major changes.

During the planning process I combined cyclical iteration using each model with external research. Meaning, I would expand more detail using one model, research its output, refine manually, and then have the next model do a review of the previous model's output. This process led me to landing on a high-low sequence allocator, Snowflake IDs as a core part of the scanning algorithm, and most of the other bones of what the system contains today. I don’t remember how long I spent on this iteration and research process, but probably no more than a few hours total. The result of this is the PLAN.md file in the [initial commit](https://github.com/bitdriftlabs/blob-stream/blob/64eaab857d5311b44a90abd45e6f6222429e7561/PLAN.md).

## Initial code output

Over a period of 2 days, January 18th and 19th, I used some combination of these models (unfortunately I don’t remember which) to output the first large batch of code. I then dropped the project until February 28th where I likely had a spare weekend day to work on it. Only a month later the world had moved on to Gemini 3.1 Pro, Opus 4.6, and GPT 5.3 Codex. I used 5.3 Codex to generate the next large batch of code. The weekend after that I again used 5.3 Codex to output the remainder of the plan which was committed on March 7th. At this point the main system was “code complete” including relatively thorough test coverage including integration and fault tests.

I should point out that during this initial code generation period I was only reviewing the outputted code in a cursory manner. Part of this was because at the time I considered this effort to be a “for fun” side project. The other reason is that I simply was not experienced yet with how to properly review and iterate on agent generated code. I will touch more on my current review and iteration process below.

## Actual deployment

Though I had worked on generating the code in my spare time, during this entire period it was difficult to justify spending any real bitdrift time on deploying it. Generating code is one thing, getting it deployed and working is another thing entirely. So again the project was dropped for a long period.

Fast forward to July, and I had a week lull before a company onsite, and I decided to see how much I could get working for real in one week.

By this past July I had been using AI assisted code production for almost 100% of my code output for some time. I had also largely unified my model usage to the GPT 5.3 Codex lineage of models. Meaning, 5.3 Codex (High) \-\> GPT 5.4 (High) \-\> GPT 5.6 Terra (High). I skipped GPT 5.5. For my own personal usage, I have found the OpenAI models in this line to be superior to Anthropic’s models by large measure. I find the price for performance to be excellent. With the right guidance the code output is very high quality and for whatever reason I find the models in this line to not be prone to “over thinking.” They just get things done very fast and mostly very well. I have dabbled in the pricier models, but I have never found a substantial improvement in performance given my working style and the type of things I work on (maybe that is my own fault but that is a topic for a different discussion and blog post).

So I settled in and developed a rollout plan:

1. As I mentioned above, historically at bitdrift we never rebalanced Kafka clusters. Too annoying. We had already built a mechanism to dual write and dual read from clusters at the application level so we could gracefully cutover to a new cluster, drain the old cluster, and delete it. I expanded this mechanism so that blob-stream producers and consumers could sit behind the same Kafka production and consumption traits we used everywhere. This allowed blob-stream to theoretically be a drop-in replacement for Kafka in our application code.
2. I then deployed blob-stream to our staging environment and attempted a cutover.

Mind you, with AI assistance I did the previous 2 items in 1-2 days. The new world is pretty awesome isn’t it?

Now, reader, what do you think happened when I actually tried to push data through blob-stream in our staging environment? You would be correct in assuming that it didn’t work at all\! 😂

## The real work starts

My general experience with AI coding tools is that their ability to engage in high level “systems thinking” is minimal to nonexistent. Meaning, while they are exceptionally good at iterating and completing well scoped tasks, as of this writing, even with the most advanced and expensive models, I have not seen any evidence that they can look at the “big picture” and iterate independently correctly against that picture to produce a low level system that is correct, highly performant, and easy to maintain and expand. Perhaps this is a skill deficit on my part and I’m sure plenty of folks will read this post and tell me so. But I also think it’s just the nature of how the models are trained and how they work. There is no complete blob-stream out there (that I know of) that the models would have been trained on that might have led to better initial code output. There are a lot of pieces of blob-stream however that independently the models are well trained on, so these pieces in isolation can be executed quickly and at high quality, when under the control of an expert human.

As I said above, what the models *are* good at is hammering relentlessly at a specific task, no matter how tedious, especially when given very specific structured guidance and very specific goal states. As has been written elsewhere, in the hands of a very experienced engineer, I believe this does lead to anywhere from 2x-10x productivity gains without any long term decrease in quality.

I won’t recount every bug I fixed, but starting at the beginning of July I fixed many, many bugs. These bugs ranged from glaring performance issues (so obvious I would have never written the code that way in the first place), to functionality/correctness issues, to observability gaps, and so on. Broadly my process looked like:

1. Old-fashioned code review of large pieces of code looking for glaring issues, and then instructing the agent to make broad structural changes.
2. Instruct the agent to add in-depth observability to various portions of the system (OTLP tracing, metrics, logging, admin state dump endpoints, and accompanying dashboards).
3. Instruct the agent to relentlessly expand integration and fault test coverage, and ensure the tests are deterministic. You can see the results of this in the history of [TEST\_AUDIT.md](https://github.com/bitdriftlabs/blob-stream/blob/main/plans/TEST_AUDIT.md) which is still in the repo.
4. Deploy the code to our staging environment and do what I do best: look at charts and see if anything looks odd based on my understanding of how the system should work, especially in the face of manual chaos interventions such as repeatedly restarting brokers, crashing brokers, restarting consumers, crashing consumers, etc.
5. Repeat.

From a token/time perspective, (1) and (2) are trivial and I was able to rapidly iterate on improvements and deploy them in staging. Many times per day. (3) is a much more expensive operation in both tokens and time, however this is a task the agent can churn on in the background while I was working on other things. And while expensive, this is a task that the models are *very* good at, and the result has an extremely high benefit-cost ratio.

So effectively in parallel I was both having the agent improve test coverage and hunt for edge cases not handled correctly while at the same time deploying the code in production and using iterative observability improvements to feed back into the agent to fix further bugs. Within about 7 days of work, by the middle of July, blob-stream was fully deployed in our staging environment and Kafka was removed.

## Code review process

I want to briefly describe my current code review process for agent assisted code since there is so much dialogue about this topic these days. For every agent assisted change I do, I broadly:

1. First plan the change, and iterate on the plan over multiple rounds until I am satisfied that there is sufficient clarity that the agent won’t get lost and off the rails during implementation.
2. Then generate the code and accompanying tests.
3. Then manually review the change, looking for “code smells.” Possibly from a long career, and a lot of open source work, I’m able to do this relatively quickly.
4. Then go through zero or more additional rounds of generation and manual review cycles.
5. When I am happy I have the agent do its own adversarial review with a prompt along the lines of: “review the entire change from scratch against origin/main, focusing on correctness, performance, and missing high value test coverage.”
6. Then go through zero or more rounds of (2)-(5) depending on the results and change complexity.
7. Finally I have GitHub code review (which I find to be very high quality) do a final review and iterate again on comments.

Obviously how many rounds I go through depends on the complexity of the change, but for some of the more gnarly consumer scanning code in blob-stream the process was very lengthy over many, many iterations.

## Deploying to production

Anyone that has worked with me knows that I am a big believer in testing in production. You can theorize all day, and write unlimited synthetic test cases, but production is where the proverbial shit hits the fan. I honestly don’t know how anyone builds anything of substance where it cannot be tested in production, and it’s the main reason that I have always enjoyed working as both an engineer and an operator. The feedback cycles are so tight and satisfying.

At the same time, I am not *completely* insane. One cannot just cutover Kafka to blob-stream in a customer cluster and hope for the best. Starting in mid-July I prepare to cut over a small production cluster. Multiple times in this post I have mentioned that bitdrift has long had the ability to cut over Kafka clusters and dual read while draining the old cluster. In preparation for cut over, I enhanced this functionality to truly dual-write one of our high volume topics to both Kafka and blob-stream. Instead of the consumer processing both inputs for real, I changed the behavior so the blob-stream data was written to parallel database tables that could be explicitly compared with the Kafka data looking for lost or corrupted data.

Controlling the app layer fully made certain types of production bugs much easier to spot. For example, I was not particularly worried about corrupted records because I would have a high likelihood of knowing right away if that was happening. My biggest worry with blob-stream was (and is) disappearing records completely due to the time synchronized consumer scanning protocol in use. Dual writing to parallel tables and comparing to the main tables gave me substantial confidence that at least in the common case no records were being lost.

By the beginning of August I was confident enough in the system that I cut over one of our smaller production clusters to blob-stream. After cutover I continued to perform regular manual chaos monkey experiments on the small production cluster to build my confidence in edge case handling.

## Larger clusters and long tail bugs

After allowing blob-stream to bake for a few weeks in the smaller production cluster, I started to plan for deployment in larger clusters. At this point it became clear that some of the cost improvements I had been planning for “the future” were going to have to get done before I could roll out further. These included:

1. Routing DynamoDB metadata reads and S3 blob reads through brokers to collapse consumer API calls.
2. Packing all topic virtual partition data into the same segment in order to reduce S3 PUTs.

Agent assisted, neither of these took very much time.

Over a period of a couple of weeks in late August I gradually started a cutover in one of our much larger production clusters. During this period I continued to improve performance, fix small bugs, and clean-up code.

By the beginning of September I had blob-stream at 100% (with the Kafka cluster on hot standby in case of emergency). My constant fear during this process was that I was disappearing customer records in edge cases due to bugs in the consumer metadata scanning engine. Mind you, this is a cluster regularly doing millions of records per second so timing based edge cases would on average happen relatively often.

I realized that while blob-stream offsets do have expected gaps when brokers restart, they do *not* have gaps at steady state. In short order I built some forensic tooling:

1. A consumer flight recorder that would efficiently keep track of scanning state in memory, and in the case of an offset gap, dump a large amount of forensic data.
2. A metadata inspector tool that allowed me (and my robot helper) to easily check the state of DynamoDB around a particular suspect offset.

Once deployed, lo and behold, across 2 different clusters in aggregate running at millions of records per second, I found that every 6-12 hours I was indeed disappearing 500ms-1s of data. However, the dumped forensic data made it easy for the agent to track down two different correctness bugs \[[1](https://github.com/bitdriftlabs/blob-stream/pull/104), [2](https://github.com/bitdriftlabs/blob-stream/pull/106)\] and both bugs have subsequently been fixed. Since then, the system has been operating at millions of records per second with no loss (that I know of).

As an aside, it should theoretically be possible for blob-stream record offsets to be contiguous in all cases other than broker crash, by having the broker give back any unused sequence block allocations on lease release. I tried to [implement this](https://github.com/bitdriftlabs/blob-stream/pull/96), but it went horribly wrong and I had to revert it. I may come back to it eventually, but given how the system works consumers must be able to deal with gaps, so it’s nice only from a monitoring perspective and I can’t justify the risk right now.

## Deployment summary

Since the successful migration of the large cluster, we have since migrated the rest of our clusters to blob-stream and no longer use Kafka anywhere at bitdrift. In total, I would estimate that it took about 2 months of calendar time to develop a full Kafka replacement tailored for bitdrift’s use case. This system is approximately 10x cheaper on a dollar for dollar basis and monumentally simpler to operate. I should add that quite a bit of this time was baking time. Myself and my robot helpers worked on many other things in parallel.

I do not feel qualified to offer any thoughts on what the future holds for computer programmers in the age of AI. I have significant concerns that we are not training the next generation properly and it’s extremely optimistic that computer intelligence is going to make up for that lost training as the rest of us retire in the coming years. At the same time, it’s clear that putting even today’s tools (let alone whatever is available in a few months) in the hands of expert programmers has the ability to create superhuman output. I have no doubt that I could have created blob-stream on my own without assistance, but there is absolutely no way I would have done it this fast.

# What comes next for blob-stream?

I quite honestly don’t know. Blob-stream is not a product. It doesn’t support the Kafka API or a bunch of Kafka-like features that people probably need like authn/authz, schema checking, and whatever else. I’m not even providing docker containers, the Rust libraries are not currently available on [crates.io](http://crates.io), and there are no libraries for other languages. This is a one person operation currently and I don’t have the resources or time to do all these things.

And in the new world we live in it’s not clear to me what open source even means anymore or if it has the value it once did. At bitdrift we fork software with abandon and in many cases no longer think that carrying patches or a fork is a burden. I’m sure many others feel the same way. AI assistance has fundamentally changed so much of how we built software previously and it’s hard to know how things are going to shake out over the next few years. I find this topic both ironic and mystifying given that the models we use today are trained in large part on the open source software that came before. 🤔

At the same time, I still believe that for certain types of software, open source *does* have value. I think it’s very likely that other people around the world will find blob-stream valuable in some form, even if it doesn’t yet have every bell and whistle, or support every blob and metadata store across every cloud provider. I think there is value in sharing this type of software and working together to improve its performance and correctness. And ultimately I think that there must be an open source blob streaming solution so that the entire industry is not beholden to the whims of individual corporations for critical infrastructure software that must be run inside our own environments.

So will anyone else use [blob-stream](https://github.com/bitdriftlabs/blob-stream)? I don’t know. Will people aim to contribute or fork? I don’t know. Will I have time to accept contributions? I don’t know. But it will be fun to find out.

---

## Frequently asked questions

### Is blob-stream a drop-in Kafka alternative or replacement?

blob-stream is an open source streaming system that works much like Apache Kafka: topics, partitions, consumer groups, commits and seeks. But it doesn't implement the Kafka API or work with existing Kafka client libraries. Producers and consumers use blob-stream's Rust libraries.

bitdrift moved off Kafka by putting blob-stream behind the same producer and consumer interfaces its application code already used, so no major application changes were needed. blob-stream is a strong Kafka alternative for high-throughput AWS workloads that can tolerate a bit higher end-to-end latency. It doesn't yet support global per-key ordering, authentication and authorization, schema checking, or clients in languages other than Rust.

### How does blob-stream’s total cost of ownership (TCO) compare to Kafka?

In bitdrift's largest production cluster, blob-stream handles about 6 million records per second at peak (about 11 Gbps uncompressed). That runs on 6 small brokers and costs about \$2,100 per month at AWS list prices, covering S3 storage, S3 API calls, DynamoDB and broker compute. The total cost of ownership is roughly 10x cheaper than bitdrift's previous Amazon MSK deployment, which makes blob-stream one of the lowest-cost Kafka alternatives for high-volume streaming. The vastly simpler operational complexity is not factored into the 10x cost reduction.

Most of the savings come from 4 places:

* **No cross-AZ traffic:** Producers and consumers only talk to brokers in their own availability zone, so there are no cross-AZ networking charges.
* **Stateless brokers:** Brokers have no local disks, so there's no storage to provision, rebalance or overprovision for peak load.
* **Cached, batched reads:** Metadata and blob reads go through brokers, which cuts DynamoDB and S3 API calls.
* **Shared S3 objects:** Data for all topics is packed into the same S3 objects, so write costs grow with throughput rather than with the number of topics and partitions.

### How is blob-stream different from WarpStream and other diskless Kafka alternatives?

Like WarpStream, blob-stream writes streaming data directly to object storage (Amazon S3) instead of broker disks. There are a few main differences between WarpStream and blob-stream:

* blob-stream is fully open source, so teams can read and change the code that runs in their clusters.
* blob-stream doesn’t need a separate control plane: brokers, producers, and consumers coordinate with standard Kubernetes service discovery and DynamoDB, and nothing else.
* blob-stream doesn't have to stay compatible with the Kafka API, and its client libraries route requests to the broker most likely to have the data cached.
* blob-stream also uses tightly synchronized clocks (the AWS Time Sync Service) and time-ordered Snowflake IDs, so segment metadata is written without transactions and consumers can scan it cheaply.
