yesterday at 4:20 PM
Object store is quickly becoming the new core data substrate. Lets build kafka, but on s3. Lets build github, but on s3. It feels like we going to see more and more "object-store first" systems in the next few years.
I am excited about this future. Give me stateless servers and a storage bucket over having to manage systems with disks any day.
I do wonder if we will see an expansion of the s3 api to support more of these use cases. S3 added a janky file append operation to their new express-one-zone bucket type, and limited to 10k total file append operations. I wonder what else we will get in the next few years.
yesterday at 7:05 PM
It makes perfect sense: it's a extremely reliable, infinitely-scalable, strongly consistent (in some implementation), cheap, and extremely simple to use key-value store. As such, it transparently solves a lot of the problems distributed systems have to be engineered around.
As long as you can run on a CSP, and can engineer around the high-ish latency (most business cases can), it's extremely expensive to try engineer around it.
yesterday at 4:31 PM
Definitely agree that every data system that doesn't need <100ms latency is moving to object storage.
> I do wonder if we will see an expansion of the s3 api to support more of these use cases
This is actually an area where I think we have a big leg up on folks building on top of S3. My team (which built K2) sits next to the R2 team, and we have the opportunity to co-evolve the products in mutually beneficial ways.
yesterday at 8:24 PM
> every data system that doesn't need <100ms latency is moving to object storage
I think the opportunity extends below 100ms too, particularly given the existence of faster object storage tiers like S3 express or more recently GCS rapid bucket (both only offering single-zone durability, so still need to do quorum writes to get region-level durability as with standard tiers).
One of the tensions of course is how long to linger before flushing to object storage - you have to trade off directly between latency and cost of your API ops for PUTs.
When building the serverless offering of s2.dev (which is in a similar space, full disclosure!), we designed around stateful backend processes capable of constantly flushing multi-tenant objects (i.e., containing records from many streams), allowing streams to offer low ack latencies (~50ms p99 from same region) without blowing up the unit economics.
Congrats on the launch btw!
yesterday at 9:57 PM
From my experience benchmarking S3 express, it is is faster, but not fast enough yet for many use cases.
An obvious disclaimer is that the word "enough" here is carrying quite the weight: I expect it to get better, and each has their own requirements. Do benchmark yourself and don't make expensive decision based on an HN comment.
yesterday at 6:03 PM
I still think S3 is underutilized.
The range of things you can do with blob storage and a (very simple) auth model are surprisingly broad.
We recently replaced our Docker container registry with S3 using a tiny tool [1] we built in-house. I think that even with current capabilities, we can still model a lot services as a very thin layer over object storage.
[1]: https://github.com/Simple-Observability/grue
yesterday at 6:11 PM
From the opposite side of things, I feel the same about OPFS. Finally having something that performant that multiple workers can operate against in a browser is such a massive boon. You could plug something like that in to it locally with markedly less bullshit than you would have had to do previously with the other APIs.
yesterday at 6:50 PM
I'm a little surprised the Docker CLI doesn't support this directly, since gcr.io works this way. Maybe there's a thin layer between the bucket and the CLI? Neat that you worked it out for the generic case.
yesterday at 7:29 PM
I was surprised too! The OCI layout for storing images is actually pretty simple. But for some weird reason you can't stream it straight into Docker. docker load only accepts tarballs.
So you need something to wrap the layout into something Docker understands. Because S3 is not a server, you have to construct the tarball on the fly as the image is pulled and stream it straight into docker load.
yesterday at 10:23 PM
I worked on OCI for quite a while, the very short and cynical answer is that Docker thought the registry was their moat for a long time and fought tooth and nail to keep it at the detriment of almost everything else. (The longer version is a bit more diplomatic.)
Even better, the OCI distribution protocol (which was Docker distribution until a few years ago) is not a static-blob-over-HTTP protocol! The blob bits are and you can route them to blob storage but you need a smart server for a few key bits of the protocol.
Back in the day I made proposals for distribution formats that didn't have these flaws and were more distributed (and previous proposals like AppC's discovery had similar ideas) but they were roundly rejected by the Docker people.
yesterday at 6:28 PM
That sounds really cool — I tend to agree with you that S3 and similar are underutilized, but I remember that essentially all of the providers charge for bandwidth measured in gigabytes and I'm like no. My consumer line is measured in megabits per second and if I have to pay for my data usage the way it's paid for in data centers it would be far more expensive. Somehow consumer ISPs, who have to pay for the lines, are cheaper than than the cloud providers.
yesterday at 6:51 PM
"I remember that essentially all of the providers charge for bandwidth measured in gigabytes" - sure, but the cost per GB is shockingly cheap. If you're doing things well, you can do a lot inside those pricing structures.
Alternatively you can stand up your own object store services, but that's not something I would like to do.
yesterday at 8:03 PM
> the cost per GB is shockingly cheap
Huh? We're talking about the same S3 right? At list pricing, 1TB is $23/TB to store for one month, and about $90/TB (plus request fees) to send it out to the internet.
While hard drive prices are roughly 3x what they were a year ago, the price per TB of a new hard disk averages around $30/TB -- assume 2x overhead for other hardware and extra space for parity etc, a disk pays for itself in less than 3 months and lasts 5 years or more.
If you assume it takes 1 month to download that 1TB (about 3Mb/s) that's $29.16/Mbps. When I first started buying internet transit in Europe back in 2008 I think I was paying under $10/month. It's now under $0.10/Mbps pretty much anywhere in the US or Europe at the big datacenters.
None of the "big" object storage services are cheap. They're somewhat reasonable if you only access the data from within the same region, but absolutely insane if you ever want to ship that data outside of that cloud vendor (or to another region etc). The pricing of storage and egress has not changed in a decade (I believe AWS last lowered the price of either in 2016) and in fact it costs even more now due to things like NAT Gateways etc.
It's definitely not "shockingly cheap". It's just cheaper than $80/TB of gp3 or $45/TB of st1, and while sc1 is $15/TB it has a baseline performance of only 12 MB/s. There's quite a few companies out there that have $6/TB/month object storage plans with similar performance, rising to about $15-$18/TB/month for SSD backed storage with far lower latency figures.
yesterday at 8:19 PM
AWS have found the ideal captive market: computer engineers who don't know how computers work or how much they should cost.
Whatever AWS is selling you - except Deep Archive - I'll figure out a way to sell you for half that price, if you want, and it'll still be 80% profit for me. Your only downside will be that I don't know what I'm doing so it might not be reliable - but neither is AWS.
yesterday at 8:26 PM
Not just AWS. IIRC, GCS and ABS are slightly cheaper for storage but more for egress. To be fair to them, their reliability is pretty much second-to-none, whereas many of those at the very cheap end pretty much just ran a Ceph/RADOS cluster and then wondered why they were having lots of problems with reliability -- but with things like RustFS becoming more and more popular, the object storage racket is long overdue a shake-up!
yesterday at 7:41 PM
It's because for providers the bottleneck isn't capacity (which is dirt cheap) but IOPS and bandwidth. To keep transfers fast they need to spread data across a lot of physical disks, meaning they often times sit mostly empty. Egress fees are a way to monetize that empty storage and make sure there is an incentive to minimize egress traffic.
We use R2 (not affiliated) which doesn't have egress fees.
yesterday at 6:47 PM
> My consumer line is measured in megabits per second and if I have to pay for my data usage the way it's paid for in data centers it would be far more expensive
Well because your provider assumes you are not using all your bandwidth constantly. Cloud bandwidth is only billed for you actually use
yesterday at 8:21 PM
It's still true that "clouds" cost much more than proper internet connections at DCs. E.g. AWS wants you to pay $90/TB, Hetzner $1.50/TB, a good deal on a contract is probably half what Hetzner pays since they need profit too.
And if you have two specific endpoints you need to transfer data between at a high rate, you can get stupidly cheap cost per GB on a leased line in exchange for making all that commitment upfront.
yesterday at 8:01 PM
Depending on your use case, exposing S3 data with CloudFront will decrease the transport cost.
yesterday at 8:16 PM
How will you prevent this system from becoming a server with extra steps?
yesterday at 7:27 PM
Came across this the other day which lets you run a etcd compatible API/Kubernetes on top of S3
Haven't tried it yet but looks nice for simpler K8s deployments
yesterday at 4:59 PM
doesn't that make egress fees egregious? or still cheaper than disks?
yesterday at 5:30 PM
You can download objects from S3 from EC2 without traversing the public internet using things like gateway endpoints [0] which avoids s3 egress fees. But doesn't avoid egress fees from EC2 to the end user.
[0] https://docs.aws.amazon.com/vpc/latest/privatelink/vpc-endpo...
yesterday at 5:10 PM
This all depends on the cloud you are building on and if the data ever leaves the datacenter.
yesterday at 8:22 PM
they are indeed egregious, like, downloading all your data one time costs about three times as much as buying disks to store it.
yesterday at 4:28 PM
People like S3 because they have hard engineering guarantees around bit rot and work well as a high level abstraction of a network filesystem with all of the low level failure recovery built-in. You don't have to worry about doing your own RAID configs. The bigger question is whether non-AWS services can offer the same level of guarantees. I have heard horror stories for example when it comes to downtime on Hetzner's S3 object store.
I am currently using Cloudflare R2 right now and if you see their forums, there's always the occasional post about objects going missing.
yesterday at 8:24 PM
I find that Ceph is pretty annoying to operate but if you do, it works fine. Keeps redundant copies of data on multiple disks, across multiple racks if your diversity is that wide. Supports erasure coding for effective redundancy less than 2x.
Don't go for the "object gateway" compatibility layer - just use raw Ceph if you're writing your own app.
yesterday at 5:50 PM
Meanwhile, just using disks and regular servers is still reliable and more so than ever, especially with some redundancy.
Consider your cloud costs.
yesterday at 7:44 PM
If you there's no SLAs to engineer for, just set the target at zero, and drive your cost to zero as well. So you have to design for _some_ defined value of reliability.
Backblaze's famous reports put HD failure rate is ~1.39%, so for the 11 9s you get as a guarantee from S3. Assuming nothing else fails, you'd need at least 6 independent copies to get that, plus all the effort to engineer recovery, and constant upkeep.
Suddenly, S3, even when considering bandwidth costs, seems like a steal.
yesterday at 8:27 PM
The usual SLA is not expressed numerically but as a feeling. And a normal server suffices to deliver that feeling. I'd bet 99% of apps are used by under a hundred people and if they have to take a day off due to a head crash, it's annoying but not catastrophic.
If you have two sets of hardware the server can run on (cold standby), and RAID, and are competent at physical IT work, you can have faulty hardware replaced in an hour. Drive fails - replace it. Anything else fails - swap the drives to the other machine, boot it up and then troubleshoot the original.
Most likely you don't even need that. If the app server runs on a standard platform like Windows you can shuffle it over to some spare tower PC. You hear horror stories of dusty towers that nobody knows what they're doing - the horror there isn't from running server software on a tower, but from unmaintained servers no matter the form factor
today at 1:30 AM
The pricing is surprising, Data Produced at $0.04/GB seems reasonable (when compared to other clouds event streams) but Data Consumed being at the same $0.04/GB is rather steep...
This means actual usage is $0.08/GB in the simplest case (one consumer) but fan-out consumer strategies get very expensive very fast.
yesterday at 3:59 PM
I'm the author of the post and tech lead for K2. Happy to answer any questions!
yesterday at 7:25 PM
I'm doing some testing, and while my writes are <1s, I'm seeing E2E delivery latency of p95=2.5s / p99=7.5s. Is this roughly expected?
yesterday at 6:09 PM
After reading the post, still don't fully understand when you would use CF Queues vs K2. Can you help to elaborate a bit more?
yesterday at 11:03 PM
My understanding: Distributed queues are generally good for when you have multiple workers processing chunks of work. Each queue item usually needs to be processed to completion exactly one time, so the queue provides the mechanism for the workers to coordinate state of each item at the item level (which allows for time-outs and re-tries if, say, a worker dies during processing, like if you were using spot instances for your worker pool)
Event streams are for multiple consumers, and can offer different guarantees. As far as I can tell, K2 is designed to ensure all consumers receive all events at least once (it's unclear to me whether this means they're continuously storing all events from the stream origin, or if older events age out at some point, or are dropped when they've been consumed by all known consumers).
Other types of guarantees with streams might be "at-least-once", "at-most-once", and "exactly-once" delivery, for different needs. Redis streams used to be at-least-once but it looks like they support all 3 use cases now. Some relational DBMSes also have the option to replicate by streaming their transaction logs to all servers in the cluster so each node maintains its own understanding of the database state (though stale reads can also occur in some/all? DBMSes that replicate this way, when a server is queried before receiving an update)
yesterday at 7:45 PM
Sure! There are definitely some overlapping use cases, and we've seen folks using/abusing queues for use cases that are more appropriate to something like k2.
Queues are great when you have a unit of work that needs to be completed, retried, and tracked individually. For example, a shop might need to call a payment processor API that can fail or timeout, and retry it until it succeeds, while polling on the frontend for the state of that particular message. With a queue, you can insert a message tracking that payment, and have a queue processor that keeps getting sent it until it succeeds or has failed too many times.
In a queue each item is its own thing that's important to someone, and queues give you APIs to interact with that particular item.
K2 is for moving large volumes of data around. Pricing is per GB, not per message. Records are produced and consumed in bulk, and what matters is that all records are processed, but no one is querying the state of a particular record. K2 also supports multiple consumers for the same record, and long term retention. For example, all of your applications may emit events when things happen, and those events need to be read by an alerting system, a system that durably stores them, and a system that uses them to build ML features.
yesterday at 9:32 PM
One thing thing I am curious about. You mentioned polling for the state of that particular message (CF queue world). Is that really possible? I would have assumed it requires the developer to track the message using D1
yesterday at 8:52 PM
Queues are for actions (do this), K2/event streams are for events (this happened).
yesterday at 11:35 PM
How does this compare to s2.dev?
yesterday at 9:31 PM
Have I understood the pricing correctly? There's charges for data published and consumed per GB?
yesterday at 4:27 PM
Great product, congratulation on launch. Are you using this internally in any way?
yesterday at 4:35 PM
Yes! We originally built K2 to serve as the ingestion layer for Basin Pipelines [0], our stream processing product. We have a number of other teams building new products on top of it at the moment which I can't talk about yet :)
[0] https://developers.cloudflare.com/basin-pipelines/
yesterday at 5:35 PM
Congrats on the launch! Stream/event-based system are really powerful, but they are also just pretty complex, in no small part because the modeling of streams for most people today is really modeling Kafka topic/partitions which has a whole bunch of foot-guns and complexity. Making the individual stream really cheap and easy is a big simplification, especially if flexibly consuming a stream for both ordered and unordered use-cases is made simple. This looks to work for unordered, be curious to see how the managed dividing the work for the mentioned key-based ordering.
yesterday at 8:30 PM
Kafka works the way it does for a reason. I don't see how Cloudflare can avoid the same fundamental limits that Kafka faces.
yesterday at 8:25 PM
To support the ordered use-case, put events on the same topic-partition. To support the unordered use-case, don't.
yesterday at 7:38 PM
I wonder how many of data infra startups are wrapper on top of S3.
Also, the boundary between OLTP and OLAP is blurring every day.
For folks who want an off shelf version of this you may be interested in https://github.com/viggy28/streambed
(Disclaimer: I am one of the committers)
yesterday at 8:28 PM
Hasn't it always been blurred? SQL supports OLAP and OLTP queries on the same database in the same language. They're different workload types in the way that "servers" and "GUIs" are two different types of software.
yesterday at 8:03 PM
LinkTree is a billion dollar business and it's basically a key value store.
yesterday at 3:53 PM
If I was a serious Cloudflare customer I would be seriously concerned about the security of my infrastructure with them. Yes LLMs can code fast but this is an almost frenetic pace of releasing new products, all with fewer staff.
yesterday at 5:22 PM
They've always been this fast. The trick is that they release beta products a lot. Kind of classic "lean" (does anyone remember that?)
My bigger concern would be their increasing grip on a lot of the market, and eventually becoming a monopoly (or part of the big tech "duopoly")
yesterday at 6:00 PM
> eventually becoming a monopoly (or part of the big tech "duopoly")
At this point this has already effectively happened. The average person doesn't realize the extent of it, because the products Cloudflare builds are inherently transparent to the average consumer.
yesterday at 7:04 PM
Not when you use firefox!
yesterday at 4:03 PM
They are working on a set of primitives that make building this kind of software easier. With AI, runtimes matter more than ever and languages matter less.
yesterday at 5:36 PM
I trust that this is not the case. Check out both the authors. Micah's company was acquired by Cloudflare, and both of these guys actually bring good expertise around this area.
The question is less about "vibe coding" the product, and more about how their acquired teams function. For most products they have released so far, they are backed by a company they acqui-hired. They will usually rebrand the product and absorb the team.
You would not be surprised if this release was from a product for a startup company rather than a large enterprise company.
Cloudflare hired a bunch of folks for sure, but they also fired a lot of folks. What they are doing these days is building products by buying out entire companies, giving them nearly independent authority to build a product like they would build a company.
yesterday at 4:08 PM
I find this release relieving because it fills a hole that was really needed. Another hole would be a proper database and improvements to D1.
yesterday at 4:08 PM
That's just perception. At the same time they were hiring 1 k. People and hired 2 k. People the year before ( interns).
There was just more news about it than with other companies ( I think they let go about 2 k. People a year ago).
( Not saying it's good, just a little perception balance)
yesterday at 7:07 PM
It hasn’t even been 6 months since they fired 20%
yesterday at 9:59 PM
Rather than requiring consumers to ack the batch, why not just have them submit the ID of the batch tail on `consume` requests?
yesterday at 10:11 PM
This is a great suggestion, it has the same semantics but avoids the extra round trip. We’ll add this.
yesterday at 9:38 PM
Cloudflare is catching up to be a full AWS/GCP/Azure. At start it was just a few services around but they are adding everything else
yesterday at 7:58 PM
Has anyone built anything on top of D1 + Durable objects for multi tenancy? How has been the experience?
yesterday at 5:16 PM
How does this solution differ from AutoMQ and WarpStream? I’ve worked with one of them, and it is indeed a serverless solution built on top of S3.
As far as I know, Kafka itself already supports offloading some data to S3 for long-term storage.
Based on the articles—which I didn't fully grasp—I’m wondering if there are additional benefits mentioned, such as multi-region distribution (though I find it hard to imagine how that would be implemented).
yesterday at 8:00 PM
It’s built for Cloudflare ecosystem, IMO using K2 on its own doesn’t make sense. I’ve worked with a startup that uses Cloudflare for everything except for Kafka. If K2 were to exist back then, I’m pretty sure they would have used it.
yesterday at 8:40 PM
Kafka also has inkless with diskless topics that run off object storage!
yesterday at 6:45 PM
Congratulations on the launch! Looks very interesting indeed!
As someone who enjoys writing Kafka streams applications I am also looking forward to the day you support the Kafka APIs.
Having a cost efficient fully serverless Kafka compatible service would be great, and something I think many businesses would find useful.
Great work!
yesterday at 7:31 PM
i think nats jetstream is superoverlooked in this space. the amount of throughput is more then enough for any medium scale enterprise as long as you know what you are doing
today at 1:11 AM
NATS is my goto message broker for most projects.
yesterday at 6:37 PM
Not a single mention of AWS Kinesis? That seems like the closest competitor.
yesterday at 9:11 PM
This is specifically targeting edge. Kinesis is single region AFAIK, so I think Kinesis and Kafka are very similar and K2 is built on their S3 equivalent to make it globally distributed.
yesterday at 6:35 PM
Minor correction: it is not true that no major object stores support appends--Azure Blob Storage has from the beginning.
yesterday at 4:45 PM
This sounds like a super interesting product, but I'm always reluctant to build on anything that isn't a portable industry standard.
yesterday at 4:50 PM
Presumably the Kafka-compatible API that they say is in the works will address this?
yesterday at 4:54 PM
Yep. I'm not personally a huge fan of the kafka API — I think it's simultaneously too low level for normal users and too high level to deeply integrate into other systems (like stream processing engines), and requires a complex client library to use effectively.
We went with a simpler and more user friendly consume API, that also allows much higher levels of read parallelism (particularly important if you're using something like Workers, which parallelize well but aren't very powerful individually).
But we know many companies are invested in the Kafka ecosystem, and we want to provide an easy on (and if necessary, off) ramp for them.
yesterday at 10:37 PM
Kafka libraries outside Java (and maybe Go) are such a disaster. I remember getting paged because all my NodeJS Kafka consumers were crashlooping from a segfault in Confluent's official Javascript client (wrapper around rdkafka C++). Why? Because Confluent changed some sort of telemetry setting on the broker. I also get segfaults in Rust using the most popular client out there (also an rdkafka wrapper) without any `unsafe` (although there were threads).
yesterday at 4:07 PM
What's the benefits of this over traditional GKE pub/sub, kafka or another queuing service?
yesterday at 4:27 PM
Google PubSub is a great product. The primary benefit of K2 is cost, particularly for longer retention periods. Being backed by object storage means that we can store data extremely cheaply compared to disk backed solution, and we pass that on in our pricing.
Compared to self-hosted or cloud-hosted Kafka (e.g., Amazon MSK or Confluent), K2 is much cheaper, and fully serverless. There are no clusters to manage or scale, and consistent performance even as you vastly increase the amount of data.
The main downside is produce (and end to end) latency is higher (around 1s p99) than systems that rely on local disk replication, like Kafka.
So it's great if you're trying to move a huge amount of data around, or for use cases where cost is more important than latency.
yesterday at 6:34 PM
Having built a hobo version of something similar (serving a minimal subset of the Kafka API on top of CosmoDB): there’s also the hybrid scenario where the cheap serverless streams are used for scalability and broadcast while a low-latency core is maintained on sharply reduced compute resources. An 80/20 approach that saves a lot and, in our case, reduced cluster (mis)management risks at the same time.
In our case BLOB and large document transfers were handled in parallel, merging them together through object storage is a highly appealing package. Great work!
yesterday at 4:18 PM
Ha setup and management is costly. K2 relieves you off that by charging you 0.04/0.04/GB read/write and 0.02/GB/month storage. Pretty useful if your volume is not into multi-GB's a month. GKE pub/sub seems pretty costly by comparison.
today at 12:09 AM
Nice.
Our Monolog is similar in spirit.
Instead of building Monolog on object storage (R2, S3), we built it on our Dip, thus achieving extreme low latency and parallelism for ingestion and consumption.
Our novel architecture enables scaling to infinite consumers without upfront partitions (no magic, different tradeoff). This plays nicely with Slyp's data locality and application architecture.
Monolog is built on Rust, is lightweight, and runs on mobiles and servers alike.
The original reason for building Monolog was comically outlandish, Kafka was slow and required JVM, Redpanda was eating too much RAM - 2GB min which is absurd for our use case. Redpanda was also consuming so much CPU that the disgusting CPU fan noise had us feel emotional pain.
We would very much like to build object storage on Dip, but we are currently preoccupied and hence don't have any immediate plans to build one in the short term. Hence, it is R2 for now.
today at 1:27 AM
You can use redpanda with interrupts just fine (no fan noise). You don’t need to busy poll for low resource envs. Same for sys alloc
yesterday at 4:43 PM
i was trying to understand why i would use this over their current offerings of queues, and im really sick so my brain isn't working. so i ran it through ai
You need... | Use
-------------------------------------------------------|-------
“Make sure this job gets done” | Queue
retries / dead-letter handling | Queue
delayed jobs | Queue
distribute jobs among workers | Queue
“Record that this event happened” | K2
multiple independent systems reading the same events | K2
replay old events | K2
ordered event streams | K2
Kafka-like architecture | K2