The Manuscript Is Done. The Code Is Public.

In August I wrote that you should build the thing first and let the code tell you what the chapters say. That post was a bet made in the middle of the work. This one is the report.

The manuscript is finished: 452 pages, 18 chapters, 72 diagrams, 79 code listings. Not a draft I am still circling. Finished.

More useful to you: the reference implementation it was built on is now public. It lives at github.com/ThomasJaeger/event-sourcing-cqrs, MIT licensed, .NET 10 and C# 14, tagged v1.0.0, which is the release the book describes.

What is actually in the repository

It models an order-management domain across five bounded contexts, event-sourced end to end. Five aggregates rebuilt from events. Two process managers that are themselves event-sourced on their own streams, with compensation branches. Eight projections maintaining read models over a mix of relational tables and JSONB. Four hosts: a Blazor Server UI, a JSON API, a workers host carrying projections and the outbox, and an admin console. Role-based authorization and tenant isolation run through all of them.

Fifty-three architecture decision records carry the reasoning. If you only read one part of the repository, read those. They are where the arguments are, including the ones I lost.

It is licensed so you can lift pieces of it into commercial work without asking me.

The part that cost the most

Four event stores sit behind one interface as first-class peers: PostgreSQL hand-rolled over Npgsql, SQL Server hand-rolled over Microsoft.Data.SqlClient, KurrentDB over gRPC, and DynamoDB with conditional writes. All four pass the same contract suite. Switching between them is a configuration change.

That constraint was the most expensive decision in the project and the one I would make again. The moment a second store has to satisfy the same contract, every convenience that leaked out of the first one becomes visible. Optimistic concurrency stops being a paragraph and becomes four different mechanisms that have to produce the same guarantee. You cannot hand-wave a contract test suite.

It also answers the question I get most often, which is some version of “do I have to adopt a new database for this?” No. The domain model does not have to change when that answer changes.

Where the book stands

It is in submission. No publisher is named here, there is no publication date, and I am not going to invent one. When there is a contract to announce, I will announce it.

I am saying that plainly because the alternative is the thing I find tiresome in this corner of the industry: announcements engineered to imply more than they say. The manuscript being finished is a fact. The code being public is a fact. A publication date is not yet a fact.

What you can do with this today

Clone it. Run it. The README is honest about what docker compose up does and does not start, which is the kind of detail that usually costs a reader an afternoon.

Then tell me where it is wrong. Not where it is unfamiliar, where it is wrong. Open an issue. The book is finished, but the implementation continues on main past the tag, and a reference implementation that nobody argues with is not much of a reference.

One thing it will not do is convince you to event-source your whole system. A good chunk of the book is about where these patterns do not belong, and the honest answer for a lot of software is a boring table you update in place. Chapters that only sell the pattern are how teams end up with an event store full of StatusUpdated.


I run Legacy to Modern LLC, where I modernize MS-DOS, Visual Basic, Delphi and .NET applications. If your .NET application is on a current runtime but was designed as CRUD over tables, that is a specific problem with a specific path out.

The Delete Nobody Approved” / “How Much Does Your Database Forget?”

There is a moment in Greg Young’s Code on the Beach talk from 2014 that I keep coming back to. He asks the room how many people have an UPDATE or DELETE statement in their system. Every hand goes up. Then he asks how many of them sat down with the CEO and the board to talk about how that data had no value. The hands come down.

I have watched that talk dozens of times, and I find something new in it every time. This point deserves far more attention than it gets: most teams have no idea how much information their CRUD systems destroy, because the destruction is invisible by design.

Two carts.

Picture two shopping carts in an online store. Customer A adds three items and checks out. Customer B adds four items, winces at the total, removes one, and checks out. In a typical CRUD schema, both orders now look the same. Three line items, one address, one total. Byte for byte identical. Two different customers, two different stories, one row.

"Two shopping-cart event histories. Customer A: CartCreated, ItemAdded times three, CheckedOut. Customer B: CartCreated, ItemAdded times four, ItemRemoved, CheckedOut. The ItemRemoved event falls out of the flow to a crossed-out marker labeled no row, no trace. Both histories converge into the same database row with three line items. Callout: the difference was a future sale."
Two different histories, one identical row. The removed item is the information the DELETE destroyed.

Greg tells the second story on himself. When his cart total climbs too high, he removes an item or two before checkout. He still wants those items. He wants them next month, after the credit card has cooled off. A removed item and a never-added item are different business facts: deferred desire on one side, no desire on the other. The DELETE statement collapses both into the same fact. Nothing.

A table stores conclusions.

This is the mechanism, and it applies to every structural model no matter how carefully designed. Current state is a derived value. It is what remains after applying every insert, update, and delete the record ever saw. Store only the result and you have thrown away the inputs. Whenever two different histories produce the same result (added then removed, changed then changed back, applied then reverted), the difference between them is gone. No backup will recover it. It was never written down.

Greg’s claim in the talk is blunt: event sourcing “is the only model that does not lose information.” Store the facts themselves (CartCreated, ItemAdded, ItemRemoved, CheckedOut) and current state becomes a function you run over them. You can delete and rebuild state whenever you like. You can never rebuild discarded facts from state.

The report that arrives a year late.

Here is where the loss stops being philosophical and starts costing money. A product manager comes to you and says: I want to know which items customers remove from their carts within five minutes of checkout, because I believe they buy those items later, and I want to remarket them.

In a CRUD system, the answer is a schema change. You add a removed_items table, deploy it, and start collecting. The report runs on day one and shows nothing, because it counts forward from the deploy. The first useful version of that report is a year away. The years of behavior your customers have already given you are gone for good.

"Two timelines from system launch to a year past today. In the CRUD system the removals happened but years of them were destroyed, so a removed_items table deployed today covers only from today forward. In the event-sourced system ItemRemoved events were kept since launch, so a new projection deployed today covers the entire history. Callout: same question, answered back to launch day."
The same report request, deployed the same day, on two architectures.

In an event-sourced system, ItemRemoved events have been accumulating since the day the system went live. You write a projection over the event log and the report covers the entire life of the system by Monday morning. Better still, you can answer what the report would have shown on any date in the past. The business asks a brand-new question and the system answers it retroactively. In my experience, the moment a stakeholder understands this is the moment the architecture conversation changes. They stop asking what the system stores and start asking what the business could learn.

Who approved the forgetting?

Every UPDATE overwrites a value that will never be seen again. Every DELETE removes a fact that was once worth recording. Which data a company remembers is a business decision. In CRUD systems that decision is made by a developer, inside a migration script, without a meeting. Nobody weighed those removed cart items against a future remarketing campaign and ruled them worthless. The schema had nowhere to keep them, and nobody noticed that a decision was being made at all.

That is the force of Greg’s question about the CEO and the board. The room laughs because the conversation he describes has never happened anywhere.

Older industries settled this centuries ago.

Ask your bank for your balance and it will add up your transactions. The balance is a derived number; the transactions are the record. Accountants correct a wrong entry by appending a reversing entry, and an erased line in a ledger is evidence of tampering. Your doctor appends to your chart rather than replacing last year’s picture of you. When Greg points out that no mature industry runs on current state, he is describing institutions that learned the hard way that history is the asset. Software is the young industry here. We made overwriting the default because storage was expensive in 1980, and we kept it as the default long after that argument died.

Count both costs.

Event sourcing is not free, and I will not pitch it as free. Event schemas become contracts you maintain for years. Read models bring eventual consistency you have to design for, and teams need time to stop thinking in rows. Weigh all of that. Then weigh the other side with the same rigor: a CRUD system spends every day quietly deciding, on your behalf, which questions your business will never be able to ask. You get to choose which bill you pay. You do not get to choose zero.

Watch the talk. It is an hour that will change how you read every UPDATE statement you have ever shipped: https://www.youtube.com/watch?v=JHGkaShoyNs

I am writing a book on Event Sourcing and CQRS, with a complete production-grade reference implementation in C# and .NET. The argument above gets a much fuller treatment there. More on that soon.

When People Mix Up Commands and Events

I can usually tell within the first ten minutes of a design conversation whether a team has really understood Event Sourcing. And no, it has nothing to do with which event store they picked, or how far along they are, or whether anyone in the room has read Greg Young.

The tell is much simpler than that.

Can they keep commands and domain events apart?

When they can’t, I stop asking about their architecture. Because whatever they tell me next isn’t going to mean much yet.

“But they’re both just records!”

Yes! On the surface they look identical. Both are small records with a name and a payload. Both get passed around the system. In C# both are usually a record. If all you had to go on was the shape of the type, you’d be forgiven for wondering what all the fuss is about.

The difference isn’t in the shape. It’s in what each one is allowed to do.

A command is a request. One recipient. It can be refused. And once it’s been handled, successfully or not, it’s gone. CancelOrder is a command. That order might already be on a truck somewhere, and the aggregate is fully entitled to tell you no.

A domain event is a fact. Any number of subscribers, none of whom get a vote, and it lives forever because it is the system of record. OrderCancelled is an event. You don’t get to argue with it. It already happened.

So here’s the ten-second check I use, and it sorts almost every confused message I’ve ever run into:

Who is allowed to say no?

Nobody? It’s an event.

Exactly one thing? It’s a command.

A subscriber? Congratulations, you’ve built a distributed bug, and you’re going to meet it at 3am on a Saturday.

What the mix-up actually looks like

It never shows up with a label on it. It shows up as a handful of small decisions that each seemed perfectly reasonable at the time.

An event that’s really an instruction. OrderShipped goes onto a queue that exactly one consumer is expected to act on, and the publisher cares whether that consumer succeeded. That’s a command wearing past tense as a disguise. WTF is that?

An event named after the command that caused it. Something like CancelOrderCommandProcessed. Now your log is coupled to a command that might get renamed, split in two, or deleted entirely next quarter. The cancellation is history. History doesn’t get renamed because you refactored a handler.

One marker interface over both. One bus, one pipeline, everything symmetrical. It feels tidy for about four months. Then nobody can say which records are the system of record and which ones were transient, and the answer is buried somewhere in dispatch code.

Commands written into the event store. Right next to the events. Replay now happily reissues intent that was already satisfied two years ago. Have fun with that one.

Every one of those is a symptom. The disease is that whoever wrote it hadn’t yet internalized that one of these things is a request and the other one is history.

Here’s what doesn’t work: explaining it again

This is the part I had to learn the hard way, and honestly it took me longer than it should have.

My first instinct used to be to explain the distinction more carefully. Better definitions. A sharper example. A nice diagram with arrows. Surely if I just say it clearly enough?

Nope. People nod, they agree it makes sense, and then they go right back to their editor where the two types still look identical and the compiler still doesn’t care one bit. A definition is cheap and it’s easy to agree with. Agreement is not understanding!

What actually changes someone’s mind is being made to sort real messages from their own business, out loud, standing up, in front of somebody who knows that business better than they do.

Which is exactly why I reach for Event Storming

If you’ve never run one: a long wall, a box of good sticky notes, no tables, no laptops, and the people who actually know how the business works. Alberto Brandolini invented it and it’s still the best two hours you can spend at the start of a project.

Domain events go on orange notes, in past tense, arranged left to right in the order they happen. Commands go on blue notes. And here’s the bit that does the teaching for me:

A blue note never gets a spot on the timeline.

Commands sit off to the side of the event they produced, next to the actor who issued them. They can’t go into the sequence, because the sequence is a story of what happened, and a request isn’t something that happened.

Somebody tries it anyway in the first hour of every single workshop I’ve run. They write “Cancel Order” on an orange note, walk up, and stick it on the wall. And within about a minute a domain expert says something like “well, that’s not what happened, that’s what the customer asked us for.”

That’s it. That’s the whole lesson, delivered by the right person, in the room, about their own business. It costs me nothing and it sticks in a way my careful explanation never did.

And the wall keeps paying after that. When the group can’t agree whether something is a fact or a request, that argument goes up as a red note, and red notes are the most valuable things in the room. Every single one is a question the team would otherwise have discovered as a production bug six months later.

Then when you’re done, the translation is almost mechanical. Orange notes are your candidate domain events. Blue notes are your candidate commands. The yellow ones are the actors your task-based UI is going to serve. You didn’t just teach the distinction, you walked out with the first draft of the model.

So

Mixing up commands and events isn’t a style thing you fix in code review. It’s a signal that the core idea hasn’t landed yet, and no amount of patient re-explaining is going to land it.

Put the team at a wall with somebody who knows the business, and let the business tell them the difference.

Have you run into this on your own teams? Which shape did it take for you? Drop a comment below, I’d love to hear about it.

Write the code first. Let it tell you what the book says.

Back in December 2024 I asked here whether anyone would be interested in a book about Event Sourcing and CQRS. I had a six page table of contents and a lot of doubt about whether I wanted to commit the time.

Well, I committed the time. Two years of it.

And I did one thing differently from what I planned, which turned out to matter more than anything else: I built the implementation first and let the code tell me what the book should say.

Sounds like a small decision, right? It is not. It rewrote chapters.

Let me show you one.

Snapshotting looks simple until you build it

Everybody describes snapshotting the same way. Your event stream gets long, replaying it on every load gets slow, so you periodically save the aggregate’s state and on the next load you restore that and replay only what came after.

That description is correct. It is also useless if you are the one writing the code.

I found four things while implementing it that no diagram tells you. Three of them ended up in the chapter.

The trigger is boundary math, not a modulus

You want to snapshot every fifty events. So you check whether the version is a multiple of fifty, right? version % 50 == 0. Done.

No. And here is the part that should worry you: it fails silently.

One command appends three events. Your stream goes from version 49 to version 52. It steps right over 50 and never lands on it. The check never fires, the snapshot never happens, and nobody tells you anything, because a missing snapshot looks exactly like an aggregate that has not hit the threshold yet. The load still works. It is just slow. You find out six months later when somebody asks why one aggregate takes 400 ms to load.

What you want is to ask whether the append moved the stream into a new bucket:

postVersion / interval > preVersion / interval

Integer division, evaluated across the append instead of at a single point. A multi-event append that jumps a boundary captures once, at the post-append version.

Simple fix. But I did not see it until I wrote the test that appends three events at once.

Do not upcast your snapshots

Events get upcast. Event shapes change, you write an upcaster to lift the old shape to the new one, and you maintain that chain forever, because events are immutable and you cannot go back and rewrite history.

So when the snapshot shape changes, you upcast the snapshot too. Right?

Wrong, and I am glad I found this in code rather than in print. A snapshot is a cache. The aggregate can always rebuild it from events. So when the shape changes, let the old snapshot read as a miss, rebuild from history, and capture a fresh one at the current shape the next time you cross a boundary.

Upcasting snapshots buys you nothing that a discard does not already give you, and it costs you a second chain of upcasters to keep correct next to your event upcasters. Two lineages that have to agree with each other, forever, where zero would do.

Put the schema version in the WHERE clause, by the way. Do not read the row and compare in memory. A snapshot at a shape you cannot read should never cross the wire in the first place.

Guard the restore or corrupt your writes

This one bit harder.

RestoreFrom(snapshot, version) seats state onto your aggregate and sets its version. That version is the concurrency token your next append checks against.

Now let somebody call that on an aggregate that has already applied events, or is holding uncommitted work. You just overwrote the token. The next append goes out with an expected version describing a state the aggregate is not in, and a stale write lands as if it were current.

In an event-sourced system! Where the whole promise is that the log tells you the truth!

So the restore throws unless the aggregate is pristine. Loud failure right at the seam, instead of silent corruption three layers downstream that you find out about in production.

Your speedup test cannot be a stopwatch

Obvious test: load with a snapshot, load without, assert the first is faster.

I have the scar on this one. The same code passed five out of five on a 32-core machine and starved on a shared 4-core CI runner. The scheduler widened the window and the timing budget stopped meaning anything at all.

Snapshotting does not save you milliseconds. It saves you replays. So test replays. Seed a stream past two boundaries, load through the snapshotting repository over a store that records where the read started, and assert two things: the read began at the snapshot’s version, and it replayed strictly fewer events than the full stream.

Machine independent, and it measures what the pattern is actually for.

This happened across the whole book

The snapshotting chapter is not the chapter I outlined. Neither is the one on correlation tracing. Neither is the one on event versioning, and that one moved the most.

Every single time the implementation and the manuscript disagreed, the implementation won and the chapter got rewritten. Not once did I look at working code and decide my prose had been right all along.

And I think this is the thing nobody tells you about writing on architecture. Any pattern you can draw on a whiteboard has a layer underneath it where the real decisions live. Interval math. Concurrency tokens. What your test can honestly assert on a CI runner you do not control. You cannot write about that layer if you have not been down there, and readers can tell.

So here is my advice to anybody thinking about writing a technical book: build it first. All of it. Ship the code, run it, break it, and let it tell you what your chapters should say.

The manuscript is at 449 pages across 18 chapters now, with a production-grade reference implementation in .NET running on PostgreSQL, SQL Server, KurrentDB and DynamoDB behind one contract.

What would you want to see from a book like this? I am still listening.