Rendered at 21:33:26 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ramon156 11 hours ago [-]
Why even have a blog when you can't be arsed to write the posts. This is so obviously LLM-written. I have a positive view of Shopify engineers, but this kind of made a dent in that confidence.
davidron 5 minutes ago [-]
I have a gemini gem that recompresses the expansion that AI causes. It seems to have worked great on this article!
Happy to share the gem with anybody who's interested in reading articles like this without having to muddle through all the AI bulk.
kimos 3 hours ago [-]
I too have a positive view of Shopify engineers. But no longer Shopify management. It went from bottom up trust basted to top down AI first and for everything based. It’s very hard to express how deeply and completely the internal workings and culture of the company has changed. (Source: I am a long time former developer there.)
briga 10 hours ago [-]
What exactly makes it "obvious" that this is written by AI? I could totally believe that AI was used to generate parts of it, but I really don't get the sense that the whole thing was written that way. I've seen way worse examples on this site.
As software engineers we are constantly told that we need to heavily use these tools for our daily work. So is it surprising that software engineers use the same tools as writing aids? Using AI does not mean no human effort was involved.
energy123 10 hours ago [-]
Subheading and dot point spam, low density writing (the opposite of standard technical english), including useless detail (like enumerating stats on Shopify's scale), using contrastive parallelism, and other llm-isms. Even if it's not AI it's bad writing done by someone who has picked up AI's worst ticks.
For example this subheading:
> "The real bottleneck: connections, not CPU"
That's two AI smells. AI likes to say vacuous punchy statements like "the final takeaway" or "here's the rub". Then contrastive parallelism "connections, not CPU". Contrastive parallelism should scarcely exist in technical writing, regardless of whether it's AI generated.
samudrijan 6 hours ago [-]
Say the thing, not contrastive parallelism. Thanks for putting a name to that horrible LLM habit.
I really wish that slop writing was disincentivized in whatever RLHF they do. I don’t want to read the weird LinkedIn pop-sci tone for the rest of my life in such amounts.
I wonder if the average reader is also getting annoyed like this or whether they just don’t care - especially seeing what seems to get upvoted on your run of the mill social media sites. They probably collectively shape things more than I do.
DenisM 7 hours ago [-]
I wish they don’t stop. Make it easier to skip past.
Espressosaurus 3 hours ago [-]
I don’t know how you stop the slop writing. What the LLM writes has the voice of whatever has been RLHF’d into the weights. There will be a voice. It will have ticks.
They can change it but it will be there.
I do like that I have a term for that “this, not that” phrasing that’s nails on the chalkboard for me after several years of reading slop.
giancarlostoro 6 hours ago [-]
The LinkedIn “techbro” writings are the worst. They talk like a cryptobro “APPLE JUST PUBLISHED A CODE REPOSITORY ON THEIR HAND CRAFTED LLM ANYONE CAN USE THIS TO COMPETE WITH OPENAI AND ANTHROPIC” and its some random simple LLM on github that handles very little compared to OpenAI and Anthropic.
stasomatic 6 hours ago [-]
That's two AI smells. AI likes to say vacuous punchy statements like "the final takeaway" or "here's the rub". Then contrastive parallelism "connections, not CPU". Contrastive parallelism should scarcely exist in technical writing, regardless of whether it's AI generated.
Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of? I am also not a fan of the fluff words that Claude pollutes the context with, it's a bit much, I could live 30% less but it doesn't want to read its instructions in the .md
If the author wanted, they could ask their "AI" to write the article as a caveman, no? I, simply for fun, slop coded up a skill so that I can ask Claude to reply as Samual L. Jackson, or Kramer, etc.
I also wonder if there is an "AI language barrier", likely not in this particular article, but let's say there was a published article by a researcher whose native language isn't English. What happens with the translation? I realize I am asking naive things :)
zdragnar 5 hours ago [-]
There are distinctive patterns of language use that can be powerful when used sparingly. LLMs trained on a massive copus of human communication, picked effective patterns, and overuses those patterns to the point that it feels both artificial and underwhelming.
DetroitThrow 5 hours ago [-]
>Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of?
I think you might be interested in reading in training data generation, training, and post training papers/articles. I think you might be surprised at how much intervention there is on some of these levels.
mattmanser 5 hours ago [-]
There's a post training.step where they get humans to interact with it.
As I understand it, there is (or was) a step where they ask people what's the 'better' response.
These linguistic forms sound good the first time you hear themz even if they are rare in real speech, so rapidly got trained in.
Now they distill off previous models, I imagine these weird linguistic forms are quite hard to get rid of.
drstewart 4 hours ago [-]
>if it's not AI it's bad writing
Is this post AI generated?
thunky 9 hours ago [-]
The good news:
Your worst complaint has nothing to do with the overall content or accuracy of the post.
Just style bashing.
asG123k 8 hours ago [-]
Just? With all the Claudisms it is hard to distill what really happened. My guess:
- Our oversell protection was a gross hack that broke ACID.
- MySQL has a new feature that allows us to remove parts of the gross hack.
- Question: To what extent does the hack still exist?
Instead you get garbage like "The answer is often in the plumbing, not the engine." and "Crucially, this wasn't about making reservations fast. It was about making them safe neighbors."
But this crap is what Lütke wants, so they deliver.
thunky 8 hours ago [-]
Did you read TFA? Because all of your questions are answered clearly there.
Semaphor 4 hours ago [-]
Maybe they should show us their prompt so we don't have to read all the bullshitting to get to those answers
thunky 4 hours ago [-]
The article is fine. LLMs are just turning everyone into divas.
muragekibicho 8 hours ago [-]
Taken verbatim "But the hardest lesson wasn't about database design. It was discovering that the real bottleneck wasn’t what we were observing and measuring. "
cataphract 5 hours ago [-]
In every AI-written post, we have the script: some top post complaining it's AI-written, someone cluelessly saying they don't see it, and another person pointing out the claudisms. Sometimes someone posts a link to pangram.
I think we need the moderation team to set up some policy. In which direction idk.
watrr 5 hours ago [-]
The EU helps. We can report Spotify for not labeling this content, since a substantial number here think this obviously AI assisted post is genuine:
Spotify's defense will then be that this post is obviously AI assisted garbage.
derangedHorse 3 hours ago [-]
Shopify*
12387128 10 hours ago [-]
Because the post is just rambling without a clear intent or direction. Why do you need "oversell protection" if you have "transactions". Isn't the whole point of a "transaction" that it handles concurrency and disk failures?
resonious 9 hours ago [-]
I work with Shopify often and the whole oversell protection thing is like a huge joke.
Oversell happens because Shopify doesn't decrement inventory until payment is confirmed. And they apparently would rather die than change that invariant.
ethbr1 7 hours ago [-]
It's funny the organizational rules that get written into a company's DNA.
I could see a historical moment where Shopify, sans that invariant, massively fucked up inventory counts from uncompleted transactions... then had to unwind all that when a bunch failed to clear payment.
Ergo, now there's an invariant.
dessimus 4 hours ago [-]
Would a retail business want the inventory to decrement just because users put items into carts that get abandoned? Seems like that would really mess things up more in the long run.
derangedHorse 3 hours ago [-]
An unconfirmed payment might also refer to a credit card charge that is 99% guaranteed to go through but hasn’t completed yet
randomperson321 5 hours ago [-]
> Scale here is not abstract:
This is a really strange, clunky way of writing.
ChrisMarshallNY 7 hours ago [-]
Well, one thing to remember: AI learned from us (See what I did, there?).
The training these LLMs got, was from endless human-slop, on sites like LinkedIn, and marketing copy, everywhere. In fact, don't be surprised, if we start learning from AI; reversing the process.
But I think that it's only a matter of time, before almost everything will be at least touched by AI.
I posted this, yesterday[0]. It wasn't a particularly popular comment, but I stand by it.
I read your linked comment and I'm bummed out because I agree.
The thing is that I LIKE writing code. It's fun.
I don't LIKE coding with an LLM. It isn't fun.
We can't go back to a place where the professional code is all hand written. But at the same time I'm sad and tired. The joy has been utterly sapped and I don't think I'll be able to do this as a career for much longer.
ChrisMarshallNY 3 hours ago [-]
Well, I am sorry to hear that, but it hasn’t been my experience.
That may have something to do with the way that I use an LLM.
I use it as a “pair partner,” not as an agent.
Most of the code in my apps is mine, but I do ask the LLM to take care of some of the “overhead” tasks, such as minor utility functions.
In fact, I just got done, ripping out a bunch of code the LLM wrote, in my current app, and replacing it with my own work. It seems that I can’t really ask the LLM to do much beyond function level, without it going sideways.
I simply can’t understand people shipping entire, vibe-coded apps. If I let the LLM write my whole app, it would be hot garbage.
cindyllm 2 hours ago [-]
[dead]
matwood 6 hours ago [-]
Yeah, people act like this style is new, when it absolutely is not. There's this weird romanticism of the time pre-LLMs where people imply all writing/code was perfect, and only now it's slop.
EDIT
I also read your linked comment and agree 100%.
lelanthran 36 minutes ago [-]
> Yeah, people act like this style is new, when it absolutely is not.
Every time this claim is made I ask for a link to a pre-2022 blog or article that has all those AI tells.
I'm still waiting for it. I don't think I will ever receive it.
ChrisMarshallNY 23 minutes ago [-]
It's likely because no one could be bothered to respond. There's plenty there, but it's not their job to fetch it for you.
People on the Internet (especially here), seem to think that it's OK to make rude demands of complete strangers, and expect them to spend considerable time, meeting them.
I'd rather some rando think they "won the Internet," than spend a bunch of my time, working for free, just to satisfy them. I have better things to do with my time, and I don't really have a lot of investment in what complete strangers think of me.
Sometimes, the fox ain't worth the chase.
jpease 6 hours ago [-]
AI slop: Trained on 100% natural human slop.
askj18 7 hours ago [-]
[flagged]
ChrisMarshallNY 7 hours ago [-]
> But enjoy your intellectual wheelchair, which for some reason you feel the need to advertise after 40 years of programming. I'm sure it is just in the best interest of future generations ...
I love this place. We are such class acts.
Have a great day!
filcuk 10 hours ago [-]
People see an — dash and kneejerk. Writers are deciding to stop using this valid punctuation because of this exact reaction, it's ridiculous.
briga 10 hours ago [-]
A true casualty of AI. Em-dashes used to be one of my favorite punctuation devices, now I find myself consciously editing them out of my writing.
vikramkr 8 hours ago [-]
There are some claudeisms here (it's not x, it's y) but it does read like they did at least one human pass on it. Or maybe this is opus 5 writing which is a bit less distinct idk. It doesn't feel as obviously generated slop as some of the stuff posted here with every other sentence a fragment and load bearing and all that nonsense
drowsspa 6 hours ago [-]
Using AI is a pretty good indicator that the reader will need to exert more effort than the writer did. Unless the reader also decides to use AI to summarize it, but then what's the point?
gsich 8 hours ago [-]
The pictures.
curt15 7 hours ago [-]
"No waiting on the same row, less contention."
jssmith 7 hours ago [-]
I agree, I too found this barely readable.
Interestingly, ChatGPT was able to parse and explain it to me. I got a lot out of its presentation.
lp4v4n 10 hours ago [-]
I agree with you, it doesn't take long to identify the itemized style that LLMs use and it's always incredibly annoying.
I would rather read the original thoughts of the engineer, grammar mistakes and stylistic imperfections included, than prompt generated AI slop.
jgalt212 7 hours ago [-]
I think they're just following the orders of the CEO.
Culonavirus 5 hours ago [-]
> This is so obviously LLM-written.
Ok? Who cares. This was very interesting to me. IDGAF if AI was used (or not) to put the words in the article, the bits of information in those words is why I'm reading it. It's almost 2027, people need to accept that articles, art, videogames, code, everything will be more and more "AI touched" and you can't stop it.
I mean I share all the people's worry about AI taking jobs and a handful of AI (and AI adjacent) corps sucking out massive amounts of resources, but reading under every 2nd or 3rd article how it's AI is beginning to be more annoying than the AI writing itself.
It's like now every interesting or weird or exceptional or controversional video of anything has (what seem to be) 80-100 IQ people spamming "AI" in the comments instead of making interesting observations, posting their own anecdotes, cracking jokes etc.
tonyhart7 9 hours ago [-]
unfortunately its new normal now
13 minutes ago [-]
j1elo 1 hours ago [-]
[flagged]
cowboylowrez 10 hours ago [-]
I just wanted to drop a note that I almost never go into an article or blog wondering if its AI or not. I am trying to be a more picky reader and try to actually do something occasionally but I do browse database articles and this is a good one. So I wonder, am I becoming insensitive about AI writing? Am I being hypnotized into taking whatever color pill thats had a color representing a surrender to the AI "hive mind" whatever that is?
So I'm sort of curious, did you not like the AI writing style or just rejecting AI in general? I myself am very conflicted, I think AI in its current form is the wrong tech at the wrong time yet I don't mind reading AI text and sort of mooch off of googles free tier.
Also I'm starting to notice lots of AI smearing, I read folks online describing someone elses contribution as "obviously AI" and I'm suspecting in some cases these could be false accusations.
I'm thinking about starting a blog, so when my AI gf writes posts, do I ask her to try to "not read like an AI"? Anybody try that? I'm gonna try that. Hehe for all you know I already did hehe
ethbr1 7 hours ago [-]
> whatever color pill thats had a color representing a surrender to the AI "hive mind"
Obviously the blue pill.
sureglymop 11 hours ago [-]
Mostly unrelated but shopify is incredibly annoying. They introduced this delivery tracking app called "shop" and it has become unavoidable when buying electronics from china. Recently looked at it with mitmproxy and it ships home more than gets shipped to me.
malfist 7 hours ago [-]
That and they helpfully share your email with a company if you add something to your cart. You don't even have to checkout or submit a form or anything. I can't tell you the number of spam emails I've gotten from companies because they auto add everything they get from shopify to their mailing list and then nag you about not checking out
sureglymop 6 hours ago [-]
Wow that's even worse! I use a catch all email and I recently even contacted a company I bought something from about potentially being compromised because I got a weird email. This here would explain what happened.
doublerabbit 11 hours ago [-]
Mood. It's not just over-seas. With a domestic courier they still use dark-ui to hide the tracker link via "shop".
manbash 20 hours ago [-]
> Instead of one row per item with a quantity column, we use one row per sellable unit. An item with 10 units has 10 rows.
> But one row per unit for all inventory would break down at scale—an item with 50,000 units across 10 locations would mean 500,000 rows, and the reserve query would slow as it scans through them. Instead, we maintain a bounded pool of available rows, capped at 1,000 per item/location combination. Reservations consume rows from this pool; a replenishment process refills it from the inventory ledger.
Shouldn't I feel uncomfortable with such approach? It seems to create a backoff (pool) for lowering the chance of having a synchronization issue.
solatic 3 hours ago [-]
Yeah, I'd be uncomfortable with that approach. One of their key design goals was to minimize underselling, recognizing that it results in lost revenue. But if a seller has 5k inventory in one location, has a spike of 2k orders, but only 1k of the orders can successfully reserve inventory, then isn't that an argument that you lost the revenue of the 2nd 1k orders that error out before the replenishment process succeeds?
I'm skeptical of this approach. Sure, row contention means that you cannot have a database transaction per customer order attempting to decrease inventory count by 1 each time. But you can have a batch transaction whereby the transaction decreases inventory by 100 (thus touching the high-contention inventory row once) and credits each of 100 different customer cart database rows (which are not under heavy contention and can be on a different disk entirely). Attempted customer orders are submitted to a reservation system put in charge of assembling the batches. Customers wait some short period of time - say, 15 seconds - for the reservation attempt to be batched and to be notified that they successfully locked a reservation. Arguing that "slow reservations trigger throttling and a worse buyer experience", without an actual number for what counts as "slow" to serve as an SLO and as a design target, is a cop-out inviting over-engineering.
soontimes 2 hours ago [-]
> But if a seller has 5k inventory in one location, has a spike of 2k orders, but only 1k of the orders can successfully reserve inventory, then isn't that an argument that you lost the revenue of the 2nd 1k orders that error out before the replenishment process succeeds?
They explicitly cover this in the article, saying they do reservation inline. It does increase latency for these orders, but it doesn’t result in an error
> But you can have a batch transaction whereby the transaction decreases inventory by 100 (thus touching the high-contention inventory row once) and credits each of 100 different customer cart database rows (which are not under heavy contention and can be on a different disk entirely).
They mention this as well, checkout batching increased implementation complexity.
> Arguing that "slow reservations trigger throttling and a worse buyer experience", without an actual number for what counts as "slow" to serve as an SLO and as a design target, is a cop-out inviting over-engineering.
True. The article would’ve been better if they included such numbers. However the fact that they didn’t mention this doesn’t imply they haven’t done research. I haven’t found anything related to checkout specifically, however there are in general articles, indicating that increased latency correlates with revenue drop.
fauigerzigerk 11 hours ago [-]
I agree, it does seem awfully complicated and there are quite a few pieces missing for this to be a complete solution.
I'm a bit surprised about the scalability case against a simpler solution. This is not about Shopify's scale. We're talking about contention for a specific SKU of a specific seller at a specific warehouse location.
How many shopping carts are competing for a single SKU at the payment stage at peak hours? Can this really be too much lock contention for a single database row?
I realise Shopify engineers are neither stupid nor inexperienced. Hence my surprise. I would have liked to hear more about that specific problem.
sgarland 6 hours ago [-]
> Can this really be too much lock contention for a single database row?
When you consider how many other (often poorly-crafted) DB queries and service calls are being performed while holding the row locked, yes, it can rapidly add up. If the entire pipeline takes 500 msec, congratulations, you can’t sell more than 2 units per second of that SKU. If the item is popular and being actively hyped, then yes, that can be a problem.
kevincox 10 hours ago [-]
Flash sales are a huge scaling issue for Shopify. There are celebrities who want to sell thousands of items in a few minutes window at the end of an advertised countdown.
Basically this is an incredibly rare case but a feature that they want to support.
cowsandmilk 9 hours ago [-]
There were woot offs 20 years ago. Flash sales are not some new scaling issue.
risyachka 7 hours ago [-]
Its also not a solved one considering basically no platforms support them.
malfist 7 hours ago [-]
It's not? Shopify does it, woot does it, amazon does it, yes style does it, royal Caribbean does it.
Seems like a lot of places can do it for it not to be a solved problem.
fauigerzigerk 10 hours ago [-]
Makes sense.
misiek08 8 hours ago [-]
We observe 200+ (can’t say closer number) purchases per second for single SKU with good marketing and price.
The other thing that bothers me - why not real stable, but maybe „too old, medieval” solution with Redis as the main source of through - without any sync with SQL at all in terms of stock… it worked in my previous job with much higher traffic (1000s/s). Yup, we ended up with app-side sharding, but it was stupid-simple.
theptip 6 hours ago [-]
You need a sql DB for the actual purchase transaction, you can’t keep financial records in Redis.
As they say in TFA they used Redis for the cart reservation but then you need to sync the two stores.
sandeepkd 19 hours ago [-]
Comes down to type of items, when you have physical inventory the number is limited so more manageable and interestingly enough the problem only applies to physical inventory.
You are just spending some more disk space to avoid synchronization issues. Denormalization for performance is a really common pattern, just that people do not start with it in the first place itself
esjeon 16 hours ago [-]
I would call this one-row-per-contract-type, and this is the most general model for the problem (e.g. the model cannot be further broken down into finer level), thus, the most scalable model given storage is dirt cheap.
LunaSea 13 hours ago [-]
Storage might be cheap but search and RAM isn't
ponector 12 hours ago [-]
RAM may be expensive for a person, but still is quite cheap for the company.
RobotToaster 10 hours ago [-]
> an item with 50,000 units across 10 locations would mean 500,000 rows
I don't get it, Wouldn't that still only be 50,000 rows, just divided across locations?
every (item, location) combo gets its own row, and then they moved to the smarter thing.
bijowo1676 18 hours ago [-]
you should, their design is not the best. There is middle ground between "one row per SKU" and "1000 rows per SKU".
Its called one row per shopping cart*SKU combo.
if two people order 100 and 500 items of the same SKU, respectively, the table should have only two rows: for order1 and order2. Not 600 rows.
codedokode 12 hours ago [-]
The problem is that in this case you have to do splits/merges. And while there are products that are sold by 100 units at a time, I think in most cases people by 1-2 items so the hassle might be not worth it.
Also you might not understand the original problem. Imagine if 100 customers want to buy product A. One thread starts a transaction, searches for amount of product A and UPDATE's it and goes searching for other products. The database locks the row until the end of transaction and other 99 treads cannot continue until first transaction commits (they can read but cannot update the rows).
This is why they made a row per item. In this case, transaction 1 hopefully locks only several rows with items of product A. Transaction 2 instead of waiting for lock release skips them (due to SKIP LOCK) and locks several next rows. And so on.
Obviously you do not need to make a row per item - if the available amount is really large (10 000 items), you could have for example 100 rows having 100 items each. In this case each transaction locks the whole row (100 items) even if it wants to reserve just one item. The problem though is that now every row might have different amount of available items and you have to do more work to reserve the amount you want.
bijowo1676 6 hours ago [-]
ok, lets model situation of 100 customers and one last remaining item. Who will get the last item?
in shopify's design, it is a user who was the first to lock the row and have successful payment. Sounds good, but how often does it happen ? It's a rare and extreme case and they model their entire system after the rare even, and incur the overhead of 1000 rows per SKU per shop for all combination of SKU and shop_id for all the normal items that are not sold out in flash sale.
the same outcome could be achieved without locking and without creating 1000 rows:
1. keep track of all active carts at the checkout in a table
2. for each cart, record the timestamp in nanoseconds when user clicked Pay (but I would prefer timestamp of clicking Checkout)
3. that timestamp will decide who gets the last available item.
4. in a shopping cart, have explicit field for each SKU: inventory_reserved.
5. this decision mechanism is now explicit via global monotonic non-decreasing counter. It is no longer tied to payment processing gateway timeouts, not opaque and implicit mechanism relying on database internals and quirks of how DB engine locks and releases some placeholder rows.
codedokode 5 hours ago [-]
On a large scale "rare" events happen every day. That is why people use locks and transactions or other measures.
In case with shopify, they want to decide whether the user may place order or not, at the moment when the user clicks "Pay" or some other button. If the user cannot place an order, they are shown the error, if they can, the items are reserved and the user is redirected to the payment page. So payment is processed only after successful reservation, and reservation is made only if the user wants to pay. The similar system works for buying train tickets online in my country, for example.
In you case, when user A clicks a button, following happens (as I understand):
1 the server increments the counter
2 the server calculates available amount as (amount_in_stock - amount reserved by carts with time < counter)
3 if the amount is large enough, the server updates the "time" field for user's cart thus reserving the item
Imagine that at step 2 the user A sees that there is one item left. However before user A does step 3, another user B might reserve the item (complete all 3 steps), and proceed to the payment. Then user A then completes step 3 and proceeds to the payment too. Now we end up with both user A and B paying for the last remaining item which doesn't solve the stated problem. Shopify's solution doesn't have such issues.
This is a classical TOCTTOU situation. There were exploits against Linux kernel based on similar issues.
bijowo1676 4 hours ago [-]
shopify is wrapping their entire dance with locking and moving rows inside a transaction. if you wrap step 1-3 inside transaction you will get same atomicity guarantee
but again, my idea was:
1) do not use throwaway placeholder rows to imitate a single item
2) do not rely on db engine to decide which transaction gets committed first (which customer gets the last item)
3) model queue explicitly by introducing counter field that sorts and prioritizes customers' orders and decides which order gets fulfilled and which customers gets the last item
codedokode 32 minutes ago [-]
No, using transactions won't change anything here.
> do not use throwaway placeholder rows to imitate a single item
The point of using multiple rows for one product is to distribute the locks.
jakewins 17 hours ago [-]
Can you explain how that works? With the row-per-item I can see how you’d use locking primitives etc easily to deal with multiple concurrent shopping carts claiming available inventory.. but how does your solution solve contention? There’d need to be some “number of items in inventory” row, wouldn’t there be contention on that?
The point of one row per item is that thousands of concurrent shoppers don’t need to block each other as they can each claim as many free rows as they need for themselves?
sroussey 14 hours ago [-]
One other advantage is item serial numbers. Or something else that makes an item that seems the same but actually be unique (perhaps the warehouse it’s in?)
I guess it depends on how the replenishment process works. Unless you're ordering over 1000 of an item, I doubt it would be a problem.
bijowo1676 18 hours ago [-]
replenishment is an unnecessary cludge that only exists due to poor design. an "algorithmical smell" if you wish
szundi 11 hours ago [-]
[dead]
e12e 11 hours ago [-]
Maybe the example numbers are just bad - but now you expect your system to fall down if you scale from 10 to 100 locations?
onion2k 11 hours ago [-]
"Number of locations" is an input so if the system has been designed to handle up to 10, and not 100, then yes I would absolutely expect it to fail with the higher value.
Developers (and everyone else really) need to think about systems, with the system taking inputs like "number of locations", and producing outputs like "available inventory", and when the input parameters change outside of the designed scope, without the system itself changing, then you should expect things to break.
jghn 18 hours ago [-]
Depends on the scale. Most companies don't approach the scale where this matters.
dbbk 15 hours ago [-]
I'm familiar with the reserved row approach (I use SELECT FOR UPDATE SKIP LOCKED) and yeah this replenishing idea terrifies me.
jhhh 6 hours ago [-]
My main takeaway from this post is that in 2026 we haven't developed enough technology to scalably and durably handle concurrently decrementing a single number. This has caused multiple organizations to develop database hacks (the multiple rows) or complex architectural solutions (redis) which destroy the atomicity of the process.
winrid 54 minutes ago [-]
The benefit of per row is that you can tag additional info like reserved user id etc, which you would have to track somewhere anyway.
isignal 19 hours ago [-]
It seems there could be a simpler solution.
1. Deduct the reservation from the inventory when the user starts to order, but in the same txn also maintain a separate row for the in progress order flow.
2. If the order flow is aborted or times out have a background process that returns these to the inventory.
That seems simpler than this approach and involves no locking. Though their presented approach is also reasonable, there must be some reason not to choose a simpler flow. It is not that difficult to have a gc service that scales, but may be they didn't want to separate that.
firasd 19 hours ago [-]
My understanding is: your proposal is not very different from what Shopify is doing except they are tracking 'reserved units' (one per row) and you are proposing tracking 'orders' as the temporary state to then reconcile back with inventory quantities.
isignal 18 hours ago [-]
Yes, at a high level. It doesn't rely on skip locked, which is not cheap at DB level. DB has to still typically run query and keep going until it finds an unlocked item. Deducting and checking inventory counts are simpler ops inside the DB.
treis 18 hours ago [-]
This seems like what triggers are for and how we do similar type things. Update trigger on order does select for update on the inventory and increases/decreases it as appropriate.
I don't think you really need that even. An indexed lookup is fast and you don't need to store a computed quantity generally.
sandeepkd 19 hours ago [-]
The moment you added a background process you just replaced the complexity.
1. Backgrounds process can back up
2. They need context of the user and need to switch context per user
3. What if they fail, you create some DLQ or another process to handle the failure
4. Who looks on those failure and how do they act
TLDR; there is always a cost
0x696C6961 18 hours ago [-]
The design in the shoppify post already had a background process for the item replenishment.
soontimes 19 hours ago [-]
Can you clarify why this involves no locking? There can still be 2 actors fighting for the same row.
isignal 18 hours ago [-]
Two concurrent deductions of inventory do contend but only during the actual DB update. That is just normal DB locking for SQL isolation levels. The blog refers to explicit locking by the app, which is where skip locked comes in.
soontimes 18 hours ago [-]
Yes, the point is to spread contention across multiple rows. They also mention this in the beginning of the article
vxxzy 19 hours ago [-]
now you have two problems. what happens when your reservation system backs up?
dbbk 14 hours ago [-]
You don't sell stuff I guess
sieabahlpark 19 hours ago [-]
[dead]
stillpointlab 18 hours ago [-]
I was investigating Durable Objects (DO) and had Fable walk me through where in my app they might be appropriate. One place had a dependency with billing (where I use a transaction now) and the proposed re-work to allow for concurrent editing with DO looked very much like this, reservations with idempotency keys. And if you add hierarchical allotments then it scales pretty well.
I disagree with the other posters about the bg process, if you have any bg processing already you should be able to handle the few edge cases without too much trouble.
lordmoma 3 hours ago [-]
I don’t get Why AI written technical article is a big thing here. Did we care if an article is typed or handwritten at a point?
giovannibonetti 8 hours ago [-]
I wonder how we could handle that in a simpler way with durable workflows (e.g. Temporal, Restante, DBOS) – which are similar to Erlang processes but with persistent disk storage. This could avoid the need to maintain the 1000 row inventory.
Perhaps each shopping cart would have its own workflow, and the inventory item would have one as well. Then, whenever a customer put an item in their cart, their cart workflow would send a signal to the inventory item workflow and wait for the response. The inventory item workflow would maintain a ledger controlling to which cart each unit goes, and it could batch the writes to this table. This way, even if 100k customers try to purchase the same item in the same second, it should handle the load.
After the batch is written to the ledger, the inventory item workflow would reply signals to each cart workflow confirming that the reservation was completed. The end-to-end latency from the consumer point of view would be a fraction of a second, without needing the 1000-row hot-inventory heuristic.
azuanrb 8 hours ago [-]
Durable workflows are different, not necessarily simpler, imo. Unless the team is already familiar with them, I wouldn’t introduce one just for this. You also have to account for the infrastructure needed to run and manage the durable workflow itself, which adds complexity.
kandros 8 hours ago [-]
simpler
bijowo1676 18 hours ago [-]
not the best design to have 1000 rows for each shop*SKU combination. If a candidate proposed this solution during Shopify's System Design interview, i doubt he would be vetted for Senior+ position.
Instead of having 1000 rows per shop*SKU, why not just have one row per shopping cart*SKU?
That way a single row would represent a single cart, and will hold info of multiple items of the same SKU.
No need a cludge with 1000 rows limit and replenishment process. Instead of dealing with N rows, you always deal with a single row.
idoubtit 15 hours ago [-]
> not the best design [...]
So those engineers at Shopify worked hard for months on a more performant system, but they missed the obvious structure? They chose a complex denormalization for no good reason?
It may be true, but I think it's presumptuous to belittle their work when we have only partial information. My guess is that they had good reasons to think that the more obvious ways would not scale.
And from reading your comments in this thread, I believe your structure would fail at their scale. A SQL query that uses 2 sub-queries with "group by" is probably too heavy. From the post, at peaks there would be millions of active shopping carts.
BTW, I suspect most orders are just for 1 or 2 of each item, so the denormalization is not as heavy as it seems.
bijowo1676 15 hours ago [-]
i also work in big tech and know that a lot of bullshit design creeps into system design and prod, because everyone is overworked, overstressed, wants to just get things done for the quarterly performance review as to not get shitcanned with severance
re concurrency, it is not a big issue at all. stock exchanges deal with HFT traders and can easily deal with concurrency of orders. Same can be implemented with shopify, but I doubt they face the same level of concurrency as stock exchange anywhere near
JamesSwift 9 hours ago [-]
Famously, stock is settled on a delay (and generally doesnt involve physical products that are not fungible). Im sure theres a lot to glean from how they handle concurrency but Im not sure they are solving the same problems.
bijowo1676 6 hours ago [-]
settlement is a different process, what exchanges are doing is they match Buy and Sell orders.
you have an open Sell 1 APPL for $100.0. Millions of other HFT orders rush to scalp your single order. How do you think exchange matches your Sell to HFT's Buy orders? which Buy order gets fulfilled first?
soontimes 9 hours ago [-]
> re concurrency, it is not a big issue at all
I would really appreciate it if you could write this up as an article. It would be an extremely interesting and valuable read
bijowo1676 4 hours ago [-]
Martin Fowler's overview of LMAX Disruptor is one of the best reads on this subject re high-load system design
If anything this article shows that concurrency is a big issue. It is such a big issue you have to write in-memory single threaded processor with custom journaling. If you have established workflows with MySQL and a team knowing how to work with it, throwing all that to do LMAX is not cost efficient. While there are domains where such approach is suitable and even required due to strict transaction ordering, Shopify case doesn’t look like one of them.
post_below 14 hours ago [-]
You might not have noticed that essentially the entire blog post was AI written.
There's even this bit where they discover a remarkable trick:
> Each round trip to the database has a cost. For carts with multiple line items, we batch reservation queries using UNION ALL so we fetch all needed units in one round trip
Insights like that really don't read like senior level output, and of course, it's LLM output. I'm not sure it's presumptuous to question it.
Chyzwar 12 hours ago [-]
There is now new type of comment in HN, if you disagree with article you attack the fact that ai was used in post editorial process. People here now dismiss anything that have em dash.
maxrev17 11 hours ago [-]
Yeah ai doesn’t mean bad content. It does make for unreadable and unbearable articles though
atomicnumber3 17 hours ago [-]
I have never worked anywhere where describing how their system actually works would pass the company's own system design interview
maxrev17 11 hours ago [-]
Hahaha best comment on the whole thing! So true!
matwood 14 hours ago [-]
Others are almost never as dumb as you hoped, and you’re rarely ever as smart as you think.
soontimes 17 hours ago [-]
> Instead of having 1000 rows per shopSKU, why not just have one row per shopping cartSKU?
At what point that row is inserted?
bijowo1676 17 hours ago [-]
per my reading of the article, the protection is only needed for a few seconds, while payment is being processed by the payment system.
so the row is inserted when Payment is initiated, and row is deleted when Payment succeeds
What is oversell protection?
Reserve: When payment starts, we mark items as reserved (a short hold, e.g. several minutes).
Claim: When payment succeeds, we permanently deduct quantity from the inventory ledger (source of truth).
but that system could be easily improved to reserve item when user Adds item to a cart, to prevent scenario when user adds item to a cart, goes through checkout, and after initiating payment gets "soldout error":
1. Let user add item to a cart by default (happy path)
2. Initiate async check in the background for SKU and quantity
2a. The check sums up rows for all SKUs and compares to Inventory table (very cheap check since its done to only active shopping carts)
3. After few seconds the check comes back, and we let user know that item is soldout, before/the moment user goes to Checkout.
soontimes 17 hours ago [-]
Ok, but before inserting you must ensure that inventory is not depleted, which means you need to know the count and you need to lock the row. So you still have contention on that item. Them having a 1k buffer allows not to take a lock on a single row every time, and only do it when buffer is empty
bijowo1676 17 hours ago [-]
there is no need to lock the row, since you a dealing with a shopping cart, not individual item piece. when you run aggregate functions, lock is no needed, it is actually better to run it with SET TRANSACTION ISOLATION LEVEL READ UNCOMMITTED; for aggregation
the check for oversold items is extremely cheap:
with current_order as (
select $SKU1, $q2 as quantity
union
select $SKU2, $q2 as quantity
),
with carts as (
select sku, sum(quantity) as reserved
from active_carts
group by sku
),
with warehouse as (
select sku, available_units
from inventory
group by sku
)
select * from current_order
inner join carts using (sku)
inner join warehouse using (sku)
where warehouse.available_units - carts.reserved < current_order.quantity
assuming there are indexes on sku field in both, results in efficient index seek and agg over 2 tables
codedokode 11 hours ago [-]
The item is reserved when the user decides to place an order, but before paying for it. Not when a product is added to the cart because the user can keep it there for a month and end up not buying.
You reserve the product by creating an "active_cart" entry. Your solution has a problem, that when you run the check, it might say the product is available, but before you create an "active_cart" to reserve it from thread A, another thread B reserves it and you end up reserving a product that is not available anymore. You end up with SUM(active_cart.quantity) > inventory.available_units.
That is exactly why the database has locks - to prevent this situation. With locks, thread A decrements inventory.available_units and that row is locked until the end of transaction. Other threads (if they do SELECT FOR UPDATE instead of SELECT) cannot see the old, invalid value until thread A either commits and the value is updated or rollbacks. However, locks cause performance issues and that is why shopify uses the architecture from the article - instead of 100 users fighting for the lock on the same row with available amount, each user locks only rows with units they plan to buy.
I don’t understand how this should prevent oversold. You have a check that reports empty or oversold inventory. But how does that check prevent 2 concurrent actors fighting for the last item from inserting 2 rows?
bijowo1676 17 hours ago [-]
how does current design resolve concurrent actors fighting for the last item ?
there is ultimately needs to be some global mechanism resolving this conflict. Currently it is an order in which db engine processes transactions by locking rows for a transaction, whoever got the first lock, wins the last remaining items.
my design is the same, except it does not need this dance with moving rows between tables, locking them, and the cludge with replenishment process.
in the simplest form, run the sum() over active non-finished orders and compare to inventory. you get the same result: whoever got the first to run sum() and get positive answer will get the last remaining items.
but the problem as formulated, imho, is not even correctly defined.
Shopify incorrectly formulated the very problem they are trying to solve.
Trying to solve it at the payment time is too late, its better to resolve it earlier, before the checkout.
the "PAY" button should only do one thing: deduct money from cc and that's it. Resolving inventory availability must be solved way earlier, the moment user clicks Checkout, not when user clicks Pay.
So ideally, the error for oversold items should be shown to a user when he clicks Checkout, not when he click PAY
admax88qqq 16 hours ago [-]
> Shopify incorrectly formulated the very problem they are trying to solve.
That’s a bold overconfident statement. Cart abandonment is real. People never clear their carts they just walk away
Shopify purposefully chooses to do it at payment time because doing it earlier results in lost sales as people “reserve” items and then walk away causing other to see out of stock and then also walk away
Whoever puts up the money first gets the item
That’s the design constraint they chose you can’t just say “their solution is wrong because they solved the wrong problem”. Each design is a different user experience and I think it’s safe to say they chose which experience they want consciously.
bijowo1676 15 hours ago [-]
that's why I mentioned active carts in my post, there are ways to define active cart to get rid of abandoned carts ( ignore carts where last user action was > N seconds ago).
Ok, let's accept the design goal that whoever paid first wins. You can use the same metric (how many milliseconds ago did user click PAY) and impose a global monotonic non-decreasing counter to distribute the scarce inventory. This is how order matching engines work at stock exchanges with HFT orders (FIFO logic).
the goal is to know with 100% certainty, before sending payment request to payment processor, who will have item and who won't, and you dont need to move mountains of rows for that.
the payment processor should be just a binary answer: payment succeeded or not, but currently it combines Inventory availability check & payment processing, which is the root cause of confusion. For clarity it is better to make that stage of order processing an explicit separage stage, instead of coupling it with payment stage.
some stores split payment into two stages: Payment and Final order confirmation. at the Payment stage you can pre-authorize money at cc and do inventory availability, and at final confirmation you capture $$
hanikesn 14 hours ago [-]
Most payment methods in the world don't support separate authorization and capture.
bijowo1676 14 hours ago [-]
i dont know about the world, by authorize.net and Stripe, which work globally and work with global credit cards, they do support separate authorize and separate capture, which seems to be part of PCI standard
You'll quickly realize PCI mainly applies to the credit card industry and not to something like Europe's psd2 and sepa instant.
7 hours ago [-]
14 hours ago [-]
neerajsi 16 hours ago [-]
Clearly this is for high concurrency cases where there are many people racing to get all the available items. It's not clear that it's in shopifys or the sellers interest to let items get sequestered in people's shopping carts, which is a spot where there isn't a strong commitment to complete the purchase. At payment time, you can be more assured that the item will actually be purchased.
Still I think their solution is a bit weird. I'd want to commit the reservation transaction with inventory decrement along with a payment key and then use a different transaction to drop the reservation when the transaction completes. If the transaction does not complete in a timely manner you probably need to query external systems anyway to resolve whether the payment actually occurred or not.
They talk about lock contention in this case, but I also wonder about latch contention since these rows are adjacent. If it's a small transaction that's not interactive, does mysql resolve it with just the latches on the needed tables?
neerajsi 13 hours ago [-]
I was curious about what Tiger Beetle does. It has two phase transfers, which appears tailor built to handle this case. But maybe Tiger Beetle isn't the right database to track all your product stocks.
jorangreef 2 hours ago [-]
TB was designed to track stock inventories (as another form of double-entry accounting).
10 hours ago [-]
soontimes 17 hours ago [-]
> how does current design resolve concurrent actors fighting for the last item ?
It resolves with skip locked. Assuming we have only 1 item left. First query scans the buffer table, locks as many rows as needed (1 in our case), and moves rows to another table. Second query scans the table, finds no rows (even if first one hasn’t finished yet, the row is locked and ignored), checks if it can increase buffer, finds out that it’s fully sold and aborts. Db guarantees that you can’t oversold.
> my design is the same, except it does not need this dance with moving rows between tables, locking them, and the cludge with replenishment process.
I can’t evaluate whether it’s the same or not, because you still haven’t clarified when exactly you’re going to insert the row. In the article they’re inserting in the same transaction. Would you also do it in the transaction? Because if you’ll introduce a separate global mechanism to resolve conflicts, on a high level it would be the same as their approach with redis (you need to have 2 systems)
EDIT: wording
bijowo1676 16 hours ago [-]
think about for a moment what that skip locked actually means, all these 1000 rows per SKU are logically equivalent to a Inventory table with a single row where available_units=1000 per SKU.
now let's think again, do we need to lock 900 rows to place order on 900 items? or can we insert a single row where order_quantity=900 ?
shopify's design relies on DB to lock rows for transaction as a way to "decrement the counter" of available units. What I am suggesting, is you can just decrement counter by updating a single row, no need to lock 900 rows. Shopify moved from one extreme (single global variable in redis) to another extreme (1000 rows in db) and forgot about the middle ground.
The dance with moving rows per each item between tables is completely unnecessary, it's like counting numbers one by one in a for loop, when you can just substract number directly.
if I were to solve the problem, I would have solved it differently, at the Checkout state, before user clicks PAY. This removes the race condition at the user UI level, before any request lands in backend/db:
1. Have a table with active shopping carts (cart_id, cart_status, sku, quantity)
2. when cart_status changes to 'Checkout' run inventory availability check
3. If inventory availability check fails, show error to user (before he clicks Pay) and suggest replacement items.
4. If inventory availability succeeds, proceed to charge cc
availability check is the SQL above: inventory-sum(active_carts.quantity)-current_order must be > 0
soontimes 13 hours ago [-]
> if I were to solve the problem, I would have solved it differently, at the Checkout state, before user clicks PAY. This removes the race condition at the user UI level, before any request lands in backend/db:
In order to avoid races you need to insert reservation and decrement availability atomically. Your proposed approach is not atomic. For it to be atomic you will need to lock whole range, to make sure no new rows appeared between the points “check for availability” and “record reservation”. Actors will be effectively competing for the single aggregate row. This is the same as having a single inventory row with quantity field, which they rejected in the beginning of the article
> now let's think again, do we need to lock 900 rows to place order on 900 items? or can we insert a single row where order_quantity=900 ?
In the proposed schema nobody is waiting for these locks, they’re skipped by concurrent queries. In your schema actors would have to wait before they can insert without breaking invariants.
pas 15 hours ago [-]
assuming their "reserve item" function is just "update the table set N rows to reserved=true where reserved==false"
more transactions can commit at the same time, but with one counter they would conflict (as it did in the Redis case)
they should use CRDT (and trying to model that with this 1000 row workspace, no?)
still, eventually at some point they need to do the math
Godsend69 16 hours ago [-]
[dead]
edoceo 16 hours ago [-]
Thanks! I don't uSe 'with' enough
zer00eyz 9 hours ago [-]
Your mental model here is mapping too close to an actual cart in a retail, at a in person, setting.
The assumption that a SKU maps 1 to 1 to a cart item is flawed.
If the first item in the cart is a bundle of SKU-A and SKU-B, the second item is a bundle of SKU-A and SKU-C and the third item is 5xSKU-B where do you do you keep the re-agregation of the SKU-X's to track them?
This is without accounting for item location in the reservation - and rules that may apply around that.
You haven't even gotten to the part where different customers will have different rules around shipping from different locations - because that can eat into margins.
You're also making a bunch of other assumptions around transaction flow and where carts are actually stored (and how they get converted to an invoice, with payment attached) that likely do not hold true.
Could you do it more like what you're sugesting -- maybe -- but only in a single tenant system.
bijowo1676 6 hours ago [-]
Modeling after physical shopping process is actually the proper way to design a eshop system.
Imagine you are at Walmart store and go to checkout stage, you would have to pick item from the shelf and take it from availability for other shoppers, before you pay for the item.
What shopify did, is customer enters the store, heads straight to checkout and retail workers races back to shelves to pick up items for client. Sometimes it says: sorry bud, item is sold out, frustrating customer experience, who is already mentally prepared to pay and own an item.
Re SKU storage, you will have multiple row entries per SKU, if I order two items, there will be two rows corresponding to the items in your Purchase order.
The sum() aggregation check will run across all active orders per sku
Re single tenant: shopify creates 1000 rows per SKU per shop(tenant!!). As long as tenant is on the same DB you can run it, just add shop_id to the group by field
zer00eyz 3 hours ago [-]
If Walmart (in person) behaved like Walmart (on line) the isles would be so littered with half full carts and items that others could not buy you would be unable to move in the store.
70 percent of carts are abandoned. You dont want your inventory sitting in carts, when other people want to buy it. It only comes off the shelf (and gets put in a box) when you have money in your hand.
> Sometimes it says: sorry bud, item is sold out, frustrating customer experience, who is already mentally prepared to pay and own an item.
This is better than
A) taking their money and then telling them you dont have it.
B) Them not being able to buy it because someone has it in their cart and is NEVER going to check out with it.
raverbashing 14 hours ago [-]
And that's why these interviews can be stupid, you can mention the real solution and interviewers might reject because it's not the textbook solution
But the real world is different
codedokode 13 hours ago [-]
Could not they shard the inventory table by shop_id? As I understand, the order includes only items from one store, so there is no need to keep all the stores in a single table.
Also, I wonder why they could not have a row status (available/reserved) and UPDATE it instead of deleting the rows.
soontimes 9 hours ago [-]
They never said they don’t shard it, however this doesn’t solve the problem they were facing. Even if they have a single store (therefore a single shard), the burst demand may be high for the item in that shop, which creates contention for “remaining item quantity” resource. Their solution spreads this contention across several rows.
> Also, I wonder why they could not have a row status (available/reserved) and UPDATE it instead of deleting the rows.
This requires a row per item unit, doesn’t it? If you have 50k units you’ll have to track status of every item, meaning 50k rows. They also mention this as a rationale to use at most 1k rows, and treat it as a buffer.
codedokode 6 hours ago [-]
I now thought that "updating a row" might be more expensive than simply deleting because UPDATE is implemented as "mark row deleted" + "insert new version of a row" in a table which support multiple versions of a row (MVCC). So maybe using DELETE is actually faster - it just marks a row as "deleted in transaction X". Unless I forgot something.
sgarland 6 hours ago [-]
That's how Postgres' (and perhaps others) MVCC works, yes. MySQL / InnoDB, however, updates tuples in-place [0], and uses the undo log to recreate older versions as needed.
I came here to make a similar comment. It seems like TigerBeetle was built for exactly the type of transactional processing this post is about?
progx 11 hours ago [-]
Why is it so hard for many people to accept, that this is a solution for a specific problem of shopify? They did not say that Redis is bad and MySql is good. They only a solve their problem.
gregoriol 11 hours ago [-]
Even the good teams may make bad choices, that's why it is interesting to read their ideas and discuss them
newsoftheday 5 hours ago [-]
The high contrast dark theme made my eyes squint and I started getting a headache within 60 seconds of trying to read the page. The war on light themes needs to end.
arichard123 11 hours ago [-]
I had a client and they weren't to bothered if they sold the last item twice, they would call the customer, apologise, and offer a discount on an alternative and keep the sale.
zhivota 20 hours ago [-]
"But the hardest lesson wasn't about database design. It was discovering that the real bottleneck wasn’t what we were observing and measuring."
Horffupolde 20 hours ago [-]
But was it load bearing?
CoastalCoder 19 hours ago [-]
Even better.
It's web-scale.
ares623 18 hours ago [-]
load = bearing
gun = smoking
insight = key
gap = closed
summary = executived
18 hours ago [-]
KingMob 18 hours ago [-]
belt = suspended
jtbaker 17 hours ago [-]
boot = strapped
tweakimp 15 hours ago [-]
Until now I thought it was boots-trapped. I am not a native speaker :)
crabmusket 13 hours ago [-]
The idiom is to "pull oneself up by one's boot straps":
It's honestly weird Claude converges on this language because it's incredibly wordy and hard to parse.
One would think semantic density would win out in training.
grey-area 10 hours ago [-]
Claude doesn’t think or have a goal like correctly summarise a topic, it just generates text based on the corpus and training. So not really weird that it is sometimes vague and sometimes incorrect.
Training could not address semantic density unless it was for the very specific pattern you need for this very specific article.
I guess the people generating the article don’t really care if is correct or easy to read, as long as it gets indexed in Google and linked by prominent sites like HN it has done its job (seo).
true_religion 17 hours ago [-]
Why? This is a common transition that people use in speech and text.
Close out previous paragraph. Segue to completely different topic.
How else are you supposed to go on a tangent?
Foobar8568 12 hours ago [-]
I feel that Claude converge to more claudism while Chatgpt sounds more natural.
Both sucks in French but at least chatgpt prose is readable while Claude is awful.
peyton 19 hours ago [-]
Who knows. I wish ant harshly penalized speaking litotically because it’s essentially reward hacking as it can often be read multiple ways.
It’s also annoying as a human because Claude et al rate their own writing very highly, putting human<>LLM interactions at a disadvantage to human->LLM<>LLM interactions.
Jolter 14 hours ago [-]
Thanks for teaching me the word ”litotically”!
sroussey 14 hours ago [-]
Each version minor version of Claude has its own preferences for vocabulary.
MagicMoonlight 11 hours ago [-]
[dead]
jasonlotito 18 hours ago [-]
[flagged]
nozzlegear 18 hours ago [-]
It's not hard to parse, but it's a dense pair of sentences that say nothing. It just pads the length of the article and gives readers mental fatigue trying to read between the lines to figure out what the point is.
jasonlotito 3 hours ago [-]
> It's not hard to parse
That's what I said. Full stop.
> but...
a bunch of stuff that just pads the comment agreeing with me but adding nothing else of value.
CoolestBeans 18 hours ago [-]
I actually don't think this article was LLM generated but these two sentences suck. I think they were moved from another part of the article without being modified.
First, "the hardest lesson". What lesson? It is out of context. Nobody was talking about lessons before this.
Second, "the bottleneck wasn't what we were measuring and observing". Of course the bottleneck itself wasn't that. They couldn't discover what the bottleneck was using the information in their measurements and observations.
It is a clunky and frankly incorrect passage in an otherwise well written article.
jasonlotito 3 hours ago [-]
> but these two sentences suck.
Sure. But that's not what's being discussed here. People are poor writers. People write, and they barely edit. But hard to parse? No, not at all.
ares623 17 hours ago [-]
Sure, it's not technically hard to read.
But it suuucks, making it hard to read, the same way (some) fast/junk food is hard to swallow.
They have access to a trillion dollar writing machine god, and they choose to publish that.
jasonlotito 3 hours ago [-]
Nope. Hard to parse. And it's not hard to parse.
> fast/junk food is hard to swallow.
Lots of people easily eat it so this doesn't make sense. Someone, like our friend above, might suggest this makes what you just said hard to parse.
> They have access to a trillion dollar writing machine god, and they choose to publish that.
Do they? Please, share with me the trillion dollar writing machine god. None of the LLMs I know of would be considered a writing machine god.
firasd 19 hours ago [-]
Makes sense... if you are counting something in MySQL and now your counter is in Redis that's already strange
But I guess the point is that even in the MySQL scenario the 'reserved_quantities' is almost like a temporary table so either way is not the 'Real' inventory
mrloopex 17 hours ago [-]
This is absolutely fascinating. I enjoy real life stories like this. I went to a Node meetup in 2013 when Target had just switched to Node from PHP and it was a similar experience to see their metrics and hear their strategy.
azuanrb 10 hours ago [-]
Pretty interesting read. One thing I’m curious about is the DB size trade off. Going from a quantity in Redis to one row per reservable unit seems like it could create a lot more rows, even with the 1,000 row cap per item/location.
pjmlp 11 hours ago [-]
I never spent much time with the whole NoSQL movement, it always seemed something out of people that don't get how to optimise SQL queries, or suffer from SQL allergy, only to reinvent it badly in custom languages.
williebeek 7 hours ago [-]
There are many good reasons to use NoSQL instead of a "full SQL database". OTOH I can relate to your experience, I remember colleagues switching to MongoDB because they couldn't get good performance on the (MS) SQL database. They didn't know enough about proper (multi-column) indexes, use proper isolation levels, etc.
szundi 11 hours ago [-]
[dead]
pythonRon 10 hours ago [-]
Tobi Lütke also made some rather controversial statements lately, too, agreeing with a retired TD Bank CEO that more votes should be given to the rich. On the surface, this sounds awful, but what Eric Thor actually said was that the number of votes should be tied to the amount of income tax a person pays. Considering that (from what I've heard) billionaires pay no income taxes, I'd say it's not a bad idea. No tax: no vote.
srcreigh 18 hours ago [-]
It’s fascinating that in order to do this, they had to remove 50% of reads and 33% of transactions from the main DB.
kennywinker 20 hours ago [-]
Shopify’s founder and their coo both fund far-right extremism, and its founder thinks only rich people should be able to vote. But anyway, they switched databases.
The link you shared is just a podcast and does not contain even contain “far right”. Can you provide specific concerns, otherwise I don’t see anything wrong with a CEO of the most successful tech company should not be concerned about a horribly performing country from a GDP perspective.
kennywinker 16 hours ago [-]
I included a link to a podcast because that's a good high level overview of the topic. If you want to dig into it further, the podcast episode page I linked to has links to more reporting and plenty of keywords you can type into google.com if that's not enough.
Anyway, if you think a country having a low gdp per capita is how you measure if it should suspend voting rights for disabled people and stay at home parents, then I suspect you're not actually reading any of this.
croes 17 hours ago [-]
You should be concerned what they spot as the problem and what they propose as the solution
The problem he was reacting to (local decisions dominated by people with locked-in benefits and low future exposure) is real. His idea there is silly but I think treating a couple of provocative X replies as some far right agenda is its own kind of overreach.
stiltzkin 20 hours ago [-]
[dead]
hdndjsbbs 20 hours ago [-]
Yeah it's an awful place to work unless you're a far-right bro. My old director used to use slurs and vape in the office. The founder hires pro gamers with no technical expertise because he thinks they're cool.
derwiki 20 hours ago [-]
Are you implying that vaping is far right?
kennywinker 19 hours ago [-]
I think that was part of the “bro” bit, not the far right bit
nozzlegear 18 hours ago [-]
Sounded to me like they were saying people vape in the office, which would make a bad work environment on top of the far right bros.
hamdingers 8 hours ago [-]
> people vape in the office
It's famously a fully remote company.
nozzlegear 17 minutes ago [-]
Shopify has employee "ports" (company jargon for employee workspaces/offices) in Toronto, Ottawa and New York.
kennywinker 5 hours ago [-]
Do you think that means they don’t have offices?
chucksmash 18 hours ago [-]
Don't get shy now. Which slurs?
hdndjsbbs 11 hours ago [-]
Let's say some racial ones (including hard-r n-word) and some ableist ones? Guy also talked about jerking off a lot and joked about rape.
KingMob 18 hours ago [-]
Why are you trying to get someone to repeat slurs?
chucksmash 17 hours ago [-]
Because "slurs" is vague and covers a wide variety of utterances that run the gamut from unprofessional to unemployable and I had a tingle of Spidey sense that OP might have chosen the vague phrasing specifically to inflate the sins of the nameless Shopify director in the mind of the reader.
kennywinker 17 hours ago [-]
Which slurs are ok, and which aren't?
chucksmash 16 hours ago [-]
Great question. It's a vague term and a moving target that not everyone agrees on.
For instance, retarded is now considered a slur. If OP is saying "my director said the R word," then I would question whether OP spared the gory details out of concern for polite company or if they're being oblique in service of their point.
And if managers at Shopify are dropping N-bombs, I want to know that too.
This section is badly written. For example, it refers to different table names than those previously introduced.
The slop shows. While I appreciate the post, I wonder why they didn't bother using an LLM in a way that would at least ensure internal consistency.
eviks 12 hours ago [-]
What way would that be?
jdw64 13 hours ago [-]
Is it really the right choice to drop Redis and go back to a disk based relational database just to wrap transactions into a single unit?
Redis handles tens of thousands of concurrent connections in a single event loop, while MySQL uses one thread per connection. No matter how I look at it, that seems like a step backward.
Of course, performance isn't everything. And if performance isn't a problem, having everything in one place does make it easier to reason about. But I'm worried that under spike traffic, this approach might actually cause more problems.
I think putting a scheduling layer in front of the DB would be a better approach. The application server could handle concurrent connections and only write to MySQL when correctness is actually needed. That seems like a cheaper way to do it. but is it different for large-scale enterprise distributed systems?
codedokode 12 hours ago [-]
Redis doesn't have transactions and persistence.
No persistence means the data gets lost if machine shuts down or process crashes. Furthermore, after restart you will need to regenerate the data which can take time. That's why Redis is a cache and not a database. You can fix the persistence issue (Redis can write WAL log, don't remember if it does fsync or not), but then Redis won't be able to handle those thousands of concurrent connections.
Redis (and other NoSQL storages) don't have some magic architecture that gives them advantages over SQL databases. They just cut corners on ACID guarantees and skip fsync. Once you start doing fsync, your transaction throughput will drop to SQL database level.
Redis also doesn't have transactions which means every app error damages the data. You will spend engineer hours investigating and fixing the problems. Transactions save so much time and worries.
el1s7 12 hours ago [-]
Everything you said is incorrect. Redis does have transactions, as well as data persistence. All cloud providers provide managed redis instances with automatic backups as well.
codedokode 10 hours ago [-]
I mentioned that Redis can write changelog (called AOF in the docs [1]). However, if you tell it to do fsync on every update (like SQL databases do), it stops being that fast and spends time waiting for the filesystem.
Furthermore, the RDB snapshot mechanism (when Redis forks and forked process writes the snapshot) can double memory consumption and cause thousands of page faults in Redis process if there are many writes happening.
The docs contains corresponding warnings. "Cloud backups" are marketing terms and not ACID guarantees.
As one more disadvantage, Redis has no SQL and you cannot easily view the data.
As for transactions, indeed it seems to have them, but their execution is serialized, i.e. when MySQL can prepare 100 transactions in parallel, Redis will execute them sequentially.
Does Redis become that slow when you enable both AOF and RDB? Sure, there's a write cost, but it doesn't lose its ability to maintain tens of thousands of connections. Redis supports AOF and lets you choose the fsync policy.
But I think using only MySQL is unnecessarily expensive, just to get single transaction tracking for bug tracing. So the article's argument seems to be:
'Use only MySQL as a solution to the distributed transaction consistency problem between two different storage systems, Redis and MySQL!'
But I think using Redis is much more elegant. It's easier to scale. I'd even argue that something like Saga would be a better approach. Of course, we might just have different opinions. But in my experience, reducing layers always ends up making things more complicated in the long run.
p.s. We have different views, but I do think some of your points are valid, so I upvoted your comment
codedokode 10 hours ago [-]
The fsync policy equivalent to SQL database would be "fsync on each change before reporting successful update to the app". Redis (as many NoSQL databases) also doesn't have SQL and is a pain to view the data, you need to write extra tools when investigating the problems.
RDB snapshots can cause multiple page faults due to use of fork() and CoW.
> It's easier to scale
The company in question manages online stores and they could easily scale by allocating a separate database for each store (sharding).
> But I think using Redis is much more elegant.
I cannot agree because I think using a single database for all the data is more elegant, than multiple different databases and there are less problems to deal with. I dislike microservice-style architecture strongly and believe it is mostly good for wasting company's money.
> 'Use only MySQL as a solution to the distributed transaction consistency problem between two different storage systems, Redis and MySQL!'
I read it as "do not create unnecessary work by using a single database".
jdw64 10 hours ago [-]
It's interesting that our views differ. I think it's because of our different experiences. I believe MSA is the right approach. But this doesn't seem like a debate that can be resolved through discussion. It's something that needs to be implemented and tested.
Still, I respect your perspective and your experience. We clearly have different values, but I think you have a mature engineering mindset. Ultimately, I think only real measurements can settle this. Have a great day.
nurettin 15 hours ago [-]
At that revenue, why not make your own filesystem, database and index structure? There is no way mysql is the best possible software for this use case. Why stop innovation and hand everything over to ops?
znpy 14 hours ago [-]
Most likely? Time.
Using off the shelf software means you mostly design how to plumb things together and how to make them correct , safe and scalable.
The things you mention, on the other hand, carry the same requirements but are also much complex to develop AND to maintain.
kgeist 11 hours ago [-]
The social network VK internally uses highly specialized database engines per business domain. They don't use stock DBs. They have a DB engine for posts, a DB engine for likes, etc. They have a team of DB engineers. Their DB load was around 250 mln RPS 3 years ago. Stock DBs were harder to scale for them. I guess if you have immense highload, having a team of DB engineers can be cheaper because you can save a lot on servers. I reviewed their code. A DB engine's source code is pretty compact and simple (relatively speaking) because they deal with very specific domain entities, so they don't have to account for all the possible user query combinations that a general-purpose DB would have to support. It was mostly shards+binlog+snapshots+views in RAM. Considering that Telegram was founded by former VK engineers, I suspect they have something similar.
Pair programming and forced AI, that sounds like absolute hell. Glorification of Lütke who didn't do that much in open source and now props up his ego by thinking "AI can do it so it wasn't all that difficult all along."
I don't think he ever worked on complex parts of Ruby. The people he now oppresses did.
Ruby should note that this company is actively repelling people from using the language. I really want to switch, but then I see Claude contributions in Ruby core, the influence of this slop company, and think it isn't worth it.
Oh, and they bought DHH in 2024 for his 180° turnaround on AI. He is now an AI booster, so Rails is out of the question as well.
culi 19 hours ago [-]
They were really so proud of that AI image that they just had to tack it on at the end? Did nothing but make the blog post feel like cheap mass produced slop
nozzlegear 18 hours ago [-]
This is Shopify, the leadership is full steam ahead on AI in a big way and they review employee performance based on AI usage.
throwatdem12311 17 hours ago [-]
And Lutke is a fascist.
All the biggest proponents of AI seem to be fascists.
Weird.
tayo42 18 hours ago [-]
The blog probably was.shopify was pretty early and publicly all in on using AI for everything
xdotcommer 7 hours ago [-]
[flagged]
welcomezhangjun 11 hours ago [-]
[flagged]
tailscaler2026 20 hours ago [-]
[dead]
cloudie78 12 hours ago [-]
[dead]
skullone 20 hours ago [-]
[flagged]
trueno 20 hours ago [-]
so this is interesting to me, im in retail i work closely with platforms ive used shopify ive used magento ive used smaller players ive helped implement various pieces of all of them.
and i was excited to get some insight, then i realized that this whole thing was written by AI and im going to guess the idea and implementation were probably very AI driven.
> The solution: SKIP LOCKED
> Core idea: one row per unit, bounded by design
cool, thanks claude.
Now I'm wondering what the engineering culture is even like at shopify.
Here's the thing. I like databases, I think there's a lot of shit in this space that went and smoked a shit ton their own good stuff to come up with these pure event driven designs that lock you into event workflows with no isolation and remove the ability to do broader bulk-functions.. and then do something even stupider and say "all you need for the interface is graphql" and such service/platform doesn't give you any other way to reconcile or do reporting for your org you have to warehouse from graphql.. this is crap. So seeing a headline where shopify says they want to kinda get behind a unified database strat behind the scenes even if it's not necessarily customer facing, like that's good imo. SQL is many decades of relational algebra that makes insane computations acrossed vast sets of data pure magic and one of the best query dml interfaces of all time.
..however i dont even agree with the claim their making here that redis isnt the tech for a reservation system. redis when used correctly feels like an insanely awesome way to do a reservation system, i lurv redis for stuff like that.
I'm just gonna go forward with the assumption that current and future shopify updates are pure vibeslop. I already hate their data interfaces, but compared to other saas offerings i appreciate that they do have bulk-features.
akamaka 20 hours ago [-]
I found Shopify’s post very easy to read, and learned about some features of MySQL. On the other hand, I didn’t get any value from reading your comment. You seem to have a bunch of opinions about how things should be done, but haven’t given any details about how you came to these conclusions.
trueno 17 hours ago [-]
I've done multiple large scale implementations with shopify paired with many flavors of order managmeent systems as well as competing offerings in the space. The only one i haven't touched that i'd like to get my feet wet with is commerce tools.
I really just disagreed with the assessment that redis is not good enough for the job for a reservation system. I use sql database all the time, I prefer them. But I'm seeing a claude written article here that seems to heel turn on a proven technology, it would at most be insightful if there was human content in here from actual engineers at shopify who want to vouch for and explain the challenges they were up against with redis rather than just expect me to take claudes word for it. Anyone who's been dabbling with AI knows damn well that you can convince claude to write up a dissertation on any hill you want to die on.
dawnerd 17 hours ago [-]
I found it really hard to read, the llm-isms are just too distracting. Does no one proof blog posts anymore?
benmmurphy 19 hours ago [-]
You should be able to do these increments/decrements in a database at the rate you can write WAL to the disk. But the problem is in a lot of these databases the transaction will hold locks until the WAL hits the disk which causes a massive serialisation problem when you have lots of writes to the same row.
For example if it takes 20ms to write a batch to the WAL then if you do 5 updates to the same row then that is a minimum of 100ms. But without waiting on locks if you can batch all the WAL writes together then this could be just 20ms.
I don’t think holding locks while waiting for WAL is strictly necessary. There is definitely some anomalies that can happen if you don’t wait for WAL to be durable because transactions that don’t write WAL can observe non-durable writes in some situations. So for example conditional updates that don’t perform work. But I assume this can be fixed by making these wait on the commit for dependent transactions to become durable if they are empty. There is also the problem of failing writes that reveal information about non-durable writes which is more tricky. For example you try to insert into a unique index and it fails, but the duplicate was due to a non-durable write that is lost.
Pure reads should be fine when using MVCC because you just show the latest durable version of the DB. I know some other replication systems will run all transactions including reads through the WAL/replicated log in order to not have anomalies.
They do heavily use AI, but you haven’t refuted their point that if inventory is in SQL, storing reservation in a second storage system increases complexity.
trueno 17 hours ago [-]
I mean I'm all for everything collapsing into sql. SQL all the things. Not really against it, I'd just rather not-AI write the challenges they were up against. It seems like these are all very behind-the-scenes scaling issues they faced, so it'd be cool to hear from them. Redis has great qualities, I don't use it often but I've also for years now understood why Redis was put in front of these use cases to handle them. Complexity be damned generally you're trying to enforce a first come first served or some level of idempotent behavior, so if sqls doing that now then hell yeah. It's just super off-putting to try and upend an important design pattern with an AI written article.
jbird99 20 hours ago [-]
The lengths companies will go to avoid running different pieces of software...
matwood 14 hours ago [-]
Most companies would be best served picking MySQL or PG and only adding something else if absolutely necessary. Every piece of software added increases complexity.
anonymars 20 hours ago [-]
It can be easier and cheaper to solve problems via technology changes than operations and people
Now you only need MySQL expertise and maintenance rather than Redis and MySQL
kirici 14 hours ago [-]
The default should be that every additional piece needs to be justified
https://share.gemini.google/WxEi6Satj7mY
Happy to share the gem with anybody who's interested in reading articles like this without having to muddle through all the AI bulk.
As software engineers we are constantly told that we need to heavily use these tools for our daily work. So is it surprising that software engineers use the same tools as writing aids? Using AI does not mean no human effort was involved.
For example this subheading:
> "The real bottleneck: connections, not CPU"
That's two AI smells. AI likes to say vacuous punchy statements like "the final takeaway" or "here's the rub". Then contrastive parallelism "connections, not CPU". Contrastive parallelism should scarcely exist in technical writing, regardless of whether it's AI generated.
I wonder if the average reader is also getting annoyed like this or whether they just don’t care - especially seeing what seems to get upvoted on your run of the mill social media sites. They probably collectively shape things more than I do.
They can change it but it will be there.
I do like that I have a term for that “this, not that” phrasing that’s nails on the chalkboard for me after several years of reading slop.
Aren't the LLMs trained on a massive corpus of human written texts? If that stands, then they are doing what they were asked, kind of? I am also not a fan of the fluff words that Claude pollutes the context with, it's a bit much, I could live 30% less but it doesn't want to read its instructions in the .md
If the author wanted, they could ask their "AI" to write the article as a caveman, no? I, simply for fun, slop coded up a skill so that I can ask Claude to reply as Samual L. Jackson, or Kramer, etc.
I also wonder if there is an "AI language barrier", likely not in this particular article, but let's say there was a published article by a researcher whose native language isn't English. What happens with the translation? I realize I am asking naive things :)
I think you might be interested in reading in training data generation, training, and post training papers/articles. I think you might be surprised at how much intervention there is on some of these levels.
As I understand it, there is (or was) a step where they ask people what's the 'better' response.
These linguistic forms sound good the first time you hear themz even if they are rare in real speech, so rapidly got trained in.
Now they distill off previous models, I imagine these weird linguistic forms are quite hard to get rid of.
Is this post AI generated?
Your worst complaint has nothing to do with the overall content or accuracy of the post.
Just style bashing.
- Our oversell protection was a gross hack that broke ACID.
- MySQL has a new feature that allows us to remove parts of the gross hack.
- Question: To what extent does the hack still exist?
Instead you get garbage like "The answer is often in the plumbing, not the engine." and "Crucially, this wasn't about making reservations fast. It was about making them safe neighbors."
But this crap is what Lütke wants, so they deliver.
I think we need the moderation team to set up some policy. In which direction idk.
https://www.theguardian.com/technology/2026/jul/31/ai-labels...
Spotify's defense will then be that this post is obviously AI assisted garbage.
Oversell happens because Shopify doesn't decrement inventory until payment is confirmed. And they apparently would rather die than change that invariant.
I could see a historical moment where Shopify, sans that invariant, massively fucked up inventory counts from uncompleted transactions... then had to unwind all that when a bunch failed to clear payment.
Ergo, now there's an invariant.
This is a really strange, clunky way of writing.
The training these LLMs got, was from endless human-slop, on sites like LinkedIn, and marketing copy, everywhere. In fact, don't be surprised, if we start learning from AI; reversing the process.
But I think that it's only a matter of time, before almost everything will be at least touched by AI.
I posted this, yesterday[0]. It wasn't a particularly popular comment, but I stand by it.
[0] https://news.ycombinator.com/item?id=49219990
The thing is that I LIKE writing code. It's fun.
I don't LIKE coding with an LLM. It isn't fun.
We can't go back to a place where the professional code is all hand written. But at the same time I'm sad and tired. The joy has been utterly sapped and I don't think I'll be able to do this as a career for much longer.
That may have something to do with the way that I use an LLM.
I use it as a “pair partner,” not as an agent.
Most of the code in my apps is mine, but I do ask the LLM to take care of some of the “overhead” tasks, such as minor utility functions.
In fact, I just got done, ripping out a bunch of code the LLM wrote, in my current app, and replacing it with my own work. It seems that I can’t really ask the LLM to do much beyond function level, without it going sideways.
I simply can’t understand people shipping entire, vibe-coded apps. If I let the LLM write my whole app, it would be hot garbage.
EDIT
I also read your linked comment and agree 100%.
Every time this claim is made I ask for a link to a pre-2022 blog or article that has all those AI tells.
I'm still waiting for it. I don't think I will ever receive it.
People on the Internet (especially here), seem to think that it's OK to make rude demands of complete strangers, and expect them to spend considerable time, meeting them.
I'd rather some rando think they "won the Internet," than spend a bunch of my time, working for free, just to satisfy them. I have better things to do with my time, and I don't really have a lot of investment in what complete strangers think of me.
Sometimes, the fox ain't worth the chase.
I love this place. We are such class acts.
Have a great day!
Interestingly, ChatGPT was able to parse and explain it to me. I got a lot out of its presentation.
I would rather read the original thoughts of the engineer, grammar mistakes and stylistic imperfections included, than prompt generated AI slop.
Ok? Who cares. This was very interesting to me. IDGAF if AI was used (or not) to put the words in the article, the bits of information in those words is why I'm reading it. It's almost 2027, people need to accept that articles, art, videogames, code, everything will be more and more "AI touched" and you can't stop it.
I mean I share all the people's worry about AI taking jobs and a handful of AI (and AI adjacent) corps sucking out massive amounts of resources, but reading under every 2nd or 3rd article how it's AI is beginning to be more annoying than the AI writing itself.
It's like now every interesting or weird or exceptional or controversional video of anything has (what seem to be) 80-100 IQ people spamming "AI" in the comments instead of making interesting observations, posting their own anecdotes, cracking jokes etc.
So I'm sort of curious, did you not like the AI writing style or just rejecting AI in general? I myself am very conflicted, I think AI in its current form is the wrong tech at the wrong time yet I don't mind reading AI text and sort of mooch off of googles free tier.
Also I'm starting to notice lots of AI smearing, I read folks online describing someone elses contribution as "obviously AI" and I'm suspecting in some cases these could be false accusations.
I'm thinking about starting a blog, so when my AI gf writes posts, do I ask her to try to "not read like an AI"? Anybody try that? I'm gonna try that. Hehe for all you know I already did hehe
Obviously the blue pill.
> But one row per unit for all inventory would break down at scale—an item with 50,000 units across 10 locations would mean 500,000 rows, and the reserve query would slow as it scans through them. Instead, we maintain a bounded pool of available rows, capped at 1,000 per item/location combination. Reservations consume rows from this pool; a replenishment process refills it from the inventory ledger.
Shouldn't I feel uncomfortable with such approach? It seems to create a backoff (pool) for lowering the chance of having a synchronization issue.
I'm skeptical of this approach. Sure, row contention means that you cannot have a database transaction per customer order attempting to decrease inventory count by 1 each time. But you can have a batch transaction whereby the transaction decreases inventory by 100 (thus touching the high-contention inventory row once) and credits each of 100 different customer cart database rows (which are not under heavy contention and can be on a different disk entirely). Attempted customer orders are submitted to a reservation system put in charge of assembling the batches. Customers wait some short period of time - say, 15 seconds - for the reservation attempt to be batched and to be notified that they successfully locked a reservation. Arguing that "slow reservations trigger throttling and a worse buyer experience", without an actual number for what counts as "slow" to serve as an SLO and as a design target, is a cop-out inviting over-engineering.
They explicitly cover this in the article, saying they do reservation inline. It does increase latency for these orders, but it doesn’t result in an error
> But you can have a batch transaction whereby the transaction decreases inventory by 100 (thus touching the high-contention inventory row once) and credits each of 100 different customer cart database rows (which are not under heavy contention and can be on a different disk entirely).
They mention this as well, checkout batching increased implementation complexity.
> Arguing that "slow reservations trigger throttling and a worse buyer experience", without an actual number for what counts as "slow" to serve as an SLO and as a design target, is a cop-out inviting over-engineering.
True. The article would’ve been better if they included such numbers. However the fact that they didn’t mention this doesn’t imply they haven’t done research. I haven’t found anything related to checkout specifically, however there are in general articles, indicating that increased latency correlates with revenue drop.
I'm a bit surprised about the scalability case against a simpler solution. This is not about Shopify's scale. We're talking about contention for a specific SKU of a specific seller at a specific warehouse location.
How many shopping carts are competing for a single SKU at the payment stage at peak hours? Can this really be too much lock contention for a single database row?
I realise Shopify engineers are neither stupid nor inexperienced. Hence my surprise. I would have liked to hear more about that specific problem.
When you consider how many other (often poorly-crafted) DB queries and service calls are being performed while holding the row locked, yes, it can rapidly add up. If the entire pipeline takes 500 msec, congratulations, you can’t sell more than 2 units per second of that SKU. If the item is popular and being actively hyped, then yes, that can be a problem.
Basically this is an incredibly rare case but a feature that they want to support.
Seems like a lot of places can do it for it not to be a solved problem.
The other thing that bothers me - why not real stable, but maybe „too old, medieval” solution with Redis as the main source of through - without any sync with SQL at all in terms of stock… it worked in my previous job with much higher traffic (1000s/s). Yup, we ended up with app-side sharding, but it was stupid-simple.
As they say in TFA they used Redis for the cart reservation but then you need to sync the two stores.
You are just spending some more disk space to avoid synchronization issues. Denormalization for performance is a really common pattern, just that people do not start with it in the first place itself
I don't get it, Wouldn't that still only be 50,000 rows, just divided across locations?
Its called one row per shopping cart*SKU combo.
if two people order 100 and 500 items of the same SKU, respectively, the table should have only two rows: for order1 and order2. Not 600 rows.
Also you might not understand the original problem. Imagine if 100 customers want to buy product A. One thread starts a transaction, searches for amount of product A and UPDATE's it and goes searching for other products. The database locks the row until the end of transaction and other 99 treads cannot continue until first transaction commits (they can read but cannot update the rows).
This is why they made a row per item. In this case, transaction 1 hopefully locks only several rows with items of product A. Transaction 2 instead of waiting for lock release skips them (due to SKIP LOCK) and locks several next rows. And so on.
Obviously you do not need to make a row per item - if the available amount is really large (10 000 items), you could have for example 100 rows having 100 items each. In this case each transaction locks the whole row (100 items) even if it wants to reserve just one item. The problem though is that now every row might have different amount of available items and you have to do more work to reserve the amount you want.
in shopify's design, it is a user who was the first to lock the row and have successful payment. Sounds good, but how often does it happen ? It's a rare and extreme case and they model their entire system after the rare even, and incur the overhead of 1000 rows per SKU per shop for all combination of SKU and shop_id for all the normal items that are not sold out in flash sale.
the same outcome could be achieved without locking and without creating 1000 rows:
In case with shopify, they want to decide whether the user may place order or not, at the moment when the user clicks "Pay" or some other button. If the user cannot place an order, they are shown the error, if they can, the items are reserved and the user is redirected to the payment page. So payment is processed only after successful reservation, and reservation is made only if the user wants to pay. The similar system works for buying train tickets online in my country, for example.
In you case, when user A clicks a button, following happens (as I understand):
1 the server increments the counter
2 the server calculates available amount as (amount_in_stock - amount reserved by carts with time < counter)
3 if the amount is large enough, the server updates the "time" field for user's cart thus reserving the item
Imagine that at step 2 the user A sees that there is one item left. However before user A does step 3, another user B might reserve the item (complete all 3 steps), and proceed to the payment. Then user A then completes step 3 and proceeds to the payment too. Now we end up with both user A and B paying for the last remaining item which doesn't solve the stated problem. Shopify's solution doesn't have such issues.
This is a classical TOCTTOU situation. There were exploits against Linux kernel based on similar issues.
but again, my idea was:
> do not use throwaway placeholder rows to imitate a single item
The point of using multiple rows for one product is to distribute the locks.
The point of one row per item is that thousands of concurrent shoppers don’t need to block each other as they can each claim as many free rows as they need for themselves?
Developers (and everyone else really) need to think about systems, with the system taking inputs like "number of locations", and producing outputs like "available inventory", and when the input parameters change outside of the designed scope, without the system itself changing, then you should expect things to break.
1. Deduct the reservation from the inventory when the user starts to order, but in the same txn also maintain a separate row for the in progress order flow. 2. If the order flow is aborted or times out have a background process that returns these to the inventory.
That seems simpler than this approach and involves no locking. Though their presented approach is also reasonable, there must be some reason not to choose a simpler flow. It is not that difficult to have a gc service that scales, but may be they didn't want to separate that.
I don't think you really need that even. An indexed lookup is fast and you don't need to store a computed quantity generally.
1. Backgrounds process can back up
2. They need context of the user and need to switch context per user
3. What if they fail, you create some DLQ or another process to handle the failure
4. Who looks on those failure and how do they act
TLDR; there is always a cost
I disagree with the other posters about the bg process, if you have any bg processing already you should be able to handle the few edge cases without too much trouble.
Perhaps each shopping cart would have its own workflow, and the inventory item would have one as well. Then, whenever a customer put an item in their cart, their cart workflow would send a signal to the inventory item workflow and wait for the response. The inventory item workflow would maintain a ledger controlling to which cart each unit goes, and it could batch the writes to this table. This way, even if 100k customers try to purchase the same item in the same second, it should handle the load.
After the batch is written to the ledger, the inventory item workflow would reply signals to each cart workflow confirming that the reservation was completed. The end-to-end latency from the consumer point of view would be a fraction of a second, without needing the 1000-row hot-inventory heuristic.
Instead of having 1000 rows per shop*SKU, why not just have one row per shopping cart*SKU?
That way a single row would represent a single cart, and will hold info of multiple items of the same SKU.
No need a cludge with 1000 rows limit and replenishment process. Instead of dealing with N rows, you always deal with a single row.
So those engineers at Shopify worked hard for months on a more performant system, but they missed the obvious structure? They chose a complex denormalization for no good reason?
It may be true, but I think it's presumptuous to belittle their work when we have only partial information. My guess is that they had good reasons to think that the more obvious ways would not scale.
And from reading your comments in this thread, I believe your structure would fail at their scale. A SQL query that uses 2 sub-queries with "group by" is probably too heavy. From the post, at peaks there would be millions of active shopping carts.
BTW, I suspect most orders are just for 1 or 2 of each item, so the denormalization is not as heavy as it seems.
re concurrency, it is not a big issue at all. stock exchanges deal with HFT traders and can easily deal with concurrency of orders. Same can be implemented with shopify, but I doubt they face the same level of concurrency as stock exchange anywhere near
you have an open Sell 1 APPL for $100.0. Millions of other HFT orders rush to scalp your single order. How do you think exchange matches your Sell to HFT's Buy orders? which Buy order gets fulfilled first?
I would really appreciate it if you could write this up as an article. It would be an extremely interesting and valuable read
https://martinfowler.com/articles/lmax.html
> LMAX Disruptor
If anything this article shows that concurrency is a big issue. It is such a big issue you have to write in-memory single threaded processor with custom journaling. If you have established workflows with MySQL and a team knowing how to work with it, throwing all that to do LMAX is not cost efficient. While there are domains where such approach is suitable and even required due to strict transaction ordering, Shopify case doesn’t look like one of them.
There's even this bit where they discover a remarkable trick:
> Each round trip to the database has a cost. For carts with multiple line items, we batch reservation queries using UNION ALL so we fetch all needed units in one round trip
Insights like that really don't read like senior level output, and of course, it's LLM output. I'm not sure it's presumptuous to question it.
At what point that row is inserted?
so the row is inserted when Payment is initiated, and row is deleted when Payment succeeds
but that system could be easily improved to reserve item when user Adds item to a cart, to prevent scenario when user adds item to a cart, goes through checkout, and after initiating payment gets "soldout error":the check for oversold items is extremely cheap:
assuming there are indexes on sku field in both, results in efficient index seek and agg over 2 tablesYou reserve the product by creating an "active_cart" entry. Your solution has a problem, that when you run the check, it might say the product is available, but before you create an "active_cart" to reserve it from thread A, another thread B reserves it and you end up reserving a product that is not available anymore. You end up with SUM(active_cart.quantity) > inventory.available_units.
That is exactly why the database has locks - to prevent this situation. With locks, thread A decrements inventory.available_units and that row is locked until the end of transaction. Other threads (if they do SELECT FOR UPDATE instead of SELECT) cannot see the old, invalid value until thread A either commits and the value is updated or rollbacks. However, locks cause performance issues and that is why shopify uses the architecture from the article - instead of 100 users fighting for the lock on the same row with available amount, each user locks only rows with units they plan to buy.
Interestingly, MySQL docs has the documentation page with a similar case: https://dev.mysql.com/blog-archive/mysql-8-0-1-using-skip-lo...
there is ultimately needs to be some global mechanism resolving this conflict. Currently it is an order in which db engine processes transactions by locking rows for a transaction, whoever got the first lock, wins the last remaining items.
my design is the same, except it does not need this dance with moving rows between tables, locking them, and the cludge with replenishment process.
in the simplest form, run the sum() over active non-finished orders and compare to inventory. you get the same result: whoever got the first to run sum() and get positive answer will get the last remaining items.
but the problem as formulated, imho, is not even correctly defined.
Shopify incorrectly formulated the very problem they are trying to solve.
Trying to solve it at the payment time is too late, its better to resolve it earlier, before the checkout.
the "PAY" button should only do one thing: deduct money from cc and that's it. Resolving inventory availability must be solved way earlier, the moment user clicks Checkout, not when user clicks Pay.
So ideally, the error for oversold items should be shown to a user when he clicks Checkout, not when he click PAY
That’s a bold overconfident statement. Cart abandonment is real. People never clear their carts they just walk away
Shopify purposefully chooses to do it at payment time because doing it earlier results in lost sales as people “reserve” items and then walk away causing other to see out of stock and then also walk away
Whoever puts up the money first gets the item
That’s the design constraint they chose you can’t just say “their solution is wrong because they solved the wrong problem”. Each design is a different user experience and I think it’s safe to say they chose which experience they want consciously.
Ok, let's accept the design goal that whoever paid first wins. You can use the same metric (how many milliseconds ago did user click PAY) and impose a global monotonic non-decreasing counter to distribute the scarce inventory. This is how order matching engines work at stock exchanges with HFT orders (FIFO logic).
the goal is to know with 100% certainty, before sending payment request to payment processor, who will have item and who won't, and you dont need to move mountains of rows for that.
the payment processor should be just a binary answer: payment succeeded or not, but currently it combines Inventory availability check & payment processing, which is the root cause of confusion. For clarity it is better to make that stage of order processing an explicit separage stage, instead of coupling it with payment stage.
some stores split payment into two stages: Payment and Final order confirmation. at the Payment stage you can pre-authorize money at cc and do inventory availability, and at final confirmation you capture $$
https://docs.stripe.com/payments/place-a-hold-on-a-payment-m...
https://support.authorize.net/knowledgebase/Knowledgearticle...
Still I think their solution is a bit weird. I'd want to commit the reservation transaction with inventory decrement along with a payment key and then use a different transaction to drop the reservation when the transaction completes. If the transaction does not complete in a timely manner you probably need to query external systems anyway to resolve whether the payment actually occurred or not.
They talk about lock contention in this case, but I also wonder about latch contention since these rows are adjacent. If it's a small transaction that's not interactive, does mysql resolve it with just the latches on the needed tables?
It resolves with skip locked. Assuming we have only 1 item left. First query scans the buffer table, locks as many rows as needed (1 in our case), and moves rows to another table. Second query scans the table, finds no rows (even if first one hasn’t finished yet, the row is locked and ignored), checks if it can increase buffer, finds out that it’s fully sold and aborts. Db guarantees that you can’t oversold.
> my design is the same, except it does not need this dance with moving rows between tables, locking them, and the cludge with replenishment process.
I can’t evaluate whether it’s the same or not, because you still haven’t clarified when exactly you’re going to insert the row. In the article they’re inserting in the same transaction. Would you also do it in the transaction? Because if you’ll introduce a separate global mechanism to resolve conflicts, on a high level it would be the same as their approach with redis (you need to have 2 systems)
EDIT: wording
now let's think again, do we need to lock 900 rows to place order on 900 items? or can we insert a single row where order_quantity=900 ?
shopify's design relies on DB to lock rows for transaction as a way to "decrement the counter" of available units. What I am suggesting, is you can just decrement counter by updating a single row, no need to lock 900 rows. Shopify moved from one extreme (single global variable in redis) to another extreme (1000 rows in db) and forgot about the middle ground.
The dance with moving rows per each item between tables is completely unnecessary, it's like counting numbers one by one in a for loop, when you can just substract number directly.
if I were to solve the problem, I would have solved it differently, at the Checkout state, before user clicks PAY. This removes the race condition at the user UI level, before any request lands in backend/db:
availability check is the SQL above: inventory-sum(active_carts.quantity)-current_order must be > 0In order to avoid races you need to insert reservation and decrement availability atomically. Your proposed approach is not atomic. For it to be atomic you will need to lock whole range, to make sure no new rows appeared between the points “check for availability” and “record reservation”. Actors will be effectively competing for the single aggregate row. This is the same as having a single inventory row with quantity field, which they rejected in the beginning of the article
> now let's think again, do we need to lock 900 rows to place order on 900 items? or can we insert a single row where order_quantity=900 ?
In the proposed schema nobody is waiting for these locks, they’re skipped by concurrent queries. In your schema actors would have to wait before they can insert without breaking invariants.
more transactions can commit at the same time, but with one counter they would conflict (as it did in the Redis case)
they should use CRDT (and trying to model that with this 1000 row workspace, no?)
still, eventually at some point they need to do the math
The assumption that a SKU maps 1 to 1 to a cart item is flawed.
If the first item in the cart is a bundle of SKU-A and SKU-B, the second item is a bundle of SKU-A and SKU-C and the third item is 5xSKU-B where do you do you keep the re-agregation of the SKU-X's to track them?
This is without accounting for item location in the reservation - and rules that may apply around that.
You haven't even gotten to the part where different customers will have different rules around shipping from different locations - because that can eat into margins.
You're also making a bunch of other assumptions around transaction flow and where carts are actually stored (and how they get converted to an invoice, with payment attached) that likely do not hold true.
Could you do it more like what you're sugesting -- maybe -- but only in a single tenant system.
Imagine you are at Walmart store and go to checkout stage, you would have to pick item from the shelf and take it from availability for other shoppers, before you pay for the item.
What shopify did, is customer enters the store, heads straight to checkout and retail workers races back to shelves to pick up items for client. Sometimes it says: sorry bud, item is sold out, frustrating customer experience, who is already mentally prepared to pay and own an item.
Re SKU storage, you will have multiple row entries per SKU, if I order two items, there will be two rows corresponding to the items in your Purchase order.
The sum() aggregation check will run across all active orders per sku
Re single tenant: shopify creates 1000 rows per SKU per shop(tenant!!). As long as tenant is on the same DB you can run it, just add shop_id to the group by field
https://baymard.com/lists/cart-abandonment-rate
70 percent of carts are abandoned. You dont want your inventory sitting in carts, when other people want to buy it. It only comes off the shelf (and gets put in a box) when you have money in your hand.
There is an entire ecosystem around Shopify to reach out to abandon cart holders and attempt to convert them: https://apps.shopify.com/categories/marketing-and-conversion...
> Sometimes it says: sorry bud, item is sold out, frustrating customer experience, who is already mentally prepared to pay and own an item.
This is better than A) taking their money and then telling them you dont have it. B) Them not being able to buy it because someone has it in their cart and is NEVER going to check out with it.
But the real world is different
Also, I wonder why they could not have a row status (available/reserved) and UPDATE it instead of deleting the rows.
> Also, I wonder why they could not have a row status (available/reserved) and UPDATE it instead of deleting the rows.
This requires a row per item unit, doesn’t it? If you have 50k units you’ll have to track status of every item, meaning 50k rows. They also mention this as a rationale to use at most 1k rows, and treat it as a buffer.
0: https://dev.mysql.com/doc/refman/8.4/en/innodb-multi-version...
It's web-scale.
gun = smoking
insight = key
gap = closed
summary = executived
https://en.wikipedia.org/wiki/Bootstrapping#Etymology
One would think semantic density would win out in training.
Training could not address semantic density unless it was for the very specific pattern you need for this very specific article.
I guess the people generating the article don’t really care if is correct or easy to read, as long as it gets indexed in Google and linked by prominent sites like HN it has done its job (seo).
Close out previous paragraph. Segue to completely different topic.
How else are you supposed to go on a tangent?
Both sucks in French but at least chatgpt prose is readable while Claude is awful.
It’s also annoying as a human because Claude et al rate their own writing very highly, putting human<>LLM interactions at a disadvantage to human->LLM<>LLM interactions.
That's what I said. Full stop.
> but...
a bunch of stuff that just pads the comment agreeing with me but adding nothing else of value.
First, "the hardest lesson". What lesson? It is out of context. Nobody was talking about lessons before this.
Second, "the bottleneck wasn't what we were measuring and observing". Of course the bottleneck itself wasn't that. They couldn't discover what the bottleneck was using the information in their measurements and observations.
It is a clunky and frankly incorrect passage in an otherwise well written article.
Sure. But that's not what's being discussed here. People are poor writers. People write, and they barely edit. But hard to parse? No, not at all.
But it suuucks, making it hard to read, the same way (some) fast/junk food is hard to swallow.
They have access to a trillion dollar writing machine god, and they choose to publish that.
> fast/junk food is hard to swallow.
Lots of people easily eat it so this doesn't make sense. Someone, like our friend above, might suggest this makes what you just said hard to parse.
> They have access to a trillion dollar writing machine god, and they choose to publish that.
Do they? Please, share with me the trillion dollar writing machine god. None of the LLMs I know of would be considered a writing machine god.
But I guess the point is that even in the MySQL scenario the 'reserved_quantities' is almost like a temporary table so either way is not the 'Real' inventory
https://www.techwontsave.us/episode/340_shopifys_leaders_are...
But here's another link if you need something that includes the phrase "far right": https://pressprogress.ca/shopify-executives-right-wing-media...
Anyway, if you think a country having a low gdp per capita is how you measure if it should suspend voting rights for disabled people and stay at home parents, then I suspect you're not actually reading any of this.
https://www.ctvnews.ca/business/article/shopify-ceo-draws-cr...
It's famously a fully remote company.
For instance, retarded is now considered a slur. If OP is saying "my director said the R word," then I would question whether OP spared the gory details out of concern for polite company or if they're being oblique in service of their point.
And if managers at Shopify are dropping N-bombs, I want to know that too.
This section is badly written. For example, it refers to different table names than those previously introduced.
The slop shows. While I appreciate the post, I wonder why they didn't bother using an LLM in a way that would at least ensure internal consistency.
Redis handles tens of thousands of concurrent connections in a single event loop, while MySQL uses one thread per connection. No matter how I look at it, that seems like a step backward.
Of course, performance isn't everything. And if performance isn't a problem, having everything in one place does make it easier to reason about. But I'm worried that under spike traffic, this approach might actually cause more problems.
I think putting a scheduling layer in front of the DB would be a better approach. The application server could handle concurrent connections and only write to MySQL when correctness is actually needed. That seems like a cheaper way to do it. but is it different for large-scale enterprise distributed systems?
No persistence means the data gets lost if machine shuts down or process crashes. Furthermore, after restart you will need to regenerate the data which can take time. That's why Redis is a cache and not a database. You can fix the persistence issue (Redis can write WAL log, don't remember if it does fsync or not), but then Redis won't be able to handle those thousands of concurrent connections.
Redis (and other NoSQL storages) don't have some magic architecture that gives them advantages over SQL databases. They just cut corners on ACID guarantees and skip fsync. Once you start doing fsync, your transaction throughput will drop to SQL database level.
Redis also doesn't have transactions which means every app error damages the data. You will spend engineer hours investigating and fixing the problems. Transactions save so much time and worries.
Furthermore, the RDB snapshot mechanism (when Redis forks and forked process writes the snapshot) can double memory consumption and cause thousands of page faults in Redis process if there are many writes happening.
The docs contains corresponding warnings. "Cloud backups" are marketing terms and not ACID guarantees.
As one more disadvantage, Redis has no SQL and you cannot easily view the data.
As for transactions, indeed it seems to have them, but their execution is serialized, i.e. when MySQL can prepare 100 transactions in parallel, Redis will execute them sequentially.
[1] https://redis.io/docs/latest/operate/oss_and_stack/managemen...
But I think using only MySQL is unnecessarily expensive, just to get single transaction tracking for bug tracing. So the article's argument seems to be:
'Use only MySQL as a solution to the distributed transaction consistency problem between two different storage systems, Redis and MySQL!'
But I think using Redis is much more elegant. It's easier to scale. I'd even argue that something like Saga would be a better approach. Of course, we might just have different opinions. But in my experience, reducing layers always ends up making things more complicated in the long run.
p.s. We have different views, but I do think some of your points are valid, so I upvoted your comment
RDB snapshots can cause multiple page faults due to use of fork() and CoW.
> It's easier to scale
The company in question manages online stores and they could easily scale by allocating a separate database for each store (sharding).
> But I think using Redis is much more elegant.
I cannot agree because I think using a single database for all the data is more elegant, than multiple different databases and there are less problems to deal with. I dislike microservice-style architecture strongly and believe it is mostly good for wasting company's money.
> 'Use only MySQL as a solution to the distributed transaction consistency problem between two different storage systems, Redis and MySQL!'
I read it as "do not create unnecessary work by using a single database".
Still, I respect your perspective and your experience. We clearly have different values, but I think you have a mature engineering mindset. Ultimately, I think only real measurements can settle this. Have a great day.
Using off the shelf software means you mostly design how to plumb things together and how to make them correct , safe and scalable.
The things you mention, on the other hand, carry the same requirements but are also much complex to develop AND to maintain.
https://www.shopify.com/careers/disciplines/engineering-data
Pair programming and forced AI, that sounds like absolute hell. Glorification of Lütke who didn't do that much in open source and now props up his ego by thinking "AI can do it so it wasn't all that difficult all along."
I don't think he ever worked on complex parts of Ruby. The people he now oppresses did.
Ruby should note that this company is actively repelling people from using the language. I really want to switch, but then I see Claude contributions in Ruby core, the influence of this slop company, and think it isn't worth it.
Oh, and they bought DHH in 2024 for his 180° turnaround on AI. He is now an AI booster, so Rails is out of the question as well.
All the biggest proponents of AI seem to be fascists.
Weird.
and i was excited to get some insight, then i realized that this whole thing was written by AI and im going to guess the idea and implementation were probably very AI driven.
> The solution: SKIP LOCKED > Core idea: one row per unit, bounded by design
cool, thanks claude.
Now I'm wondering what the engineering culture is even like at shopify.
Here's the thing. I like databases, I think there's a lot of shit in this space that went and smoked a shit ton their own good stuff to come up with these pure event driven designs that lock you into event workflows with no isolation and remove the ability to do broader bulk-functions.. and then do something even stupider and say "all you need for the interface is graphql" and such service/platform doesn't give you any other way to reconcile or do reporting for your org you have to warehouse from graphql.. this is crap. So seeing a headline where shopify says they want to kinda get behind a unified database strat behind the scenes even if it's not necessarily customer facing, like that's good imo. SQL is many decades of relational algebra that makes insane computations acrossed vast sets of data pure magic and one of the best query dml interfaces of all time.
..however i dont even agree with the claim their making here that redis isnt the tech for a reservation system. redis when used correctly feels like an insanely awesome way to do a reservation system, i lurv redis for stuff like that.
I'm just gonna go forward with the assumption that current and future shopify updates are pure vibeslop. I already hate their data interfaces, but compared to other saas offerings i appreciate that they do have bulk-features.
I really just disagreed with the assessment that redis is not good enough for the job for a reservation system. I use sql database all the time, I prefer them. But I'm seeing a claude written article here that seems to heel turn on a proven technology, it would at most be insightful if there was human content in here from actual engineers at shopify who want to vouch for and explain the challenges they were up against with redis rather than just expect me to take claudes word for it. Anyone who's been dabbling with AI knows damn well that you can convince claude to write up a dissertation on any hill you want to die on.
For example if it takes 20ms to write a batch to the WAL then if you do 5 updates to the same row then that is a minimum of 100ms. But without waiting on locks if you can batch all the WAL writes together then this could be just 20ms.
I don’t think holding locks while waiting for WAL is strictly necessary. There is definitely some anomalies that can happen if you don’t wait for WAL to be durable because transactions that don’t write WAL can observe non-durable writes in some situations. So for example conditional updates that don’t perform work. But I assume this can be fixed by making these wait on the commit for dependent transactions to become durable if they are empty. There is also the problem of failing writes that reveal information about non-durable writes which is more tricky. For example you try to insert into a unique index and it fails, but the duplicate was due to a non-durable write that is lost.
Pure reads should be fine when using MVCC because you just show the latest durable version of the DB. I know some other replication systems will run all transactions including reads through the WAL/replicated log in order to not have anomalies.
Now you only need MySQL expertise and maintenance rather than Redis and MySQL