First up, who the heck am *I* to have an opinion on this?
Well, I am a computer scientist specializing in deterministic sequences used in randomized algorithms.
Read my peer reviewed published papers on blue noise, read my blog posts on irrational numbers and the benefit of using them as pseudorandom number sources, read what I’ve written about using light weight cryptography to shuffle and group objects in constant time, with random access and inverse lookups using the Chinese remainder theorem. Watch my YouTube videos of the same. I find these things fascinating, so if you are interested, here are some of my favorites:
Beyond White Noise For Real Time Rendering (YouTube): https://www.youtube.com/watch?v=tethAU66xaA
Constant Time Stateless Shuffling And Grouping (EA Blog Post): https://www.ea.com/seed/news/constant-time-stateless-shuffling
Irrational Numbers (Blog Post): https://blog.demofox.org/2020/07/26/irrational-numbers/
Ok so i know about determinism, but do I know about AI and ML? I stick to the bare metal of ML and don’t like to go into the highly abstracted levels that most modern ML people do, but yes. I know backprop, dual numbers (forward mode AD), ADAM, and other random things. During the AI winter, I kept myself warm implementing neural networks in excel (when I wasn’t working on game dev).
Machine Learning For Game Developers (YouTube): https://www.youtube.com/watch?v=sTAqWRsEiy0
I also have an ML patent (pending) I won’t describe here, that is pretty interesting. Unfortunately, I fear it would help something like terminator’s skynet to exist. I tried to park the patent at a game company, thinking that would be safe, but now the American Fascist Party’s very own Jared Kushner and the Saudi prince who ordered the murder and dismemberment of Jamal Khashoggi own EA, and my patent, so RIP to humanity i guess.
Funny thing – a core feature of that patent is deterministic ML.
So yes, I do lift, bro.
The Determinism Argument
Ok so it annoys me that people complain that LLMs aren’t deterministic. I am annoyed because it isn’t that they want it to be deterministic, it’s that they want it to be RIGHT, and they don’t even realize it.
Let’s look at a 2×2 permutation matrix.
Considering each square, “When asked the same question multiple times”… :
- #1 – it gives random answers and is wrong.
- #2 – It gives the same answer every time and is wrong.
- #3 – It gives random answers, but they are correct.
- #4 – It gives the same answer every time, and is correct.
The argument that determinism is the problem with LLMs says that you want option #2 or #4 and are fine with either. You are saying “I don’t care if it’s right, I just want it to do the same thing every time”. If you are thinking this, good job: you are now breathing manually. (Note this feeling: This is what it feels like to be wrong. Do you feel it often?)
What people really want is #3 and #4. They don’t care if it’s random or not, they just want it to be CORRECT.
Any “random variation” within being correct doesn’t make it any less correct. If it did, you wouldn’t consider it correct, and it wouldn’t be part of box #3.
Yes i even am talking about LLMs generating code or executing some set of steps that a python script is probably better suited to. If you are in box #4, you are 100% happy. If you are in box #3, you are still fully happy because what it is doing is CORRECT. If you say “yeah but i need zero variation in the code generation or the utility execution”, ok so that’s part of your definition of CORRECT and you get that. If you don’t get what you consider correct, it is in a different box.
Another Way of Looking At It
You can run an LLM deterministically – both during training, and during inference. Set the seed of all pseudorandom number generators used, make sure nothing runs on CPU or GPU threads in a way that is non deterministic. This is fully possible and not even a big stretch. If you only know vibe coding you might think this is magic, but unfortunately it’s mundane as heck.
Would that make you happy? Of course not. It puts you into box 2 or 4 for every question you ask, but you don’t get to choose which box you end up in. The only variation in the network output comes from your input.
(Oh no, did i forget rounding modes? Go learn how an RTS works before you have an opinion about the feasibility of program determinism across different machines and compilers.)
Yet Another Way of Looking At It
I can hear you saying it:
“But bro – you are mixing up terms. The network represents a distribution in high dimensional space and the output is a single sample. You taking one sample using a a single pseudorandom value as input doesn’t make it deterministic.”
Bro, you are so correct.
Think about square number 3 in my lovely MS Paint ™ diagram. It gives random but correct answers when asked the same question repeatedly. How is that possible? BECAUSE IT’S A PROBABILITY DISTRIBUTION WITH WIDER SUPPORT.
Square #4 has a ~dirac delta as a distribution. That is the only difference.
Yes, I know it’s a distribution, so what.
Yet Another Way of Looking At It
There’s a debate about whether the universe is deterministic or not.
As best as we can tell, it is absolutely deterministic, except perhaps at the quantum level, which somehow seems to get laundered into determinism at large scales. Maybe through the central limit theorem, for all we know.
If that’s true, it is deterministic already so you can relax. Everything is exactly as you say you want it to be. Still not happy? THAT’S BECAUSE YOU DON’T CARE ABOUT DETERMINISM, YOU WANT IT TO BE RIGHT.
Also, maybe take a couple steps back from the kool-aid bowl. That stuff is bad for your health.