Hardware RNG: RDRAND, Entropy Pools, and Why Software Randomness Isn't Enough

Give me 624 consecutive outputs from a Mersenne Twister and I can tell you every number it will ever produce afterwards.

That is not an exploit. It is the documented behaviour of MT19937, the generator sitting behind Python's random module, PHP's mt_rand() and a large fraction of the game code written in the last twenty-five years. Its internal state is 624 32-bit words. Observe the state, reconstruct it, and the sequence is yours in both directions.

For particle effects and procedural terrain this is completely fine. For anything where somebody has a financial reason to predict your next value, it is a catastrophe, and the gap between those two cases is wider than most developers assume.

Deterministic by design

A pseudorandom generator is a state machine. Seed it identically and it produces identical output, which is exactly why reproducible builds, deterministic replays and seeded world generation work at all.

The weaker classical generators fail even without state recovery. Linear congruential generators of the sort still lurking in older C libraries have obvious lattice structure in higher dimensions, and their low-order bits are notoriously poor. Modern non-cryptographic generators like PCG and the xoshiro family fix the statistical problems handsomely. They do not fix predictability, because they were never trying to.

The distinction that matters is not "good randomness" versus "bad randomness." It is whether an adversary who sees output can infer state. And that raises an awkward question: who actually checks?

The one industry that audits this properly

Outside cryptographic key generation, the most rigorous third-party RNG testing anywhere happens in online casino software. The standards are stricter and considerably more specific than most engineers expect.

Laboratories including GLI, BMM Testlabs and iTech Labs test casino games and the platforms serving them against published criteria, with GLI-19 covering interactive systems specifically. The scope goes well beyond running a statistical battery. An audit examines the entropy source and its failure modes, the seeding procedure, the generator's period, the mapping from raw output to game outcomes, and how the system behaves when the source degrades. It runs against the deployed implementation rather than a reference build, and it must be repeated whenever that implementation materially changes.

Certification, rather than any claim about the underlying technology, is therefore the substantive content of any credible guide to safe online casinos. Two casino operators can run the same certified algorithm, one seeded properly and one not, and the output looks identical from the outside. No amount of inspecting the client tells a player which of the two they are connected to. Only the audit trail does, which is an uncomfortable property for a technical reader to sit with, because it means the interesting engineering is invisible at the point of use.

The mapping step is where real defects concentrate, and it is the part developers most often get wrong. Converting a uniform integer to a range with a modulo operation introduces bias whenever the source range is not an exact multiple of the target. Take a 16-bit value reduced modulo 52 to draw a card in an online blackjack game: 65536 divided by 52 leaves a remainder of 16, so sixteen of the fifty-two cards come up roughly 0.08% more often than the rest. Invisible across ten thousand hands. Very visible across the hundreds of millions a busy casino deals in a year, and precisely the sort of finding a laboratory exists to catch. The same arithmetic applies to slot reel positions, where the target ranges are often smaller and the bias correspondingly larger. Rejection sampling fixes it in about four lines.

Jurisdictional requirements then sit on top of the lab work. The UK Gambling Commission's remote technical standards, the Malta Gaming Authority's framework, Ontario's registrar and New Jersey's regulator each impose their own testing and reporting obligations, which is why the same casino platform gets certified repeatedly rather than once. Live dealer casino games sidestep the question entirely, incidentally, since outcomes come from a physical wheel or shoe read by optical recognition, and the generator plays no part in determining them.

So that is the checklist a laboratory works through. The rest of this is what each item on it actually means.

What a CSPRNG adds

A cryptographically secure generator guarantees two properties a fast generator does not.

Given all previous output, the next bit is computationally infeasible to predict. And given the current state, previous output cannot be reconstructed. Those are usually called forward and backward secrecy, and they are what let you keep using the same generator after a partial compromise.

The constructions are standardised. NIST SP 800-90A specifies HMAC-DRBG, Hash-DRBG and CTR-DRBG, the last typically instantiated with AES. Outside that document, ChaCha20 has become the dominant choice in operating system generators, largely because it is fast in software without AES-NI and has a very clean security margin.

None of this helps if the seed is guessable. Which is the harder problem, and the one that occupies most of an audit.

The entropy problem

A CSPRNG stretches a small amount of genuine unpredictability into an arbitrarily long stream. Everything depends on that initial input.

The Linux kernel collects entropy from interrupt timings, and historically from disk seek latency and human input device timing. Every one of those sources degrades in the environment most server software now runs in. Solid state storage removed the mechanical jitter. Headless machines have no keyboard. Virtual machines see interrupt timing that is partly an artefact of the hypervisor's scheduler rather than physical noise.

The result is boot-time entropy starvation, a genuine operational problem rather than a theoretical one. A freshly cloned VM image asking for cryptographic keys within the first second of uptime has historically been able to get output before the pool was meaningfully seeded.

The interface history reflects the struggle. /dev/random once blocked whenever the kernel's entropy estimate ran low, which produced hangs, which produced developers switching to /dev/urandom for the wrong reasons. The getrandom() syscall arrived in Linux 3.17 with the correct semantics: block until the pool is initially seeded, then never block again. Jason Donenfeld's substantial rewrite of the random driver, landing across the 5.17 and 5.18 kernel releases, modernised the internals around BLAKE2s and ChaCha20 and largely retired the old entropy accounting arguments.

For virtualised guests, virtio-rng passes host entropy through. On bare metal, rngd can feed a hardware source into the pool.

RDRAND and the trust question

Intel shipped RDRAND with Ivy Bridge in 2012, adding RDSEED with Broadwell in 2014. The distinction is worth knowing: RDRAND reads from an on-die DRBG, while RDSEED sits closer to the raw entropy source and is intended for seeding other generators. ARM added equivalent instructions in ARMv8.5-A.

One instruction, high throughput, no drivers. It should have ended the discussion. Instead it started a better one.

In 2007 Shumow and Ferguson demonstrated that Dual_EC_DRBG, another NIST-standardised generator, had a structure permitting a backdoor if the curve constants were chosen maliciously. The 2013 Snowden disclosures made that possibility look considerably less hypothetical, and NIST withdrew the algorithm in 2014.

The fallout reached the Linux kernel almost immediately. A public petition demanded RDRAND be removed on the grounds that an on-die generator is unauditable by anyone outside Intel. Torvalds rejected it in characteristically direct terms, and the technical reason he was right is the important part: Ted Ts'o's driver never trusted RDRAND as a sole source. It mixes hardware output into the pool alongside everything else. A compromised or backdoored instruction cannot reduce the pool's quality below what the other sources already provide.

That design decision was vindicated for unrelated reasons. Some AMD processors shipped with a firmware bug causing RDRAND to return all-ones after suspend and resume, silently, with the instruction still reporting success. Systems treating it as authoritative were generating keys from a constant.

This is exactly the failure mode a casino certification audit is looking for when it asks how a system behaves as its entropy source degrades, and it is why the labs test seeding procedure rather than just output. Do not trust a single source. Not because vendors are adversaries, but because they ship bugs.

What to actually do

Call getrandom(), or BCryptGenRandom on Windows, or SecRandomCopyBytes on Apple platforms. Let the operating system own entropy collection, because it has visibility into interrupt timing that your process does not.

Never seed a security-relevant generator from a timestamp. Never use a fast generator where an adversary benefits from prediction. If you need range reduction, use rejection sampling rather than modulo. And if you are running in a container or a cloned VM image, verify the pool is seeded before asking it for anything that matters.

Statistical batteries like NIST SP 800-22, Dieharder and TestU01's BigCrush are worth running, with one caveat that gets forgotten: MT19937 passes most of them comfortably. Passing a test suite demonstrates the absence of detectable structure. It says nothing whatsoever about whether somebody watching your output stream can work out what comes next.

ЦЕНТР ОБСУЖДЕНИЙ

@GameGPU_com

Присоединяйтесь к обсуждению, делитесь результатами своих бенчмарков и спорьте о производительности железа в X.

Обсудить в X