Reducing undefined behavior in the C language
signa11
79 points
47 comments
October 09, 2026
Related Discussions
Found 5 related stories in 90.6ms across 8,906 title embeddings via pgvector HNSW
- Reducing undefined behavior in the C language chmaynard · 27 pts · September 28, 2026 · 99% similar
- Tail-call optimization in C is relatively recent (2025) prakashqwerty · 88 pts · August 10, 2026 · 54% similar
- An ambiguity in c89 which will never be fixed runningmike · 15 pts · August 10, 2026 · 51% similar
- Anecdotally, Programmers Dislike "Reduce" praptak · 12 pts · September 14, 2026 · 50% similar
- Anecdotally, Programmers Dislike "Reduce" vinhnx · 12 pts · September 14, 2026 · 50% similar
Discussion Highlights (10 comments)
elais-dev
i've seen static analyzers catch many ub patterns, but guaranteeing zero ub needs whole‑program analysis that blows up compile time and still produces false positives that drown developers
p1necone
The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me, and I'm yet to read anything about it that has made it seem any less so.
jdw64
If we follow this video, the real constraints essentially mean POSIX and ABI. This implies that the contract a programmer must understand is distributed across multiple layers—not just the language specification, but effectively the operating system and ABI as well. Ultimately, this leads to the conclusion that even if the language itself is fundamentally free, constraints are inherently necessary at its lower layers. If we were to bloat the compiler—that is, if we restricted freedom like Rust does with its borrow checker—then the freedom available to the programmer would vanish. If that happens, people might grow weary of a language that is supposed to be free. In the end, any single layer inherently restricts freedom. In other words, I feel there is a need to transfer the complexity that a programmer must manage over to a mechanical management system, but which layer would be best for that? Currently, based on experience, this level of complexity is categorized into the language layer, and that level of complexity into the operating system layer. But in the future, won't there be some sort of complexity theorem that determines which layer minimizes complexity the most, and won't systems be completely rewritten based on that?
BeaverGoose
Make signed overflow defined please.
chasil
"Some Honeywell machines, for example, had nine-bit bytes." OS 2200 has 36-bit words. It is still a supported platform. https://en.wikipedia.org/wiki/UNIVAC_1100/2200_series This platform was the first SMP UNIX implementation: "Any configuration supplied by Sperry, including multiprocessor ones, can run the UNIX system." https://www.nokia.com/bell-labs/about/dennis-m-ritchie/other...
1vuio0pswjnm7
1790620504 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/b9325731ea9b61e0/ | https://news.ycombinator.com/item?id=49882419 | 15 comments 1790673092 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/efcdbcf080cfa4c6/ | https://news.ycombinator.com/item?id=49890290 | 0 comments
hn_submit
I believe this is barking up the wrong tree since IMHO C is just "high level assembly" for systems programming. As soon as you add runtime behavior to combat Undefined Behavior (UB) you're blowing up execution times. And static analysis can only go so far without blowing up compile times. C is "the right tool for the right job" which is operating systems and its code which is called thousands of times per second. You cannot afford even one iota of runtime checks in that code. The developer must know what he's doing or he should get out of the kitchen. We should discourage the usage of C in application programming and prod developers towards memory safe languages like Rust or Go. And I'm not even sure if Rust solves this case as far as UB is concerned.
strenholme
I recently had a heated but productive discussion about undefined behavior here. As per the linked article: “There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them.” I wonder how they handle the specific case of uninitialized but allocated memory. Let’s look at something which will result in undefined behavior in C99: [1] #include<stdio.h> #include<stdint.h> #include<stdlib.h> #define b(z) for(c=0;c<z;c++) uint32_t c,e[42],f[42],g=19,h =13,n[45],i,j,k;void m(){j=0; b(12)f[c+c%3*h]^=e[c+1];b(g){ i=c*7%g;k=e[i++];k^=e[i%g]|~e [(i+1)%g];j=j+c;n[c]=n[c+g]=k >>j%32|k<<-j%32;}for(i=39;i-- ;f[i+1]=f[i])e[i]=n[i]^n[i+1] ^n[i+4];b(3)e[c+h]^=f[c*h]=f[ c*h+h];*e^=1;}int main(int c, char**v){char*q=malloc(2);if( q==0)return 0;q[0]&=31;q[0]|= 64;q[1]=0;for(;;m()){b(3){ for(j=0;j<4;){f[c*h]^=k=(*q? 255&*q:1)<<8*j++;e[c+16]^=k; if(!*q++){b(18)m();b(2){j=c; b(4)printf("%02x",(e[1+j%2] >>8*c)&255);c=j;if(c%2)m();} puts("");return 0;}}}}} The key part of the above brick of code is this: char *q=malloc(2); if(q==0)return 0; q[0]&=31; q[0]|=64; q[1]=0; Here, we see that q[0] is an allocated but undefined byte. As per C99, this results in undefined behavior, however 20 years ago this was a good trick to get kinda-randomish bytes to use as a possible entropy source. Someone claimed that the above brick of code will compile in newer versions of clang such that, since the complex cryptographic pseudo random number generator code depends on uninitialized but allocated memory, the entire cryptographic operation isn’t performed. So I tested it against multiple versions of GCC and clang; I also tested it against TCC for good measure. In all cases, with all levels of optimization, the cryptographic routine ran. I even ran it against clang 23. In cygwin, it was a randomish but consistent byte (except for clang at a higher level of optimization, at which point the uninitialized byte had a value of 0); in Ubuntu 26, the uninitialized memory consistently had a value of 0 (in tcc/gcc/clang). I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values). Naturally, I have updated my code to no longer use uninitialized memory as a source of entropy. 20 years ago, MacOS didn’t support clock_gettime() with nanosecond resolution, so that wasn’t a portable way to get pseudo-random bits; these days clock_gettime() is universal across modern development environments, and it provides pretty good entropy (along with using /dev/urandom in *NIX, which isn’t in POSIX but is widely supported, as well as CryptGenRandom() in the legacy Win32 port). [2] [1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative. [2] I don’t blindly trust /dev/urandom to always make really hard to guess pseudo-random bits, because my code is open source, and, as such, doesn’t just compile in Linux. It often times will be compiled in embedded systems, and even Linux has had at times issues with /dev/urandom on Raspberry Pis. [3] I would also like to see uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, and int64_t mandated. They exist in C99, but aren’t mandated, even though every real world compiler from this century supports all of the above types. Yes, I know about _BitInt(8/16/32/64/128/etc.) but a compiler from 2004—and yes I still use one to make win32 binaries—doesn’t support these new C23 datatypes.
pizlonator
It’s cool that this mentions Fil-C but it also undersells it. TFA also undersells CHERI. Fil-C doesn’t just “find a lot of temporal-safety” bugs. It closes off memory safety bugs (special and temporal) for exploit writers and ascribes a tight semantics to the whole language. CHERI makes some different trade offs but also gives a tight enough semantics that memory safety exploits aren’t going to work. Both CHERI and Fil-C are more comprehensive than Rust, since they attack the problem at the ABI level (and so you don’t get the problem that the protection only applies to the parts that were rewritten in the safe subset of a new language). Rust could be claimed to be better in that its compile time, but that doesn’t make a significant difference if you’re worried about the definedness of semantics or exploitability.
fithisux
We need to learn to use assembly and call it from C. C was never meant to be the end of programming. Reducing undefined behavior is the least the standards body can do. Fil-C is also an extremely useful tool. The real problem I see is the inability of independent compiler writers upgrading to the latest standard. I think this is also an issue that the standards body should pay attention. Help implementers.