[+] Proceed to the article
As a professor of biomedical engineering, Martin Uecker perhaps does not
fit the profile of a typical presenter at Kernel Recipes. He is,
however, a longtime Linux user, and works on free software for controlling
magnetic resonance imaging (MRI) scanners. He was at the conference to
talk about the C programming language, the specific problem of undefined
behavior in C, and whether it can eventually be made into a memory-safe
language.
Why bother with C in 2026? It is, he said, still a great language. C is
portable, stable over the long term, offers fast compilation, and the
resulting binary code is fast. "What you see is what you get
"; it
is easy to look at C code and have some idea of what the computer will
actually do. There are a lot of tools for working with the language, and C
gets out of the way when necessary.
C does have a long history, and that affects the language as we see it
today, he said. The C89 standard had to cope with a wide variety of
hardware, including machines with signed-magnitude or one's-complement
integer representations, segmented memory, exotic pointer representations,
and surprising sizes for types. Some Honeywell machines, for example, had
nine-bit bytes. That greatly complicated the task of writing a standard
that would enable the writing of portable code.
The approach that was taken was to define the semantics of the language in
terms of an abstract machine. All operations are to be executed as if they
had run on that abstract machine, which may not exactly match the actual
hardware. The observable behavior of the program must be what the abstract
machine would have done. The "observable" part matters: access to
volatile variables, being defined as observable, must happen
exactly according to the abstract machine; everything else just has to
produce the same eventual result.
The standard gives a lot of freedom to compiler implementers; only the
observable behavior has to be preserved. There are many aspects of that
behavior that are either undefined or implementation-defined. These are
not observable behavior, and thus do not constrain what compiler
implementers can do. There are, of course, other specifications that
can constrain compiler developers where the C standard does not;
these include ABI requirements, standards like POSIX, or the need for
backward compatibility.
Undefined behavior comes about when a program does something that is either
not portable or not defined by the standard at all. In such cases, the C89
standard states that it "imposes no requirements
" on the
implementation. Undefined behavior exists for a number of reasons.
It allows implementations to support extensions, manage
interactions with hardware-based safety mechanisms, and perform aggressive
optimization, all while allowing difficult-to-detect errors to be ignored.
It explicitly gives the compiler the right to ignore whole classes of
hard-to-detect errors.
Nasal demons
The problem, Uecker said, is that the standard allows a compiler to do
anything in response to undefined behavior, up to the point of
invoking nasal
demons. If a program contains any undefined behavior at all, according
to compiler writers, then it has no expected semantics. The C++23 standard
goes further to explicitly state that the standard imposes no requirements
for these programs. That has led to widespread disagreements between
developers about what can be expected from the language.
For example, if you zero an entire structure (perhaps with a call to memset()),
then write to specific fields, what will happen if you read from any
padding bytes in that structure? Might they contain security-relevant
data? A
2015 survey showed that there was no consensus on what should happen in
that case. Or consider this simple code:
extern int x;
int f(int a, int b)
{
x = b ? 42 : 43;
return a/b;
}
If b is zero, then the return statement is a division by
zero, which is undefined behavior. In this case, is the compiler entitled
to omit the test entirely and just execute x = 42? After all, the
b = 0 case has no expected semantics, and can thus be ignored.
There are compilers that will do exactly that. In the undefined-behavior
case, the store to x is not observable behavior. But now consider
this case:
extern void g(int x);
int f(int a, int b)
{
g(b ? 42 : 43);
return a/b;
}
This might seem to be the same situation, with the compiler being entitled
to remove the test and just pass 42 to g(), and some compilers
have treated that way — but that compiler behavior was a bug. Imagine a
definition of g() that calls exit() if b is
zero. In that case, the division will never happen and the program's
behavior is not undefined. So eliding the test and simply passing 42 to
g() is incorrect.
One more interesting case:
volatile int x;
int foo(int a, int b, bool store_to_x)
{
if (! store_to_x)
return a/b;
x = b;
return a/b;
}
The question here is: can the compiler hoist the final division operation
above assignment to x? If there are no semantics associated with
the b = 0 case, then there is no change in observable behavior.
This, too, is something compilers have done, but the C23 standard added a
"no time travel" stipulation to disallow it. In C++, instead, hoisting
must be explicitly prevented by inserting a call to
std::observable_checkpoint().
Time-travel bugs should eventually go away, but there are a lot of other
situations where, even if the standard is clear, compiler writers often
disagree. These include reading of uninitialized variables (which is
almost always defined), and equality comparisons of pointers, which
is always defined, but is also miscompiled by both Clang and GCC.
Fighting undefined behavior
To try to address all of these problems and more, the C committee runs
three study groups focused specifically on the memory object model, memory
safety, and undefined behavior. There are currently about 100 instances of
undefined behavior in the C standard, but the in-progress C2y draft has
removed 45 of them. The situation is indeed getting better.
There is an increasingly rich set of tools aimed at finding issues:
compiler warnings, static analyzers, sanitizers, LLM-based tools, formal
verification, and more. The number of situations where a compiler will
emit a warning where possible undefined behavior is detected is growing;
recent examples include better warnings for integer overflows and potential
use-after-free situations. Static analyzers are available as standalone
tools, but are also increasingly being built into the compilers themselves;
GCC can now warn about a number of potential buffer-overflow situations,
for example. Sanitizers work by inserting run-time checks; they can catch
a lot of undefined behavior and, in trapping mode, be used for hardening as
well.
Memory safety has never been one of C's strong points, but Uecker wanted to
make the point that it can be improved. That problem breaks down into
three sub-problems: type safety, spatial memory safety, and temporal memory
safety.
C, he said, has a strong type system, and the remaining problems are
fixable. Tagless unions, for example, can create type confusion, but the
compiler can enforce types with some additional annotations. New
diagnostics can catch unsafe casts from void. Type checking
across translation units is traditionally not a huge problem in C, since
header files are used to ensure consistent types, but the situation could
be improved with a link-time checker.
Spatial memory safety — bounds checking — is a partially solved problem;
the compilers can perform array-bounds checking in many situations now. In
some cases, some code changes are needed to fully benefit from this
checking. Use of the counted_by attribute
can enable checking for flexible array members, for example.
Temporal memory safety — avoiding use-after-free bugs and the like — is
harder, Uecker said, and Rust definitely has an advantage there. Still,
better temporal memory-safety enforcement is possible. Architectures like
CHERI
can help here is well. Fil-C can find a
lot of temporal-safety bugs.
Can all of these tools and language changes get us to full memory safety?
Completely solving the problem will require either expensive run-time
checking or formal verification, he said. In the near future, the most
complete results will be had with the combination of a restricted language
and formal verification tools.
Overall, he concluded, C is still a living language and is still improving.
The C23 standard removed a number of problematic features, including
old-style (K&R) function definitions, support for sign-magnitude and
one's-complement machines, and trigraphs. It added bit-precise integer
types, checked integer operations, and more. C2y will go further, adding
case ranges, named for loops, the _Countof() macro to
determine array lengths, and a lot of "demon removal
". It will not
achieve full memory safety for C, but that is an eventual possibility, and
will become more practical over time. He ended by encouraging interested
people to participate in the working groups.
The raw
video and slides from this
talk are available.
[Thanks to the Linux Foundation, LWN's travel sponsor, for supporting my
travel for this event.]
Did you like this article?? Subscribe now at the special discounted rate to get a lot more like it.