C's Flexible Integer Sizes Were Not a Design Mistake
C's integer types were never meant to be fixed-width. In the 1970s, machines had 12-, 18-, 36-, and 60-bit words, and C needed to run efficiently on all of them. The flexible sizes of char, int, short, and long were a deliberate portability feature, not an oversight. This article explains why judging that 1970s decision by today's 64-bit monoculture misses the point.
A 'plain' int object has the natural size suggested by the architecture of the execution environment.
- adrian_b
I agree that the C flexible integer sizes were still necessary at the time of its creation, when some important computers still had word sizes that were not powers of two.
Nonetheless, I started to use C for programming only in 1990, when I got access to the Microsoft C and Borland Turbo C compilers.
At that time, 36 years ago, the C flexible integer sizes were already obsolete.
Since that time until now, while using C on a great variety of computers, from servers and workstations to the smallest microcontrollers, I have seen plenty of portability problems created by the existence of the flexible integer sizes.
The only programs that had no portability problems were those that never used the flexible integer sizes, but only integers with a definite size, e.g. 8-bit, 16-bit, 32-bit or 64-bit.
While sizeof solves the problems of memory allocation or copying, it does not help in preventing unexpected integer overflows, because even the size of "char" may be unknown, and even if the size of "char" is known, writing code with multiple paths that would check or prevent overflow for different integer sizes is very cumbersome.
Flexible integer sizes would work well only on the old computers, where integer overflow generated a hardware exception, so installing an overflow handler would have been sufficient to make the C code work correctly regardless of the size of the native integers.
- layer8
> Language types such as char, int, short, and long do not come with a guarantee of how many bytes they occupy in memory.
Char is actually guaranteed by C to occupy exactly 1 byte in memory. It’s just that a byte can have more than eight bits in C. “Byte” is simply the smallest unit of memory addressable by a pointer.
Further down the article acknowledges that “C requires char to have at least 8 bits (CHAR_BIT >= 8), not exactly 8” and mentions the Honeywell 6000 as an example of a C implementation with 9 bits (and 36-bit ints).
Historically in computing, the size of a byte was hardware-dependent and not standardized. The Wikipedia article on “byte” cites Knuth’s 1968 TAOCP where byte denotes a unit which “contains an unspecified amount of information […] capable of holding at least 64 distinct values […] at most 100 distinct values. On a binary computer a byte must therefore be composed of six bits”.
- InvisibleUp
Where the flexible integer sizes break the most is when dealing with ABIs, which weren't really a concern before dynamic linking existed but are very much a concern today. We've also, for some reason, decided that the standard way of defining a library ABI is with a C header. That means that everyone has to worry about precisely defining integer sizes, as well as more esoteric types like size_t or intmax_t. Good writeup on all that here: https://thephd.dev/to-save-c-we-must-save-abi-fixing-c-funct...
- nayuki
Some pieces of code in the article look suspicious:
#define LUAI_IS32INT ((UINT_MAX >> 30) >= 3)
If unsigned int is just 16 bits wide, then `65535 >> 30` shifts by more positions than the width of the type, which should be undefined behavior, right?
#define NB CHAR_BIT
#define MC ((1 << NB) - 1)
If we have a DSP machine (as mentioned in the article) where CHAR_BIT is 16 and int is 16 bits, then `1 << NB` is `1 << 16` which shifts as much as the width of the type, which I believe is undefined behavior as well.
Regardless of whether you think C's flexible integer sizes are good or bad, and whether it helped propagate the language, there's no denying that if you want to write portable code across machines, you have to put in more effort compared to a language with fixed-size integers. Whether this matters or not is a matter of situation and opinion.
Other than that, I made a simulator where you can set the bit width of each integer type, and then it shows you how `x operation y` gets promoted to some output integer type. https://www.nayuki.io/page/summary-of-c-cpp-integer-rules section "Conversion rules simulator"
- codedokode
I think it didn't work out well, because "int" being different size makes programming difficult. For example, a system must manage up to 100 000 records. Can I use int for record number? What if it is 16 bits? What if I need to send data between machines, how can I use "int" if it can be different size?
Probably someone noticed that it is inconvenient, and on 64-bit machine ints are still 32-bit and not 64.
The computers with 16-bit ints or 9-bit bytes are long gone, but the language still has to carry that legacy.
- habitue
It was intentional, sure. It was an attempt to solve a particular kind of problem.
In hindsight though, it was a mistake.
Evidence: when the world moved to 64 bit, we didnt just let int mean 8 bytes on amd64. That's a clear acknowledgement that the design was not correct once we understood things better.
- stkdump
The problem begins when you start mixing the traditional types and (u)intN_t, because the latter are merely aliases for the internal types, and it messes up overload resolution. All relevant platforms have pretty much agreed the size of char, short (int), int and long long (int). They have different opinions about long (int) and thus an int64_t might use either long (int) or long long (int).
So the best solution for nowadays is to use just char, short, int and long long (and make strong assumptions that these are exactly 8, 16, 32 and 64 bits wide respectively), never use long or long double. Never use (u)intNN_t. Then you are good.
Those caveats of the past (but int might be 16 or 36 bits), are exactly that. An artifact of the past. A historical curiosity. Not relevant for today or the future. No, I don't believe for a second that any future platform will change their size.
Platforms also still disagree on the signedness of char, so when an 8 bit numeric type (as opposed to an ascii character type) is needed, one should always explicitly specify signed char or unsigned char, both of which are separate types from char.
Further things of note: platforms also have agreed on little endian (so called "network byte order" is dead and should never be used in new protocols, because it forces everyone to convert) and on IEEE memory representation of float and double. Contrary to popular belief the main floating point operations (+,-,*,/,==,<,>,<=,>=) are also precisely defined and al […]
- quelsolaar
Good article.
C would probably not have survived unless it had this flexibility.
But its not justa historical thing. Today there are modern platforms like DSPs that have 32bit sized char, because that is the smallest addressable type. These platforms depend on C for tool chains, even if most "portable" C wont run correctly on them. The fact that you can build hardware like that, and not have to invent a new language / dialect to program them is a huge win for the world.
<edit> I didnt see the footnote about DSPs at first read </edit>