Proven C Book한국어 GitHub

27 Integers — a world of finite numbers

What to know first

chapter 26, The families of types · where the integer and basic types sit
chapter 7, Representing integers · sign and overflow
chapter 23, Declaring variables · a type is the shape of the vessel

Looking back

Chapter 7 said that overflow of unsigned integers is “defined wrap-around” while signed overflow is outside the contract. But is the int we made in chapter 23 signed or unsigned — which world have our variables been in all along?

A. int is a signed integer — that is, our variables have been in the world of “outside the contract if it overflows.” Fortunately the examples so far never came near the ±2.1 billion container, but it is time to draw that boundary exactly. The kinds and sizes of containers, where the boundary lies, and the rule when it is crossed — that is this chapter.

The need for this chapter, and its context

The representation of integers learned in chapter 7 turns here into C’s types. The twenty-chapter gap between them was needed to erect the skeleton of the language; now the bit-level story can be laid on top of syntax. That is also why part 6 begins with forging values: know the values exactly before controlling how they flow.

By the end of this chapter

Part VI forges values and governs flow. Its first chapter faces C’s integers head on — how the world of representation learned in chapter 7 appears as C’s family of types, what face overflow wears in practice, and the modern C practice (<stdint.h>).

The questions this chapter answers

  1. What happens if the value fits nothing in the list?
  2. It seems strange that char is a member of the integer family — is it not a character?

This chapter digs into one cell of what chapter 26 laid out — the family passed over there in the single phrase “the signed and unsigned integer types”.

27.1 The family of integer types

C’s integer types are not int alone but a family. The basic members on the signed side, in order of size, are char (1 byte — the “byte = character” story of chapter 4 survives in the name), short, int, long and long long, and putting unsigned in front of each gives its unsigned partner (unsigned int and so on). The standard fixes only each type’s minimum size — int at least 16 bits, long at least 32 — so the exact size is the platform’s business (on mainstream systems int is 32 bits).

If “it differs by machine” sounds uneasy, that is an accurate instinct — which is why modern C practice is to use types with the size fixed in the name. The <stdint.h> toolbox’s int32_t (exactly 32 bits, signed), uint8_t, int64_t and the like. Wherever size is part of the contract — file formats, communication, everywhere chapter 5′s endianness matters — that side is standard practice. C23 added _BitInt(N) on top: an integer whose bit width you write yourself. This book’s examples use int by default for simplicity on the page, but switch to the <stdint.h> family in scenes where size matters.

27.2 The basic types at a glance — minimum and actual ranges

Here we gather in one place exactly what C’s basic types are and how large each is. This section is a place to look things up.

Integer types come in five tiers, each with a signed and an unsigned version — char, short, int, long, long long. To these are attached bool (C99) and the character-specific types (char16_t, char32_t, wchar_t). Floating types are three — float, double, long double.

What the standard pins down is not size but minimum range. The common saying “int is 4 bytes” is not a sentence of the standard but an observation that most platforms are like that today.

typeminimum range guaranteedminimum widthlimit macros
signed char−127 to +1278 bitsSCHAR_MIN SCHAR_MAX
unsigned char0 to 2558 bitsUCHAR_MAX
charone of the two above (implementation-defined)8 bitsCHAR_MIN CHAR_MAX
short−32767 to +3276716 bitsSHRT_MIN SHRT_MAX
unsigned short0 to 6553516 bitsUSHRT_MAX
int−32767 to +3276716 bitsINT_MIN INT_MAX
unsigned int0 to 6553516 bitsUINT_MAX
long−2147483647 to +214748364732 bitsLONG_MIN LONG_MAX
unsigned long0 to 429496729532 bitsULONG_MAX
long longabout ±9.2×101864 bitsLLONG_MIN LLONG_MAX
unsigned long long0 to about 1.8×101964 bitsULLONG_MAX

Table 27.1

Your eye will go to the minimum range being −127 (not −128). That is because the old standard permitted all three sign representations (chapter 7); now that C23 has pinned two’s complement down, it is effectively −128.

The standard also fixes the order of sizes. Widths cannot run against this order.

char  ≤  short  ≤  int  ≤  long  ≤  long long

And sizeof(char) is always 1 — because a byte is by definition the size of a char (chapter 4). How many bits are in a byte, though, is told by CHAR_BIT, and the standard guarantees only 8 or more.

The limits of floating types are the business of <float.h>. Picking only those in frequent use:

macromeaningfor IEEE 754 double
FLT_DIG DBL_DIGtrustworthy decimal digits6 / 15
FLT_MAX DBL_MAXlargest representable valueabout 1.8×10308
FLT_MIN DBL_MINsmallest normalised valueabout 2.2×10−308
FLT_EPSILON DBL_EPSILONsmallest difference distinguishable from 1.0about 2.2×10−16
FLT_RADIXthe base of the exponent2

Table 27.2

DBL_EPSILON is met again in chapter 50 when comparing floating-point numbers — it is the value used to set the criterion for judging “equal”.

examples-en/ch27/limits.c

#include <stdio.h>
#include <limits.h>
#include <stdint.h>
#include <float.h>
#include <stddef.h>

int main(void)
{
    printf("--- sizes of the basic integer types (on this machine) ---\n");
    printf("char=%zu short=%zu int=%zu long=%zu long long=%zu\n",
           sizeof(char), sizeof(short), sizeof(int), sizeof(long), sizeof(long long));
    printf("size_t=%zu ptrdiff_t=%zu void*=%zu\n",
           sizeof(size_t), sizeof(ptrdiff_t), sizeof(void *));

    printf("\n--- the real ranges (<limits.h> macros) ---\n");
    printf("CHAR_BIT  = %d\n", CHAR_BIT);
    printf("SCHAR_MIN = %d, SCHAR_MAX = %d, UCHAR_MAX = %u\n",
           SCHAR_MIN, SCHAR_MAX, (unsigned)UCHAR_MAX);
    printf("SHRT_MIN  = %d, SHRT_MAX  = %d, USHRT_MAX = %u\n",
           SHRT_MIN, SHRT_MAX, (unsigned)USHRT_MAX);
    printf("INT_MIN   = %d, INT_MAX   = %d, UINT_MAX  = %u\n",
           INT_MIN, INT_MAX, UINT_MAX);
    printf("LONG_MIN  = %ld, LONG_MAX = %ld\n", LONG_MIN, LONG_MAX);
    printf("LLONG_MAX = %lld, ULLONG_MAX = %llu\n", LLONG_MAX, ULLONG_MAX);

    printf("\n--- the fixed-width types (<stdint.h>) ---\n");
    printf("INT8_MAX=%d INT16_MAX=%d INT32_MAX=%d\n",
           (int)INT8_MAX, (int)INT16_MAX, (int)INT32_MAX);
    printf("INT64_MAX=%lld UINT64_MAX=%llu\n",
           (long long)INT64_MAX, (unsigned long long)UINT64_MAX);
    printf("SIZE_MAX=%zu\n", SIZE_MAX);

    printf("\n--- the real types (<float.h>) ---\n");
    printf("float : %zu bytes, %d digits, max %g\n", sizeof(float), FLT_DIG, (double)FLT_MAX);
    printf("double: %zu bytes, %d digits, max %g\n", sizeof(double), DBL_DIG, DBL_MAX);
    printf("DBL_EPSILON = %g (the smallest difference distinguishable from 1.0)\n", DBL_EPSILON);
    return 0;
}

Output

--- sizes of the basic integer types (on this machine) ---
char=1 short=2 int=4 long=8 long long=8
size_t=8 ptrdiff_t=8 void*=8

--- the real ranges (<limits.h> macros) ---
CHAR_BIT  = 8
SCHAR_MIN = -128, SCHAR_MAX = 127, UCHAR_MAX = 255
SHRT_MIN  = -32768, SHRT_MAX  = 32767, USHRT_MAX = 65535
INT_MIN   = -2147483648, INT_MAX   = 2147483647, UINT_MAX  = 4294967295
LONG_MIN  = -9223372036854775808, LONG_MAX = 9223372036854775807
LLONG_MAX = 9223372036854775807, ULLONG_MAX = 18446744073709551615

--- the fixed-width types (<stdint.h>) ---
INT8_MAX=127 INT16_MAX=32767 INT32_MAX=2147483647
INT64_MAX=9223372036854775807 UINT64_MAX=18446744073709551615
SIZE_MAX=18446744073709551615

--- the real types (<float.h>) ---
float : 4 bytes, 6 digits, max 3.40282e+38
double: 8 bytes, 15 digits, max 1.79769e+308
DBL_EPSILON = 2.22045e-16 (the smallest difference distinguishable from 1.0)

What matters is that this output belongs to the machine that made this book. Run it on another machine and different numbers may appear — and that is exactly this section’s point: do not assume sizes; ask.

27.3 The type of an integer constant — the same value, typed by its notation

Chapter 20 showed the four bases and the suffixes for writing an integer constant. What was deferred there — which type the compiler gives that constant — can be faced now that integers have been met.

The rule is “walk a list and take the first type that fits”. But the list differs by base.

Unsuffixed constantThe candidate list, in orderThe point
decimal (4294967295)intlonglong longunsigned types are not candidates
octal, hex, binary (0xFFFFFFFF)intunsigned intlongunsigned longlong longunsigned long longunsigned types are interleaved

Table 27.3

Adding a suffix narrows the list by hand — u walks only the unsigned ones, l starts at long, ll at long long.

The difference shows up for real.

WrittenMeasured sizeofType
0xFFFFFFFF4unsigned int
42949672958long

Table 27.4

The same number, a different type. And once the type differs, everything downstream differs — promotion and the usual arithmetic conversions (chapter 29) apply differently, and comparisons can come out reversed.

Counter-example. Writing a bit mask in decimal

x & 4294967295      /* becomes an operation with a long */
x & 0xFFFFFFFFU     /* visibly an unsigned 32-bit mask */

Practice writes masks in hexadecimal not only because it reads better. The type differs, and above all the number of bits is visible0xFFFF is 16 bits and 0xFFFFFFFF is 32, right there in the digit count. Adding U to pin the signedness as well is the convention (shifting a signed integer is chapter 28′s grey area).

Q. What happens if the value fits nothing in the list?

A. If it does not fit even the widest candidate, it goes into an extended integer type the implementation provides, or, if there is none, it is a constraint violation and gets diagnosed. Meeting this in practice has one answer — use the fixed-width types and their macros (UINT64_C(…), <stdint.h>). “Do not leave the type to the compiler” is the same discipline as the rest of this chapter.

27.4 Where size is the contract — fixed-width types

There are places where the width must be exactly determined: file formats, network protocols, hardware registers. In such places use the fixed-width types of <stdint.h>.

kindexamplesmeaning
exact widthint8_t uint16_t int32_t uint64_texactly that many bits. not provided if unavailable
minimum widthint_least16_tthe smallest type that is at least that wide
fastestint_fast32_tat least that wide and fastest for the machine to handle
pointer-sizedintptr_t uintptr_tan integer able to hold a pointer (mind chapter 14′s provenance)
largestintmax_t uintmax_tthe widest integer
size and differencesize_t ptrdiff_tthe types of a size and of a pointer difference

Table 27.5

Each has its limit macros too — INT32_MAX, UINT64_MAX, SIZE_MAX, PTRDIFF_MAX and so on. Their format specifiers come from <inttypes.h>‘s PRId32 family (appendix B).

The rule is simple. If the meaning is “this machine’s natural integer”, int; if “a size or an index”, size_t; if “a width fixed by a format”, a fixed-width type.

27.5 Seeing the boundary with our own eyes

The edge of a container is told by the <limits.h> toolbox. And the wrap-around of the unsigned world — the one learned on the page in chapter 7 — can now be run and shown.

examples-en/ch27/wrap.c

#include <limits.h>
#include <stdio.h>

int main(void)
{
    printf("INT_MAX  = %d\n", INT_MAX);
    printf("UINT_MAX = %u\n", UINT_MAX);
    printf("UINT_MAX + 1 = %u\n", UINT_MAX + 1u);  /* defined wrap-around */
    return 0;
}

Output

INT_MAX  = 2147483647
UINT_MAX = 4294967295
UINT_MAX + 1 = 0

One new format has joined — unsigned integers are printed with %u, not %d (chapter 22′s format contract growing along with the family of types). And the third line is exactly chapter 7′s promise: add 1 to the maximum and you get 0 — the clock has gone round once, and this is defined behaviour.

The signed side is another matter. Code that computes INT_MAX + 1 is outside the contract (undefined behaviour), so it cannot even be put into this book’s verification pipeline — the UBSan equipped in chapter 17 catches exactly this kind of code at run time. That the same “+1” can be demonstrated on one side and not on the other is itself eloquent about the difference between the two worlds.

A common misconception. “Overflow is a problem for special programs that handle large numbers”

Plausible, but the list of real accidents says the opposite. The commonest overflows happen in perfectly ordinary places — taking the average of two indices as (a + b) / 2 and having the intermediate sum overflow (the famous binary search bug, which hid in standard libraries for nearly twenty years); a millisecond timer wrapping after 49 days (the Windows 95 49.7-day hang); a size calculation count * size overflowing and taking an absurdly small piece of memory (a classic of security incidents). What they share is that the intermediate calculation overflows, not the final value — worrying about the container is done for every intermediate step of an expression, not for the result, and that is why checked arithmetic of the kind proven provides exists (chapter 43).

Q. It seems strange that char is a member of the integer family — is it not a character?

A. In C a character and a small integer are the same thing — as chapter 8 taught, a character is a number, and char is merely a one-byte integer container holding that number. The value of the character literal 'A' is simply 65. One trap in advance — whether char is signed or unsigned differs by platform (it is a third type, neither of the two). So char is not used for numeric work; when a byte is needed, modern practice uses uint8_t. The proper story of strings and char is in chapter 42.

The containers of integers are sorted. The next chapter takes the operations on those containers — the truth about division, deferred in chapter 20, and the operators of the bit world learned in chapter 7, joining C’s syntax.