Proven C Book한국어 GitHub

20 Expressions and constants — the things that become values

What to know first

chapter 7, Representing integers · the representation of integers
chapter 19, The structure of a program · where statements and expressions sit

Looking back

Chapters 7 and 8 taught how to hold numbers in the machine (two’s complement, IEEE 754). So when a number is written in source code — when you write 12, say — when does it become the machine’s bits?

A. At compile time. The 12 in the source is just two characters (in the terms of chapter 9, the characters 1 and 2), and the compiler reads it, translates it into a two’s-complement bit pattern and plants it in the executable. A notation for a value written in source is called a literal — meaning “as the letters say.” That the bridge between the world of letters and the world of bits is the compiler — chapter 16′s relay is at work here too.

The need for this chapter, and its context

The first step inside the skeleton is how to write a value down. Making this chapter the gathering place for constant notation is deliberate — scatter the integer, character, floating and string notations and there is nowhere to look them up later. The integer representation of chapter 7 turns here, for the first time, into characters written in source.

By the end of this chapter

We go inside the statement. How to write values in source code is half of it — the notations for integer, character, floating and string constants are gathered here in one place, so this is where to come back when “what was that notation again?” arises — and the other half is weaving values into calculations (operators and expressions) and the order of calculation (precedence and parentheses). As this book’s principle demands, only the minimum set of operators: addition, subtraction, multiplication and parentheses. Even that goes surprisingly far.

The questions this chapter answers

  1. Then if I write 0.1 in the source, does the machine hold exactly 0.1?
  2. Why the advice never to use a lowercase l?
  3. Why must a hex float have the p exponent, and why use one at all?
  4. Why was division put off? Leaving one of the four operations out feels odd.
  5. Is printf("hello") a value too? A function call is written in the position of an expression.

20.1 Literals — values written in source

We have met literals several times already. The 0 of hello world is an integer literal, and "Hello, world!\n" is a string literal — the characters inside the quotation marks become a value “as they are.” Add a decimal point to an integer literal (3.5) and it becomes a floating-point value from chapter 8.

Q. Then if I write 0.1 in the source, does the machine hold exactly 0.1?

A. It does not — exactly as chapter 8 taught. The compiler translates the literal 0.1 into its nearest neighbour (3FB9 9999 9999 999A). A literal is “a notation writing the value as it is”, but if that value is not on the grid of representable numbers, the approximation has already happened at translation. Writing it in the source does not make it exact — chapter 8′s lesson holds in the world of source code too.

20.2 Writing constants — all of them at a glance

Having met literals in the previous section, this is the place to gather every notation for writing a value in source at once. The standard scatters them through its lexical clauses (§6.4.4, §6.4.5), and what trips people in practice is nearly always one of three things: a prefix, a suffix or an escape.

This section is for reference. There is no need to memorise it now — take away the sense that these notations exist, and come back when you need them. Read the listing in that spirit too: skimming the output and matching notation to result is enough.

examples/ch20/constants.c

/* 상수를 적는 방법들 — 표기가 값과 타입을 어떻게 정하는가. */
#include <inttypes.h>
#include <stdio.h>
#include <string.h>
#include <uchar.h>

int main(void)
{
    puts("[integer constants - four bases and digit separators (C23)]");
    printf("  1234        = %d\n", 1234);          /* 10진 */
    printf("  0755        = %d   <- a leading 0 means octal\n", 0755);
    printf("  0xFF        = %d\n", 0xFF);
    printf("  0b1010      = %d   <- binary, new in C23\n", 0b1010);
    printf("  1'000'000   = %d   <- digit separators, new in C23\n", 1'000'000);
    printf("  0b11'10'11'01 = %d\n", 0b11'10'11'01);

    puts("\n[접미어가 타입을 정한다]");
    printf("  sizeof 1    = %zu, sizeof 1L  = %zu, sizeof 1LL = %zu\n",
           sizeof 1, sizeof 1L, sizeof 1LL);
    printf("  sizeof 1U   = %zu, sizeof 1wb = %zu  ← wb 는 C23 의 _BitInt\n",
           sizeof 1U, sizeof 1wb);

    puts("\n[문자 상수 — 접두어가 타입을 정한다]");
    printf("  'a'   크기 %zu, 값 %d      ← C 에서 문자 상수는 int 다\n",
           sizeof 'a', 'a');
    printf("  u8'a' 크기 %zu (char8_t)\n",  sizeof u8'a');
    printf("  u'a'  크기 %zu (char16_t)\n", sizeof u'a');
    printf("  U'a'  크기 %zu (char32_t)\n", sizeof U'a');
    printf("  L'a'  크기 %zu (wchar_t)\n",  sizeof L'a');
    /* 'ab' 같은 다중 문자 상수는 값이 구현 정의라 -Wmultichar 가 경고한다.
       여기서는 경고를 켠 채 두고, 값은 본문의 실측 표로만 보인다. */

    puts("\n[이스케이프 — 8진과 16진]");
    printf("  '\\101' = %d, '\\x41' = %d   ← 둘 다 'A'\n", '\101', '\x41');
    printf("  \"\\x41\" \"1\" = \"%s\"        ← 16진은 가장 긴 열을 먹는다.\n",
           "\x41" "1");
    puts("    그래서 \"\\x411\" 이 아니라 문자열을 쪼개 이어 붙인다.");

    puts("\n[부동소수점 상수]");
    printf("  3.14  1e3=%g  1.=%g  .5=%g\n", 1e3, 1., .5);
    printf("  0x1p-3 = %g            ← 16진 부동소수점. 지수부 p 는 필수다\n", 0x1p-3);
    printf("  sizeof 1.0 = %zu, 1.0f = %zu, 1.0L = %zu\n",
           sizeof 1.0, sizeof 1.0f, sizeof 1.0L);
    printf("  0.1 == 0.1f ? %s   ← 접미어가 다르면 값도 다르다\n",
           (double)0.1f == 0.1 ? "예" : "아니오");

    puts("\n[문자열 리터럴]");
    printf("  sizeof \"abc\" = %zu        ← NUL 이 한 칸 더 붙는다\n", sizeof "abc");
    printf("  \"hello, \" \"world\" = \"%s\"  ← 인접한 것은 하나로 이어진다\n",
           "hello, " "world");
    printf("  \"a\\0b\": strlen = %zu, sizeof = %zu  ← 안에 NUL 을 넣어도 배열은 남는다\n",
           strlen("a\0b"), sizeof "a\0b");
    printf("  sizeof u8\"\" = %zu, sizeof u\"\" = %zu, sizeof U\"\" = %zu\n",
           sizeof u8"\uAC00", sizeof u"\uAC00", sizeof U"\uAC00");
    return 0;
}

Output

[integer constants - four bases and digit separators (C23)]
  1234        = 1234
  0755        = 493   <- a leading 0 means octal
  0xFF        = 255
  0b1010      = 10   <- binary, new in C23
  1'000'000   = 1000000   <- digit separators, new in C23
  0b11'10'11'01 = 237

[접미어가 타입을 정한다]
  sizeof 1    = 4, sizeof 1L  = 8, sizeof 1LL = 8
  sizeof 1U   = 4, sizeof 1wb = 1  ← wb 는 C23 의 _BitInt

[문자 상수 — 접두어가 타입을 정한다]
  'a'   크기 4, 값 97      ← C 에서 문자 상수는 int 다
  u8'a' 크기 1 (char8_t)
  u'a'  크기 2 (char16_t)
  U'a'  크기 4 (char32_t)
  L'a'  크기 4 (wchar_t)

[이스케이프 — 8진과 16진]
  '\101' = 65, '\x41' = 65   ← 둘 다 'A'
  "\x41" "1" = "A1"        ← 16진은 가장 긴 열을 먹는다.
    그래서 "\x411" 이 아니라 문자열을 쪼개 이어 붙인다.

[부동소수점 상수]
  3.14  1e3=1000  1.=1  .5=0.5
  0x1p-3 = 0.125            ← 16진 부동소수점. 지수부 p 는 필수다
  sizeof 1.0 = 8, 1.0f = 4, 1.0L = 16
  0.1 == 0.1f ? 아니오   ← 접미어가 다르면 값도 다르다

[문자열 리터럴]
  sizeof "abc" = 4        ← NUL 이 한 칸 더 붙는다
  "hello, " "world" = "hello, world"  ← 인접한 것은 하나로 이어진다
  "a\0b": strlen = 1, sizeof = 4  ← 안에 NUL 을 넣어도 배열은 남는다
  sizeof u8"가" = 4, sizeof u"가" = 4, sizeof U"가" = 8

20.3 Integer constants (§6.4.4.1)

BaseNotationExampleNote
decimalstarts with 1~91234
octalstarts with 00755 = 493the commonest trap08 is an error
hexadecimal0x or 0X0xFF = 255
binary0b or 0B0b1010 = 10added in C23

Table 20.1

Letters can follow too — a suffix, as in 10L, 1U, 3ULL. It marks not the value but the type.

SuffixWhat it marks
u Uan unsigned type
l Llong or wider
ll LLlong long or wider
wb WB_BitInt (C23)

Table 20.2

For now, that a suffix changes the type is all you need. The type names above, and the exact rule for which type the compiler picks when there is no suffix, come in chapter 27 once integers have been met properly, in “The type of an integer constant”. One thing in advance — writing the same value in decimal or in hexadecimal can give it a different type.

Q. Why the advice never to use a lowercase l?

A. Because in many fonts 1 (one) and l (ell) look almost identical. 10l being read as 101 has really happened. Suffixes in capitals10L, 1UL — is the long-standing convention, and this book follows it. The 0x prefix is conventionally lowercase, which looks like the opposite rule, but the reason is the same: whichever is easier to tell apart.

20.4 The C23 digit separator '

A notation for breaking up long numbers arrived in C23. It has no effect on the value — in the standard’s words it is ignored when determining the value of the constant.

The standard’s own example shows the traps along with the feature. Measured:

WrittenResultWhy
12'341234a separator goes only between digits
0b11'10'11'01237binary takes it too
0x1'2'3'4AB'C'D305441741so does hexadecimal
0x'FFerror — “digit separator after base indicator”it may not follow 0x
'1'2errorread as the character constant '1' followed by 2

Table 20.3

The last row is this notation’s one danger. A separator is a separator only between two digits; at the front it is read as a single quote — the start of a character constant.

20.5 Character constants (§6.4.4.4)

The type names read fully only after chapter 9′s character sets and chapter 27′s integers — for now just see that the prefix settles the type.

NotationTypeMeasured sizeNote
'a'int4★in C a character constant is not a char
u8'a'char8_t1C23. Must be one UTF-8 code unit
u'a'char16_t2one UTF-16 code unit
U'a'char32_t4one UTF-32 code unit
L'a'wchar_t4 (Linux), 2 (Windows)the wide literal encoding (chapter 9)
'ab'int4value implementation-defined. GCC warns with -Wmultichar

Table 20.4

A common misconception. 'a' is a char, so its size is 1”

True in C++ and false in C. Measured, the very same sizeof('a') is 4 in C and 1 in C++.

The standard’s sentence is “an integer character constant has type int” (§6.4.4.4p11). Its value is “what results when an object of type char holding that character is converted to int”, so the value is what you expect and only the type is wider.

Where it shows is mostly sizeof, _Generic, and overloading on the C++ side. Write only C and the practical harm is near zero, but in a header used from both languages it must be known.

The escapes are exactly these (§6.4.4.4).

KindNotationNote
must be escaped\' \\the single quote and the backslash must take this form
optional\" \?inside a string \" is needed
non-graphic characters\a \b \f \n \r \t \vtheir meanings are defined in §5.2.3
octal\ + octal digitsat most three digits
hexadecimal\x + hex digitsthere is no digit limit
universal character names\uXXXX \UXXXXXXXXnaming characters outside the basic set

Table 20.5

Counter-example. "\x411" — a hex escape eats the letter after it

The standard nails it: each octal or hexadecimal escape sequence is the longest sequence of characters that can constitute the escape sequence (§6.4.4.4p7). Octal stops at three digits; hexadecimal does not stop.

"\x411"      /* not 'A'(0x41) then '1', but a request for 0x411 */

Measured, GCC warns hex escape sequence out of range. The fix is to split the string and let it join — adjacent string literals concatenate (below), so "\x41" "1" is exactly "A1".

20.6 Floating constants (§6.4.4.3)

The decimal form must have either a decimal point or an exponent part. So 1., .5 and 1e3 are all valid, while 1 is an integer constant.

Floating constants take suffixes too (what the types are comes in chapters 8 and 50).

SuffixTypeMeasured size
(none)double8
f Ffloat4
l Llong double16 (x86-64 Linux)
df dd dl_Decimal32 / _Decimal64 / _Decimal128 (C23)4 / 8 / 16

Table 20.6

There are also hexadecimal floating constants (C99) — 0x1p-3 is exactly 0.125.

Q. Why must a hex float have the p exponent, and why use one at all?

A. Because e is unavailable: in hexadecimal e is the digit 14 and cannot start an exponent. So p, meaning a binary exponent, was given its own place, and p cannot be omitted — without it there is no telling where the significand ends. 0x1p-3 is “1 × 2−3”.

The reason to use one is exactness. As chapter 8 showed, decimal 0.1 does not sit exactly in binary, whereas the hex notation transcribes the binary representation itself, so no rounding happens in translation. Hence its use in floating-point tests’ expected values, in the standard library’s tables of constants, and in papers about floating point. printf’s %a prints in the same notation (appendix B).

Worth noting too that the suffix changes the value. Measured, (double)0.1f == 0.1 is false0.1f is the nearest value on the float grid and 0.1 the nearest on the double grid, and those are different numbers (chapters 8 and 50).

20.7 String literals (§6.4.5)

NotationElement typeEncodingMeasured sizeof
"가"charthe literal encoding (chapter 9)4 (3 UTF-8 bytes + NUL)
u8"가"char8_talways UTF-84
u"가"char16_tUTF-164 (1 code unit + NUL)
U"가"char32_tUTF-328
L"가"wchar_tthe wide literal encoding8 (Linux)

Table 20.7

Four properties go together.

  1. A NUL is appended. sizeof "abc" is not 3 but 4.
  2. Adjacent literals join into one (translation phase 6). "hello, " "world" is one string. It is the standard way to split a long string across lines, and the way out of the \x trap above. But the prefixes must not be mixedu"a" U"b" is a compile error (a constraint violation).
  3. A NUL inside does not cut the array short. "a\0b" has strlen 1 and sizeof 4 — the string functions stop, the data is all there.
  4. Modifying one is undefined behaviour. Measured, it usually dies at run time (it is placed in a read-only section). So take string literals as const char *.

Counter-example. char *s = "abc"; s[0] = 'X';

It compiles (in C the type of a string literal is char[N], not const), and it dies when run — SIGSEGV in the measurement.

C++ closed this off entirely (a string literal is const char[N] there, so the assignment is an error). C left it open for compatibility with old code, so the habit has to close it — take it as const char *s = "abc"; and the compiler catches it. Turning on -Wwrite-strings is another way.

20.8 Things that look like constants

NotationWhat it really isMore
RED (an enumeration constant)an integer constant, of type intchapter 55 — it lives in the ordinary-identifier yard
nullptra keyword of type nullptr_t (C23)chapter 36
true falsekeywords yielding bool values (C23)chapter 30
(int[]){1,2,3} a compound literalnot a constant but an object — you can take its addresschapter 47
#define N 100not a constant but token replacementchapter 57
constexpr int n = 10;C23′s real constant — usable in a constant expressionchapter 23

Table 20.8

The last two rows pay off in practice. A macro has neither type nor scope (chapter 57), and a const int is not a constant expression in C — the place where C and C++ part. But “cannot be used” is less accurate than where it cannot be, so here it is, measured.

Given const int n = 10;, writingResult (GCC, C23)
int a[n]; inside a blockaccepted — but as a variable length array, not a constant one
int a[n]; at file scopeerror — variably modified 'a' at file scope
static int a[n]; inside a blockerror — storage size of 'a' isn't constant
case n:error — case label does not reduce to an integer constant

Table 20.9

That is, it fails wherever a real constant expression is required. C23′s constexpr came in to fill that place.

20.9 Expressions — the things that are evaluated into values

Weave literals together with operators and you have an expression, as in 2 + 3 * 4. The definition takes one sentence — that which is calculated (evaluated) into a value. This property of “becoming a value” is the whole of an expression, and half of the eye for reading C from here on. Wherever a value is needed in code, an expression may go there — a single literal is the simplest expression of all.

This chapter has only four operators — plus +, minus -, times *, and parentheses ( ). Following this book’s spiral principle, the remaining operators join in the chapters where they become necessary (comparison in chapter 30, division in chapter 28 — why division was put off, in a moment).

Here is the demonstration. The %d in the example below is a mark meaning “print an integer value here in decimal”; its formal explanation is in chapter 22 — for now it is only a window for seeing results.

examples-en/ch20/expr.c

#include <stdio.h>

int main(void)
{
    printf("%d\n", 2 + 3 * 4);    /* the multiplication is computed first */
    printf("%d\n", (2 + 3) * 4);  /* the parentheses state the order */
    return 0;
}

Output

14
20

20.10 Order — precedence, and the practice of parentheses

The first line of the demonstration is 14 because the multiplication was calculated before the addition. Exactly the convention of the mathematics lesson, and C has fixed such a ranking of “who goes first” — precedence — for every operator.

Here we state this book’s recommendation in advance. Do not memorise the precedence table; use parentheses. C has dozens of operators and the ranking table runs to more than fifteen rows — even people who have memorised it all get confused and cause accidents. Make the order explicit with parentheses, as in the second line, and there is nothing to memorise and nothing to misread. The full table is in the appendix as reference material; the text goes with “arithmetic by the mathematical convention, parentheses everywhere else.”

Q. Why was division put off? Leaving one of the four operations out feels odd.

A. Because integer division differs from division in mathematics. In C 7 / 2 is not 3.5 but 3 — division between integers is the quotient with the fractional part discarded, and it only makes sense understood as a set with its partner operator % (remainder). Including the exact rule of that discarding (which way it goes for negative numbers), it is a subject worth treating properly on top of chapter 7′s world of integers, so it has been given a place in chapter 28. Better to treat it squarely, all at once, than to introduce it half-heartedly now and create the misconception “I thought 7/2 was 3.5.”

Q. Is printf("hello") a value too? A function call is written in the position of an expression.

A. Good eye — it is. Calling a function is itself an expression, and therefore becomes a value. What that value is and where it goes is exactly the next chapter’s subject. That one question has already half-opened the next chapter’s door.

20.11 One caution planted early — the order in time may differ

About precedence there is one caution to plant as a seed. Precedence is a rule of binding — who is whose material — not of what the machine calculates first in time. Which of the left and right materials of an addition is calculated first, for instance, is not fixed by the standard and may differ between compilers. In an expression like 2 + 3 it does not matter at all, but when calculating a material leaves a trace such as output, the order of those traces may differ.

Be reassured on one point, though — the order between statements is guaranteed absolutely. Chapter 19′s “one statement at a time, top to bottom” is a contract. What can waver is only the inside of a single statement, and the exact rules there (side effects, the notion of sequence points) are faced head on in chapter 33, once more material is in place. For now one practical rule suffices: do not cram order-sensitive work into one statement; split the statements.

To summarise — a notation writing a value in source is a literal, that which is evaluated into a value is an expression, and the order of binding is governed by arithmetic convention and parentheses, while the order in time inside one statement may differ. In the next chapter we learn the most important device that consumes and produces values — how to call a function. Hello world’s heart, printf(...), is finally treated head on.