Skip to content

Incorrect "Fix bad UTF-8 char " step in Dockerfile actually corrupts the input #33

Description

@alamb

I spent a large amount of time confused about this and wanted to file an issue

I found it in the context of

The Dockerfile in this repo implies that the DSGen data generator for TPCH has incorrect input:

# Fix bad UTF-8 char
RUN iconv -f ISO-8859-14 -t UTF-8 tpcds.dst > tpcds.dst2
RUN mv tpcds.dst2 tpcds.dst

However, what that command actually does is corrupt the file

For example,t he originl data has a CÔTE in it: (\x43 \xc3 \x94 \x54 \x45 in utf8)

grep -n "IVOIRE" tpcds.dst | head -1 | xxd | head -3
echo "---"…)
⎿  00000000: 3634 393a 6164 6420 2822 43c3 9454 4520 649:add ("C..TE
00000010: 4427 4956 4f49 5245 223a 3129 3b0a D'IVOIRE":1);.

iconv -f ISO-8859-14 -t UTF-8

Tells iconv "treat the input as ISO-8859-14 single-byte data." But the file was already UTF-8. So iconv took each existing UTF-8 byte (e.g. C3 and 94) as if each were a separate Latin character, and re-encoded each one into UTF-8.

This results in

tpcds.dst (original) — already correct UTF-8:

  • C3 94 = the valid UTF-8 encoding for Ô (U+00D4 LATIN CAPITAL LETTER O WITH CIRCUMFLEX)

tpcds.dst2 (after iconv) — corrupted (double-encoded):

  • C3 83 C2 94 = two characters: Ã (U+00C3) + a control character (U+0094)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions