SP (Space) is ASCII code point 32 (0x20), Unicode U+0020, and a printable character — though it prints nothing visible. It is the most common whitespace character, used to separate words in virtually every written language that uses the Latin, Cyrillic, or Greek alphabets. In programming, space is both trivial and treacherous. Most languages use it to separate tokens, but Python makes it syntactically meaningful — indentation with spaces (or tabs) defines code blocks. In HTML, multiple consecutive spaces collapse into one unless you use ` ` or `white-space: pre`. URLs encode it as `%20` or `+`, a distinction that has quietly broken countless web forms. It originated on the telegraph, where operators needed a way to signal "advance the carriage but print nothing." Typewriters inherited this as the space bar, and ASCII formalized it at position 32 — the first code point after the 33 control characters (0–31 and 127). UTF-8 encodes it as a single byte 0x20, identical to its ASCII value. Unicode classifies it as Zs (Separator, Space) with the name SPACE. It is not the same as the non-breaking space (U+00A0), the em space (U+2003), or the zero-width space (U+200B) — Unicode defines over a dozen space-like characters, each with subtly different typographic behavior.
const char *s = "a b c";int words = 0; // count transitions into non-spacefor (int i = 0; s[i]; i++) if (s[i] != ' ' && (i == 0 || s[i-1] == ' ')) words++;
Space (ASCII 32) is the bridge between control characters and printable ones — it's the first character you can actually 'see' (as blank area). It's the most typed character in any language and the primary token separator in programming, natural language, and data formats.
Space (32) allows line breaks; Non-Breaking Space (160 or ) prevents them. They look identical but are treated differently by browsers and parsers. NBSP is crucial for keeping units with numbers (e.g., '10 kg') but can cause confusing bugs in code if accidentally used instead of a normal space.
HTML's whitespace collapsing rule treats any run of spaces, tabs, and newlines as a single space in rendered output. This was a design choice to let authors format source code freely. Use or CSS white-space: pre to preserve exact spacing.
Both represent a space, but in different contexts. %20 is the universal URL encoding. The + sign only means space in application/x-www-form-urlencoded data (like form submissions). Using + in a URL path is technically incorrect — always use %20 there.
Calling `split()` (no arguments) splits on *any* run of whitespace and removes empty strings, which is usually what you want. Calling `split(' ')` splits strictly on single spaces, preserving empty strings. This subtle difference causes many data parsing bugs.
Use the `\s` character class. It matches space (32), tab (9), newline (10), carriage return (13), form feed (12), and often vertical tab (11). It's safer than typing a literal space because it catches all the invisible separators at once.