String Title Case Conversion – From Intuition to Precision
You have a sentence like "the quick brown fox jumps". It looks plain. Now imagine you want it to look like a book title: "The Quick Brown Fox Jumps". Every important word starts with a capital letter. That transformation is title case conversion.
The intuition is simple: you want the first letter of each word to be uppercase, and the rest of the letters in that word to be lowercase. But the moment you try to do this for real text, you run into edge cases. What about "don't"? Should it become "Don'T" (ugly) or "Don't" (natural)? What about "the" in the middle of a title — some style guides keep it lowercase ("The Lord of the Rings"), others capitalise it. And what about a word like "iPhone" — should it become "Iphone" (destroying the brand name) or stay as is?
So the intuitive idea is clear, but the precise definition depends on what rules you adopt. In programming, the most common version is the simplest one: capitalise the first character of every word, and make every other character lowercase. That is the definition used in most standard library functions (like Python's str.title() or JavaScript's String.prototype.toTitleCase() — though the latter isn't built-in).
Title Case (simple definition)
For a string S consisting of words separated by whitespace, the title-case version T is:
T=join({word[0].upper()+word[1:].lower()∣word∈split(S)})
Let's break that down with an example.
Step-by-step with "hello WORLD"
- Split the string into words:
["hello", "WORLD"]
- For each word:
"hello": first character 'h' → uppercase 'H'. Rest "ello" → lowercase "ello". Result: "Hello"
"WORLD": first character 'W' → uppercase 'W'. Rest "ORLD" → lowercase "orld". Result: "World"
- Join the results with a space:
"Hello World"
That's it. The entire operation is: first letter up, rest down.
Where the simple definition breaks
Consider "they're here". Splitting by spaces gives ["they're", "here"]. Applying the rule:
"they're" → "They'Re" (because the rest after the first character is "hey're", and lowercasing that gives "hey're" — wait, that's wrong. Let's do it carefully.)
Actually, "they're" has 7 characters: t h e y ' r e. First character 't' → 'T'. The remaining substring is "hey're". Lowercasing that gives "hey're". So the result is "They're". That's correct! The problem is with words like "don't" where the apostrophe is inside the word — the simple split-by-space rule actually handles that fine. The real trouble is with punctuation attached to words: "hello, world" splits into ["hello,", "world"], and "hello," becomes "Hello," (first char 'h' → 'H', rest "ello," lowercased → "ello,"). That's fine too.
The genuine edge case is words that are intentionally mixed-case, like "McDonald" or "iPhone". The simple rule would turn them into "Mcdonald" and "Iphone", which is probably not what you want. But for a first encounter, that's a detail you handle later. The core idea is rock-solid.
The simple title-case conversion destroys intentional capitalisation inside a word. "iPhone" becomes "Iphone". If you need to preserve brand names or proper nouns, you need a smarter algorithm (often called "true title case" or "intelligent title case").
Why does this matter for exams? …