Phone Number Regex: Why You Should Not Write Your Own
A few years ago I reviewed a signup form that validated phone numbers with this pattern: ^\(\d{3}\) \d{3}-\d{4}$. It accepted exactly one format: (555) 123-4567. It rejected 555-123-4567, +1 555 123 4567, and every phone number on Earth outside North America. The company had just launched in the UK. I wish I were exaggerating.
Phone numbers look regular. They are not. Every country runs its own numbering plan with its own digit counts, its own trunk prefixes, and its own rules about which ranges are actually assigned. A regex can check shape. It cannot know numbering plans. That gap is where hand written phone validation goes to die.
What goes wrong with the homemade regex
The typical evolution goes like this. Version one handles your home country. Version two adds an optional country code. Version three tries to handle international numbers and becomes ^\+?[1-9]\d{7,14}$, which accepts +11111111111 and rejects nothing meaningful. Each version is a little more permissive and a little less useful, converging on a pattern that verifies the string contains digits, which you already knew.
The specific traps:
- Trunk prefixes differ by country. A UK number written 07700 900123 becomes +447700900123 internationally: the leading zero is dropped. Italy keeps its leading zero for landlines. Guess wrong and every digit after shifts onto a different subscriber.
- Length is not uniform. E.164 caps numbers at 15 digits total, but valid lengths vary wildly within that. A pattern with a fixed digit count is wrong for most of the world.
- Format is not identity. Parentheses, spaces, and dashes are presentation. +44 7700 900123 and +447700900123 are the same number. Storing the formatted version as the identity breaks deduplication and lookups.
- Numbering plans change. Countries add area codes, split ranges, and reassign prefixes. Your regex is a snapshot; the world keeps moving.
What to use instead: libphonenumber
Google's libphonenumber is the de facto standard, and it exists precisely because this problem is a database problem, not a pattern problem. It ships with per country metadata: valid length ranges, trunk prefix rules, and which prefix ranges are actually assignable. It distinguishes a number that is well formed from a number that is actually possible in a given country.
In JavaScript, the lightweight port is libphonenumber-js:
import { parsePhoneNumberFromString } from 'libphonenumber-js';
const phone = parsePhoneNumberFromString('+1 213 373 4253');
phone.isValid(); // true
phone.formatInternational(); // +1 213 373 4253
phone.formatNational(); // (213) 373-4253
The default metadata build is about 80KB, with smaller 65KB and fuller 145KB variants. That is the entire world's numbering plans in less space than a hero image. When a country changes its plan, you update the library, not your regex.
Where regex still has a role
I am not saying banish regex from phone handling. It has two legitimate jobs:
- Input sanitization: strip everything that is not a digit or a leading + before parsing.
value.replace(/[^\d+]/g, '')is honest work. - Normalizing presentation: reformatting a parsed number for display is a fine use of string replacement on a number you have already validated.
What regex should not do is decide whether a number is real. It should not pretend to be a numbering plan database.
The practical pattern I recommend
- Client side: light format check for instant feedback, then parse with libphonenumber-js. Pair it with a country selector so local numbers like 07700 900123 get parsing context instead of guesswork.
- Server side: parse and validate with the full libphonenumber port for your language. Store the E.164 form (+447700900123) as the canonical identity. Format for humans at display time.
- When you need to know if the number is active: neither regex nor libphonenumber can tell you. That requires a verification SMS or a lookup API that checks carrier and line status. Format validation and existence validation are different problems; do not confuse them.
Normalize for machines. Format for humans. And let the library that tracks 200+ numbering plans do the job your regex never could.