Regex Generator · Guides

Why Most Email Regex Patterns Are Wrong (And What to Use Instead)

I have watched the same email regex bug ship in three different companies. The pattern looked responsible, something like ^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$, and it quietly rejected real customers. A woman named O'Brien could not register because of the apostrophe. A German user with an umlaut in his domain was told his address was invalid. A plus-addressed Gmail user, the kind who uses name+shopping@gmail.com to track who sells their data, got bounced by a signup form. Each time, the fix was the same: stop trying to be clever, and understand what an email regex can and cannot do.

The two failure modes

Email regexes fail in two opposite directions. The strict ones reject valid addresses. The loose ones accept garbage. Most hand written patterns are strict in the wrong places and loose in the wrong places simultaneously, which is quite an achievement.

Too strict: patterns that whitelist specific TLDs, or forbid plus signs, or require the TLD to be 2 to 4 letters. That last one was defensible in 2003. Today there are TLDs like .technology and .international, and country codes are still 2 letters, so {2,4} rejects real domains on both ends. Apostrophes are valid in the local part. So are many characters people never expect. If your pattern does not accept o'brien@example.com and user+tag@example.technology, it is wrong.

Too loose: the classic ^.+@.+\..+$ accepts a@b.c and @@@-shaped nonsense. Worse, patterns without anchors match substrings, so a "validator" built on an unanchored pattern will happily accept not an email user@example.com definitely not because it found an email shaped thing inside.

The pragmatic pattern most validators actually use

Here is what experienced teams converge on, and what JSON Schema's email format check looks like in practice:

^[^\s@]+@[^\s@]+\.[^\s@]+$

Read it as: something that is not whitespace or @, then @, then something that is not whitespace or @, then a dot, then something that is not whitespace or @. It is deliberately permissive. It will not reject O'Brien, umlauts, plus addressing, or new TLDs. It catches the actual typos users make: missing @, missing domain, spaces pasted from a spreadsheet.

This is a willful violation of RFC 5322, and that is the point. The HTML5 spec's email validation does exactly this kind of deliberate simplification, because the RFC grammar allows comments, whitespace folding, and quoted strings in ways no real signup form should accept. Full RFC compliance in a regex is thousands of characters long and still cannot tell you whether the inbox exists.

What a regex cannot do (and what to do instead)

A regex answers one question: does this string have the shape of an email address? It cannot answer the question you actually care about: can I reach this person? Only the mail server knows that.

The correct validation stack, in order:

If you are validating emails server side for security purposes, use a mature RFC parsing library, not a hand rolled pattern. The edge cases that matter there, header injection via newlines, quoted local parts, address literals, are exactly the ones a regex gets wrong.

My rule of thumb

Be liberal in what your pattern accepts and strict in what your confirmation flow requires. Every false rejection is a customer you turned away at the door. Every false acceptance costs you one bounced verification email, which your system already handles. The asymmetry is enormous, and it points in one direction: when in doubt, accept the shape and verify by sending.

Frequently asked questions

Should I use the full RFC 5322 email regex?
No. It is enormous, slow, and still cannot verify that an inbox exists. It also accepts formats like comments and quoted strings that no real user will ever type into your form. Use a permissive shape check plus a confirmation email.
Why do so many patterns require the TLD to be 2 to 4 characters?
It is a fossil from the early web, when TLDs were .com, .org, .net, and two letter country codes. Modern TLDs like .technology or .photography are much longer. Use {2,} with no upper bound, or skip the length check on the TLD entirely.
Is the plus sign valid in email addresses?
Yes. Plus addressing (user+tag@example.com) is valid and widely used, especially by Gmail users who track which services share their address. Any pattern that rejects + is rejecting real users. The pragmatic pattern above accepts it.

Try it on the generator

Theory is nice, but patterns earn trust against real strings. Build the pattern on the homepage, paste your own test cases into the live tester, and watch both sides: what should match and what should not.

Keep reading

Phone Number Regex: Why You Should Not Write Your Own