Regex Generator · Guides

Regex Lookahead and Lookbehind Explained: The Zero-Width Trick

You have a string: 42kg, 7lb, 150kg. You want the numbers, but only the kilogram ones, and you want 42 and 150 without the kg attached. A capture group gets you there with post-processing: match (\d+)kg, then throw away group zero and keep group one. It works, but it is annoying, and it gets worse the more context you need to check.

There is a cleaner way. It lets you say "match this, but only when that is next to it" without including "that" in the match at all. It has an intimidating name: lookaround assertions. Two directions, two polarities, four forms total. That is the whole topic.

The core idea: checking without consuming

Most regex tokens consume characters. \d+ eats digits. A lookaround checks a condition at the current position and consumes nothing. Think of it as a bouncer who checks your ID and waves you through without stamping anything: the position in the string does not move, but the match only proceeds if the condition holds.

Lookahead checks what comes next. Lookbehind checks what came before. Each comes in positive ("must be there") and negative ("must not be there") flavors:

The kilogram problem, solved

Back to the opening string. \d+(?=kg) matches 42 and 150, skipping 7 because lb follows it instead of kg. The kg never appears in the match. No groups, no post-processing, no stripping.

The mirror image: (?<=\$)\d+ against $42, €15, $99 matches 42 and 99 but not 15, because 15 is preceded by €. The dollar sign stays out of the result. This is the pattern I reach for whenever the marker is noise and the value is signal: currency symbols, units, prefixes in log lines.

Negatives work the same way, inverted. \d+(?!%) matches numbers that are not percentages: against 25%, 30, 40% you get 30 (and, watch out, the 2 in 25% can match too, since 2 is not directly followed by %; anchor with \b or match the full number first if that bites you).

The case that made lookaheads famous: password rules

Here is the one real-world pattern most developers meet first. You need a password with at least one lowercase letter, one uppercase letter, one digit, and a minimum length, in any order:

^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$

Why it works: each (?=.*...) anchors at the start of the string (because of the ^) and scans the entire string for its condition, consuming nothing. All three checks run against the same full string. Then .{8,}$ does the actual consuming and enforces the length. Without lookaheads you cannot express "contains X and Y in any order" in a single positional pattern, because regex reads left to right and the characters can appear anywhere.

My honest opinion, though: this pattern existing does not mean you should build password rules out of it. Length beats complexity every time, and the zxcvbn comparison shows why character-class checklists are a weak proxy for real strength. Use the lookahead to implement the rule if you must have one. Prefer fewer, longer-is-better rules.

Where it breaks: engine support

Lookahead works everywhere. Lookbehind does not. JavaScript only gained lookbehind in ES2018, so anything running on an older runtime chokes on (?<=...). Python's re module requires lookbehind patterns to be fixed-width: (?<=\$)\d+ is fine, (?<\$\d+) is not, because the engine needs to know exactly how far back to look. Several other engines have the same restriction.

If your pattern throws on a lookbehind, check two things: the engine's age, and whether anything inside the lookbehind has variable length. The fix is usually to restructure with a lookahead or a capture group instead of fighting the engine.

One more trap, and it is a sharp one: quantifiers inside lookarounds can backtrack catastrophically. ^(?=.*a.*a.*a).*$ looks innocent and can explode on long non-matching strings. Lookarounds do not exempt you from backtracking math. The full story of how that explosion works is in Catastrophic Backtracking: The Regex That Took Down Cloudflare for 27 Minutes.

Frequently asked questions

What is the difference between lookahead and lookbehind?
Direction. Lookahead asserts something about the text after the current position (\d+(?=kg)). Lookbehind asserts something about the text before it ((?<=\$)\d+). Neither consumes characters.
Do lookaheads include the matched context in the result?
No. That is the entire point: they are zero-width. \d+(?=kg) returns 42, not 42kg. If you want the context in your result, use a capture group instead.
Why does my lookbehind throw an error?
Usually one of two causes: your engine predates lookbehind support (JavaScript before ES2018), or the pattern inside the lookbehind has variable length, which Python's re and several other engines forbid. Check the engine version first, then simplify what is inside.
Can I use lookahead in JavaScript?
Yes, lookahead has worked in JavaScript since the beginning. Lookbehind requires ES2018 or later, which covers all modern browsers and Node.js releases from the last several years.
When should I use a lookahead instead of a capture group?
When the surrounding context is a condition, not data. Validating "contains a digit somewhere" is a lookahead job. Extracting "the digits before kg" is also a lookahead job. Extracting both the digits and the unit for later use is a capture group job.

One practical regex guide a week. Get it by email.

Try it on the generator

Lookarounds are easier to trust when you can see them fire. Build the pattern on the homepage, paste your own test strings into the live tester, and watch what matches and what does not.

Keep reading

Catastrophic Backtracking: The Regex That Took Down Cloudflare for 27 Minutes

Why Most Email Regex Patterns Are Wrong (And What to Use Instead)