Regex Lookahead and Lookbehind Explained: The Zero-Width Trick
You have a string: 42kg, 7lb, 150kg. You want the numbers, but only the kilogram ones, and you want 42 and 150 without the kg attached. A capture group gets you there with post-processing: match (\d+)kg, then throw away group zero and keep group one. It works, but it is annoying, and it gets worse the more context you need to check.
There is a cleaner way. It lets you say "match this, but only when that is next to it" without including "that" in the match at all. It has an intimidating name: lookaround assertions. Two directions, two polarities, four forms total. That is the whole topic.
The core idea: checking without consuming
Most regex tokens consume characters. \d+ eats digits. A lookaround checks a condition at the current position and consumes nothing. Think of it as a bouncer who checks your ID and waves you through without stamping anything: the position in the string does not move, but the match only proceeds if the condition holds.
Lookahead checks what comes next. Lookbehind checks what came before. Each comes in positive ("must be there") and negative ("must not be there") flavors:
(?=...)positive lookahead: followed by(?!...)negative lookahead: not followed by(?<=...)positive lookbehind: preceded by(?<!...)negative lookbehind: not preceded by
The kilogram problem, solved
Back to the opening string. \d+(?=kg) matches 42 and 150, skipping 7 because lb follows it instead of kg. The kg never appears in the match. No groups, no post-processing, no stripping.
The mirror image: (?<=\$)\d+ against $42, €15, $99 matches 42 and 99 but not 15, because 15 is preceded by €. The dollar sign stays out of the result. This is the pattern I reach for whenever the marker is noise and the value is signal: currency symbols, units, prefixes in log lines.
Negatives work the same way, inverted. \d+(?!%) matches numbers that are not percentages: against 25%, 30, 40% you get 30 (and, watch out, the 2 in 25% can match too, since 2 is not directly followed by %; anchor with \b or match the full number first if that bites you).
The case that made lookaheads famous: password rules
Here is the one real-world pattern most developers meet first. You need a password with at least one lowercase letter, one uppercase letter, one digit, and a minimum length, in any order:
^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$
Why it works: each (?=.*...) anchors at the start of the string (because of the ^) and scans the entire string for its condition, consuming nothing. All three checks run against the same full string. Then .{8,}$ does the actual consuming and enforces the length. Without lookaheads you cannot express "contains X and Y in any order" in a single positional pattern, because regex reads left to right and the characters can appear anywhere.
My honest opinion, though: this pattern existing does not mean you should build password rules out of it. Length beats complexity every time, and the zxcvbn comparison shows why character-class checklists are a weak proxy for real strength. Use the lookahead to implement the rule if you must have one. Prefer fewer, longer-is-better rules.
Where it breaks: engine support
Lookahead works everywhere. Lookbehind does not. JavaScript only gained lookbehind in ES2018, so anything running on an older runtime chokes on (?<=...). Python's re module requires lookbehind patterns to be fixed-width: (?<=\$)\d+ is fine, (?<\$\d+) is not, because the engine needs to know exactly how far back to look. Several other engines have the same restriction.
If your pattern throws on a lookbehind, check two things: the engine's age, and whether anything inside the lookbehind has variable length. The fix is usually to restructure with a lookahead or a capture group instead of fighting the engine.
One more trap, and it is a sharp one: quantifiers inside lookarounds can backtrack catastrophically. ^(?=.*a.*a.*a).*$ looks innocent and can explode on long non-matching strings. Lookarounds do not exempt you from backtracking math. The full story of how that explosion works is in Catastrophic Backtracking: The Regex That Took Down Cloudflare for 27 Minutes.
Frequently asked questions
What is the difference between lookahead and lookbehind?
\d+(?=kg)). Lookbehind asserts something about the text before it ((?<=\$)\d+). Neither consumes characters.Do lookaheads include the matched context in the result?
\d+(?=kg) returns 42, not 42kg. If you want the context in your result, use a capture group instead.Why does my lookbehind throw an error?
re and several other engines forbid. Check the engine version first, then simplify what is inside.Can I use lookahead in JavaScript?
When should I use a lookahead instead of a capture group?
One practical regex guide a week. Get it by email.