Algorithms

An algorithm generates plausible variants of a name. There are 32, each one modelling a specific way a name goes wrong — the error taxonomy behind them is chapter 3.

urlinsane typo --list algorithms      # what this build has

The full set

Applies to is blank where an algorithm binds by capability rather than by type: those run on any nameable node, domain or package or handle alike.

ID Name Applies to What it does
aci Adjacent Character Insertion any Insert a character adjacent on the keyboard: googhle
acs Adjacent Character Substitution any Replace a character with a keyboard neighbour: ezample
afx Affix Squatting package, repo, username Add a plausible prefix or suffix: node-acme, acme-js
bf Bit Flipping any Flip one bit of a character — bitsquatting
cb Combo Squatting any Append or prepend a common keyword: acme-login
cm Common Misspellings any Apply a curated misspelling for the language
cns Cardinal Substitution any Swap a number for its cardinal word and back: file2filetwo
co Character Omission any Drop a character: gogle
cr Character Repetition any Double a character: gooogle
cs Character Swapping any Transpose two adjacent characters: examlpe
dhs Dot Hyphen Substitution any Swap dots and hyphens: my-acme.commy.acme.com
di Dot Insertion any Insert a period: exa.mple
do Dot Omission any Remove a period: wwwacme.com
fsd Delegated Subdomain domain Put the name under a host that gives subdomains away: paypal.duckdns.org
gi Grapheme Insertion any Insert a grapheme from the language’s alphabet
gr Grapheme Replacement any Replace a grapheme with another from the alphabet
hi Hyphen Insertion any Insert a hyphen: ex-ample
ho Hyphen Omission any Remove a hyphen: oneforall
hr Homoglyph Replacement any Replace a character with one that looks the same
hs Homophone Substitution any Replace a word with one that sounds the same
nsc Namespace Confusion package, repo Move a name between namespaces or scopes
ons Ordinal Substitution any Swap a number for its ordinal word and back
rar Repetition Adjacent Replacement any Double a character, then replace the double with a neighbour: gppgle
sep Separator Substitution package, repo, username Swap the separator a registry allows: -_.
si Subdomain Insertion domain Insert a subdomain label
sld Wrong Second-Level Domain domain Swap the second level under a ccTLD: bbc.co.ukbbc.org.uk
sp Singular Pluralise any Make a word singular or plural
tld Wrong TLD domain Substitute a different public suffix
tos Token Order Swap any Reorder the words: shop-onlineonline-shop
tli TLD Insertion domain Append a suffix, making the whole name a subdomain: example.com.br
vs Vowel Swapping any Swap one vowel for another: ixample
xhs Cross-language Homophone any Swap for a spelling that sounds the same in another language: youtubeyutup

Selecting them

urlinsane typo -a cs acme.com              # only transpositions
urlinsane typo -a cs,co,acs acme.com       # three of them
urlinsane typo -a '^bf' acme.com           # everything except bit flipping
urlinsane typo acme.com                    # everything (the default)

^id excludes. You can check what a selection actually compiled to without running anything:

$ urlinsane typo --explain -a cs,co acme.com | sed -n '/operators/,$p'
operators
  co             on nameable           where in-seed-closure
  cs             on nameable           where in-seed-closure
  decompose.domain on domain
  …

Note the where in-seed-closure condition. Variant operators only fire on nodes inside the seed closure, which is what stops the scan from generating variants of every nameserver it happens to discover. See Limits.

Which ones to run

Running all 32 is the default and is right for a one-off audit where you have time. For anything repeated, pick by threat model.

A consumer-facing brand. Human error dominates, and reading errors dominate over typing errors:

urlinsane typo -a cs,co,acs,aci,cr,vs,cm,hs,sp,hr,tld acme.com

A package or a repo. Convention exploitation dominates; keyboard geometry barely matters, because the name is usually copied rather than typed:

urlinsane typo -a afx,nsc,sep,cb,co,cs,hr npm:acme-utils

A quick triage pass. The four highest-yield generators against domains, and fast:

urlinsane typo -a co,cs,acs,tld -d 1 acme.com

Everything except the noisy one. bf produces the largest and least human-plausible set:

urlinsane typo -a '^bf' acme.com

Notes on individual algorithms

bf (bit flipping) does not model a human. Its output looks like nonsense because the error source is memory corruption in a resolver, not a typist — see chapter 3. It generates a lot. Exclude it unless you specifically want bitsquat coverage.

hr (homoglyphs) is the algorithm you cannot replace with careful reading. Its output often looks identical to the seed in a terminal. Pair it with the idn operator’s punycode output to see what is really there.

xhs is the one generator that is not about a single language. hs swaps a word for one that sounds like it in the same language — base and bass. xhs swaps for one that sounds like it to a speaker of a different language: youtubeyutup, boutiqueboetiekbutik, babybebi. It catches the person who heard a name in a language whose orthography spells that sound differently and wrote down what they heard, which a curated English homophone list cannot reach by construction. X-squatter (ACM TOPS 2024) measured these and found ~15% carry TLS certificates against 7% for other squatting types — they are provisioned, not just registered.

fsd is tld narrowed by risk rather than by shape. tld swaps the suffix for any of the ~7,000 ICANN entries, where taking a name costs money and leaves a WHOIS trail. fsd uses only the ~3,000 private entries of the public suffix list — hosts that delegate subdomains to anyone, like duckdns.org, github.io, vercel.app — where it costs nothing, takes a minute and leaves no registration record. Ask for fsd when you want the cheap attacks first.

tos is the only generator that moves whole words. cs transposes two adjacent characters, turning shop-online into shpo-online; tos turns it into online-shop. It matters most where multi-word names are the convention — every package registry — because a developer who half-remembers node-fetch remembers the words, not their order.

cm, hs, vs, gi, gr, cns, ons are language-driven: they read vocabulary out of the dataset database, and what they generate depends on which languages that database carries. An empty or stale dataset makes them generate little or nothing, silently. If these look quiet, check Datasets.

acs, aci, rar are keyboard-driven, and the keyboard is geometry rather than a grid of rows — e’s neighbours on a US layout come out as w r d 3 4 s, in distance order. On QWERTZ, AZERTY or Dvorak the answers differ entirely. Keyboards has the model.

tld substitutes from the public suffix list, which is large. It is the main reason a default scan produces so many candidates against a domain target.

si (subdomain insertion), nsc (namespace confusion) and tli (TLD insertion) model structural attacks rather than errors: the attacker is not imitating a typo but exploiting how a name is parsed.

tli is the one with no misspelling in it at all. example.com.br contains the target name in full, spelled correctly, and is a subdomain of somebody else’s registration — so it defeats the check most people actually perform, which is “does the address contain the name I expect”. It is also what a truncating mobile address bar shows first.

sld exists because tld cannot produce it. tld replaces the whole public suffix, so bbc.co.uk becomes bbc.de — a different country. sld keeps the country and changes the category: bbc.org.uk, bbc.ac.uk. A reader checking “is this the UK site?” gets the right answer and still lands on the wrong name.

--language and --keyboard are registered flags but do not yet reach the plan, so language and keyboard selection is currently whatever the dataset holds rather than something you choose per run.

Cost

Generation is free; observation is not. One algorithm at depth 1 against example.com took 43 seconds in the example in chapter 6 — essentially all of it network. Doubling the algorithm count roughly doubles the candidate set, and the candidate set is what drives the number of DNS and WHOIS calls.

The cheap way to explore is to generate without observing much: -d 1 keeps the scan on the variants themselves and their immediate records, rather than walking out through nameservers and address space.


Next: Observation and depth.


Back to top

URLInsane is licensed under the GPLv3. Copyright © 2024-2026 Rangertaha.

This site uses Just the Docs, a documentation theme for Jekyll.