Customizing the parser

Every piece of nameparser configuration sorts into one of three places by asking what it varies with: vocabulary varies by language (Lexicon), behavior varies by data source or application (Policy), and presentation varies by output destination (a rendering argument). See How the parser works for why the split is drawn there.

Vocabulary: Lexicon

Adding and removing words

>>> from nameparser import Lexicon, Parser
>>> lex = Lexicon.default().add(titles={"dean"})
>>> Parser(lexicon=lex).parse("Dean Robert Johns").title
'Dean'

add() and remove() both return a new Lexicon — the one you started from (here, Lexicon.default()) is never mutated. Every field accepts a plain set of lowercase words, keyword by field name (titles above; particles, suffix_words, and the rest work the same way) — see API reference for the full field list.

The default word lists themselves — TITLES, PARTICLES and the other frozensets in nameparser.config — are frozen, so a runtime addition belongs on a Lexicon as above, or on a private Constants if you are still parsing through HumanName. REGEXES and CAPITALIZATION_EXCEPTIONS are the two members the freeze does not cover — they are still plain dicts. Editing one at runtime is not a supported override, and it is not a clean no-op either: the edit reaches a freshly built Constants, while the shared CONSTANTS (copied at import) and the cached default() never see it. That is the same inconsistent reach the freeze removed for the word lists, so these overrides belong on a config object too.

Five of these lists were renamed in 2.2 to match the field names used here: PREFIXES, NON_FIRST_NAME_PREFIXES, BOUND_FIRST_NAMES, FIRST_NAME_TITLES and SUFFIX_NOT_ACRONYMS became PARTICLES, NON_GIVEN_NAME_PARTICLES, BOUND_GIVEN_NAMES, GIVEN_NAME_TITLES and SUFFIX_WORDS, and two modules moved with them. Names outside that list, TITLES among them, are unchanged. Every 1.x name still imports, with a DeprecationWarning, until 3.0 — see Migrating from HumanName for the full mapping.

Vocabulary entries are matched one word at a time, with two exceptions, so a multi-word entry like titles={"grand moff"} can never match; the constructor warns when it sees one (capitalization_exceptions keys included — they are looked up per word too). The exceptions are given_name_titles, looked up as the space-joined run of words already read as titles or as that run’s last word — several titles written together are one form of address and the last one does the addressing, so "Her Majesty Queen Elizabeth" is read by queen — and maiden_markers, matched by lookahead over the words as written: maiden_markers={"z domu"} matches the pair and neither word alone, which is how the shipped Polish entry works. The words have to stand together — a bracketed clause or a comma between them ends the run, and the first word is then an ordinary name word. Where a phrase entry and a word entry starting with it are both configured, the phrase wins where it matches and the word matches everywhere else. No warning is raised for a multi-word entry in either of these two fields, since there it is not a mistake.

The limit is on storage, not on the shape a name can have. Adjacent suffix words are reassembled after they match, so a multi-word credential is reachable as its component words even though the phrase itself cannot be stored:

>>> from nameparser import parse
>>> parse("John Smith, MD PhD").suffix
'MD PhD'

That has held since 1.4.0. A credential whose words are not in the default vocabulary is reached by adding those words, not the phrase:

>>> lex = Lexicon.default().add(suffix_acronyms={"leed", "ap"})
>>> Parser(lexicon=lex).parse("John Smith, LEED AP").suffix
'LEED AP'

Removing works the same way, and drops the word from recognition:

>>> lean = Lexicon.default().remove(titles={"professor"})
>>> Parser(lexicon=lean).parse("Professor Robert Johns").title
''

A few fields mark a subset of another — given_name_titles over titles, particles_ambiguous over particles, suffix_acronyms_ambiguous over suffix_acronyms, conjunctions_ambiguous over conjunctions, and honorific_tails over suffix_words. Entries belong in the base field too, so add to both and remove from the marker first. Three of them enforce that — particles_ambiguous, suffix_acronyms_ambiguous and honorific_tails raise ValueError naming the orphans, because an orphan in each of those does real harm rather than nothing. The other two are deliberately unchecked because an orphan there is inert: given_name_titles matches a title run as one space-joined string, or by that run’s last word, so a legitimate entry like "sir and dame" is no single word in titles; and a conjunctions_ambiguous entry is only ever read for a word that is a conjunction, so remove(conjunctions={"e"}) simply works and the stale marker entry is never consulted.

Turning title detection off

The subset rule matters most when clearing a field wholesale. Emptying titles alone orphans every given_name_titles entry, so the two go together:

>>> d = Lexicon.default()
>>> lean = d.remove(titles=set(d.titles),
...                 given_name_titles=set(d.given_name_titles))
>>> Parser(lexicon=lean).parse("Hon Solo").given
'Hon'

Emptying the title vocabulary does not switch titles off entirely, though. A word ending in a period, standing at the front of the part that carries the given name, is read as a title structurally, without consulting titles at all — that is what lets unfamiliar ranks and abbreviations work (see Titles you didn’t configure):

>>> bare = Parser(lexicon=lean)
>>> bare.parse("Professor John Smith").title      # vocabulary gone
''
>>> bare.parse("Dr. John Smith").title            # structural, stays
'Dr.'

Combining two lexicons

Whole lexicons compose with |, which unions field by field — handy for keeping a shared house vocabulary separate from a per-source one and combining them at parser construction:

>>> house = Lexicon.empty().add(titles={"dean"})
>>> per_source = Lexicon.empty().add(titles={"provost"})
>>> sorted((house | per_source).titles)
['dean', 'provost']

Fixing the case of a particular word

capitalization_exceptions is the one pair-valued field — each entry maps a lowercase key to a case mask: the key’s own letters and digits, each in the case it should take ("phd" → "PhD"). Case repair lays the mask over the word as it was written and keeps every other character where it stood, so it recases a word and never re-spells it. A value that spells anything else raises ValueError when the lexicon is built. Being pair-valued, the field isn’t a fit for add()/remove(). Change it with dataclasses.replace() instead, and pass the result to capitalized():

>>> import dataclasses
>>> from nameparser import parse
>>> str(parse("jane smith dphil").capitalized())
'Jane Smith Dphil'
>>> default = Lexicon.default()
>>> lex = dataclasses.replace(
...     default,
...     capitalization_exceptions=tuple(default.capitalization_exceptions)
...     + (("dphil", "DPhil"),))
>>> str(parse("jane smith dphil").capitalized(lex))
'Jane Smith DPhil'
>>> str(parse("JANE SMITH D.PHIL.").capitalized(lex))
'Jane Smith D.Phil.'

Note the tuple(...) + ...: assigning a bare (("dphil", "DPhil"),) would replace the default exceptions rather than extend them, so the shipped masks (phd, bsc, psyd and the rest) would be lost and those words fall back to the all-capitals acronym repair: john smith phd would give John Smith PHD rather than John Smith PhD.

A mask also outranks the lowercase that case repair gives a surname particle (de la Vega). The shipped map uses that for the Irish particles, which are written capitalized:

>>> str(parse("SEÁN Ó MURCHÚ").capitalized())
'Seán Ó Murchú'
>>> str(parse("JUAN DE LA VEGA").capitalized())
'Juan de la Vega'

How a key matches a word

The key is matched against the token with punctuation normalized away, not against the raw text, so one "phd" entry covers phd, PHD and Ph.D. alike, and you don’t need a separate key for each way a source might punctuate it. (A mixed-case Phd matches too, but mixed case is left as written unless you pass force=True; see Presentation: rendering arguments.)

Each word keeps its own punctuation: Ph.D. repairs to Ph.D., not PhD. Punctuation in the value is never written into the word; it only marks which of the mask’s letters are joined. That matters for one case: a single letter the writer split off beside a full stop is an initial, and is capitalized even where the mask keeps that letter inside a longer run:

>>> str(parse("john smith ph.d.").capitalized())
'John Smith Ph.D.'
>>> str(parse("john smith p.h.d.").capitalized())
'John Smith P.H.D.'

Words that need no entry

An acronym already listed in suffix_acronyms, plain or dotted, and a roman numeral are written in capitals by case repair on its own. Most listed acronyms whose usual spelling is not all capitals already carry a shipped mask (DSc, PsyD, PharmD, MDiv and others):

>>> str(parse("john smith md iv").capitalized())
'John Smith MD IV'
>>> str(parse("john smith psyd").capitalized())
'John Smith PsyD'

The exception is an acronym that is also a name word, such as meng or edd. A mask applies wherever its word stands, so it would re-spell a person called Meng or Edd, and these get none: they repair in plain capitals. The reasoning is the Excluded block for CAPITALIZATION_EXCEPTIONS under R4 in docs/design/decisions.md.

>>> str(parse("john smith edd").capitalized())
'John Smith EDD'

Acronyms the vocabulary doesn’t list

An acronym suffix_acronyms doesn’t list repairs according to how it is written:

  • Plain (dphil), it parses as an ordinary name word and repairs as one: Dphil, as in the first example above.

  • Dotted (d.phil.), it is a suffix by shape alone (see Credentials the vocabulary doesn’t list) and repairs in all capitals the same as a listed acronym does: D.PHIL..

Either way, give it a suffix_acronyms entry, or a mask of its own, rather than relying on this fallback.

Words that are also ordinary names

Some vocabulary words are also ordinary name words: an acronym suffix that is also borne as a surname, a particle that doubles as a given name, a connective letter that doubles as an initial. Three fields mark them — suffix_acronyms_ambiguous, particles_ambiguous and conjunctions_ambiguous, one each for suffix_acronyms, particles and conjunctions. They add no vocabulary by themselves; they narrow how an existing entry is read when it appears alone, and the parse reports the fork as an ambiguity.

Credentials that are also surnames

ma is a shipped example of suffix_acronyms_ambiguous. It is both a credential and a common surname, so it is listed there. Written with its periods it is a suffix outright; written bare, the capitalization decides:

>>> parse("Jack Ma").family
'Ma'
>>> parse("Jack M.A.").suffix
'M.A.'
>>> parse("Jack MA").suffix
'MA'

The full reading of a bare marked acronym, in order:

  • ALL CAPITALS in a mixed-case name leans to the credential: it is a suffix even with no words to spare ("Jack MA").

  • Title-case in a mixed-case name leans to the surname: it stays the family name even WITH words to spare ("John Smith Ma").

  • No signal — all lowercase in a mixed-case name, or any spelling in a name written wholly in one case — falls back to the count: a suffix with two or more words before it, not counting a title or a nickname ("John Smith ma"), the family name otherwise ("JACK MA").

  • After a comma the count of name words before the comma decides FIRST, and the case is read only where the count leaves the word a name. "John Smith, Ba" reads suffix Ba on the count alone, and "Smith, BA" reads suffix BA on the capitals lean. What still reads as the given name is "Smith, Ba" (one word, and Title-case has no credential lean) and "smith, ba" (one word, one case, no lean at all).

  • In brackets, "John Smith (BA)" falls through to nickname parsing.

Mark a word, or leave it out

When you add a word that is both a credential and a name, weigh how often it is one against the other. There are three choices, and none is free:

  • Unambiguous (suffix_acronyms alone): the credential reading is taken silently, and a wrong claim can lose a real person’s surname.

  • Marked ambiguous (suffix_acronyms_ambiguous too): the reading depends on the writing as above, and every bare reading reports the fork.

  • Left out of suffix_acronyms altogether: right when the name reading is far more common, as for an acronym whose credential is tenuous or specialized beside a common surname. That is what the default vocabulary did with rai and cha in 2.3; a caller who needs one adds it back with Lexicon.default().add(suffix_acronyms={"cha"}).

The same conservatism is why dean above isn’t in the default vocabulary in the first place: “Dean” is also a common given name, and a default that swallowed it as a title would misparse “Dean Martin” for everyone.

Particles that are also given names

particles_ambiguous is the same idea for surname particles. A particle listed there may also be a given name, which is what makes a leading one a decision to take; a particle not listed there never is, so there is nothing to decide. That shows up in what a particle standing alone at the front of a name does: a listed one is a name part in its own right, while an unlisted one pulls the rest of the name into the surname and leaves no given name at all. Which field a listed particle lands in is name_order’s question, covered below; an unlisted one opening the name is the surname under every order, because a word that can never be a given name leaves the order nothing to decide.

>>> parse("van Gogh").given          # 'van' may be a given name
'van'
>>> parse("de Mesnil").given         # 'de' may not
''
>>> parse("de Mesnil").family
'de Mesnil'

A comma forestalls the question rather than answering it. Writing the surname before the comma has already said which words are the surname, so a particle at the front of them decides nothing, and whatever follows the comma is the given name as usual:

>>> parse("de Mesnil, Juan").given   # the comma named the surname
'Juan'

If your data never uses Van as a given name, take it out of the ambiguous set: a leading van is then no decision at all, so no ambiguity is recorded and it becomes part of the surname — under any name_order, since that is what taking the word out asserted:

>>> lex = Lexicon.default().remove(particles_ambiguous={"van"})
>>> Parser(lexicon=lex).parse("van Gogh").family
'van Gogh'

For particles_ambiguous the default runs the other way from credentials: a particle that is not borne as a given name belongs outside the set, which is where mc and ste were moved (#360).

One-letter connectives that are also initials

conjunctions_ambiguous is the same idea for one-letter connectives. A single letter written against the name’s own case is an initial and one written with it is the connective — but a name written wholly in one case, all upper or all lower, says nothing either way, and this is the set that decides it there. e and i are the two entries shipped: a bare E or I initial is common where those letters between two surnames are rarer, and y runs the other way, so y joins even written as a bare capital.

>>> parse("jose e maria santos").middle       # 'e' reads as an initial
'e maria'
>>> parse("JUAN GARCIA Y LOPEZ").family       # 'y' joins
'GARCIA Y LOPEZ'
>>> parse("Jose e Maria Santos").given        # mixed case decides itself
'Jose e Maria'

A member also reports the fork, so a caller can see which reading was taken:

>>> [a.kind for a in parse("jose e maria santos").ambiguities]
[<AmbiguityKind.CONJUNCTION_OR_INITIAL: 'conjunction-or-initial'>]

If your data is Portuguese, where e links surnames the way y does in Spanish, take it out and the connective reading comes back:

>>> lex = Lexicon.default().remove(conjunctions_ambiguous={"e"})
>>> Parser(lexicon=lex).parse("jose e maria santos").given
'jose e maria'

If your data is Catalan or Polish, where i links two surnames the way y does in Spanish, take that one out instead and the link joins in a one-case name too:

>>> lex = Lexicon.default().remove(conjunctions_ambiguous={"i"})
>>> Parser(lexicon=lex).parse("josep carod i rovira").family
'carod i rovira'

If your data is Dutch, where a bare single letter is an initial and never a connective, add the other one instead:

>>> lex = Lexicon.default().add(conjunctions_ambiguous={"y"})
>>> Parser(lexicon=lex).parse("juan garcia y lopez").middle
'garcia y'

Bound given names

bound_given_names holds given-name prefixes that attach to the following word to form one given name — abdul, abu, umm and their Arabic-script spellings (عبد, أبو, أم) among them:

>>> parse("abdul salam ahmed salem").given
'abdul salam'

Add your own, or empty the set to switch the behavior off entirely:

>>> lex = Lexicon.default().add(bound_given_names={"mohamad"})
>>> Parser(lexicon=lex).parse("mohamad salam ahmed salem").given
'mohamad salam'
>>> d = Lexicon.default()
>>> off = d.remove(bound_given_names=set(d.bound_given_names))
>>> Parser(lexicon=off).parse("abdul salam ahmed salem").given
'abdul'

Behavior: Policy

When your data source or application needs different parsing behavior — a different name order, stricter suffix rules, extra delimiters — set it on Policy, a small, closed set of fields, listed below.

Field

Type

Effect

name_order

one of the three exported order constants

Assigns positional (no-comma) input to given/middle/family in this order. Use the exported GIVEN_FIRST (default), FAMILY_FIRST, or FAMILY_FIRST_GIVEN_LAST constants. Ignored when a comma separates family from given (“Thomas, John” puts the family name first); a comma that only sets off suffixes (“John Smith, Jr.”) leaves it governing the name part. See Family-first name order.

script_orders

pairs of Script and an order

Assigns a name in one of these scripts in the paired order, whatever name_order says. Defaults to family-first for a name wholly in Han or Hangul, or in kanji and kana other than katakana alone. See East Asian defaults, and turning them off.

segment_scripts

frozenset[Script]

Scripts whose unspaced names are split into surname and given name. Defaults to Hangul. See East Asian defaults, and turning them off.

patronymic_rules

frozenset[PatronymicRule]

Reorders patronymic-shaped names via opt-in detectors — East Slavic formal order (EAST_SLAVIC) or Turkic reversed order (TURKIC) — but stands down under a declared FAMILY_FIRST or FAMILY_FIRST_GIVEN_LAST name_order. Defaults to empty.

middle_as_family

bool

Folds middle into family instead of splitting them — for naming systems with no middle-name concept. Defaults to False.

nickname_delimiters

frozenset[tuple[str, str]]

Routes content enclosed by these delimiter pairs to nickname. Defaults to DEFAULT_NICKNAME_DELIMITERS — straight quotes and parentheses plus the typographic conventions (smart quotes, guillemets, CJK brackets, …). See Nicknames, maiden names, and brackets.

maiden_delimiters

frozenset[tuple[str, str]]

Routes content enclosed by these delimiter pairs to maiden instead of nickname. Needed only for a clause that does not announce itself: "(née Jones)" is a maiden name without it. Defaults to empty. See Nicknames, maiden names, and brackets.

extra_suffix_delimiters

frozenset[str]

Adds separators that split suffix groups, e.g. " - " for "Jane Smith, RN - CRNA". Additions only — the comma always splits suffix groups and cannot be replaced. See Suffixes not separated by commas.

lenient_comma_suffixes

bool

Reads an initial-shaped suffix word after a comma as a suffix: "John Smith, V" is John Smith the fifth when True (default); False reads V as a given-name initial instead. Multi-letter suffixes (III, MD) are unaffected.

unlisted_dotted_suffixes

bool

Reads an unlisted dotted acronym such as X.Y.Z. as a credential where the name’s shape allows it. Defaults to True. See Credentials the vocabulary doesn’t list.

unlisted_caps_suffixes

CapsSuffixes

Where an unlisted all-caps word such as XYZ reads as a credential: after a comma (AFTER_COMMA, the default), at the end of a name as well (EVERYWHERE), or nowhere (OFF). See Credentials the vocabulary doesn’t list.

strip_emoji

bool

Excludes emoji from tokenization — they appear in no field or rendered view, though original keeps them. Defaults to True. See Keeping emoji and control characters.

strip_bidi

bool

Excludes bidirectional control characters the same way. Defaults to True. See Keeping emoji and control characters.

To apply a PolicyPatch directly – without going through a locale pack – call Policy.patched():

>>> from nameparser import Policy, PolicyPatch
>>> Policy().patched(PolicyPatch(middle_as_family=True))
Policy(middle_as_family=True)

Family-first name order

name_order is the one most likely to matter for data that is not in Western order. Positional input is assigned in the order you declare — with the two vocabulary exceptions noted under Where the vocabulary answers first — so a name written family-first — Hungarian, here — parses as written instead of needing to be rearranged afterwards:

>>> from nameparser import Parser, Policy, FAMILY_FIRST, parse
>>> parse("Nagy Laszlo Peter").family            # default GIVEN_FIRST
'Peter'
>>> family_first = Parser(policy=Policy(name_order=FAMILY_FIRST))
>>> name = family_first.parse("Nagy Laszlo Peter")
>>> name.family, name.given, name.middle
('Nagy', 'Laszlo', 'Peter')

An explicit comma still wins, on the reasoning that someone who wrote one meant it — so the same parser reads "Thomas, John" as family-then-given regardless of the configured order:

>>> family_first.parse("Thomas, John").family
'Thomas'

A Vietnamese full name needs a third order. It is written family, then middle, then given — the name a person is actually called by is the last word, not the second. Family-first order gets the family name right and then reverses the remaining two, so FAMILY_FIRST_GIVEN_LAST exists for the names that read this way:

>>> from nameparser import FAMILY_FIRST_GIVEN_LAST
>>> family_first.parse("Tran Quoc Toan").given       # FAMILY_FIRST
'Quoc'
>>> given_last = Parser(policy=Policy(name_order=FAMILY_FIRST_GIVEN_LAST))
>>> viet = given_last.parse("Tran Quoc Toan")
>>> viet.family, viet.middle, viet.given
('Tran', 'Quoc', 'Toan')

Nothing keys this order to a script the way the East Asian defaults below do — Vietnamese is written in the Latin alphabet, which carries no order of its own — so it applies only where you set it, and there is no vn locale pack.

Declaring the order settles where a surname ends

A surname particle joins forward, onto the word after it. Where a particle ends the name there is nothing ahead of it to join, and what it belongs to is decided by what the writing says rather than by the word. Two things say it, and both amount to someone stating that the family name came first — a family comma, and a declared family-first order:

>>> parse("Jong, Anke de").family                  # the comma says so
'de Jong'
>>> family_first.parse("Jong Anke de").family      # the order says so
'de Jong'

The comma and the order are one shape

That pair is not a coincidence but one shape written two ways: form 4 (Title Family Given Middle Middle [Particle] [, Suffix]) is form 2 (Family [Suffix], Title Given (Nickname) Middle Middle[,] Suffix [, Suffix]) with the comma removed and the family folded inline — titles included, so a title that form 2 writes after the comma leads the name in form 4 instead. If your records are family-first without commas, Policy(name_order=FAMILY_FIRST) reads them the way the comma format is already read. Measured over the whole particle vocabulary — every particle nameparser ships, crossed with three families and three given names, 630 generated pairs in all — 603 of 630 agree (2026-08-30); the executable form of this correspondence is tests/v2/test_order_correspondence.py.

Where the correspondence stops

Three limits keep that statement honest:

  • It covers one shape, not comma deletion in general. A name whose shape changes when the comma is removed — a title or suffix crossing to a different position — parses as the shape it becomes, not as a disagreeing reading of form 2.

  • A word that is both particle and suffix reads differently in the two writings. The shipped words in both vocabularies are do, mc and vd (2026-10-02). The particle attachment outranks the suffix reading on the comma side alone — that is the scope the rule is stated in — so parse("Ménil, Christophe vd") reads family vd Ménil, while family_first.parse("Ménil Christophe vd") reads family Ménil and suffix vd. It is that asymmetry, not a precedence, that breaks the correspondence, and a listing ending in one of those words is the whole of the 27 disagreeing pairs measured above, not scatter.

  • Only FAMILY_FIRST is reached. It is the only order that puts a trailing piece in the middle, where a particle means nothing. FAMILY_FIRST_GIVEN_LAST puts it in the given slot, where your own declaration already says it is the given name, and no comma format writes the given name last, so form 5 has no comma twin to correspond to in the first place:

>>> given_last.parse("Nguyen Thi Van").given
'Van'

Where a leading particle run stops

The declaration also bounds how far a leading particle run reaches. With no order declared, nothing marks where the surname ends and a particle followed by several words really can be all surname, so the whole name is read as one. Declaring family-first asserts that what follows the family is not more surname, which settles it:

>>> parse("de Mesnil Jean").family                 # nothing says where it ends
'de Mesnil Jean'
>>> family_first.parse("de Mesnil Jean").family    # the order does
'de Mesnil'
>>> family_first.parse("de Mesnil Jean").given
'Jean'

In the default order, write the comma for that reading. Under a family-first order the run stops after one name word rather than one token, so it cannot cut inside a conjunction-joined run or a bound given-name pair — "de la Vega y Santos Juan" keeps family de la Vega y Santos. Where two or more words are left over, the two family-first orders differ from each other:

>>> family_first.parse("de la Cruz Juan Carlos").middle
'Carlos'
>>> given_last.parse("de la Cruz Juan Carlos").middle
'Juan'

Where the vocabulary answers first

In two places the vocabulary layer answers before name_order is consulted at all:

  • A middle word that is also a shipped particle is claimed by the vocabulary. That is why the Vietnamese example above is not the more obvious "Nguyen Van Minh": Van is the Dutch particle van, so that name reads family Nguyen with Van Minh given under both family-first orders, and the choice between them makes no difference.

  • A never-given particle opening the name overrides the declared order outright: where a particle that can never be a given name stands alone as the opening piece, the whole name is the surname, in every name_order. "de Mesnil" is family de Mesnil under both family-first orders exactly as it is by default, not family de with Mesnil given — a word that can never be a given name leaves the order nothing to decide. Only the never-given set does this: "van Gogh" reads family van, given Gogh under a family-first order, because van can be a given name and so leaves a real question to answer.

Words that are also ordinary names covers dropping a word from a vocabulary, or moving one between those two sets.

East Asian defaults, and turning them off

Two defaults key on the script a name is written in rather than on anything you set: a name written wholly in Han or Hangul — or in kanji and kana, unless it is katakana alone — is assigned family-first (script_orders), and an unspaced hangul name is split into surname and given name against the shipped Korean census list (segment_scripts). East Asian names explains the naming conventions both rest on; this section is how to switch them off. Two further behaviors are not policy fields at all, and are covered last.

Switching off order and splitting

The two defaults switch off separately:

>>> parse("김민준").family                    # both defaults on
'김'
>>> positional = Parser(policy=Policy(script_orders=()))
>>> positional.parse("김민준").family         # still split
'민준'
>>> unsplit = Parser(policy=Policy(segment_scripts=frozenset()))
>>> unsplit.parse("김민준").family            # one token, not split
'김민준'

The two switches interact, and clearing only script_orders produces a third behavior rather than the old one: the split still runs, so 김민준 still becomes two tokens, and the positional default then assigns them given-first — the surname lands in given. To restore nameparser 2.0’s reading exactly, clear both fields.

Note

Every field here is annotated with its canonical storage type rather than with everything the constructor accepts — the same as capitalization_exceptions, and for the same reason: the annotation is what you get back when you READ the attribute, which is the commoner operation.

The constructor is deliberately wider. It takes any mapping for script_orders, any iterable of Script for segment_scripts, and plain strings wherever a Role is wanted (Role is a StrEnum precisely so that works). A dataclass cannot express those two types separately, so the examples in this guide use the spellings that check clean under mypy — () and frozenset(...) rather than {} and a bare set literal. The wider spellings parse identically; they just need a # type: ignore[arg-type] if you run a type checker.

Teaching the splitter a surname

To teach the splitter a surname it doesn’t ship with, add it to the surnames vocabulary like any other word:

>>> lex = Lexicon.default().add(surnames={"김민"})
>>> Parser(lexicon=lex).parse("김민준").family
'김민'

Chinese surnames are deliberately absent from that default set, because splitting Han text requires knowing Chinese from Japanese; Locale packs covers the opt-in zh pack that supplies them.

Japanese names

The Japanese behaviors ride these same two fields, so they need no switches of their own. script_orders=() clears the kana-licensed entry along with the Han and Hangul ones, so a name in kanji and kana reads given-first again:

>>> parse("山田 エミ").family
'山田'
>>> positional.parse("山田 エミ").given
'山田'

segment_scripts=frozenset() deactivates every script at once, which also stops a parser consulting whatever segmenter it was given. The segmenter has an off-switch of its own as well: Parser(segmenter=None), which is the default. See Segmenters for what a segmenter is expected to do with text it does not handle.

Behaviors no policy field controls

Two behaviors apply however both fields are set:

  • The katakana middle dot ・ separates tokens the way a space does, decided in tokenization.

  • A glued honorific, a listed CJK honorific at the end of a name token, is split off it. The honorific vocabulary carries its own license rather than borrowing a script’s: every entry is a word that can never end a name, so there is no per-script question for segment_scripts to answer.

>>> both_off = Parser(
...     policy=Policy(script_orders=(), segment_scripts=frozenset()))
>>> both_off.parse("マイケル・ジャクソン").family
'ジャクソン'
>>> both_off.parse("田中さん").suffix
'さん'

The honorific vocabulary is the peel’s off-switch. Removing an entry from honorific_tails leaves that honorific glued, while the spaced form still reads it as a suffix:

>>> no_san = Parser(
...     lexicon=Lexicon.default().remove(honorific_tails={"さん"}))
>>> no_san.parse("田中さん").family
'田中さん'
>>> no_san.parse("田中 さん").suffix
'さん'

honorific_tails marks a subset of suffix_words, so dropping a word from it alone orphans nothing and needs no matching suffix_words edit. Emptying it is also how to opt out of what the peel costs a non-ASCII parse: an empty honorific_tails stops the peel at its first check, whereas segment_scripts never gated it and so cannot turn it off.

Nicknames, maiden names, and brackets

Maiden markers

A maiden name is usually announced by a marker word, and then it needs no brackets and nothing configured:

>>> parse("Jane Smith née Jones").maiden
'Jones'

Nicknames and maiden names in the usage guide covers how the marked forms read. The marker words themselves are the maiden_markers vocabulary, a Lexicon field, shipped in several languages (see nameparser.config.maiden_markers). Add your own like any other word; a marker you add works bare and in brackets alike:

>>> parse("Jane Smith formerly Jones").maiden       # not a marker
''
>>> marked = Parser(
...     lexicon=Lexicon.default().add(maiden_markers={"formerly"}))
>>> marked.parse("Jane Smith formerly Jones").maiden
'Jones'
>>> marked.parse("Jane Smith (formerly Jones)").maiden
'Jones'

A multi-word entry such as the shipped z domu is matched as a phrase, not word by word; see Vocabulary: Lexicon above.

How a bracketed clause is read

A delimiter pair carries no meaning of its own, so what a clause enclosed in one reads as is settled in steps, and the pair is asked last:

  1. Suffix-shaped content is taken first. The brackets are dropped and what was inside parses as if it had been written bare, which is not the same as the clause becoming the suffix.

  2. A clause that announces itself, opening with a recognized maiden marker and carrying a word after it, is a maiden name inside any configured pair, with nothing configured (since 2.2). The marker is dropped from the value.

  3. Everything else is decided by the pair: content in a nickname_delimiters pair is a nickname, and content in a maiden_delimiters pair is a maiden name.

>>> name_jr = parse("Jane Smith (née Jr.)")         # step 1
>>> name_jr.family, name_jr.suffix
('née', 'Jr.')
>>> parse("Jane Smith (née Jones)").maiden          # step 2
'Jones'
>>> parse('Jane Smith "née Jones"').maiden          # step 2, any pair
'Jones'
>>> parse("Jane (Jones) Smith").nickname            # step 3
'Jones'

Routing a pair to maiden names

Step 3 is what maiden_delimiters is for. It reaches two kinds of clause, not one: content with no marker word in it, such as a birth surname written bare in parentheses, and a lone marker word. Listing a pair in maiden_delimiters drops it from the effective nickname_delimiters set automatically, and the one-liner is the whole recipe:

>>> policy = Policy(maiden_delimiters=frozenset({("(", ")")}))
>>> Parser(policy=policy).parse("Jane (Jones) Smith").maiden
'Jones'

A marker is dropped from the value only where it stands as its own word with a name word after it. A lone marker is kept as the value — and so is a multi-word one filling the clause, such as z domu — and so is a marker written against the name it marks, which is one token with it, so 旧姓 stays in the value too:

>>> parse("Jane Smith (Nee)").nickname
'Nee'
>>> Parser(policy=policy).parse("Jane Smith (Nee)").maiden
'Nee'
>>> Parser(policy=policy).parse("Jane Smith (z domu)").maiden
'z domu'
>>> cjk_parens = frozenset({("(", ")")})       # full-width
>>> cjk = Parser(policy=Policy(maiden_delimiters=cjk_parens))
>>> cjk.parse("山田花子(旧姓佐藤)").maiden
'旧姓佐藤'

Adding a delimiter pair

To add a delimiter pair rather than reroute one, build on the exported default — assigning a bare set replaces the built-in pairs instead of extending them, the same trap as capitalization_exceptions:

>>> from nameparser import DEFAULT_NICKNAME_DELIMITERS
>>> parse("Benjamin {Ben} Franklin").middle        # not a pair by default
'{Ben}'
>>> policy = Policy(
...     nickname_delimiters=DEFAULT_NICKNAME_DELIMITERS | {("{", "}")})
>>> Parser(policy=policy).parse("Benjamin {Ben} Franklin").nickname
'Ben'

Suffixes not separated by commas

extra_suffix_delimiters handles sources that separate post-nominals with something other than a comma. Undeclared, the separator is a word. Since 2.4 such a name still parses when a credential opens the part after the comma, because that makes the whole part the suffix, but the separator stays in the suffix as written. Declared, it splits the suffix into groups the way a comma does:

>>> name = parse("Jane Smith, RN - CRNA")
>>> name.given, name.family, name.suffix
('Jane', 'Smith', 'RN - CRNA')
>>> policy = Policy(extra_suffix_delimiters=frozenset({" - "}))
>>> name = Parser(policy=policy).parse("Jane Smith, RN - CRNA")
>>> name.given, name.family, name.suffix
('Jane', 'Smith', 'RN, CRNA')

Through 2.3 the undeclared reading was given RN, family Jane Smith and suffix CRNA, so the delimiter was the only way to get the name right.

Credentials the vocabulary doesn’t list

No suffix list holds every post-nominal, so two Policy fields, both new in 2.4, read an unlisted word as a credential from how it is written: in periods (X.Y.Z.), or in capitals (XYZ). The writing cannot settle it alone, because a surname can be written either way, so both fields also decide by position: the word has to stand behind a name that can spare it.

Both fields share one exception. Two letters alone right after a comma are how a person’s initials are written, and the words before the comma may be one surname of two words, so "García Márquez, G.J." and "García Márquez, MJ" keep given G.J. and MJ. The dotted spelling reports the bare reading as a fork; the capitals do not.

The exception gives way to evidence that the letters are a credential: an unambiguous post-nominal in front of them that is not also a title, or another unlisted dotted or all-caps word beside them that its own field reads as a credential. A title in front of them says the opposite:

>>> parse("John Smith, PhD X.Y.").suffix
'PhD X.Y.'
>>> parse("John Smith, PhD MJ").suffix
'PhD MJ'
>>> parse("García Márquez, MJ XYZ").suffix
'MJ XYZ'
>>> cred = parse("García Márquez, Ms G.J.")
>>> cred.title, cred.given
('Ms', 'G.J.')

Dotted acronyms

unlisted_dotted_suffixes is on by default. It reads a token of two or more period-separated chunks as a credential at the end of a name, right after a comma behind two or more name words, at the end of the given part after a family comma, and at the end of a maiden marker’s clause. Case is irrelevant; the periods are the signal. With nothing to spare in front of it, the word stays a name. At the end of a name, given part or clause, either reading reports the fork as a suffix-or-name ambiguity, so a record that reads wrong can still be found; right after a comma behind a full name, the credential reading is taken silently:

>>> cred = parse("John Smith X.Y.Z.")
>>> cred.family, cred.suffix
('Smith', 'X.Y.Z.')
>>> parse("Jack X.Y.Z.").family
'X.Y.Z.'
>>> cred = parse("Doe, John X.Y.Z.")
>>> cred.given, cred.suffix
('John', 'X.Y.Z.')
>>> parse("Doe, X.Y.Z.").given
'X.Y.Z.'
>>> cred = parse("Jane Doe nee Smith X.Y.Z.")
>>> cred.maiden, cred.suffix
('Smith', 'X.Y.Z.')
>>> cred = parse("John Smith, X.Y.Z.")
>>> cred.suffix, cred.ambiguities
('X.Y.Z.', ())

Four things are not this shape. A token the vocabulary already knows (M.A., Ph.D.) is read by the vocabulary. A single trailing period is how any word is abbreviated, so "John Smith Xyz." keeps family Xyz.. Every chunk must be alphabetic, so a digit anywhere refuses it ("John Smith 1.4" keeps family 1.4). And a script with no period abbreviations of its own refuses it too ("John Smith 田.中." keeps family 田.中.).

Setting the field to False reads such a token as name material everywhere, and still reports the fork. It does not bring back the pre-2.4 reading of a chunk that is one ASCII character, a roman numeral or the digit 2, as a credential. That reading was retired outright, not put behind this switch, so "Jack X.Y.I." and the version string "John Smith 1.4.2" keep their last word as the family name either way:

>>> dotted_off = Parser(policy=Policy(unlisted_dotted_suffixes=False))
>>> dotted_off.parse("John Smith X.Y.Z.").family
'X.Y.Z.'
>>> dotted_off.parse("Jane Doe nee Smith X.Y.Z.").maiden
'Smith X.Y.Z.'

All-caps words

unlisted_caps_suffixes reads an unlisted word of two or more capital letters, with no period in it. Its value is a CapsSuffixes, because where capitals mean a credential depends on a second convention: many records write the SURNAME in capitals ("Jean DUPONT", "DUPONT, Jean").

The capitals only mean something against a name that does not use them. The name must hold a word of its own with a capital and a lowercase last letter (Smith, DiCaprio) that the vocabulary does not claim as a title, particle or credential. A record written wholly in capitals, or wholly in lowercase, keeps every word a name word.

AFTER_COMMA (the default)

CapsSuffixes.AFTER_COMMA reads the word only in the part right after a comma with two or more name words before it. The all-caps surname convention never writes capitals there:

>>> from nameparser import CapsSuffixes
>>> cred = parse("John Smith, XYZ")
>>> cred.family, cred.suffix
('Smith', 'XYZ')
>>> parse("Smith, XYZ").given
'XYZ'
>>> parse("JOHN SMITH, XYZ").given
'XYZ'
EVERYWHERE

CapsSuffixes.EVERYWHERE also reads the end of a name, the given part’s last word after a family comma, and the word ending a maiden marker’s clause. Those are the places an all-caps surname IS written, which is why it is not the default:

>>> caps_everywhere = Parser(
...     policy=Policy(unlisted_caps_suffixes=CapsSuffixes.EVERYWHERE))
>>> caps_everywhere.parse("John Smith XYZ").suffix
'XYZ'
>>> caps_everywhere.parse("Doe, John XYZ").suffix
'XYZ'
>>> cred = caps_everywhere.parse("Jane Doe nee Smith XYZ")
>>> cred.maiden, cred.suffix
('Smith', 'XYZ')
>>> cred = caps_everywhere.parse("Jean Pierre DUPONT")
>>> cred.family, cred.suffix
('Pierre', 'DUPONT')
OFF

CapsSuffixes.OFF reads none of them and reports nothing, as 2.3 did. It is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential:

>>> parse("García Márquez, GABRIEL").suffix
'GABRIEL'
>>> caps_off = Parser(
...     policy=Policy(unlisted_caps_suffixes=CapsSuffixes.OFF))
>>> caps_off.parse("García Márquez, GABRIEL").given
'GABRIEL'

Keeping emoji and control characters

The strip flags keep characters the parser removes by default. Note what happens to an emoji you keep — it becomes a token like any other, and lands in the middle name:

>>> str(parse("Sam 😊 Smith"))                      # stripped by default
'Sam Smith'
>>> kept = Parser(policy=Policy(strip_emoji=False)).parse("Sam 😊 Smith")
>>> str(kept), kept.middle
('Sam 😊 Smith', '😊')

strip_bidi=False does the same for invisible bidirectional control characters, which is occasionally what you want when round-tripping right-to-left text verbatim.

Presentation: rendering arguments

Once a name is parsed, how it’s displayed is a separate decision made at the point of output, not baked into the parse. Three methods on ParsedName cover it — see API reference for full signatures:

  • render() fills a format spec from the seven role fields.

  • initials() is the same idea narrowed to first letters, with its own delimiter/separator arguments.

  • capitalized() returns a new, case-fixed ParsedName instead of a string. It only touches a name whose words are written in a single case (all lower, all upper) unless you pass force=True — mixed case is left alone by default on the assumption that someone already capitalized it on purpose. The suffixes don’t count toward that: III or PhD written the usual way says nothing about how the name was cased. A suffix written in more than one case is the writer’s spelling and is kept as written (john smith EdD gives John Smith EdD) unless you pass force=True.

>>> from nameparser import parse
>>> name = parse("Dr. Juan Q. Xavier de la Vega III")
>>> name.render("{family}, {given} {middle}")
'de la Vega, Juan Q. Xavier'
>>> name.initials(spec="{given}{middle}{family}", delimiter="", separator="")
'JQXV'
>>> str(parse("DR. JUAN DE LA VEGA").capitalized())
'Dr. Juan de la Vega'
>>> str(parse("JuAn DE LA vEGA").capitalized())
'JuAn DE LA vEGA'
>>> str(parse("JuAn DE LA vEGA").capitalized(force=True))
'Juan de la Vega'
>>> str(parse("juan garcia III").capitalized())
'Juan Garcia III'

Looking for v1’s string_format? It’s the render(spec) argument now — pass your own format string per call instead of setting it once on a shared config object.

A spec chooses what the output is for. The default is written for display and does not survive a reparse — it parenthesizes the maiden name, which reads back as a nickname. When the rendered string will be parsed again, spell the marker out (née {maiden}) so the field round-trips; see the round-trip note in the tour.

Sharing a configured parser

A Parser is a frozen value, so the way to share one configuration across a codebase is the same way you’d share any other constant: build it once at module level and import it wherever you parse.

# myapp/names.py
from nameparser import Lexicon, Parser, Policy

lex = Lexicon.default().add(titles={"dean"})
policy = Policy(strip_emoji=False)
parser = Parser(lexicon=lex, policy=policy)

# elsewhere
from myapp.names import parser
name = parser.parse(raw_name)

Because Parser and its lexicon/policy are immutable and hashable, parser is safe to import and call from multiple threads with no locking — there is no shared mutable state to protect, unlike v1’s module-level CONSTANTS.