Release Log

  • 2.4.0 - Unreleased

    nameparser 2.4 is under development.

    Behavior Changes

    • Fix a particle surname before a comma being split when a credential follows it. HumanName("van der Berg, PhD") gives last van der Berg, suffix PhD, where 1.4.0 through 2.3.0 gave first van, last der Berg. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads as Berg, PhD does, and Abu Bakar, PhD gives last Abu Bakar where 1.4.0 through 2.3.0 gave first Abu. The same count keeps a particle surname whole in front of a credential that is also a name: De La Cruz, Ed gives first Ed, last De La Cruz, as 2.3.0 read it, and van der Berg, MA gives last van der Berg, suffix MA, the reading Smith, MA gets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name: Van Johnson, Dr. gives last Van Johnson, where 2.2 and 2.3 gave first Van. A given name in front still makes two words, so John van Buren, Ed reads suffix Ed as John Smith, Ed does. A connective surname is not counted as one: Ortega y Gasset, PhD still gives first Ortega, middle y, as Ortega y Gasset reads on its own. See the C1 entry of docs/design/decisions.md (closes #575)

    • Fix a surname particle that is also a credential (vd, mc) being split away from the surname or the credentials around it. HumanName("SMITH VD MA, JOHN") gives last SMITH VD MA, where 2.2 and 2.3 gave last SMITH MA, suffix VD, and HumanName("Doe, Jane PhD vd MA") gives suffix PhD vd MA, where 2.3 gave middle MA, suffix PhD vd. A vd with nothing behind it still joins the last name: HumanName("Doe, Jane PhD vd") gives last vd Doe. (closes #573)

    • Fix a one-letter connective joining a name that gives no sign it is a connective. HumanName("jose e maria santos") gives first jose, middle e maria, last santos, where 1.4.0 through 2.3.0 gave first jose e maria; and JUAN GARCIA Y LOPEZ gives last GARCIA Y LOPEZ, where every release since 1.4.0 read the bare capital as an initial and gave middle GARCIA Y. A single letter is an initial where the writing says so – a bare Latin capital in a name that is not written wholly in one case – and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there: e reads as an initial and y joins. Mixed-case input is untouched in both directions: Jose e Maria Santos still gives first Jose e Maria and Jose E Maria Santos still gives middle E Maria. Short names move in the derived views rather than the fields, P3’s three-word carve-out being unchanged: parse("john e smith").initials() is j. e. s. where 2.3.0 gave j. s., and HumanName("john e smith").capitalize() gives John E Smith where 2.3.0 gave John e Smith; JUAN Y GARCIA moves its capitalize() the same way in reverse, giving Juan y Garcia, while its initials do not move at all: parse(...).initials() is J. Y. G., what 2.3.0 gave and what 1.4.0’s own view gave, the Y holding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461). HumanName.initials() agrees with the core on both – see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (Хосе И Мария Сантос still gives first Хосе И Мария), and Arabic و never enters the rule, having no case to be written against. A Lexicon knob decides which letters are marked, so the reading is configurable rather than fixed. See the P3 entry of docs/design/decisions.md (closes #383, closes #479)

    • Fix HumanName.initials() reading a one-letter connective by vocabulary and written shape instead of by the parse. HumanName("john e smith").initials() gives j. e. s., where every release from 1.4.0 through 2.3.0 gave j. s.; JUAN GARCIA Y LOPEZ gives J. G. L. where 2.3.0 gave J. G. Y. L.; JUAN Y GARCIA does not move at all this cycle, giving J. Y. G. on both surfaces as 2.3.0 and 1.4.0 did, since the connective-initials fix further down this list (closes #461) gives its Y an initial again. That first name read 1.4.0’s way at 2.3.0 and only there: 2.0.0 through 2.2.0 already gave today’s answer, by the unrelated bug the 2.3.0 note below records as fixed (the facade dropping a bare capital that is also a one-letter conjunction, #462), so against those three releases it does not move at all. The v1 facade decided whether a word was the connective by looking the word up and checking its shape, while parse(...).initials() read the tag the parse recorded – so the change above, which reads a single letter in a one-case name from the vocabulary rather than from its case, moved one view and not the other. Both views of a parse now give the same answer. Mixed-case names are untouched on both, the writing having decided the letter: John E Smith is still J. E. S. and Scott E. Werner still S. E. W.. So is a one-case name whose letter is outside the marked set – maria y lopez is still m. l., y having joined before this release and after it. Two costs, and both match what capitalize() has always done: editing C.conjunctions after a name is parsed no longer changes its initials until full_name is assigned again, and a name restored from a pickle, copied with copy.copy/copy.deepcopy (the same state hooks), or built from keyword fields (HumanName(first=..., middle=..., last=...)) carries no tags, so its initials come from the vocabulary and can differ from a fresh parse of the same string. One private break, stated because a v1 subclass can hit it: an override of _process_initial written to v1’s (name_part, firstname=False) signature now raises TypeError the first time initials() runs, since initials() passes the part’s tokens. Such an override has to accept a tokens keyword and pass it on – return super()._process_initial(name_part, firstname, tokens=tokens) – to receive this fix. Widening the signature without forwarding still works, but on the pre-#528 STRING path: the token call hands the override the group’s own text as name_part rather than an empty placeholder, so john e smith initials j. s. under such an override, not the j. e. s. above. A subclass overriding one of the public first_list, middle_list or last_list properties keeps working too: that member takes the pre-2.4 vocabulary reading instead of the change above, while an un-overridden member still moves. See the R3 entry of docs/design/decisions.md (closes #528)

    • Fix a credential acronym that is also a surname being read by position alone. HumanName("Jack MA") gives suffix MA where 2.0 through 2.3 gave last MA, and John Smith Ma gives last Ma where they gave suffix Ma. In a name written in more than one case, an ambiguous acronym written in capitals is written the way a credential is written and is read as one even where removing it leaves no surname; one written in any other cased form that is not wholly lower is written the way a surname is written and stays one even where there are words to spare (John Smith ma and John Smith ed – all lower, no contrast – give suffix ma/ed instead). A name written wholly in one case says nothing either way and keeps the reading it had: JOHN SMITH MA is still a credential, ANH DO still a surname, jack ma still a surname. The same reading reaches the comma forms, where the words-to-spare count is now a count of NAME words: Smith, MA gives last Smith, suffix MA; Smith Jr., MA keeps last Smith; and John Smith, MA, John Smith, Ed, john smith, ma and JOHN SMITH, MA all give a suffix again, which is what 1.4.0 read and 2.0 through 2.3 did not. Jack Ma and Anh Do are unchanged. The LEAN is inert on a caseless script, but the comma count above is not – it asks name-word count, not case – so 마틴 킹, MA and 田中 太郎, MA also give a suffix again (1.4.0 parity on the suffix, two pre-comma name words each) while the single-token 毛泽东, MA does not move, having no case to write a contrast in either way. See the S2 entry of docs/design/decisions.md (closes #289)

    • Fix a bare trailing Meng or Lac being read as a credential and losing the family name: meng and lac are now acronyms that are also ordinary names. HumanName("wang meng") gives first wang, last meng, and parse() reports a suffix-or-name ambiguity, where every release from 2.0.0 through 2.3.0 gave suffix meng and no last name; 1.4.0 read last meng, so this is 1.4.0’s answer plus the flag. li meng and tran lac move the same way, Wang, Meng gives first Meng, last Wang again, and Parser(policy=Policy(name_order=FAMILY_FIRST)).parse("Wang Meng") gives given Meng where 2.0.0 through 2.3.0 gave family Wang, suffix Meng and no given name. With a full name in front the credential reading stays: john smith meng and nguyen van lac keep suffix meng and lac, now flagged. But a Title-case Nguyen Van Lac gives last Van Lac where every release gave last Van, suffix Lac. The cost is the marking’s own, and it falls on the conventional spellings: a MEng or LAc written that way, in a name written in more than one case, is read the way John Smith Ma is (above), so John Smith MEng gives middle Smith, last MEng, where every release gave suffix MEng. Where one of them LEADS a credential run nothing speaks for it: John Smith MEng PhD gives middle Smith, last MEng, suffix PhD, where every release read suffix MEng PhD (MEng, PhD through 2.2); behind a credential, or written with its periods, it is read with the run (the next entry). After a comma Smith, MEng and Smith, meng give first MEng and meng (1.4.0’s reading, not the suffix 2.0 through 2.3 gave) and Smith, John MEng gives middle MEng (every release gave suffix MEng), and a bracketed John Smith (MEng) falls through to nickname, where every release gave suffix MEng. A lone credential after a comma behind a full name (John Smith, MEng) keeps the credential reading, and so does a run of them (the next entry). Meng is a common Chinese surname and given name, Lac a Vietnamese given name (Nguyen Van Lac) and a French surname; see the suffix-acronym-collisions entry of docs/design/decisions.md (closes #540)

    • Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole. HumanName("John Smith, Ed Ma") gives first John, last Smith, suffix Ed Ma, where 2.0 through 2.3 gave first Ed, middle Ma, last John Smith – 1.4.0’s reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (John Smith, MA), and parse() reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report when another credential opens it, the capitals having decided it (John Smith, PhD MA, John Smith, MS MA), while one opened by such an acronym reports on that first word, as John Smith, MA alone does (John Smith, MA PhD). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so john smith, v phd ma gives first john, last smith, suffix v phd ma, where 2.3 gave first v, middle ma, last john smith. john smith, md ma and John Smith, Ms Ma move the same way, where 2.0 through 2.3 gave title md/Ms – the second is the accepted cost, Ms read as the suffix word it also is, as John Smith, Ms alone already reads it – while one name word before the comma keeps the listing form (Smith, Ms Ma gives title Ms, first Ma). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential’s company whatever its case: John Smith PhD MEng and Doe, Jane PhD MEng give suffix PhD MEng, the fields every release gave (PhD, MEng through 2.2), now reported, and doe, jane v phd do gives suffix v phd do where 2.3.0 gave last do doe – a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: Wang Ma PhD keeps last Ma. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: Smith, PhD Ma gives last Smith, suffix PhD Ma, where 2.3 gave first PhD, middle Ma. A title that is also a credential (MD, Ms) opening that part stays a title and nothing in the part speaks, so Smith, MD PhD Ma keeps title MD, first PhD, middle Ma and Smith, Ms MD Ma title Ms MD, first Ma, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so Wang M.Eng. gives suffix M.Eng., as Wang M.A. does and as 2.0 through 2.3 did. See the #544 entry under S2 in docs/design/decisions.md (closes #544)

    • Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run. HumanName("John Smith, PhD DO DO") gives first John, last Smith, suffix PhD DO DO, where 2.2 and 2.3 gave first PhD, last DO DO John Smith and 2.0 and 2.1 gave title PhD, first DO DO – 1.4.0’s reading, restored. DO, MC and VD are credentials and surname particles at once, and two particles next to each other had joined into one particle run that took the credentials apart; the part after the comma is now read as credentials before any particle can join it. John Smith, PhD vd DO gives suffix PhD vd DO the same way, and John Smith, MD DO DO gives suffix MD DO DO where 2.3 gave title MD, first DO, middle DO. A single DO is still left to its capitals (John Smith, PhD DO gives suffix PhD DO), and with one name word before the comma a part of credentials is all suffixes (Smith, PhD DO DO gives last Smith, suffix PhD DO DO, and Smith, MA DO DO last Smith, suffix MA DO DO, where 2.3 gave first PhD and first MA, last DO DO Smith). See the #562 entry under C1 in docs/design/decisions.md (closes #562)

    • New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position. HumanName("John Smith X.Y.Z.") gives suffix X.Y.Z. where every release gave last X.Y.Z., while Jack X.Y.Z. keeps its surname, the same words-to-spare rule a listed acronym takes – and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person’s initials are written, and two words before a comma may be one surname, so García Márquez, G.J. keeps first G.J. and last García Márquez and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (John Smith, PhD X.Y. gives suffix PhD X.Y., while García Márquez, Ms G.J. keeps title Ms, first G.J.). Three letters or more read by the count, so John Smith, X.Y.Z. gives suffix X.Y.Z. – and so does García Márquez, G.J.R., the accepted cost of the line, since initials are conventionally written apart (García Márquez, G. J. R.), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so john smith x.y.z. reads the same way. Words the vocabulary does know are untouched (M.A., Ph.D., A.B.C.), a single trailing period is still not this shape (John Smith Xyz. keeps last Xyz.), and a dotted run at the FRONT of a name is untouched (J.R.R. Tolkien). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS – the roman numerals the suffix list holds, and the lone digit 2 – was reading as a generational suffix, so Jack X.Y.I. gives last X.Y.I. again, as 1.4.0 read it, while Msc.Ed., JD.CPA and Lt.Gov. are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: John Smith 1.4.2 gives last 1.4.2 where 2.3 gave suffix 1.4.2, and John Smith, 1.4.2 gives first 1.4.2, last John Smith. Such a token reports nothing at any policy – it is no acronym either, the shape reading wanting every chunk alphabetic – and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way – setting it to False reads an unlisted dotted word as name material by position instead (John Smith X.Y.Z. keeps last X.Y.Z.), the pre-2.4 reading for THAT half alone. See the S2 and suffix-acronym-collisions entries of docs/design/decisions.md (closes #516)

    • New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request. Its value is a CapsSuffixes. The default, CapsSuffixes.AFTER_COMMA, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials: HumanName("John Smith, XYZ") gives first John, last Smith, suffix XYZ, where 1.4.0 through 2.3.0 gave first XYZ, last John Smith; John Smith, LEED AP and John Smith, PhD XYZ give suffix LEED AP and PhD XYZ the same way, and John Smith, RAI gives suffix RAI again, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (Jean DUPONT, DUPONT, Jean) and never there. A word after a one-word surname stays the given name (Smith, XYZ), a two-letter word reads exactly as dotted initials do (García Márquez, MJ and García Márquez, MJ PhD keep first MJ), and the name has to contrast the capitals with a word of its own holding a capital whose last letter is lowercase (Smith, DiCaprio); a name typed with decomposed accents reads as its composed spelling. A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (GISCARD d'ESTAING, VALÉRY, LLOYD FitzGERALD, RONALD, LLOYD WEBBER, ANDREW PhD), as does a name written wholly in lowercase. CapsSuffixes.EVERYWHERE also reads the end of a name, the given part’s last word after a family comma and the word ending a maiden marker’s clause: .parse("John Smith XYZ") gives suffix XYZ, and Jean Pierre DUPONT gives last Pierre, suffix DUPONT – why it is not the default. CapsSuffixes.OFF reads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (García Márquez, GABRIEL gives suffix GABRIEL). The field reaches the core parser only, through Parser(policy=Policy(unlisted_caps_suffixes=...)); a HumanName tracks the parser’s defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field nor unlisted_dotted_suffixes has a v1 Constants manager. See the S2 and C1 entries of docs/design/decisions.md (closes #516, closes #564)

    • The comma’s own decision about an ambiguous credential is now reported. parse("Smith, MA").ambiguities names suffix-or-name, and so does every other decision at the ambiguous credential class – before or after a comma, in either direction, with no new AmbiguityKind (the family-comma attachment fork already reported this way, e.g. parse("Berg, Jan vd")). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: John Smith, X.Y.Z. and John Smith, PhD X.Y. report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: John Smith, X.Y. P.Q. gives last Smith, suffix X.Y. P.Q. (#563). One report per decision: Smith, Ma reports that the word was kept as the given name just as Smith, MA reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: John van der Berg Ma gives last van der Berg Ma and names suffix-or-name, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: Steven Hardman, MD, DO, DDS no longer reports comma-structure, on its written case. That is the whole of the losses over the differential corpora – John Smith, MD, R.A.I. is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release’s own development. The other movement an upgrader sees is a SWAP rather than a loss: Jack X.Y.I. reported given-or-family at 2.3.0 and reports suffix-or-name here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker’s clause. See the S2 and C1 entries of docs/design/decisions.md

    • Fix a credential ending the given part of a family-comma listing being read as a middle name in silence. HumanName("Doe, John MA") gives first John, last Doe, suffix MA, where 2.0 through 2.3 gave middle MA – and 1.4.0 gave the suffix, so this restores v1’s reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: Doe, John Ma keeps middle Ma, written the way a name is written, and Doe, John Ed keeps middle Ed. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields – a declined name re-rendered without its comma, John Ma Doe, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent – Doe, John MA Smith gives middle MA Smith and reports nothing – while a credential run or a trailing title is transparent to it: Doe, John MA PhD gives suffix MA PhD and Doe, John MA Prof. gives title Prof. with suffix MA. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: Doe, John Prof. MA now gives title Prof. where it gave middle Prof. MA, the trailing-title chain reaching a word the credential used to hide; and Doe, John van MA gives last van Doe with suffix MA where it gave middle van MA, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: DOE, MARY JO MA, doe, john ma, 田中, 太郎 MA and 김, 민준 MA all give a suffix. The unlisted dotted spelling moves with them without the parity claim – Doe, John X.Y.Z. gives suffix X.Y.Z. where 1.4.0 and 2.3.0 both gave a middle name – to match the comma-less John Doe X.Y.Z.. One word is carved out: do is the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry). Doe, John DO gives suffix DO, while Doe, John do, Doe, John Do, DOE, JOHN DO and doe, john do are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so SMITH, JOHN DO keeps last DO SMITH as NASCIMENTO, EDSON ARANTES DO does – right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the S2 and P6 entries of docs/design/decisions.md (closes #531)

    • Change where a maiden marker takes a maiden name: only behind a surname, and only up to the credentials and titles the name ends with. HumanName("Jane Doe nee Smith"), Mai Le née Nguyen and Doe nee Smith, Jane give maiden Smith and Nguyen as before, but a marker counts only in the name before any comma or in the surname part before a family comma, and only behind a name word. Anywhere else it is an ordinary word, as in 1.4.0: Doe, Jane nee Smith gives middle nee Smith, Smith, John, PhD née Jones gives suffix PhD née Jones, and Dr. nee Smith gives first nee, last Smith, where 2.0 through 2.3 gave maiden Smith, Jones and Smith – the last with no name at all. A credential in front of the marker ends the name (see the credential-run entry below), so Jane Doe PhD nee Smith gives suffix PhD nee Smith where 2.0 through 2.3 gave suffix PhD, maiden Smith. A particle or a lone word is a surname here: Jane van nee Smith gives last van, maiden Smith, and Jane Smith née V gives last Smith, maiden V, as 2.0 and 2.1 read it, where 2.2 and 2.3 gave last née, suffix V. The words the marker takes now end where the run of post-nominals and titles the name would end with if the clause were not written begins, and the words it gives up read as they would there: Jane Doe nee Smith MA gives maiden Smith, suffix MA, where 2.0 through 2.3 gave maiden Smith MA and said nothing, as do the one-case JANE DOE NEE SMITH MA and jane doe nee smith ma; Jane Doe nee Smith MA PhD gives suffix MA PhD where 2.3.0 gave maiden Smith MA, suffix PhD, so the two orders of the credentials now agree; Jane Doe nee Smith DO DO gives suffix DO DO; Jane Doe nee Smith Prof., Mary Smith née Jones Prof. and Jane van der Berg nee Smith Prof. give title Prof.; Jane Doe nee Smith King. gives title King.; Jane Doe nee Smith MA Prof. and Jane Doe nee Smith Prof. MA both give title Prof., suffix MA; and Jane Doe nee Smith V Prof. gives suffix V – each where 2.3.0 kept every word in the maiden name. A credential the clause gives up or keeps at that boundary is reported as suffix-or-name. The writing still decides, as it does at the end of a name with no clause: Jane Doe nee Smith Ma keeps maiden Smith Ma, and Jane Doe nee Yo-Yo Ma keeps a two-word birth surname whole. The first word after the marker is always taken, the marker having announced a name: Jane Doe nee MA and Jane Doe nee King. keep it, and Jane Doe nee Prof. Dr. gives maiden Prof., title Dr., where 2.3.0 gave maiden Prof. Dr.. Before a family comma the clause keeps every word, as before: Doe nee Smith Prof., Jane keeps maiden Smith Prof.. Delimiters settle the question outright: Jane Doe (nee Smith MA) keeps the whole span, and brackets around a clause ending in a period are dropped as in 2.3.0, so Jane Doe (nee Smith Prof.) gives title Prof.. The name #548 reported, Dr. nee Smith PhD Prof., gives title Dr. Prof., first nee, last Smith, suffix PhD, where 2.3.0 gave last PhD, maiden Smith. John Smith nee Jones R.A.I. gives suffix R.A.I., as 2.3.0 read it. The rule replaces the clause rules this cycle’s #533 and #535 had added, and the readings those changes gave here are superseded. See the M2 entry of docs/design/decisions.md (closes #601)

    • Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name. parse("John Smith, Jr., Freiherr von Richthofen").ambiguities names comma-structure alone, where 2.0 through 2.3 also named particle-or-given for von – a word the same parse had put in the suffix, so the report described a reading it never made. This fix moves no field; the title change below (#603) moves Freiherr to the title. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain’s credential-acronym report this release adds (above) keeps the same bound: John Smith, PhD Do Ma reports the comma’s decision once and not again for Ma. See the 2026-09-28 bullet of the C1 entry in docs/design/decisions.md

    • A credential after the name core starts a suffix run to the end of its part. HumanName("John Smith PhD Jones") gives suffix PhD Jones and reports suffix-or-name for the name word it took, where 2.3.0 gave middle Smith PhD, last Jones; Smith, John PhD Jones gives suffix PhD Jones where 2.3.0 gave middle Jones, suffix PhD. A title in the run is a title, so Eric H. Holder Jr. Attorney General gives title Attorney General, suffix Jr., as the comma spelling Eric H. Holder, Jr. Attorney General already read, where 2.3.0 gave middle H. Holder Jr. Attorney, last General. Only an unambiguous credential or generational word starts the run, and only behind the name core – two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, so John Smith MA Jones and Mohamed Ali Abd Allah read as before, and a bare title word starts nothing: Mary Jane King Smith keeps middle Jane King. See the S2 entry of docs/design/decisions.md (closes #602)

    • Fix a comma part that opens with a credential being read as a given name. HumanName("John Smith, PhD Jones") gives first John, last Smith, suffix PhD Jones and reports suffix-or-name for the name word it took, where 2.3.0 gave first PhD, middle Jones, last John Smith (and 1.4.0 title PhD, first Jones). With one word before the comma the whole part is suffixes: Doe, PhD Jones gives last Doe, suffix PhD Jones. A title word in the part stays a title. A word that is a title as well as a credential (MD, Ms) opens nothing, so Smith, Ms Jane keeps title Ms, first Jane. An undeclared delimiter becomes one more word of the part: Steven Hardman, RN - CRNA gives first Steven, last Hardman, suffix RN - CRNA, where 1.4.0 through 2.3.0 gave first RN, middle -, last Steven Hardman. The split spelling now counts like the joined one, so Smith, John Ph. D. Jones gives suffix Ph. D. Jones, where 2.3.0 gave middle Jones, suffix Ph. D.. See the C1 entry of docs/design/decisions.md (closes #603)

    • Fix a comma part read as credentials being taken apart again by a later join. HumanName("John Smith, Ph. D. and Mary Jones") gives first John, last Smith, suffix Ph. D. and Mary Jones, where 2.3.0 gave first Ph. D. and Mary, middle Jones, last John Smith, and parse() reports each name word the part took: the parser now decides once, after every word is recognized, whether the part after a comma holds the credentials, and nothing joins across that decision afterwards. Titles joined by and stay one title (John Smith, Mr. and Mrs. keeps title Mr. and Mrs.). A surname joined by y counts as its words, as #575 already counted it in front of a credential: Ortega y Gasset, Dr. gives title Dr., first Ortega, middle y, last Gasset, where 2.3.0 gave last Ortega y Gasset. And a surname with a suffix in it before the comma is the surname, as Smith Jr., John always read it: Smith Jr., Esq. gives last Smith, suffix Jr., Esq., where 2.3.0 gave first Smith, and Smith Jr., V gives first V, last Smith, suffix Jr., where 2.3.0 gave first Smith, suffix Jr., V. See the #613 entry under C1 in docs/design/decisions.md (closes #613)

    • Change a title word in a part after a second comma to read as a title. HumanName("Eric H. Holder, Jr., Attorney General") gives title Attorney General, suffix Jr., where 1.4.0 through 2.3.0 gave suffix Jr., Attorney General, and the part no longer reports comma-structure. A word that is also a suffix stays a suffix (John Smith, MD, Ms). A title joined by a connective is read as a title, but its part keeps the report (Secretary of State). (#603)

    • Add the Catalan and Polish surname link. parse("Josep Carod i Rovira") gives family Carod i Rovira, where every release from 1.4.0 through 2.3.0 gave middle Carod i with family Rovira; Josep Lluis Carod i Rovira gives middle Lluis with that same family; and Carod i Rovira, Josep gives it too, where they read family Carod Rovira and took the link into suffix as a generation marker. i is connective vocabulary now, the way y already was, and a connective counts as a name word wherever the three-word carve-out counts them – whatever else the vocabulary says the word is, which matters here because i is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so John Quincy Smith i keeps suffix i, Josep Lluis Carod i III keeps suffix i III, and the two-word Carod i keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: JOSEP CAROD I ROVIRA and josep carod i rovira keep the fields they had and gain a conjunction-or-initial report, which a one-case name gains wherever a bare i or I stands among the name’s own words – a letter inside a maiden clause is read by the clause’s rules and stays silent, as e already was – and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (parse("john smith i jr") gives middle smith, family i and suffix jr where it gave family smith and suffix i jr). Case repair follows the reading: a lower-case i the parse read as the generation is still title-cased by capitalize(force=True) (Carod i gives Carod I, as every release did), while one standing among the name words keeps its lower case as y always has (Carod i Rovira gives Carod i Rovira, where Carod I Rovira was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: HumanName("Jane Doe nee Puig i Soler") gives maiden Puig i Soler with last Doe, where 2.0 through 2.3 ended the birth name at the link and gave maiden Puig with middle Doe i, last Soler – and the same words would have joined into last Doe i Soler under the change above, carrying a word of the birth name into the current surname. Jane Doe née Kowalska i Nowak moves with it, as does the all-lower jane doe nee puig i soler; the y spelling always read this way and is untouched. The link still has to be joining: Jane Doe nee Puig i keeps maiden Puig with suffix i, and Jane Doe nee Puig i III suffix i III. A caller with Catalan or Polish data removes the entry from conjunctions_ambiguous and gets the join in the one-case names too; a caller who wants none of this removes i from conjunctions, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it – parse("Dr. John i Smith").capitalized(force=True) keeps i where the off switch gives Dr. John I Smith, and Carod i Rovira and Josep i Rovira are the same shape. Those two are the whole of what the switch does not undo, and tests/v2/test_properties.py states them as its invariants’ only exemptions. A delimiter the caller declares through Policy(extra_suffix_delimiters=...) parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under (" - ",), Smith, John, PhD - i Soler keeps suffix PhD, i Soler, as 2.3.0 read it. The default policy declares no such delimiter. See the P3 and M2 entries of docs/design/decisions.md (closes #397, closes #538)

    • Fix a connective contributing no initial even where it is joining nothing. parse("Juan de y").initials() gives J. y., where every release gave J. while family_base said y – two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, so John and Jane Smith gives J. J. S. where 2.0 through 2.3 gave J. a. J. S. and 1.4.0 the run-together J a J. S., Duke of Edinburgh gives D. E. where 2.0 through 2.3 gave D. o. E. and 1.4.0 D o E., and John & Jane gives J. J.. The question is asked of the whole part and never of a word count, so Jon Dough and has base Dough and and keeps J. D., and Juan Velasquez y Garcia keeps J. V. G.. HumanName.initials() moves with the core – over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it: JUAN Y GARCIA and محمد و علي both give the answer 1.4.0 gave. Parsing got cheaper by the same change – the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, so initials() and capitalize() disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See the R3 entry of docs/design/decisions.md (closes #461)

    • Change case repair’s exceptions map from replacement spellings to case masks, so md repairs to MD and phd to PhD. HumanName("john smith phd").capitalize() gives John Smith PhD and john smith md gives John Smith MD, where every release from 1.4.0 through 2.3.0 gave John Smith Ph.D. and John Smith M.D.. A capitalization_exceptions value is now the key’s own letters and digits in the case each should take, laid over the word as it was written, so the one phd entry repairs ph.d. to Ph.D. and JOHN SMITH PH.D. to John Smith Ph.D., and repair never adds or drops a character: john smith iii. gives John Smith III. where every release dropped the period. Punctuation in a value only marks which of its letters are joined and is never written into the word, which matters for a lone initial: john smith p.h.d. gives John Smith P.H.D., each letter the writer split off from the mask’s one run PhD being an initial, where every release gave John Smith Ph.D.. The shipped map holds phd → PhD, bsc → BSc and msc → MSc and fifteen more of the listed post-nominals whose usual spelling is mixed case – DSc, PsyD, PharmD, MDiv, ThD and Bt among them – so john smith bsc gives John Smith BSc and john smith psyd gives John Smith PsyD, where every release gave John Smith Bsc and John Smith Psyd; meng and edd get no mask, since a mask applies wherever its word stands and both are also names (Meng Li, Edd Smith); md, ii, iii and iv left it, a suffix md now repairing by the acronym repair listed under Additions and a suffix numeral by the numeral repair below. A mask still applies wherever its word stands (phd smith gives PhD Smith), but a word that left the map and was parsed as anything but a suffix repairs as that reading: iv smith gives Iv Smith where every release gave IV Smith, and Md Abdul Karim stays Md under force=True where every release gave M.D.. A value that does not spell its key’s letters and digits – {"jr": "Junior"} – now raises ValueError when the Lexicon is built, and at the first parse for a v1 Constants, where the raise names the fix in v1’s own spelling (constants.capitalization_exceptions['jr'] = 'JR') rather than the Lexicon() constructor call the same check offers a 2.0 caller; a value may still carry its own punctuation ({"md": "M.D."} is accepted, and repairs md to MD). Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 11 move on the default capitalize() path and 78 under force=True; no role field moves. The recipe is the R4 entry’s 2026-09-23 MEASURED bullet in docs/design/decisions.md (closes #459)

    • Fix case repair treating a suffix written in capitals as evidence that the whole name was cased on purpose. HumanName("juan garcia III").capitalize() gives Juan Garcia III, where every release from 1.4.0 through 2.3.0 returned it untouched – v1’s test for it had been a known failure since 2012 (the Google Code tracker’s issue 22) – and juan garcia PhD and JUAN GARCIA Jr. repair the same way. Repair still acts only on a name written wholly in one case, but the suffixes are left out of that test now: a credential or a generation written the way one is written says nothing about how the writer cased the name. A title still counts, so Dr. juan garcia is left alone, and so does every other word, nicknames and maiden names included (Juan garcia III and jane doe nee SMITH III are left alone too). A suffix written in more than one case is the writer’s spelling and is kept as written: john smith EdD gives John Smith EdD and juan garcia PsyD gives Juan Garcia PsyD, and so does a garbled one, juan garcia Iii giving Juan Garcia Iii; force=True repairs it (John Smith EDD, Juan Garcia III), and every release returned those names untouched. The parser’s own reading of a name’s case is unchanged and still counts the suffix, so the two can differ: john e jones III gives John e Jones III, the capitals making the e a connective to the parser, where john e jones iii gives John E Jones III. Because the test follows the parser’s suffix reading, jack MA gives Jack MA and MD, PhD gives Md PhD, where 2.3.0 left both untouched. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), the new test admits 17 and 12 of them move on the default path; force=True is unchanged. The recipe is the R5 entry’s 2026-09-23 MEASURED bullet in docs/design/decisions.md (closes #492)

    • Fix case repair capitalizing the connective inside a hyphenated compound surname. HumanName("jose ortega-y-gasset").capitalize() gives Jose Ortega-y-Gasset and maria silva-e-sousa gives Maria Silva-e-Sousa, the lowercase connective 1.4.0, 2.0.0 and 2.1.0 gave, where 2.2.0 and 2.3.0 gave Jose Ortega-Y-Gasset and Maria Silva-E-Sousa – so this restores 1.4.0’s answer after a 2.2 regression. JOSE ORTEGA-Y-GASSET gives the same Jose Ortega-y-Gasset, where every release gave Jose Ortega-Y-Gasset. A connective with a part on each side of it inside one hyphenated word keeps its lowercase, as the spaced spelling always has, in every role; at either end of the word it is ordinary name text, so juan e-f smith still gives Juan E-F Smith and juan y-garcia gives Juan Y-Garcia. A single letter marked with a period is an initial there, never the connective, so j.-e.-p. dupont keeps J.-E.-P. Dupont. The hyphen is read as the writer’s join even in a name written wholly in one case, where the same letter spaced reads as an initial, so a bare hyphenated initial that spells a connective is lowered – J-E-P DUPONT gives J-e-P Dupont where every release gave J-E-P Dupont, a recorded boundary – while the spaced maria silva e sousa gives Maria Silva E Sousa. Only connective vocabulary is read, so the Māori Te Awanui-a-Rangi Black still repairs to Te Awanui-A-Rangi Black under force=True. The shape is rare: none of the 1340 names in the differential corpora at the commit before this change (2026-09-23) carries it, the only movers being the rules document’s own example lines. See the R4 entry of docs/design/decisions.md (closes #478)

    • Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals. HumanName("john smith x.y.z.").capitalize() gives John Smith X.Y.Z. where every release gave John Smith X.y.z., the dotted word being a suffix now (the unlisted_dotted_suffixes change above) and repaired as a listed acronym is; and john smith vi gives John Smith VI where every release gave John Smith Vi, with vii, viii and ix alike. Both are keyed on the suffix role: Jack X.Y.Z., which keeps its surname, still repairs as a name word (Jack X.y.z. under force=True), and john smith xi still gives John Smith Xi, the parser reading xi as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (john smith B.Tech. gives John Smith B.Tech.), and reads all capitals under force=True (John Smith B.TECH., where every release gave John Smith B.tech.), which a capitalization_exceptions mask such as {"btech": "BTech"} undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under force=True. The recipe is the R4 entry’s 2026-09-23 MEASURED bullet in docs/design/decisions.md (#459)

    • Fix case repair breaking a name typed with decomposed accents at each accent. HumanName(unicodedata.normalize("NFD", "josé garcía")).capitalize() gives José García, where every release from 1.4.0 through 2.3.0 gave José GarcíA: decomposed text (NFD, which macOS file names and some databases hand back) writes í as i followed by a combining accent, and the letters after the accent were repaired as a separate word. A decomposed name now repairs the way its composed spelling does, Mac/Mc names and case-repair masks included, and the output keeps the form it was typed in. See the R4 entry of docs/design/decisions.md (closes #542)

    • Fix initials dropping the accent from a name typed with decomposed accents. parse(unicodedata.normalize("NFD", "émile zola")).initials() gives é. z., where every release gave e. z.: the initial was a word’s first code point, which in decomposed text is the letter without its accent. Decomposed katakana lost its voicing mark the same way and now keeps it: マイケル ジャクソン typed decomposed initials マ. ジ., where every release gave マ. シ.. HumanName.initials() moves alike, a decomposed name’s initials stay decomposed, and a composed name’s initials do not change. See the R3 entry of docs/design/decisions.md (closes #585)

    • Change the parse pipeline to copy its state without dataclasses.replace. Every stage returns a copy of its frozen state, and several also copy tokens one at a time; dataclasses.replace goes through fields() and __init__ on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline’s own three dataclasses and checks them at import. One parse of the benchmark’s reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, parse 406 to 370 and HumanName 443 to 407), and the call-count baselines move with them. Recomputable with uv run python tools/perf/call_count.py --against e0f1a2f; the counts for every interpreter are in the parse-cost entry of docs/design/decisions.md. No user-visible behavior changes (#546)

    • Fix a long given part after a family comma costing quadratic time. Since 2.3.0, parsing "Doe, Jane " + "Smith " * n took time growing with the square of the part’s length: going from 1,600 to 6,400 words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (Doe, Jane Smith Prof. Prof. ...) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553)

    • Fix a name ending in alternating credentials and titles costing quadratic time. Since 2.3.0, parsing "John Smith " + "MA Prof. " * n took time growing with the square of n: going from 400 to 1,600 pairs cost 12.9x the time, where 2.2.0 cost 4x. It costs 4x again (Python 3.11, HumanName, measured 2026-10-01). No field moves (closes #558)

    • Fix many leading titles plus many surname particles costing quadratic time. Parsing "Dr. " * n + "Jan " + "van Berg " * n took time growing with the product of the two counts in every 2.x release: going from 400 to 1,600 of each cost 12.2x the time at 2.2.0 and 2.3.0, and 12.5x at 2.0.0. It costs 4x now (Python 3.11, HumanName, measured 2026-10-01). No field moves (closes #559)

    • Fix a v1 ``Constants`` entry with stray whitespace raising or changing the parse. After c.suffix_not_acronyms.add("ma "), HumanName("John Smith", c) raised ValueError in 2.0 through 2.3, and after c.titles.add(" dean "), HumanName("dean john smith", c) gave title dean, where 1.4.0 read both as if the entry were absent. An entry that is empty or holds edge whitespace, a whitespace run or any whitespace character other than a single space, in any set or as a capitalization_exceptions key, is now ignored with a UserWarning naming it, as 1.4.0 ignored it. A Constants restored from a 1.4 pickle carries two such entries from 1.4.0’s own title list, and these are ignored without a warning, so HumanName("Actor John Smith", c) gives first Actor as on 1.4.0, where 2.0 through 2.3 gave title Actor. (closes #541)

    • Fix a v1 ``Constants`` entry that is only a CJK full stop raising at the first parse. After c.titles.add("。"), HumanName("john smith", c) raised ValueError in 2.3; 。, . and 。 in any set or as a capitalization_exceptions key are now ignored with the same UserWarning as an entry with stray whitespace. 1.4.0 through 2.2.0 accepted such an entry and applied it to a name token that is nothing but that full stop, which 2.3 and later cannot match. An entry ending in such a full stop no longer raises either: c.suffix_not_acronyms.add("ma。") made HumanName("jack ma", c) raise ValueError in 2.3, and now gives last ma, as without the entry. (closes #582)

    • Fix a Japanese name written in halfwidth katakana being read given-first. HumanName("山田 タロウ") gives last 山田, first タロウ, as 山田 タロウ does, where every release gave first 山田, last タロウ. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (高橋一郎 タロウ), the Chinese · divides between halfwidth kana (タロウ·ヤマダ gives first タロウ, last ヤマダ where it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title: タナカ. John gives first タナカ. where it gave title タナカ.. A name written wholly in katakana, halfwidth or not, still keeps the declared order, given-first by default (ヤマダ タロウ gives first ヤマダ), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add (Script.KATAKANA, FAMILY_FIRST) to Policy.script_orders (see East Asian names). See the W4 entry of docs/design/decisions.md (closes #594)

    • Fix a katakana name typed with a separate voicing mark being read family-first. HumanName("ア゙イ タロウ") gives first ア゙イ, last タロウ, as アイ タロウ does, where 2.1 through 2.3 gave last ア゙イ, first タロウ. A dakuten or handakuten after a kana that has no precomposed voiced form (ア゙, ン゙), or the spacing ゛ and ゜, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too: マイケル ア゙イ gives first マイケル where it gave last マイケル. Hiragana names and kanji-and-kana names with such a mark read as before. See the W4 entry of docs/design/decisions.md (closes #596)

    • Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes. HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffix is PhD, and MD, where 2.0 through 2.3 gave PhD - and MD; Smith, John, Puig - y Soler gives Puig, y Soler where they gave Puig - y Soler. Both are 1.4.0’s answers again. A delimiter declared through suffix_delimiter or Policy(extra_suffix_delimiters=...) now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. A maiden marker in a part the delimiter separates takes no maiden name, as no marker after a suffix comma does (see the maiden-marker entry above): Smith, John, MD - née Jones Smith gives suffix MD, née Jones Smith and no maiden, where 2.0 through 2.3 gave maiden Jones Smith. Only the parts after a suffix comma are affected – the credentials of Name, PhD or of Family, Given, PhD; in the name before the first comma, and in the given-name part of Family, Given, the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See the C1 entry of docs/design/decisions.md (closes #549)

    • Fix halfwidth corner brackets not being read as a nickname. HumanName("山田 「タロー」 タロウ") gives nickname タロー, last 山田, first タロウ, where every release gave middle 「タロー」 and first 山田 (the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth 「」 are the corner brackets of legacy JIS X 0201 data, the same punctuation as 「」, and are now a default nickname pair in both APIs: DEFAULT_NICKNAME_DELIMITERS gains ("「", "」") and the 1.x nickname_delimiters gains the key halfwidth_corner_brackets. They are not limited to Japanese text: John 「Jack」 Smith gives nickname Jack where it gave middle 「Jack」. A Constants restored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See the N1 entry of docs/design/decisions.md (closes #597)

    • Fix the Irish particles Ó, Ní and Ua and the Malay binti being read as a middle name. HumanName("Liam Ó Murchú") gives first Liam, last Ó Murchú, where 1.4.0 through 2.3.0 gave middle Ó, last Murchú; Sinéad Ní Mhurchú, Seán Ua Buachalla and Ina binti Navalamar (and the Singapore spelling binte) move the same way. Ó and Ní are never given names, so Ó Murchú alone is all last name, where every release gave first Ó. Ua and binti can be, so a leading one stays the first name and parse() reports particle-or-given: Ua Buachalla gives first Ua, last Buachalla, as before. Ó. written with a period is still an initial: Juan Ó. Pérez keeps middle Ó., while Juan Ó Pérez gives last Ó Pérez. Case repair writes the Irish particles capitalized, SEÁN Ó MURCHÚ repairing to Seán Ó Murchú, and binti in lowercase: INA BINTI NAVALAMAR repairs to Ina binti Navalamar, where 2.3.0 gave Ina Binti Navalamar. The Irish casing comes from new capitalization_exceptions entries, which now outrank the lowercase case repair gives a particle, and that holds for your own entries too: with constants.capitalization_exceptions['van'] = 'Van', ludwig van beethoven repairs to Ludwig Van Beethoven, where 1.4.0 through 2.3.0 kept van. See P7 and the #604 entries under vocabulary-collisions and R4 in docs/design/decisions.md (closes #604)

    • Fix salutations in Finnish, Estonian, Romanian, Croatian/Serbian, Icelandic, Czech, Lithuanian, Malay, Filipino and other languages being read as a first name. HumanName("Herra Väinö Johansson") gives title Herra, first Väinö, last Johansson, where 1.4.0 through 2.3.0 gave first Herra, middle Väinö. The titles gain Mr/Mrs/Miss forms such as rouva, proua, doamna, gospođa, frú, meneer, paní, ponas, encik and ginang; the word for Count in several of them (kreivi, krahv, hrabia, greve, graaf); familie (Familie Hansen gives title Familie, last Hansen); knight; and the offices commissioner, counsel and administrator, which complete titles that half-worked: Police Commissioner James Gordon gives title Police Commissioner, first James, where it gave title Police, first Commissioner. Some of the new titles are also names, nearly always the last name (Greve, Knight), and those still read as the last name there: Gladys Knight and Knight, Gladys are unchanged (Gladys Knight., with a trailing period, gives title Knight., as Mary Jane King. already does). The cost falls on a name that begins with one of them, the cost Graf already pays: Greve Anna gives title Greve, last Anna, and so does Hrabia Anna under FAMILY_FIRST; the few borne as first names (Herra, Batoni) lose them the same way. Words that commonly lead a real name stay out, among them Vietnamese Ông, Polish and Czech Pan, and Marshal and Justice. Knt (Knight) joins the post-nominals beside Kt: Sir John Smith Knt gives suffix Knt, where it gave middle Smith, last Knt. See the salutation-titles entry of docs/design/decisions.md (closes #606)

    Additions

    • Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials. A subset of conjunctions holding e and i by default; it is the knob for the change above rather than a switch. Portuguese data, where e links surnames the way y does in Spanish, takes it out: Lexicon.default().remove(conjunctions_ambiguous={"e"}) restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: Lexicon.default().add(conjunctions_ambiguous={"y"}). A v1 Constants has no manager of its own for it – deleting the word from conjunctions is what turns the marking off, the same rule the glued-honorific tails follow. See docs/customize.rst (#383, #479)

    • Add AmbiguityKind.CONJUNCTION_OR_INITIAL, reported when a one-letter connective in a name written wholly in one case is read as an initial: parse("jose e maria santos").ambiguities and parse("JOSE E MARIA SANTOS").ambiguities both name it, and detail names the letter. That is the call the behavior change above had to make. A letter outside the marked set reports nothing, its reading not being in doubt, so JUAN GARCIA Y LOPEZ is silent; so is every mixed-case name, where the writing decided it. See the P3 entry of docs/design/decisions.md (#383, #479)

    • Repair a credential acronym the case-repair exceptions map does not carry to all-caps instead of title-casing it. HumanName("JOHN SMITH MBA").capitalize() gives John Smith MBA where every release since 1.4.0 gave John Smith Mba; john smith jd gives John Smith JD. The repair is keyed on the word having parsed in the suffix role from the acronym vocabulary, so a word that is an ordinary name merely sharing a spelling with an acronym is untouched, and the exceptions map is still asked first – since the case-mask change above it holds masks, so john smith bsc gives John Smith BSc rather than BSC, and md, which left the map, is one of the acronyms this repair reaches (john smith md gives John Smith MD) – while the generational jr is unaffected (john smith jr gives John Smith Jr). The given-name half of a mixed run is unchanged, so QC MP gives Qc MP with the QC (given role) still title-cased and only the MP (suffix role) repaired. When this repair landed (PR #521, before the case-mask change), twenty-two names moved in the differential corpora on the default capitalize() path and 119 under force=True, every one a single-case name with an acronym suffix the map did not carry; no role field moves. See the R4 entry of docs/design/decisions.md (#459)

    • Remove ph from the default post-nominal acronyms. The fragment existed only so the merged Ph. D. token could pass the acronym test on its first piece, and the repair above would have read john smith ph. d. as John Smith PH. D.; the parser merges the split spelling by its own rule, so HumanName("John Smith Ph. D.") still gives suffix Ph. D., john smith ph. d. capitalizes to John Smith Ph. D., and phd/Ph.D. parse as they did. The cost is a bare ph with no D. behind it, dotted or not, alone or inside a credential run: HumanName("John Smith Ph.") gives middle Smith, last Ph., where every release since 1.4.0 gave suffix Ph., and John Smith MD Ph. gives middle Smith MD, last Ph., the MD leaving the suffix with it. A caller who needs that back adds it: Lexicon.default().add(suffix_acronyms={"ph"}). See the Excluded (SUFFIX_ACRONYMS -- ph) entry of docs/design/decisions.md (#459)

  • 2.3.0 - September 12, 2026

    nameparser 2.3 is parsing fixes and new honorific vocabulary; nothing in the API is removed or renamed.

    The fixes cluster around post-nominals and titles. A space-separated run of post-nominals keeps the spacing the writer typed, and the acronyms that are also surnames – Rai, Cha, Ba – no longer take a name’s family name. A run of titles addresses by its last, and a trailing abbreviated title reads as a title. CJK names and honorifics written with a full stop of any width now parse. The additions are renunciate and royal given-name titles, Devanagari and the first Bengali honorifics, and two AmbiguityKind members that report a reading nothing in the name decided.

    One incompatibility: a Lexicon pickled by 2.1.x or 2.2.x with a caller-added wide-stop or NFD entry no longer loads; the full-stop bullet below has the remedy.

    Behavior Changes

    • Fix HumanName.initials() dropping a middle- or family-group initial that is also a one-letter conjunction. HumanName("Scott E. Werner").initials() gives S. E. W. again where 2.0.0 through 2.2.0 gave S. W.; Juan Y. Garcia and a bare ASCII capital John E Smith likewise. v1 excluded initial-shaped words from its conjunction test and the 2.0 facade had not; parse(...).initials() was already right and is unchanged. A bare lowercase john e smith still reads the e as the connective. See the R3 entry of docs/design/decisions.md (closes #462)

    • Record a 2.0.0 change to HumanName.initials() that no release note had classified: since 2.0.0 the facade initials each WORD of a name part, where 1.4.0 initialed a joined run as one group – HumanName("Juan Velasquez y Garcia").initials() is J. V. G. and was J. V G.; Abdul Salam Hassan is A. S. H. and was A S. H.. Nothing changes in 2.3.0; the differential gate now compares initials() (#484) and this is what it found. See the differential-ledger, the initials view entry of docs/design/decisions.md

    • Fix a space-separated run of post-nominals rendering with a comma the writer never typed. HumanName("John Smith MD PhD").suffix is MD PhD and was MD, PhD at every release since 1.4.0; Kenneth Clarke QC MP gives QC MP, and the CJK honorific runs (김민준 박사 씨) follow the same rule. This is a deliberate deviation from 1.4.0, and the v1-parity suite is re-pinned to match. The comma forms are unchanged – HumanName("Smith, MD, PhD").suffix is still MD, PhD – because the separator is now the comma the writer typed; a configured suffix delimiter still parts a run, as does a name word standing between two post-nominals. Round-tripping is fixed for these runs, which the 2.2.0 note below recorded as broken: str(HumanName("Smith, MD PhD")) is Smith MD PhD and re-parses to suffix MD PhD. str() is still a rendering rather than a canonical form. See the C1 entry of docs/design/decisions.md (closes #436, closes #437)

    • Fix Parser.revise() splitting a space-separated suffix value into comma-separated entries. Parser().revise(n, suffix="MD PhD").suffix is MD PhD and was MD, PhD. A suffix value’s entries are now derived from the value’s own commas by the rule a whole name uses – a comma parts two credentials and a space joins them – so revise(n, suffix="MD, PhD") is still two entries, and revise(n, suffix="Ph. D.") renders Ph. D. where the 2.2.0 note below accepted Ph., D.. One limit: a delimiter configured through extra_suffix_delimiters parts a value only where the value’s own words read as a name with a tail segment, so in a run of post-nominals it stays a word; write a comma at the boundary instead. ParsedName.replace() is unchanged. See the C1 entry of docs/design/decisions.md (closes #511)

    • Remove rai and cha from the default post-nominal acronyms, so a trailing Rai or CHA keeps the family name. HumanName("Aishwarya Rai") gives last Rai, where 2.0.0 through 2.2.0 gave suffix Rai and no last name at all – 1.4.0’s reading, restored. The cost is that a genuine credential written after a full name is no longer recognized: John Smith RAI gives last RAI, and John Smith, RAI gives first RAI, last John Smith. A caller who needs either back adds it: Lexicon.default().add(suffix_acronyms={"cha"}). See the suffix-acronym-collisions entry of docs/design/decisions.md (closes #342)

    • Mark ba as an acronym that is also an ordinary name, so a bare trailing Ba keeps the family name. HumanName("Anna Ba") gives last Ba and reports a suffix-or-name ambiguity, where 2.0.0 through 2.2.0 gave suffix Ba and no last name. The spaced full-name form keeps the credential reading – John Smith BA still gives suffix BA, now flagged – and the dotted John Smith B.A. is an unflagged suffix. The comma forms move, which is the marking’s cost: John Smith, BA gives first BA, last John Smith, as Smith, Ed already did for the other ambiguous acronyms. Write B.A. to keep the credential reading. Ba is a real surname in Vietnamese and Senegalese Fula, the ma/Ma shape exactly (#342)

    • Fix a title run addressing by its first title rather than its last. HumanName("Her Majesty Queen Elizabeth") gives first Elizabeth with an empty last name, where every release since 1.4.0 gave last Elizabeth. Several titles written together are one form of address and the one that does the addressing is the last, so the run is now matched whole or by its last word: Reverend Mother Teresa, Dr. Sir John and Sir Sheikh abdul rahman move the same way. A run whose last word addresses by surname does not move: His Excellency Lord Duncan still gives last Duncan, lord not being a given-name title. A caller’s multi-word entry still matches as a phrase. See the H1 entry of docs/design/decisions.md (closes #489)

    • Fix the leading title peel taking a name word and leaving a post-nominal to be the name. HumanName("Dr King Jr") gives title Dr, last King, suffix Jr, where every release since 1.4.0 gave title Dr King, last Jr and no suffix at all; Dr. King MD moves the same way, and both now read as the comma spelling King, Dr Jr always has. A name that is nothing but titles or nothing but post-nominals is untouched (Marquess of Bath, MD DDS), and so is a title written as one joined unit: Prince of Wales Jr keeps title Prince of Wales. Where the only word left is the title itself, Dr Jr gives first Dr, suffix Jr and reports a title-or-name ambiguity. See the H3 entry of docs/design/decisions.md

    • Fix a trailing abbreviated title reading as a name word. HumanName("John Smith Prof.") gives title Prof., first John, last Smith, where every release since 1.4.0 gave last Prof. and lost the surname; John Smith Dr., John Smith Rev. and Andrew Perkins (Mgr.) move the same way, and the comma form Smith, John Prof. now agrees with the bare one where it gave middle Prof.. A run chains from the end (John Smith Prof. Dr. gives title Prof. Dr.), a leading title keeps its place (Dr. John Smith Prof. gives title Dr. Prof.), and a post-nominal is read through it (John Smith Jr. Prof. gives suffix Jr.). Only a listed title word wearing the abbreviation period is claimed: an unlisted abbreviation (John Smith Xyz.), a bare title word (John Smith Sir) and a post-nominal (John Smith Esq.) do not move. The reach is the whole title vocabulary, ordinary surnames in it included, so Mary Jane King. gives title King. where the bare Mary Jane King keeps last King – accepted rather than prevented, the period being evidence the bare spelling never gives. The same argument holds in a native script: 毛 泽东 Dr. gives title Dr., first 泽东, last 毛, where 2.2.0 gave last Dr.. See the H5 entry of docs/design/decisions.md (closes #316)

    • Remove esq from the default post-nominal acronyms, and assert the two post-nominal sets disjoint. HumanName("John Smith E.S.Q.") gives middle Smith, last E.S.Q., where every release since 1.4.0 gave suffix E.S.Q.; Esq, Esq., ESQ and esq are unchanged, the post-nominal word list carrying every single-token spelling. Esquire is a contraction rather than an initialism, and it was the one word in both post-nominal sets, which can now assert they do not overlap. A caller who needs the dotted spelling back adds it: Lexicon.default().add(suffix_acronyms={"esq"}). See the suffix-acronym-collisions entry of docs/design/decisions.md

    • Fix the East Slavic and Turkic patronymic rotations overriding a declared family-first name order. With patronymic_rules opted in and Policy(name_order=FAMILY_FIRST), Мицкевич Адам Юзеф gave last Адам through 2.2.0 and now gives last Мицкевич – the reading the declaration asks for – and oglu Ahmad Vali Ali with Turkic handling gave last Ahmad and now oglu. The rotations exist to restore the given-first reading a family-first listing hides, so under a declared family-first order the declaration decides. See the O1 entry of docs/design/decisions.md (closes #384)

    • Fix CJK honorifics and names written with a full stop of any width. HumanName("김민준 씨.").suffix is 씨. (the fullwidth stop a Japanese or Chinese IME produces by default), with last 김 and first 민준, where every release since 2.1.0 gave first 김, middle 민준, last 씨. and no suffix; the ideographic 씨。 and halfwidth 씨。 spellings move the same way. The vocabulary lookup now composes NFC, so a decomposed Señor or née from macOS-origin data is recognized too. A period glued to a name word no longer breaks the name: 양. 지훈 gives last 양., first 지훈 where 2.1.0 through 2.2.0 gave first 양., middle 지, last 훈, and 김민준씨. peels its honorific as the stop-less spelling does. The period stays on the word it was written with; nothing is rewritten. A Latin name written with ASCII periods is untouched (Smith. John still reads title Smith.), but a Latin or Cyrillic word wearing one of the three wider stops now reaches the vocabulary: Dr。 John Smith reads title Dr。 where it read first Dr。. This retires the 2.2.0 note below that read period strictly. One incompatibility, by decision: a Lexicon pickled by 2.1.x or 2.2.x that carries a caller-added entry the widened fold now changes – a non-ASCII entry written with a fullwidth or ideographic stop, or in NFD – no longer loads (ValueError: incompatible Lexicon pickle: entries are not normalized); the shipped vocabulary is unaffected, and the remedy is to rebuild the Lexicon from its source rather than unpickle it. See the cjk-full-stops entry of docs/design/decisions.md (closes #322, closes #323)

    Additions

    • Add the renunciate titles to the given-name title list, so a renunciate’s one name is a given name. HumanName("Swami Vivekananda") gives first Vivekananda with an empty last name, where every release since 1.4.0 gave last Vivekananda; Guru Nanak, Baba Ramdev and Lama Zopa move the same way, and so do the Devanagari and Bengali spellings added below. Two name words behind the title are unchanged – Swami Vivekananda Saraswati keeps last Saraswati – and a surname-retaining title is untouched: Rabbi Cohen still gives last Cohen. venerable is deliberately not in the list, the traditions using it splitting on whether the family name survives. See the indic-honorifics entry of docs/design/decisions.md (closes #346)

    • Add prince and princess to the given-name title list, so a royal’s one name is a given name. HumanName("Prince Harry") gives first Harry with an empty last name, where every release since 1.4.0 gave last Harry; Princess Anne and Her Royal Highness Princess Anne move the same way. Two name words behind the title are unchanged – Prince Harry Windsor keeps last Windsor – and so is Prince of Wales Jr, the joined title addressing by its last word. lord and lady are deliberately NOT in the list: they address by given name only as a courtesy style (Lord Peter, Lady Diana) and by surname for every peer and every wife (Lord Byron, Lady Thatcher), which the list cannot express. The given-name collision is untouched: Prince Fielder now gives first Fielder (#348). See the H1 entry of docs/design/decisions.md (closes #519)

    • Add trailing honorifics as post-nominal vocabulary: Latin rinpoche, Devanagari जी, साहब, साहिब, साहेब, महाराज, and Bengali সাহেব, বাবু, মহারাজ. HumanName("Lama Zopa Rinpoche") reads title Lama, first Zopa, suffix Rinpoche; नरेन्द्र मोदी जी reads last मोदी, suffix जी. They are recognized SPACED only and are deliberately absent from Lexicon.honorific_tails: Banerjee, Mukherjee and Chatterjee end in the जी substring (बनर्जी, मुखर्जी, चटर्जी), so a glued peel would cut a real family name in two, and गांधीजी staying unpeeled is the accepted cost. Bengali বাবু is trailing where Devanagari बाबू is a leading title (#344, #343)

    • Add Devanagari honorifics (#344): डॉक्टर, डा, प्रो, प्रोफेसर, प्राध्यापक, प्रा, पंडित, पं, सरदार, सुश्री, श्रीयुत, श्रीमान, सौ, बाबू, महात्मा, न्यायमूर्ति, मौलाना, जनाब and महाराजा as titles, beside the श्री/श्रीमती/डॉ that shipped in 2.1.0, and स्वामी, गुरु, बाबा, संत as given-name titles. डॉक्टर शर्मा reads title डॉक्टर, last शर्मा; स्वामी विवेकानंद reads first विवेकानंद with no last name. Dotted spellings (प्रो., पं.) match the same entries. Excluded under the collision rule: कुमारी (Kumari is a given and a family name), बेगम, शेख, आचार्य, राजा/रानी and ठाकुर, all borne as ordinary names (closes #344)

    • Add Bengali honorifics – the first Bengali vocabulary in the default lexicon (#343): ড, ডঃ, ডক্টর, ডাঃ, ডা, ডাক্তার, শ্রী, শ্রীমতী, জনাব, অধ্যাপক, প্রফেসর, বিচারপতি, মাওলানা, মুফতি, আলহাজ্ব, আলহাজ, মিঃ, মি, মিসেস, মোঃ, মো, মোসাঃ, মোসা, মোছাঃ and মোছা as titles, and স্বামী, শ্রীল, গুরু, বাবা as given-name titles. ড. মুহাম্মদ ইউনূস reads title ড., first মুহাম্মদ, last ইউনূস – the vocabulary beats the initial reading – while real initials are untouched: র. কে. নারায়ণ is unchanged. মোঃ আবদুল করিম reads title মোঃ, first আবদুল, last করিম, the mirror of Latin Md; the visarga spelling and the মো. period spelling both match, and the women’s মোসাঃ/মোসা. rides the same pair of entries. ঠাকুর stays out, being Tagore. Latin transliterations (Sri, Pandit, Mst) are not added – they collide with real given names where the native scripts cannot – and belong to the opt-in packs of #345 (closes #343)

    • Add AmbiguityKind.GIVEN_OR_FAMILY, reported when a name of one name word had nothing to decide which field it is: parse("Andrew") still gives given Andrew and now says that field was a convention rather than a reading – the library picks the given name under the default order and the family name under a declared family-first one, and detail names the field it picked. parse("Smith Jr.") reports it too, the suffix being peeled first. A name something DID decide stays silent – "Dr. Smith", "Smith, Andrew", "abdul" (bound given-name vocabulary), "J." (an initial’s shape) – and so does "毛泽东", where the writing system settles the order. No field moves anywhere. See the O5 entry of docs/design/decisions.md (closes #449)

    • Add AmbiguityKind.TITLE_OR_NAME, reported when an input that is nothing but honorifics had its last word read as the name: parse("Lord Chancellor") still gives title Lord, family Chancellor, and now says so – this is a name parser, not a title parser, so handed a string with no name in it, it reads the last title word as one. "His Holiness" and "Dr. King" move the same way, king being title vocabulary. A title with an ordinary word behind it is silent ("Dr. Smith", "King Charles"), and so is a lone title word: parse("Dr.") is a title with no name beside it. The same convention on the post-nominal vocabulary reports the existing suffix-or-name: parse("Rinpoche") gives given Rinpoche and flags it, as does "QC MP". No field moves for this change; Dr King Jr, Dr. King MD and Dr Jr gain the report from the title-peel fix above. See the H4 entry of docs/design/decisions.md (closes #491)

    Documentation

    • Document the family-first and East Asian input shapes beside the three Latin ones. The input-shapes list in usage.rst grows from three forms to seven: forms 4 and 5 for a declared FAMILY_FIRST or FAMILY_FIRST_GIVEN_LAST order, and forms 6 and 7 for the native East Asian arrangements the script carries on its own. A comma or a Latin wrapper around a CJK name is named as tolerated input – parsed best-effort, its handling changeable without notice – and the customize.rst correspondence between forms 2 and 4 is written out. Behavior is unchanged (#469)

  • 2.2.0 - August 31, 2026

    nameparser 2.2 is a rename plus about thirty parsing fixes.

    The nameparser.config word lists were still named for v1’s fields — PREFIXES, BOUND_FIRST_NAMES, FIRST_NAME_TITLES — while the Lexicon they feed has used particles and given names since 2.0. They now agree, and the lists are frozen, which retires editing one in place as a way to change a default. The rename itself changes no parse.

    The fixes cluster around surname particles, largely what a declared name_order means for Latin-script names, which this release settles; then maiden-name clauses, Arabic bound given names, and credentials after a comma. Most reach the default name order, and a bullet says so where its change is family-first only. Each names the shapes it moves, and the issue it closes carries the measurement.

    What breaks is code that writes to a default word list. Code that imports one by its 1.x name has until 3.0.

    Breaking Changes

    • Add docs/design/ contributor documentation: rules.md (the parser’s normative rules, with executable examples), decisions.md (the decision record) and mechanisms.md (the solution-pattern catalog). New tests execute every documented example and verify every code citation

    • Change every vocabulary set in nameparser.config to a frozenset. Editing one in place – TITLES.add("dean"), the old way of changing a global default – now raises AttributeError at the line that writes it. To change the defaults for HumanName, build a private Constants and pass it (c = Constants(); c.titles.add("dean"); HumanName(name, constants=c)); for the 2.0 API, build a lexicon (Parser(lexicon=Lexicon.default().add(titles={"dean"}))). Mutating the shared CONSTANTS still works, but warns and goes away in 3.0. CAPITALIZATION_EXCEPTIONS is a mapping, not a set, and is unchanged. See Migrating from HumanName and Customizing the parser (#293)

    Behavior Changes

    • Fix a title changing how the name behind it is read. "Dr. Van Johnson" gave family Van Johnson with no given name, and "Sir Van Johnson" gave given Van Johnson with no family at all; both now read given Van, family Johnson – the reading the untitled "Van Johnson" has always had. A leading word that is both a title and a particle is unchanged: "St John Smith", "Do John Smith" and "Freiherr von Richthofen" keep their readings. See the P2 entry of docs/design/decisions.md (closes #367)

    • Fix a given-name title keeping a bound given name from joining the word after it. "Sheik abdul salam" read given abdul, family salam, and now reads given abdul salam with an empty family, as "Sir John" does; "الشيخ عبد الله" reads given عبد الله. A title that addresses by family is unchanged ("Dr. abdul salam"). This also restores "Sheik Abu Bakar" to given Abu Bakar, which the fix above had regressed, and drops the PARTICLE_OR_GIVEN ambiguity that name reported through 2.1 (closes #369)

    • Fix a bound given name swallowing the family name before a single-letter generational suffix. "abdul Smith V" read given abdul Smith with no family, where "abdul Smith II" and "abdul Smith Jr" read correctly; it now reads given abdul, family Smith, suffix V, and so do I and X, for every bound given-name word. A suffix word before the numeral no longer hides it ("abdul Smith Jr V" reads family Smith). Shipped since 1.x: 1.4.0 read first abdul Smith, last V (closes #401)

    • Fix a bound given name joining a suffix as “the word after it”. "abdul Jr Smith Berg" read given abdul Jr and now reads given abdul, middle Jr Smith, where "John Jr Smith Berg" puts it. Where the suffix was a split credential the bound word joined into it – "abdul Ph. D. Smith Berg" read suffix abdul Ph. D., a 2.0 regression – and now reads given abdul, middle Smith, suffix Ph. D. (closes #421)

    • Fix a bound given name joining past a credential that the suffix rule then takes, leaving no family. "abdul Smith Jr Ma" read given abdul Smith with no family and now reads family Smith, suffix Jr, Ma, as "John Smith Jr Ma" does; "abdul Smith Ma" reads family Smith, suffix Ma. Both as 1.4.0 read them. "abdul Smith Berg Ma" keeps its join, and "Berg, abdul Sir" still reads given abdul Sir (closes #425)

    • Remove the Czech/Slovak abbreviation roz. from the default maiden markers. Marker matching is case-folded and period-insensitive, so Roz – the diminutive of Rosalind – was the same string as the marker, and a marker takes every word after it: "Rosalind Roz Smith" read maiden Smith with no family name at all. It and "Rosalind Roz Jones Smith" now read as 1.4.0 read them. The full participle is untouched – "Anna Nováková rozená Svobodová" still reads maiden Svobodová – and a caller who wants the abbreviation back adds it to their own lexicon: Parser(lexicon=Lexicon.default().add(maiden_markers={"roz"})) (found in #335’s review)

    • Add the Polish maiden marker z domu to the default vocabulary, and let a maiden_markers entry be more than one word. "Maria Kowalska z domu Nowak" now reads family Kowalska, maiden Nowak, where every earlier version read the marker as part of the name (1.4.0: middle Kowalska z domu, family Nowak). The bracketed spelling moves with it. maiden_markers and given_name_titles are now the two fields exempt from the multi-word warning. See Customizing the parser (#434)

    • Fix a bracketed maiden clause reading as a nickname because its brackets were not declared. "Jane Smith nee Jones" gave maiden Jones while "Jane Smith (née Jones)" gave nickname née Jones; the bracketed spelling now reads family Smith, maiden Jones too, and so does the Japanese "山田 花子(旧姓 佐藤)", which needed Policy(maiden_delimiters=...) through 2.1. Every delimiter pair the parser ships moves the same way, quotes included. An interior clause no longer eats the name behind it ("Jane (née Jones) Smith" keeps family Smith), and two clauses beside each other each keep their own role ("Jane "Janey" Smith (née Jones)" reads nickname Janey, maiden Jones). A clause with no marker in it is still a nickname, which is what Policy(maiden_delimiters=...) remains for. This reaches HumanName (closes #335)

    • Fix a particle chain and a maiden name taking a trailing generational numeral as a name word. "John van der Berg V" read family van der Berg V and "John née Jones Smith V" read maiden Jones Smith V, where "John Smith V" reads suffix V; both now stop before the numeral, for I and X alike. A word before the numeral that is an initial keeps its reading ("John van der J. V"). The chain also stops before a bare credential with words to spare – "John van der Berg Ma" reads suffix Ma, as 1.4.0 did – and no longer swallows the given name behind an unlisted abbreviation: "Xyz. van Johnson" and "Esq. van Gogh" read given van (closes #424)

    • Fix a name losing its given/family split when a comma is followed only by an honorific. "John Smith, Mr." returned the whole of "John Smith" as the family name and now gives given John, family Smith, title Mr.: it is "Mr. John Smith" with the honorific moved to the end, and marks no surname boundary. A comma followed by an actual name still fixes the family ("John Smith, Jones"), and a single pre-comma piece has no split to keep ("Smith, Dr." is unchanged). The pre-comma name now also picks up the declared name order – "de Mesnil Jean, Dr." keeps family de Mesnil under a family-first order – and the particle-or-given ambiguity report ("Van Johnson, Mr.")

    • Fix pure postnominals being claimed as titles: jr, junior, phd, do and se have left the default titles vocabulary, and dr/sra have left the suffix vocabulary they never belonged in, so "Smith, PhD" gives suffix rather than title PhD. Twelve words are genuine duals and keep both memberships, with position deciding – "Lt. Smith" is a title, "Smith, LT" a postnominal, bare Md before a name the Bengali and South Asian abbreviation of Muhammad, MD after it the degree. The cost is in leading position, where a dropped word now reads as a name: "PhD Smith" gives given PhD, which is what makes "Do Nguyen" parse as the Vietnamese name it is. dr and sra also stop being recognized in trailing position, so "John Smith Dr." gives family Dr.. An ambiguous credential acronym (ma, ed, jd, do) counts as a suffix only when written with its periods, so "Jack Ma." keeps family Ma. as 1.4.0 read it. Routing a trailing title word to title is a separate open question (#316)

    • Fix a credential run after a one-word family comma reading as a title or a given name. "Smith, Jr." and "Smith, PhD" now give suffix Jr./PhD where they gave title, and "Smith, Ph. D. Jr." gives suffix Ph. D. Jr. where the split credential landed in the given name – a regression from 1.4.0. The position right after a family comma is postnominal position. Vocabulary still decides which words qualify ("Smith, Dr." keeps title Dr.), the leading readings are untouched ("Sr. Garcia" is still title Sr.), and a name word in the run makes it the given-and-suffix reading it always had ("Smith, John Jr.") (closes #296, closes #325)

    • Fix a space-separated credential run after a family comma rendering with a comma the name never had. "Smith, MD PhD" gives suffix MD PhD where it gave MD, PhD, and "Smith, CBE MC", "Smith, BSc MBA" and "Smith, Dr. MD PhD" the same. The roles are unchanged; only the rendered string carried the extra comma. This reaches any family comma whose following segment holds no name word, not only a one-word family, so "John Smith, Jr. III" gives suffix Jr. III – also what 1.4.0 gave. A run written with commas keeps them ("Smith, MD, PhD"), and a name written without a comma is unaffected and still renders its run comma-joined, so re-parsing str() output does not reproduce the run (closes #429)

    • Fix a one-character suffix word after a comma being read by the wrong neighbour. "Smith, PSM I" gives suffix PSM I where it gave given PSM and suffix I, and "Smith, John V." gives middle V. where it gave suffix V.. Inside a comma part a suffix word short enough to be mistaken for an initial – I, V and 2 in the shipped vocabulary – is read by what stands before it: behind a credential it describes that credential (PSM I is Professional Scrum Master level I), and behind a name a period marks an abbreviation and so a middle initial. A numeral written bare after a name is still the generation it looks like ("Smith, John V" is suffix V), and a name with no comma is untouched (closes #430, closes #432)

    • Fix a name opening with a particle that is never a given name being split at the particle under a family-first name order. "de Mesnil" read as family de, given Mesnil and "de la Vega" as family de, given la Vega; each is now the whole surname, as it has always been in the default order, under FAMILY_FIRST and FAMILY_FIRST_GIVEN_LAST alike. A word that can never be a given name leaves name_order nothing to decide. Standing alone is the whole of it: "Juan de la Vega" under FAMILY_FIRST still reports given de la Vega. A leading particle that may be a given name is genuinely order-dependent and is untouched, so "van Gogh" still reads family van, given Gogh under both family-first orders (closes #359)

    • Fix a family name made only of particle words reporting no base, so the surname vanished from family_base and from the initials. parse("Anh Do") gave family Do with family_base '' and initials A., and under Policy(name_order=FAMILY_FIRST) "Del Toro" gave family Del the same way. A particle standing alone in a name part is not doing a particle’s work there and now reads as an ordinary name word: "Anh Do" is base Do, initials A. D.; "Juan van der" is base van der, initials J. v. d.; "Nguyen, Van Le" initials V. L. N. where the middle name used to be dropped. The parse fields themselves do not move – only the derived views and the initials. See the R2 entry of docs/design/decisions.md (closes #385, closes #402)

    • Fix case repair lowercasing the words of a family name made only of particle words, where every other view already reads them as ordinary name words. HumanName("ANH DO").capitalize() gives Anh Do where it gave Anh do, and "anh van do" gives Anh Van Do. This DIFFERS FROM 1.4.0 deliberately and does not restore it: 1.4.0 returned Anh do. The accepted cost is that a family which is nothing but particles capitalizes too, so "juan van der" gives Juan Van Der. A conjunction is untouched ("der, y van" gives y Van Der), and where the particles DO join a name word nothing changes ("juan de la vega" still gives Juan de la Vega) (closes #407)

    • Change case repair to read the parser’s own conjunction tag instead of re-deciding, from the word’s spelling, whether a word is a conjunction or an initial. Two spellings of one name disagreed because of it: "jose ortega-y-gasset" capitalized to Jose Ortega-y-Gasset while "JOSE ORTEGA-Y-GASSET" gave Jose Ortega-Y-Gasset; both give Jose Ortega-Y-Gasset now, a hyphenated token being one word to the parse whatever it contains. The spaced spelling is untouched and still repairs to Jose Ortega y Gasset. A field assigned after the parse was never classified, so repair asks the vocabulary there – today’s vocabulary, which is narrower than 1.4.0 parity: h.last = "хосе и мария сантос" gives Хосе И Мария Сантос on 1.4.0 and Хосе и Мария Сантос here. One reading changes for hand-built Tokens in the 2.0 API: an untagged token whose text is conjunction vocabulary now capitalizes as an ordinary name word. See the R4 entry of docs/design/decisions.md (closes #458)

    • Change the parse-cost benchmark to bound function calls per parse rather than wall-clock seconds. The old one-second bound failed four times on CI while the same code re-ran green on master; frame counts do not move under load. The bound is a per-interpreter band of ±2%, with a loose five-second backstop for what frame counts cannot see. Recomputable with uv run python tools/perf/call_count.py --against v2.1.0; the counts and the per-PR attribution are in the parse-cost entry of docs/design/decisions.md. No user-visible behavior changes (closes #475)

    • Fix a name that opens with a spaced Ph. D. losing its surname. parse("Ph. D. Van Johnson") read given Van Johnson with an empty family and suffix Ph. D.; it now reads title Ph., given D., family Van Johnson. A suffix never begins a name, and the split credential was the only shape that reached the defect. A family comma still opens a listing rather than a name ("John Smith Ph. D." and "Smith, Ph. D. Jr." keep their suffixes), and “the head” means the head of the string rather than of the name, so "Sir Ph. D. Van Johnson" is unchanged. This RESTORES 1.4.0. One accepted consequence: Parser.revise(suffix="Ph. D.") renders Ph., D. (closes #371)

    • Fix a trailing surname particle being stranded as a standalone middle name under a family-first name order. Under Policy(name_order=FAMILY_FIRST) the same listing written with a comma reads it as part of the surname: "Jong Anke de" gave family Jong with de left as a middle name and now gives family de Jong, given Anke – the answer parse("Jong, Anke de") has always given. FAMILY_FIRST is the only order that puts a trailing piece in a middle; FAMILY_FIRST_GIVEN_LAST puts it in the given slot, so "Nguyen Thi Van" under that order still reads given Van. "Beethoven Ludwig van" under FAMILY_FIRST now gives family van Beethoven. A particle standing alone in the given slot is no longer folded into the family either, so "Ménil de" reports given de. Nothing moves under the DEFAULT name order. See Customizing the parser for what declaring an order settles, and the P6 entry of docs/design/decisions.md for the reasoning (closes #467)

    • Fix initials() ordering a name differently from the fields of the same parse. Two rules fold words into the family and render them ahead of it – Policy(middle_as_family=True) and the tussenvoegsel attachment after a family comma – and the family field honored the fold where initials() did not: parse("der, y van") gave family van der but initials y. d. v., and now gives y. v. d.. Under middle_as_family this RESTORES v1, that option being middle_name_as_last’s successor: "Doe, Dr. John A." gives J. A. D. again where 2.0 through 2.2 gave J. D. A.. HumanName.initials() was already right and is unchanged. See the R3 entry of docs/design/decisions.md (closes #408)

    • Fix a tussenvoegsel attached to the family name after a comma deciding a genuinely uncertain reading and reporting nothing. "Van Johnson" reports a PARTICLE_OR_GIVEN ambiguity – Van is a Dutch particle and a Vietnamese given name – while "Nguyen, Thi Van" picked the same word the same way, silently. The attachment now reports the fork it decides: "Nguyen, Thi Van", "Berg, Jan van der" and "Vega, Juan de la" each gain a PARTICLE_OR_GIVEN, while a particle already read as a post-nominal reports SUFFIX_OR_NAME instead ("Berg, Jan vd"). Worth knowing before you filter on this: "Beethoven, Ludwig van" – read exactly right – now carries a report too, nothing in the input separating it from "Nguyen, Thi Van". ambiguities is the only value that grows (closes #405)

    • Fix a tussenvoegsel after a family comma being parsed as a middle name. Dutch and Belgian alphabetized listings move the particle behind the given name – "Beethoven, Ludwig van" is how "Ludwig van Beethoven" is filed – and it was read as a middle name rather than as part of the surname: "Beethoven, Ludwig van" gave middle van, last Beethoven, and "Berg, Jan van der" gave middle van der. Those now read family van Beethoven and van der Berg. Two guards bound it: a name whose only given word is the particle keeps it ("Nguyen, Van" still reads given Van), and where the word is BOTH particle and suffix vocabulary the attachment wins, so "Berg, Jan vd" reads family vd Berg where 1.4.0 and 2.1 alike gave suffix vd – as does mc. (closes #379, closes #380)

    • Add abd to BOUND_GIVEN_NAMES, so the spellings that write the article as its own word join like the others do: "abd Allah Smith" was given abd, middle Allah and is now given abd Allah. abdul, abdel and abdal were already there, and the Arabic-script عبد has covered the same word since 2.0, so only the Latin spelling was short. The word is also the postnominal ABD (“All But Dissertation”) and stays in SUFFIX_ACRONYMS: position tells the two readings apart, so "Jane Smith ABD", "Jane Smith, ABD" and "Jane Smith A.B.D." all still read the credential as a suffix (#400)

    • Change how far a leading never-given particle takes the surname when a family-first name_order is declared. Policy(name_order=FAMILY_FIRST) read "de Mesnil Jean" as family de Mesnil Jean – the whole name – and now reads family de Mesnil, given Jean. The default order is unchanged, deliberately: with no order declared nothing marks where the surname ends, and a particle followed by several words really can be all surname (von Bergen Wessels); a caller who means family de la Vega plus given Juan there writes the comma. The stop cannot land inside a conjunction-joined run or a bound given-name pair: "de la Vega y Santos Juan" reads family de la Vega y Santos, "ibn Awf abdul Rahman" given abdul Rahman. Where two or more words are left over, the two family-first orders differ from each other for the first time: "de la Cruz Juan Carlos" reads given Juan, middle Carlos under FAMILY_FIRST and the reverse under FAMILY_FIRST_GIVEN_LAST. See Customizing the parser, and the P1 entry of docs/design/decisions.md for the reasoning (closes #395)

    • Change the detail text of a PARTICLE_OR_GIVEN ambiguity to name the role the leading particle was actually given. It said “read as a given name” under every name_order, which is false under Policy(name_order=FAMILY_FIRST) – there "Van Johnson" reads family Van, given Johnson, and the report described the reading not taken. It now ends “read as a family name” in that case. The kind is unchanged and stays PARTICLE_OR_GIVEN; only the human-readable text moved, and default-order output is identical (#355)

    • Fix a maiden name being lost when a particle stood in front of the marker. "Ursula Leyen geb. Albrecht" reported maiden Albrecht correctly, but "Ursula von der Leyen geb. Albrecht" – the same words one particle chain apart – gave family von der Leyen geb. Albrecht and no maiden name at all, as did "Jane van der Berg née Jones". A suffix already stopped the particle chain; a marker now does too, so those read family von der Leyen maiden Albrecht and family van der Berg maiden Jones. Under a family-first order "de la Cruz née Vega" now reads family de la Cruz, maiden Vega. Two limits remain, both recorded in rules.md#M2: a conjunction join and a bound given-name join each still absorb a marker first (closes #399)

    • Move mc and ste into the never-given half of the particle vocabulary, and add los, las and das. The Spanish and Portuguese articles were absent from it entirely. A never-given particle opening a name folds into the family (rules.md#P1) instead of being read as a given name, so "Mc Donald" was first Mc, last Donald and is now last Mc Donald; "Ste Marie", "Los Santos", "Las Casas" and "Das Silva" move the same way, and "Mc Donald Smith" becomes last Mc Donald Smith. The PARTICLE_OR_GIVEN ambiguity goes with it for mc and ste. los, las and das were not particles at all, so for those three the ordinary particle join fires from a non-leading position too: "Maria das Neves" is now last das Neves. (closes #360)

    • Fix a bound given-name join leaving no family name when the name also carries a maiden clause, and stop the join absorbing the marker itself. "abdul Berg née Jones" read given abdul Berg with an EMPTY family, where "abdul Berg" alone reads given abdul, family Berg; it now reads given abdul, family Berg, maiden Jones. The join also declines when the piece it would absorb is a marker, so "van der Berg, abdul née Jones" reads given abdul, family van der Berg, maiden Jones where it read given abdul née. Where the bound word is ALSO suffix vocabulary, a declining join after a family comma leaves the post-nominal reading: "Berg, abd née Jones" reads family Berg, suffix abd, maiden Jones, as "Berg, abd" alone always has (closes #411)

    • Fix a maiden clause changing how the rest of the name is read, and a connective join keeping the marker in the surname. The grouping rules that count a name’s words counted the maiden clause too: "juan y garcia" reads given juan, middle y, family garcia, while "juan y garcia nee jones" read given juan y garcia with NO family name at all. A name of two or more name words now reads as it reads without its maiden clause, plus the maiden name. The connective join used to merge the marker into a multi-word piece, so "Jane van der Berg née y Jones" kept family van der Berg née y Jones and now reads family van der Berg, maiden y Jones. Two limits: a bound given-name word still never joins onto a marker standing as a word of its own ("Berg, abdul née PhD"), and a suffix-vocabulary word inside the maiden name stops the marker, so "Jane née Jr y Jones" now reads family Jr y Jones with no maiden name (closes #412, closes #417, closes #418)

    • Fix a title-plus-surname name losing its family name whenever anything stood beside it. "Dr. Smith" reads family Smith, but "Dr. Smith née Jones" read given Smith with no family at all, and so did "Dr. Smith PhD" and "Dr. "Smitty" Smith". A suffix, a nickname and a maiden name each stand beside the name rather than in it, and the rule now counts name words alone: those read family Smith with maiden Jones, suffix PhD and nickname Smitty respectively, and "Freiherr von Richthofen geb. Albrecht" reads family von Richthofen, maiden Albrecht. A given-name title still names no family: "Sir John née Jones" keeps given John, exactly as "Sir John" does. One name moves where the nickname LEADS: "'Smitty' Dr. Jones" reads family Jones where it read given (closes #410)

    • Fix a name that is a surname and a maiden clause reporting no family at all. "Smith née Jones" read given Smith with an empty family; it now reads family Smith, maiden Jones. A maiden marker announces a FORMER surname, which only means something beside a current one, so the lone name word left standing is the surname in use now. Every spelling of the shape moves, including a bracket pair you declared to mean maiden: under Policy(maiden_delimiters=frozenset({("(", ")")})), "Smith (Jones)" reads family Smith, maiden Jones. "Smith (née Jones)" RESTORES 1.4.0, which read family Smith with the clause as a nickname. Two shapes deliberately do NOT move: a word the vocabulary claims as a given name ("abd née Jones") and a word written as an initial ("J. née Jones Smith V") – this rule changes what POSITION decided and does not reach what a word already is. If you have code that reads the lone name word beside a maiden clause out of given, this is the release where it moves to family (closes #445)

    • Change what a star import of the two 2.2 vocabulary modules binds. nameparser.config.particles and nameparser.config.bound_given_names now declare __all__, which their 1.x shims already did, so from ... import * binds their vocabulary alone – it also bound the assert_normalized invariant helper, and from particles the BOUND_GIVEN_NAMES it imports only for a disjointness check. Importing a constant by name is unaffected, and no parse changes (#356)

    Deprecations

    • Rename the four vocabularies whose 1.x names described the fields they feed in v1’s words, so the data layer matches the Lexicon:

      1.x name

      2.2 name

      nameparser.config.prefixes

      nameparser.config.particles

      prefixes.PREFIXES

      particles.PARTICLES

      prefixes.NON_FIRST_NAME_PREFIXES

      particles.NON_GIVEN_NAME_PARTICLES

      nameparser.config.bound_first_names

      nameparser.config.bound_given_names

      bound_first_names.BOUND_FIRST_NAMES

      bound_given_names.BOUND_GIVEN_NAMES

      titles.FIRST_NAME_TITLES

      titles.GIVEN_NAME_TITLES

      suffixes.SUFFIX_NOT_ACRONYMS

      suffixes.SUFFIX_WORDS

      Every row above still resolves and is removed in 3.0. The two module rows are import paths: importing them still works and says nothing, since both modules are now empty shims. Reading a constant – by attribute access, by from ... import, or by from ... import * – emits a DeprecationWarning naming the module and constant to move to, once per line that reads it rather than once per process, so every place you have to edit is reported rather than only whichever one ran first. python -W error::DeprecationWarning -c "import yourapp" surfaces them; Python hides DeprecationWarning outside __main__. The CONSTANTS attribute names (prefixes, non_first_name_prefixes, bound_first_names, first_name_titles, suffix_not_acronyms) are v1 facade surface and are unchanged. See Migrating from HumanName (#293)

  • 2.1.0 - August 7, 2026

    nameparser 2.1 makes East Asian names work without configuration. A name written wholly in Han or hangul, or in kanji with kana, is read family-first. An unspaced Korean name is split against the census surname list. CJK honorifics are recognized whether they are spaced or written against the name. The two conventions that need you to declare a language, Han segmentation and kana-aware division, ship as the opt-in locales.ZH and locales.JA packs. See East Asian names for how it fits together, and Customizing the parser for the switches that turn it off.

    Most of this is default-on, deliberately: wherever nameparser acts unasked, the script itself settles the convention and no language detection is involved. Latin-script names are unaffected. To restore 2.0’s reading of CJK text, use Parser(policy=Policy( script_orders=(), segment_scripts=frozenset())).

    East Asian name support

    • Add the Chinese locale pack locales.ZH: opt-in Han segmentation for unspaced names like 毛泽东, with the surname vocabulary it needs. A pack rather than a default because a Chinese surname list corrupts Japanese names written in the same characters (高橋一郎 would split 高 + 橋一郎). Japanese data goes through locales.JA instead. See Locale packs (#271)

    • Add Lexicon.surnames, Policy.script_orders, Policy.segment_scripts, the Script enum and the DEFAULT_SCRIPT_ORDERS constant to the public API. This is the first behavior nameparser keys on the script a name is written in, allowed only where the script itself settles a convention and never as a proxy for guessing the language. See API reference (#271)

    • Add Lexicon.honorific_tails, the vocabulary the glued-honorific peel matches: entries that may be split off the end of a name token. It is a narrower set than the spaced honorific vocabulary, since a glued tail has no token boundary to lean on, and every entry must also be a suffix_words entry. Extend both in one call, or adding to honorific_tails alone raises ValueError. See Customizing the parser (#308)

    • Add AmbiguityKind.SEGMENTATION, reported when a surname split had a vocabulary-supported alternative: "남궁민수" is 남궁 + 민수 by the compound surname but 남 + 궁민수 by the single-syllable one, and longest-match had to pick. A name with only one possible split reports nothing (#271)

    • Add the Japanese locale pack locales.JA and the segmenter factory locales.ja_segmenter(), which together divide an unspaced Japanese name: parser_for(locales.JA, segmenter=locales.ja_segmenter()) reads 山田太郎 as family 山田, given 太郎. They are separate because no surname list can do this job, so the pack activates the stage and a third-party divider performs it. ja_segmenter() wraps namedivider-python, installed with the new nameparser[ja] extra; the core stays dependency-free. locales.available() is now ('ja', 'ru', 'tr_az', 'zh'). See Locale packs (closes #272)

    • Add Segmentation and the Segmenter type alias to the public API, plus the keyword-only Parser(segmenter=…) hook: any callable from a token’s text to a Segmentation or to None to decline. It is consulted only for scripts in Policy.segment_scripts, and only where the surname vocabulary declined first. Note what that ordering means when packs are stacked: a Japanese name opening on a listed Chinese surname never reaches the segmenter, so 高橋一郎 still splits 高 + 橋一郎. The packs are alternatives, one per corpus. See Segmenters (#272)

    • Add a construction-time UserWarning when a parser activates segmentation for scripts nothing can divide. parser_for(locales.JA) without segmenter= used to build a parser that behaved like a working one minus the feature, silently. It now names the dead scripts and the call to pass. Any configured segmenter or covering surname vocabulary silences it, so the default parser and the zh pack never warn. A from-scratch lexicon with no hangul surnames warns under the default policy, with Policy(segment_scripts=frozenset()) offered as the deactivation. See East Asian names

    • Add the Script members HIRAGANA and KATAKANA. Two members rather than one KANA because the parser treats them differently: hiragana never transcribes a foreign name, while a wholly-katakana name usually is one (#272)

    Breaking Changes

    • Change the pickle compatibility of Policy and Lexicon: the new fields change the field layout their guarded __setstate__ checks, so a pickle written by 2.0.0 raises ValueError naming the missing fields instead of loading. Re-pickle after upgrading. HumanName pickles are unaffected, since the facade pickles v1-shaped component state rather than these objects

    • Change the pickle compatibility of Parser the same way, for the new segmenter field. Re-pickle after upgrading. A Parser carrying a segmenter pickles only if that segmenter does, which a module-level function does and a closure or lambda does not (#272)

    • Change one thing about parse totality: parse() still never raises on any input, but a user-supplied Parser(segmenter=...) runs inside the parse and its own exceptions propagate rather than being absorbed. A failure there is a bug in your callable, not a fact about the name (#272)

    Behavior Changes

    • Fix names written wholly in Han or hangul parsing given-first. Native-script CJK now reads family-first through the new Policy.script_orders table, so "毛 泽东" gives family 毛 where 1.x gave family 泽东. A single unspaced token moves the same way, which is the easiest form of this to miss: "毛泽东" renders identically whichever field holds it, while the spaced form shows the change (str(HumanName("毛 泽东")) is now "泽东 毛"). No language detection is involved, and an explicit comma still wins. Default-on, and it reaches HumanName. See East Asian names (closes #271)

    • Fix unspaced Korean names not splitting. The census surname list now ships as default vocabulary with hangul segmentation on by default, so "김민준" parses family 김, given 민준 where 1.x returned the whole string as first. Rendering follows the split. Nothing but Korean is written in hangul and its surnames are a closed set, which is what makes this safe as a default rather than a pack. Default-on. See Customizing the parser for the two switches (#271)

    • Fix Japanese names carrying kana parsing given-first. The family-first rule extends to any name whose characters stay within kanji and kana while carrying at least one kana character, so "高橋 みなみ" gives family 高橋. The reasoning is the one hangul already uses: hiragana never transcribes a foreign name, and a transcription is kana alone, so kanji-plus-kana is a Japanese name in Japanese order. A name written wholly in katakana is excluded and stays positional, being predominantly a transcribed foreign name. Default-on. See East Asian names (#272)

    • Fix names containing 〆 (U+3006, the shime mark opening Japanese surnames like 〆木) parsing given-first. The script classifier now counts it as Han, so these names take the family-first reading like any other wholly-Han name (#303)

    • Fix the katakana middle dot ・ (U+30FB, and its halfwidth twin U+FF65) being read as part of a name rather than as a divider. It now separates tokens exactly as a space does, so "マイケル・ジャクソン" gives given マイケル, family ジャクソン where 1.x left the whole string in first. Native Japanese names never contain this character, so the separation is unconditional and the policy opt-outs do not cover it. Rendering returns the dot as a space (#272)

    • Fix 间隔号-divided transcriptions parsing as one unsplit token. U+00B7, the interpunct Chinese text divides a transcribed foreign name with, is now a token separator between characters of a classified script, and a name it divides keeps its source order and is never segmented. It divides only between classified-script characters, so the Catalan punt volat in Gal·la is untouched. The Japanese nakaguro is deliberately not a transcription marker, so 高橋・一郎 keeps its family-first reading (#298)

    • Fix spaced CJK postnominal honorifics parsing as name parts. 씨, 박사, 선생님, 교수님, 군, 양, 先生, 女士, 小姐, 博士, 教授, 様 and 氏 now route to suffix, so 王小明 先生 reads family 王小明 where the family-first default had made 先生 the given name. See East Asian names (closes #307)

    • Fix glued CJK honorifics parsing as part of the name. 田中さん, 김민준씨 and 王小明先生 now split the honorific off the end of the name into suffix. Previously it stayed in the name: 田中さん was entirely the family name, and 김민준씨 gave given 민준씨. The peel runs before the name is split or ordered, so 김민준씨 still divides into family 김, given 민준. Only entries that can never end a name peel; 양, 군, 氏, 博士 and 殿 are recognized in their spaced form only, since 김지양 is a given name and some ninety Japanese surnames end in 殿. Default-on, and a lone family name written with a glued honorific now divides where it did not. See East Asian names for the full set and Customizing the parser for the off-switch (closes #308)

    • Fix a comma or a 间隔号 stopping the glued-honorific peel. 김, 민준씨 now gives family 김, given 민준, suffix 씨, the same as the spaced 김 민준씨, and likewise 田中, 太郎さん. Previously each left the honorific inside the name. Both marks say where a name divides into surname and given, and an honorific is not part of the name in either reading. What a comma does instead is say which runs to look in: the two around a family comma. Anything past those is out of reach, so 김, 민준 지훈씨 peels while 김, 민준, 지훈씨 does not. Default-on. See East Asian names (closes #312)

    • Fix a glued honorific staying inside the name when the whole post-comma remainder is a credential. 田中さん, V. and 田中さん, Ph. D. now give up さん to suffix the way 田中さん, PhD already did. The peel had been taking the post-comma run for name text on the strength of the comma alone, walking into the credentials and abandoning the peel there. That run is now tested with the same rule that decides the comma structure, and declined where it is credentials, provided the part before the comma offers a peel site of its own. Where the credential itself lands is still the comma’s business and still differs by spelling. Default-on (closes #319)

    • Fix an ASCII period after a CJK honorific stopping it being recognized. 씨., 様., 氏., 님., 군., 양. and 殿. now route to suffix like their periodless spellings, where the trailing period had left them inside the name. The cause was v1’s initial regex, whose \w is Unicode-aware and matched a hangul syllable or Han ideograph as readily as a letter; a veto written for Latin was being asked of scripts it was never about. Alphabets keep their initials untouched, and "А. С. Пушкин" is unaffected. Read period strictly: only the ASCII full stop is covered, so "김민준 씨." written with the fullwidth stop still reads the honorific as the family name (superseded in 2.3.0, above). Default-on (#320)

    • Fix NFD-decomposed input missing the East Asian defaults entirely. Script classification now normalizes to NFC before deciding, so a Korean or Japanese name typed on macOS, where decomposed text is routine, gets the same order rule as its composed twin. Segmentation matching deliberately stays raw, so an unspaced NFD hangul name is ordered correctly but not split, rather than being split in the wrong place. One gotcha: parse output preserves the encoding it was given, so for NFD input name.family == "김" is False even though it is the same name. Compare NFC-normalized text when comparing across encodings. See East Asian names (#272)

    • Fix the Ukrainian conjunction й not joining the pieces around it. It is the euphonic alternate of і, chosen by the surrounding sounds rather than by meaning, so real Ukrainian data carries both spellings. "Олесь й Олена Коваленки" now gives given "Олесь й Олена" where the й previously landed in middle. Same treatment as the и/і entries added in 2.0.0: the conjunction joins only once the name has enough pieces, and a punctuated initial still wins, so "Й. Сліпий" is unaffected. Raised in a comment on #267

    • Add the Japanese maiden-name marker 旧姓 to the default vocabulary. "山田花子 旧姓 佐藤" now gives family 山田花子 and maiden 佐藤, where 1.4.0 left the marker in the name. It sits beside the Cyrillic урожд. and German geb. entries rather than in locales.JA, on the rule that admitted those: a native-script marker cannot collide with a Latin-script name, so it is safe as a default. Matching is whole-token, so the marker has to be a token, which for Japanese means a space or a configured delimiter must divide it from the name. The fullwidth colon does not, so "山田(旧姓:佐藤)" still returns maiden "旧姓:佐藤"; that one wants the head-peel #317 tracks. Default-on. See Customizing the parser (#309)

    • Fix a maiden marker inside bracketed content staying in the maiden value. Where a delimiter pair is routed to maiden by Policy(maiden_delimiters=...), a marker at the head of the clause is now dropped the way it always has been in the bare form, so "Jane Smith (née Jones)" gives maiden Jones, the same answer as the unbracketed spelling. Each clause loses its own leading marker, and only where the clause holds more than one token, so "Jane Smith (Nee) (Jones)" still gives maiden Nee Jones (Nee is a real surname). Scope before you count on it: Policy.maiden_delimiters is empty by default, so under the default policy brackets route to nickname and none of this applies. See Customizing the parser (closes #329)

    Documentation

    • Correct the documented scope of the period-abbreviation title rule. It was described as applying to “a leading word”, which was never true of any comma form: the rule runs at the front of the part that carries the given name, which after a family comma is the part after the comma, so "Morse, Det. Insp. Jane" gives title Det. Insp.. Behavior is unchanged and matches 1.4.0; only the description was wrong

  • 2.0.0 - July 27, 2026

    Two release candidates preceded this release (rc1 on 2026-07-23, rc2 on 2026-07-26). Please report anything the migration missed on issue #284. The notes below describe 2.0.0 as a whole.

    nameparser 2.0 adds a new parsing API alongside HumanName. parse() returns an immutable ParsedName whose seven fields are named for what they are (given, family) rather than where they sit in a Western name, configured by two frozen value objects – a Lexicon of vocabulary and a Policy of behavior – instead of a mutable global. HumanName keeps working: it is now a compatibility facade over the same pipeline, and it stays through 2.x. The removals below are the deprecations announced in 1.3.0 and 1.4.0 coming due; if your code runs warning-free on 1.4.0, most of them will not affect you. See Migrating from HumanName for the field-by-field map.

    The 2.0 API

    • Add parse(text), returning an immutable ParsedName with seven fields – title, given, middle, family, suffix, nickname, maiden – plus the derived views given_names, surnames, family_base and family_particles. Parsing is a pure function of the text, a Lexicon and a Policy: nothing global is consulted and nothing is mutated

    • Add Lexicon, the frozen vocabulary object. Lexicon.default() is the shipped vocabulary and Lexicon.empty() is a blank one; add(**entries), remove(**entries) and field-wise | all return new instances. Its fields are titles, given_name_titles, suffix_acronyms, suffix_words, suffix_acronyms_ambiguous, particles, particles_ambiguous, conjunctions, bound_given_names, maiden_markers and capitalization_exceptions

    • Add Policy, the frozen behavior object: name_order, patronymic_rules, middle_as_family, nickname_delimiters, maiden_delimiters, extra_suffix_delimiters, lenient_comma_suffixes, strip_emoji and strip_bidi

    • Add Parser, a reusable parser bound to a lexicon and policy (Parser(lexicon=..., policy=...).parse(text)), and parser_for(*locales, base=None), which folds locale packs onto a base parser

    • Add name_order and the constants GIVEN_FIRST, FAMILY_FIRST and FAMILY_FIRST_GIVEN_LAST, so a family-first name can be parsed as written rather than reordered by hand (the configuration half of #270; the locale packs below complete it)

    • Add PatronymicRule with the members EAST_SLAVIC and TURKIC. v1’s single patronymic_name_order flag enabled both detectors at once; Policy(patronymic_rules=...) lets you enable either one alone

    • Add PolicyPatch and the UNSET sentinel for partial policy deltas that compose – set-valued fields union, scalar fields override with later winning. This is the mechanism locale packs are built from, and UNSET is only needed when you must distinguish “not set” from a real False or None

    • Add Token, Span and Role: every field is backed by tokens carrying exact (start, end) offsets into the original string, reachable with tokens_for(Role.GIVEN). This replaces v1’s *_list attributes and makes it possible to highlight or re-slice the input the parse came from. Role is a StrEnum, so members compare and stringify as their field names (Role.GIVEN == "given"), matching AmbiguityKind; tokens_for() accepts a role’s string name too, and raises ValueError for an unknown role

    • Add STABLE_TAGS to the public API: the four documented Token.tags values (particle, conjunction, initial, joined)

    • Add Policy.patched(patch), applying a PolicyPatch directly without wrapping it in a locale pack

    • Add Parser.matches(a, b) and Parser.capitalized(name): the ParsedName methods of the same names fall back to the default configuration for str/omitted arguments, which is silently wrong for names parsed with a custom Parser

    • Add Parser.revise(name, **fields): ParsedName.replace() with the replacement text classified by the parser’s vocabulary, so particle/initial/suffix-join behavior survives the edit

    • Add Ambiguity and the AmbiguityKind enum, so a parse reports what it had to guess at instead of guessing silently. The kinds emitted today are particle-or-given, suffix-or-name, suffix-or-nickname, unbalanced-delimiter and comma-structure; order is reserved and not yet emitted. The two suffix kinds cover the post-nominals that are also ordinary words: "John Smith MA" reports that MA was read as a credential rather than a surname, and "JEFFREY (JD) BRICKEN" that the delimited JD was read as a nickname rather than a suffix. A reading the vocabulary settles on its own — "John Smith M.A.", "Andrew Perkins (MBA)" — is not a guess and reports nothing

    • Add ParsedName output and comparison methods: render(spec), initials(), capitalized() (which returns a new value rather than mutating in place), as_dict() (whose include_empty flag is keyword-only, unlike HumanName.as_dict()’s), replace(**fields), matches() and comparison_key()

    • Add Locale, the public pack type. Writing your own needs no registration – construct a Locale and pass it to parser_for()

    • Ship a fully typed public API (PEP 561): the core modules are checked under strict mypy settings, and nameparser 2.0 has no runtime dependencies

    Breaking Changes

    • Raise the minimum Python to 3.11 and drop the last runtime dependency, typing_extensions (#257). Python 3.10 is no longer supported

    • Remove HumanName.__eq__ and __hash__ (deprecated in 1.3.0, #223): instances now compare and hash by identity, so HumanName("John Smith") == "John Smith" is False where 1.x returned True. This changes result silently rather than raising – it is the one removal that can pass unnoticed into production. Use matches() to ask whether two names are the same, and comparison_key() as a dict key or sort key

    • Remove every v1 parsing hook and subclass extension point: pre_process, post_process, parse_pieces, parse_nicknames, join_on_conjunctions, the is_* predicates, cap_word/cap_piece, handle_firstnames, fix_phd and the rest. A subclass that overrides one gets a DeprecationWarning at construction naming the hooks, because the facade delegates to the core Parser and never calls them (closes #280). Customize through Lexicon/Policy instead

    • Remove regex configuration: CONSTANTS.regexes is now a read-only proxy. Reads still work, but CONSTANTS.regexes.bidi = False, item assignment and Constants(regexes=...) all raise TypeError. If you followed 1.3.1’s advice to keep bidi marks with CONSTANTS.regexes.bidi = False, that opt-out is now Policy(strip_bidi=False) on the 2.0 API; the same applies to regexes.emoji and Policy(strip_emoji=False)

    • Remove bytes input and the encoding argument (#245): passing bytes to HumanName or to a set manager raises TypeError with a decode hint, and SetManager.add_with_encoding() and DEFAULT_ENCODING are gone. Decode first, then use add()

    • Remove SetManager.__call__ (#243); iterate the manager or call set(manager). remove() of a missing member now raises KeyError like set.remove – discard() is the ignore-missing form – and the set operators |, &, - and ^ return a plain set

    • Remove HumanName slice access and item assignment (#258): name[1:-3] and name['first'] = value raise TypeError. String-key reads (name['first']) and iteration are unchanged; assign fields as plain attributes

    • Remove Constants.empty_attribute_default (#255): empty fields are always ''. Assigning it raises AttributeError; a pickle carrying the key still loads, with the value ignored

    • Remove constants=None (#261): both HumanName(..., constants=None) and hn.C = None raise TypeError. Use Constants() for library defaults or CONSTANTS.copy() for a private snapshot

    • Remove silent unknown-key access on the mapping managers (#256): CONSTANTS.capitalization_exceptions.typo and CONSTANTS.regexes.typo raise AttributeError naming the miss, instead of returning None/EMPTY_REGEX (capitalization_exceptions also lists the known keys). .get() remains available on both for intentional soft access

    • Remove support for Constants pickles written by nameparser 1.2.x or earlier (#279): loading one raises ValueError telling you to re-pickle under 1.3/1.4 first

    • Remove the dead regexes.no_vowels pattern (#268) and the Constants.suffixes_prefixes_titles cached union; neither was read by the parser

    • Change HumanName’s *_list attributes to read-only properties: they remain readable snapshots, but name.first_list = [...] now raises AttributeError

    • Change the Constants constructor to keyword-only; positional construction no longer works

    Behavior Changes

    • Recognize maiden-name markers – née, nee, geb., roz. and the Scandinavian participle forms – and route the following name to the new maiden field. 1.x folded them into middle/last, so "Jane Smith née Jones" parsed as middle="Smith née" (closes #274)

    • Recognize typographic nickname delimiters by default in both APIs (closes #273): smart quotes (“Jack”), German and Polish low-high quotes („Hansi“), Swedish right-right quotes (”Ann”), guillemets in either direction («Petit», »Hansi«), CJK corner brackets (「タロ」, 『ハナ』) and fullwidth parentheses. In 1.x these leaked into middle as literal text. Curly single quotes stay excluded, because U+2019 is the apostrophe in “O’Connor”

    • Fix the pre-comma piece being routed to first when everything after the comma is a suffix or title: "Andrews, M.D." now reads family Andrews / suffix M.D. where 1.x read given M.D. / family Andrews, and "Smith, Dr." moves Smith from first to family (the title was already correct in 1.x). The piece before a comma is definitionally the family name

    • Fix a lone recognized trailing suffix with no comma being routed to first/last: "Johnson PhD" and "Mr. Johnson PhD" now keep the suffix in suffix

    • Fix a split “Ph. D.” credential being read as two tokens; it now classifies as one suffix, replacing v1’s fix_phd hook. This now holds wherever the credential sits: 1.x healed the pair only when it trailed, so "Ph. D. John Smith" parsed as title Ph. / given D. with the real given name pushed to middle; it now reads given John, family Smith, suffix Ph. D.

    • Fix chargé d’affaires: shipped as one unmatchable TITLES entry since it was added, it is now two chainable entries (chargé, d'affaires), so the title is recognized; the unaccented spelling charge ships too, like attaché/attache

    • Remove seven multi-word SUFFIX_ACRONYMS entries (leed ap, nicet i–nicet iv, psm i, psm ii) that could never match in any release; splitting them would swallow real names (“John Leed”, “Smith, A.P.”), so they are dropped instead

    • Add a UserWarning when a multi-word entry is stored in a per-word Lexicon field or as a capitalization_exceptions key – such entries can never match

    • Parse an input with no alphanumeric character to an empty name in both APIs. 1.x kept pure punctuation as a name part, so "." gave first="." and bool() was True; 2.0 empties it, keeping bool(parse(x)) an honest “did I get a name?” test. The check is Unicode-aware, so names in any script are unaffected; only inputs that are entirely punctuation or symbols (".", "- -") change. Junk embedded in a name with real content – the stray dot in "John . Smith" – is still kept, since that parse is already truthy

    • Fold a leading never-given particle into the family name. Note that Lexicon.particles_ambiguous is the complement of v1’s non_first_name_prefixes, not a rename – it lists the particles that may double as a given name, where v1 listed the ones that may not. Copying a v1 customization across without inverting it silently reverses the behavior; see Migrating from HumanName

    • Add ma and do to suffix_acronyms_ambiguous, the set of post-nominals that are also ordinary surnames. An entry there is read as a credential when the name can spare it — written with periods ("John Smith M.A."), or when removing it still leaves a given and a family name ("John Smith MA" → suffix MA). With only two pieces to go around, the surname reading wins instead: "Jack Ma" and "Anh Do" keep their family names. As a side effect, a parenthesized or quoted "(MA)"/"(DO)" now falls through to nickname parsing rather than escaping to suffix, since inside delimiters the nickname reading is the plausible one

    • Change comparison_key() and matches() in both APIs to fold with str.casefold() where 1.4 used str.lower(), so Unicode case pairs compare equal – "STRASSE" matches "Straße", and a Greek final sigma matches its regular form. This is strictly more permissive: anything 1.4 matched still matches. Vocabulary normalization deliberately still uses lower(), for v1 parity

    • Change the 2.0 API’s default render()/str() spec to show every non-empty field: '{title} {given} "{nickname}" {middle} {family} ({maiden}) {suffix}'. The quoted nickname round-trips exactly; the parenthesized maiden re-parses as a nickname, a deliberate choice of presentation over lossless round-trip – use née {maiden} in a custom spec if you need it to survive a reparse. HumanName keeps v1’s own string_format default, unchanged

    • Change delimiter-overlap precedence in the 2.0 API: a pair listed in Policy.maiden_delimiters is dropped from the effective nickname set, so Policy(maiden_delimiters={("(", ")")}) alone routes parenthesized content to maiden. The default nickname set is exported as DEFAULT_NICKNAME_DELIMITERS. HumanName keeps v1’s nickname-wins precedence, so no existing behavior changes

    • Change suffix-delimiter rendering when a custom suffix delimiter is configured: for suffix_delimiter="/" and "John Smith, RN/CRNA", 1.x split the token and rendered suffix="RN, CRNA" where 2.0 keeps it whole as "RN/CRNA". Role assignment is unchanged; only rendering differs

    • Correct a long-standing typo in the shipped vocabulary that 1.x carried: actor and television had trailing spaces in TITLES, so "actor" in TITLES was False. Parse output is unchanged – the parser normalizes entries on ingest, which is exactly what let the typo go unnoticed – but code that tests membership against the exported constants directly will see corrected results. The data modules now assert their invariants at import time, alongside the ones prefixes.py already checked; entries must be stored lowercase and whitespace-free, so this class of typo now fails the build

    International name support

    • Add nameparser.locales with the first two packs: locales.RU (East Slavic patronymic order) and locales.TR_AZ (Turkic patronymic markers). Packs are pure data folded in at parser_for(locales.RU), they compose (parser_for(locales.RU, locales.TR_AZ) unions the rules), and they are never auto-detected – there is no reliable way to detect a name’s language from the name alone. locales.available() and locales.get(code) look packs up by code; loading is lazy, so importing the package imports no pack (completes #270; #271/#272/#146 stay staged for 2.x)

    • Add non-Latin vocabulary to the default lexicon (#269): Cyrillic, Greek, Arabic and Hebrew titles, conjunctions and name particles. Native-script entries cannot collide with Latin-script names, which is what makes them safe to enable by default. Deferred pending vetting: the Cyrillic мл/ст suffixes and the bare Greek κ title, which collides with the initial-plus-surname shape. Behavior note: محمد بن سلمان now chains بن onto the family name where 1.x read it as a middle name

    • Add Arabic-script bound given names (#269): عبد, the kunya pair أبو/ابو, and أم/ام join the following word into the given name exactly as their transliterations (abdul, abu, umm) always did, so عبد الرحمن محمد parses given عبد الرحمن, family محمد where 1.x split it into given plus middle

    • Add Arabic honorific titles and the conjunction و (#269): the doctor, professor, hajj, sheikha and engineer forms (الدكتور/الدكتورة/دكتور/دكتورة, الأستاذ/الأستاذة/أستاذ/أستاذة, الحاج/الحاجة, الشيخة, مهندس) as given-name titles, since Arabic honorifics precede the given name like الشيخ. Deferred under the collision rule: bare سيد/شيخ/أمير/سلطان (all common given names), the د. abbreviation (bare د would swallow initials), and the Ottoman post-nominals باشا/بك/أفندي (which survive as family names)

    • Add Hebrew honorifics and post-nominals, and Devanagari titles (#269): the Israeli honorifics גברת, פרופ'/פרופ׳, פרופסור, עו"ד/עו״ד and הרב as titles; ז"ל/ז״ל and שליט"א/שליט״א as suffixes, in both gershayim spellings; and Devanagari श्री, श्रीमती and डॉ. Latin sri/shri were deliberately not added, because they collide with real given names where the native script cannot. Deferred: bare רב (an ordinary word meaning “many”) and בר as a particle (Bar is a common modern given name)

    Compatibility layer

    • Reimplement HumanName as a facade over the 2.0 pipeline, and Constants as a shim resolving to a (Lexicon, Policy) snapshot with a shared parser cache. Fields, aggregates, mutation through name.C.titles.add(...), rendering defaults, capitalize(), matches(), comparison_key(), iteration, as_dict() and pickling are all preserved, and nameparser.parser and nameparser.config remain importable. The compatibility layer ships through 2.x and is removed in 3.0

    • Note that CONSTANTS.capitalize_name and force_mixed_case_capitalization are still honored through the facade, but the 2.0 API never capitalizes during parse() – call capitalized() when you want it

    • Add Role members (and their string values given/family) as valid HumanName subscript keys: hn[Role.GIVEN] returns hn.first

    • Add a UserWarning when assigning HumanName.given or .family: the facade spells those attributes first/last, so the assignment creates an inert stray attribute while the parse keeps the old value. The assignment still happens (v1-legal ad-hoc attributes keep working); the warning names the v1 spelling to use

    Command line

    • Rewrite python -m nameparser over the 2.0 API with a real argument parser. It prints the parse plus its capitalized form and initials by default, takes --json to emit the fields as JSON (python -m nameparser --json "Doe, John"), takes --locale CODE to apply a pack, and exits with a usage message on bad arguments

    Documentation

    • Rewrite the documentation new-API-first: a new front page and README, usage.rst as a tour of the 2.0 API, a principle-first customize.rst, a reference split into the 2.0 API and the compatibility layer, and new pages for How the parser works, Locale packs and Migrating from HumanName – the last carrying full attribute and configuration maps from v1 names to 2.0 names (#262). The 1.x documentation remains online as the readthedocs stable build

    Parsing changes were checked against a 652-name differential corpus – names harvested from the v1 test banks, plus names reported in the issue tracker – and everything not listed above parses identically between 1.4.0 and 2.0. The harness lives in tools/differential/ in the source repository (it is development tooling, not part of the installed package) – see its README to reproduce the comparison against your own names.

    Changed since 2.0.0rc1 (for anyone who tested the release candidate)

    • Role became a StrEnum; str(Role.GIVEN) is now "given"

    • ParsedName.tokens_for() raises ValueError for unknown roles instead of returning no tokens; it also accepts role-name strings

    • ParsedName.as_dict()’s include_empty is keyword-only

    • HumanName subscripting accepts Role members

    • Assigning HumanName.given/.family warns (the facade spells them first/last; the assignment was and remains an inert stray attribute)

    • The eight multi-word vocabulary entries that could never match were repaired (chargé d'affaires split; seven credential acronyms removed), and storing a new multi-word entry now warns; those eight are dropped silently, not warned about, when a restored Constants pickle carries all eight of them (the pre-2.0 signature)

    • PolicyPatch’s repr shows only the fields a patch sets, instead of all nine with UNSET sentinels

    • New since rc1: STABLE_TAGS, Policy.patched(), Parser.matches(), Parser.capitalized(), and Parser.revise() – see the API section above

  • 1.4.0 - July 12, 2026

    • Add Constants.copy(), a detached deep copy that preserves the source instance’s current customizations (unlike Constants(), which always starts from library defaults) – useful as CONSTANTS.copy() for a private snapshot of the shared config (#260)

    • Deprecate passing constants=None to HumanName (or assigning hn.C = None): it silently builds a fresh Constants(), discarding any customizations the caller may have expected to carry over from the shared CONSTANTS. Emits DeprecationWarning; will raise TypeError in 2.0. Use constants=Constants() for fresh library defaults or constants=CONSTANTS.copy() for a private snapshot instead (closes #260)

    • Deprecate assigning Constants.empty_attribute_default for removal in 2.0 (#255): once None support goes, the only legal value left is the default '', so a dial with one position isn’t configuration. Emits DeprecationWarning; reading the attribute is unaffected

    • Deprecate unknown-key attribute access on TupleManager/RegexTupleManager (CONSTANTS.regexes.typo, CONSTANTS.capitalization_exceptions.typo, etc.) for removal in 2.0 (#256): a misspelled or omitted key currently degrades silently (None/EMPTY_REGEX) with no traceback pointing at the typo. Emits DeprecationWarning naming the miss and the known keys; will raise AttributeError in 2.0. .get() remains available for intentional soft access

    • Deprecate HumanName slice access (name[1:-3]) and item assignment (name[‘first’] = value) for removal in 2.0 (#258): field access by position has no real use case, and item assignment duplicates plain attribute assignment. Both emit DeprecationWarning; string-key access (name['first']) is unaffected

    • Deprecate SetManager.add_with_encoding() itself for removal in 2.0 (#245), regardless of argument type: use add() instead (decoding bytes first). Previously only the bytes path warned; the str path was silent even though the whole method goes away

    • Deprecate loading a legacy-format Constants pickle (written by nameparser <= 1.2.x, before the 1.3.0 pickle fix) for removal in 2.0 (#279): __setstate__’s migration shim currently skips the stale computed-property key silently. Emits DeprecationWarning once per call telling users to re-pickle; will raise ValueError in 2.0

    • Fix the "Lastname, Firstname" comma format not being recognized when the input uses the Arabic comma ، (U+060C, the standard comma in Arabic/Persian/Urdu text) or the fullwidth CJK comma , (U+FF0C) instead of the ASCII comma: both variants now also split the format and no longer leak into the parsed output (closes #265)

  • 1.3.1 - July 11, 2026

    • Fix invisible Unicode bidirectional control characters (LRM/RLM/ALM, the embedding/override marks, and the isolates U+2066–U+2069) surviving parsing and sticking to first/last/etc., so a copy-pasted right-to-left name silently failed equality and dedup. They are now stripped in preprocessing like emoji; disable via CONSTANTS.regexes.bidi = False (closes #266)

    • Fix str() corrupting name text containing the substring “None” when empty_attribute_default is None (e.g. “Nonez Smith” rendered as “z Smith”): empty attributes are now substituted as '' before the format string is applied, instead of scrubbing the interpolated "None" from the output afterward (closes #254)

  • 1.3.0 - July 5, 2026

    Breaking Changes & Deprecations

    • Deprecate HumanName.__eq__ and __hash__ for removal in 2.0 (#223): the current design’s three promises — case-insensitive equality, equality with plain strings, and hashability — are mutually inconsistent (equal objects can hash differently), equality depends on string_format, and maiden is invisible to it. Both now emit DeprecationWarning naming the replacement; behavior is otherwise unchanged until 2.0 (closes #224)

    • Deprecate bytes input for removal in 2.0 (#245): passing bytes to HumanName/full_name or to SetManager.add()/add_with_encoding() now emits DeprecationWarning — decode first, e.g. value.decode('utf-8'). The encoding constructor argument is deprecated with it

    • Deprecate SetManager.__call__ for removal in 2.0 (#243): calling a manager returns the raw underlying set, so mutating the result bypasses normalization and cache invalidation; iterate the manager or copy with set(manager) instead

    • Add SetManager.discard(), and deprecate remove() of a missing member (#243): it currently does nothing but will raise KeyError in 2.0, matching set.remove; use discard() for intentional ignore-missing removal. Removing present members is unchanged and does not warn

    • Fix HumanName acting as its own iterator with a stored cursor: breaking out of a loop, iterating in nested loops, or calling len(name) mid-loop corrupted subsequent iteration; iter(name) now returns a fresh independent iterator each time. next(name) on the instance itself (undocumented) now raises TypeError — call next(iter(name)) instead (closes #225)

    • Remove the vestigial unparsable attribute: the guard that was meant to set it has been unreachable since 2013 (v0.2.9), so it has reported False for every parsed name for over a decade; check len(name) == 0 to detect an empty parse

    • Remove __ne__; Python 3 derives != from __eq__ automatically

    • Change internal initials helper __process_initial__ to _process_initial: double-underscore-both-sides names are reserved for Python special methods; subclasses overriding the old name must rename their override

    • Change REGEXES from a set of (name, pattern) tuples to a dict, so a duplicate name is a visible overwrite in the source instead of a nondeterministic winner at import time; code iterating REGEXES directly now gets keys instead of pairs — use .items() (#227)

    • Change CAPITALIZATION_EXCEPTIONS from a tuple of (key, value) tuples to a dict; code iterating it directly now gets keys instead of pairs — use .items() (#233)

    Behavior Changes (affect existing parse output)

    • Add bound_first_names set to Constants; bound Arabic given-name prefixes (abdul, abu, etc.) now join forward to form a single first name (e.g. "abdul salam ahmed salem" → first="abdul salam", middle="ahmed", last="salem"). Disable via CONSTANTS.bound_first_names.clear(). Default-on: changes parsing output for names with these prefixes. (#150)

    • Treat an unrecognized, multi-letter token ending in a period in the leading title run (before the first name is set), e.g. "Major.", as a title instead of a first name; internal-period abbreviations ("E.T.") and single-letter initials ("J.") are unaffected. Default-on: changes parsing of names with a leading unknown period-abbreviation (closes #109)

    • Fix parsing writing back into the Constants it reads (usually the shared module-level CONSTANTS): pieces derived while parsing a name — period-joined titles/suffixes like "Lt.Gov." and conjunction-joined pieces like "Mr. and Mrs." or "von und zu" — are now tracked per parse instead of being permanently add()-ed to the config, so parse results no longer depend on which names were parsed earlier in the process and parsing no longer mutates shared state across threads

    • Fix __hash__ to lowercase the name like __eq__ does, so equal HumanName instances hash equal and behave correctly in sets and dicts

    New comparison methods

    • Add matches() and comparison_key() for explicit name comparison: matches() compares parsed components case-insensitively (parsing str arguments first, so name.matches("Smith, John") and name.matches("John Smith") both match) and comparison_key() returns a hashable tuple of the seven components for dedup, dict keys, and sorting (#224)

    New name fields

    • Add a first-class maiden field and maiden_delimiters to Constants, so a delimiter (e.g. parenthesis) can be routed to maiden instead of nickname for alternate/maiden surnames, e.g. "Baker (Johnson), Jenny" (closes #22)

    • Add given_names (and given_names_list) attribute as aggregate of first and middle names, mirroring surnames (closes #157)

    • Add last_base, last_prefixes (and _list variants) for splitting last-name prefix particles (tussenvoegsels) from the core surname (#130, #132)

    New customization options

    • Add initials_separator to Constants and HumanName to control spacing between consecutive initials within a name group (#171)

    • Add suffix_delimiter to Constants and HumanName for parsing suffixes separated by arbitrary delimiters, e.g. "RN - CRNA" (#156)

    • Add nickname_delimiters to Constants for registering additional nickname-delimiter regex patterns at runtime, without subclassing (closes #110, #112)

    • Add suffix_acronyms_ambiguous to Constants for acronym suffixes that also read as given-name nicknames (e.g. "JD", "Ed"), used when disambiguating parenthesized/quoted content (#111)

    International name support

    • Add patronymic_name_order flag to Constants and HumanName for opt-in detection and reordering of Russian formal-order names (Surname GivenName Patronymic) (#85)

    • Add Turkic (Azerbaijani/Central-Asian) patronymic detection to patronymic_name_order, rotating the reversed 4-token formal shape (Surname GivenName PatronymicRoot Marker, e.g. oglu/qizi) into Western order (#185)

    • Add middle_name_as_last flag to Constants and HumanName for opt-in folding of middle names into the last name, for naming systems with no middle-name concept (e.g. Arabic patronymic chaining) (#133)

    • Add non_first_name_prefixes to Constants: a leading particle that is never a first name (e.g. "de Mesnil", "dos Santos") now parses as a surname with an empty first name, instead of treating the particle as the first name (closes #121)

    • Add international honorifics to TITLES (#187)

    • Add German/Austrian nobility and ecclesiastical titles to TITLES (closes #101)

    • Add German/Dutch last-name prefixes and title/degree suffixes; fix join_on_conjunctions() to register multi-word prefix chains (e.g. "von und zu") as prefixes, mirroring existing title handling (closes #18)

    Parsing fixes

    • Fix suffix boundary lookup for prefixed last names with a title before and after (e.g. "dr Vincent van Gogh dr" producing a corrupted middle name) (closes #100)

    • Fix a repeated prefix word in a prefix chain (e.g. “Juan de la de la Vega”) silently dropping the earlier occurrence in join_on_conjunctions(): value-based pieces.index(prefix) lookups re-found the wrong occurrence once the list had already been mutated by prior joins; prefix positions are now tracked positionally instead of re-derived by value (closes #208)

    • Fix a trailing suffix being silently dropped after an empty comma segment, e.g. "Doe, John,, Jr." losing the "Jr."

    • Fix degenerate comma input (a bare "," or an empty comma segment, e.g. "Doe,, Jr.", "John Doe, Jr.,,") leaving an empty-string member in first_list, last_list, or suffix_list; whitespace-only tokens assigned via the setters are dropped the same way

    • Fix suffix-shaped parenthesized/quoted content (e.g. "(Ret)", "(MBA)") being misclassified as a nickname instead of a suffix (closes #111)

    • Fix single-character symbol conjunctions (e.g. "&", "/") being ignored in short names (#173)

    • Fix recognition of single-letter roman numeral suffixes (e.g. "I", "V") in suffix-comma format (closes #136)

    • Fix recognition of trailing suffix_not_acronyms (e.g. "Jr.") in lastname-comma format (closes #144)

    • Fix missing comma between 'msc' and 'mscmsm' in suffix_acronyms, which silently concatenated them into a bogus 'mscmscmsm' entry (#111)

    • Fix 'apn aprn' split into separate suffix_acronyms entries so each is recognized independently (closes #155)

    Formatting and output fixes

    • Fix IndexError in initials()/initials_list() when a *_list attribute was assigned directly with an element containing unnormalized whitespace (e.g. name.middle_list = ['Q R']), bypassing the parser’s whitespace normalization (closes #232)

    • Fix initials() emitting a stray empty initial (e.g. “J. . V.”) – or raising TypeError when empty_attribute_default is None – for name parts with no initialable words, e.g. a prefix-only middle name like "de la"

    • Fix capitalization of suffix acronyms written with dots, e.g. "M.D." (closes #141)

    • Fix extra whitespace before punctuation in str() output when a string_format field is empty (closes #139)

    • Fix spurious leading space in surnames and empty token in suffix list after capitalize() with an empty middle or suffix (#164)

    API correctness and cleanup

    • Fix the five non-cached-union SetManager-backed Constants attributes (first_name_titles, conjunctions, bound_first_names, non_first_name_prefixes, suffix_acronyms_ambiguous) accepting non-SetManager assignment silently (e.g. constants.conjunctions = 'and'), degrading membership checks into substring tests with no error; assignment now raises TypeError like the four cached-union attributes already did (closes #241)

    • Fix HumanName.C accepting an invalid constants value on post-construction assignment (e.g. hn.C = 'garbage'), bypassing the constructor’s validation and failing later with an unrelated AttributeError; C is now a property that validates on assignment too (closes #239)

    • Fix TupleManager (and RegexTupleManager) accepting a bare string/bytes argument (raising a cryptic dict-internals ValueError) or an iterable of 2-character strings (silently shredding each into a key/value pair, e.g. Constants(capitalization_exceptions=['ii']) becoming {'i': 'i'}); both now raise TypeError with a clear message (closes #242)

    • Fix SetManager.__contains__ being the one operation that didn’t normalize (lowercase, strip leading/trailing periods) its operand, so e.g. 'Dr.' in constants.titles could return False even though the title was correctly configured; membership checks now normalize like add()/remove()/the constructor/the set operators (closes #244)

    • Fix a bare string passed to a set-backed Constants argument (e.g. Constants(titles='dr')), to SetManager, or as a SetManager set-operator operand (e.g. constants.titles |= 'esq') being silently split into single characters, replacing or polluting the set and producing wrong parses with no error; it now raises TypeError with the suggested fix — wrap strings in a list, decode bytes first (closes #238)

    • Fix SetManager set operators and the constructor skipping the lowercase/strip-edge-periods normalization that add() applies: constants.titles |= ['Esq.'] kept a raw 'Esq.' the parser’s lookups could never match, titles & ['Dr.'] missed 'dr', and Constants(titles=[...]) stored raw elements that silently never matched; elements and operands are now normalized everywhere, and non-str elements (bytes, None, numbers) raise TypeError instead of crashing cryptically or being coerced

    • Fix the constants constructor argument silently discarding Constants subclass instances: the exact-type check replaced them with fresh defaults, throwing away the caller’s configuration. Subclass instances are now used as given; anything that is neither None nor a Constants instance now raises TypeError instead of being silently swapped for defaults (closes #226)

    • Fix Constants customizations, singleton identity, and TupleManager subclass being lost across pickle/deepcopy round-trips (#167, #168, #169)

    • Fix is_rootname() returning stale results after add()/remove() on titles, prefixes, suffix_acronyms, or suffix_not_acronyms (#166)

    • Fix the library logger calling setLevel(logging.ERROR) on import, which silently discarded log records regardless of an application’s own logging configuration; the logger now leaves its level at NOTSET and lets the application control verbosity (closes #228)

    • Minor internal cleanups: drop a dead length check in the initials helper, simplify double-wrapped len(list(...)) calls, and other small parser tidy-ups with no behavior change (closes #229)

    • Change Constants.__repr__ to report collection sizes and non-default scalar config, replacing the uninformative <Constants() instance> (#221)

  • 1.2.1 - June 19, 2026
    • Fix initials() interpolating the literal None for empty name parts when empty_attribute_default = None (e.g. "J. None D."); empty parts now render as an empty string and a fully-empty result returns empty_attribute_default

    • Add python -m nameparser "Name String" command-line helper that prints a parsed name

    • Reorganize the test suite from a single tests.py into a tests/ pytest package

  • 1.2.0 - June 11, 2026
    • Drop Python 2 and Python < 3.10 support; Python 3.10–3.14 now required

    • Add type hints and type declarations (PEP 561 py.typed marker)

    • Migrate build tooling to pyproject.toml, drop setup.py

    • Remove dead Python 2 compatibility shims (ENCODING constant, next() aliases)

    • Modernize CI: uv-based workflow, trusted publishing to PyPI, Dependabot

  • 1.1.3 - September 20, 2023
    • Fix case when we have two same prefixes in the name ()#147)

  • 1.1.2 - November 13, 2022
    • Add support for attributes in constructor (#140)

    • Make HumanName instances hashable (#138)

    • Update repr for names with single quotes (#137)

  • 1.1.1 - January 28, 2022
    • Fix bug in is_suffix handling of lists (#129)

  • 1.1.0 - January 3, 2022
    • Add initials support (#128)

    • Add more titles and prefixes (#120, #127, #128, #119)

  • 1.0.6 - February 8, 2020
    • Fix Python 3.8 syntax error (#104)

  • 1.0.5 - Dec 12, 2019
    • Fix suffix parsing bug in comma parts (#98)

    • Fix deprecation warning on Python 3.7 (#94)

    • Improved capitalization support of mixed case names (#90)

    • Remove “elder” from titles (#96)

    • Add post-nominal list from Wikipedia to suffixes (#93)

  • 1.0.4 - June 26, 2019
    • Better nickname handling of multiple single quotes (#86)

    • full_name attribute now returns formatted string output instead of original string (#87)

  • 1.0.3 - April 18, 2019
    • fix sys.stdin usage when stdin doesn’t exist (#82)

    • support for escaping log entry arguments (#84)

  • 1.0.2 - Oct 26, 2018
    • Fix handling of only nickname and last name (#78)

  • 1.0.1 - August 30, 2018
    • Fix overzealous regex for “Ph. D.” (#43)

    • Add surnames attribute as aggregate of middle and last names

  • 1.0.0 - August 30, 2018
    • Fix support for nicknames in single quotes (#74)

    • Change prefix handling to support prefixes on first names (#60)

    • Fix prefix capitalization when not part of lastname (#70)

    • Handle erroneous space in “Ph. D.” (#43)

  • 0.5.8 - August 19, 2018
    • Add “Junior” to suffixes (#76)

    • Add “dra” and “srta” to titles (#77)

  • 0.5.7 - June 16, 2018
    • Fix doc link (#73)

    • Fix handling of “do” and “dos” Portuguese prefixes (#71, #72)

  • 0.5.6 - January 15, 2018
    • Fix python version check (#64)

  • 0.5.5 - January 10, 2018
    • Support J.D. as suffix and Wm. as title

  • 0.5.4 - December 10, 2017
    • Add Dr to suffixes (#62)

    • Add the full set of Italian derivatives from “di” (#59)

    • Add parameter to specify the encoding of strings added to constants, use ‘UTF-8’ as fallback (#67)

    • Fix handling of names composed entirely of conjunctions (#66)

  • 0.5.3 - June 27, 2017
    • Remove emojis from initial string by default with option to include emojis (#58)

  • 0.5.2 - March 19, 2017
    • Added names scrapped from VIAF data, thanks daryanypl (#57)

  • 0.5.1 - August 12, 2016
    • Fix error for names that end with conjunction (#54)

  • 0.5.0 - August 4, 2016
    • Refactor join_on_conjunctions(), fix #53

  • 0.4.1 - July 25, 2016
    • Remove “bishop” from titles because it also could be a first name

    • Fix handling of lastname prefixes with periods, e.g. “Jane St. John” (#50)

  • 0.4.0 - June 2, 2016
    • Remove “CONSTANTS.suffixes”, replaced by “suffix_acronyms” and “suffix_not_acronyms” (#49)

    • Add “du” to prefixes

    • Add “sheikh” variations to titles

    • Add parameter to force capitalization of mixed case strings

  • 0.3.16 - March 24, 2016
    • Clarify LGPL licence version (#47)

    • Skip pickle tests if pickle not installed (#48)

  • 0.3.15 - March 21, 2016
    • Fix string format when empty_attribute_default = None (#45)

    • Include tests in release source tarball (#46)

  • 0.3.14 - March 18, 2016
    • Add CONSTANTS.empty_attribute_default to customize value returned for empty attributes (#44)

  • 0.3.13 - March 14, 2016
    • Improve string format handling (#41)

  • 0.3.12 - March 13, 2016
    • Fix first name clash with suffixes (#42)

    • Fix encoding of constants added via the python shell

    • Add “MSC” to suffixes, fix #41

  • 0.3.11 - October 17, 2015
    • Fix bug capitalization exceptions (#39)

  • 0.3.10 - September 19, 2015
    • Fix encoding of byte strings on python 2.x (#37)

  • 0.3.9 - September 5, 2015
    • Separate suffixes that are acronyms to handle periods differently, fixes #29, #21

    • Don’t find titles after first name is filled, fixes (#27)

    • Add “chair” titles (#37)

  • 0.3.8 - September 2, 2015
    • Use regex to check for roman numerals at end of name (#36)

    • Add DVM to suffixes

  • 0.3.7 - August 30, 2015
    • Speed improvement, 3x faster

    • Make HumanName instances pickleable

  • 0.3.6 - August 6, 2015
    • Fix strings that start with conjunctions (#20)

    • handle assigning lists of names to a name attribute

    • support dictionary-like assignment of name attributes

  • 0.3.5 - August 4, 2015
    • Fix handling of string encoding in python 2.x (#34)

    • Add support for dictionary key access, e.g. name[‘first’]

    • add ‘santa’ to prefixes, add ‘cpa’, ‘csm’, ‘phr’, ‘pmp’ to suffixes (#35)

    • Fix prefixes before multi-part last names (#23)

    • Fix capitalization bug (#30)

  • 0.3.4 - March 1, 2015
    • Fix #24, handle first name also a prefix

    • Fix #26, last name comma format when lastname is also a title

  • 0.3.3 - Aug 4, 2014
    • Allow suffixes to be chained (#8)

    • Handle trailing suffix in last name comma format (#3). Removes support for titles with periods but no spaces in them, e.g. “Lt.Gen.”. (#21)

  • 0.3.2 - July 16, 2014
    • Retain original string in “original” attribute.

    • Collapse white space when using custom string format.

    • Fix #19, single comma name format may have trailing suffix

  • 0.3.1 - July 5, 2014
    • Fix Pypi package, include new config module.

  • 0.3.0 - July 4, 2014
    • Refactor configuration to simplify modifications to constants (backwards incompatible)

    • use unicode_literals to simplify Python 2 & 3 support.

    • Generate documentation using sphinx and host on readthedocs.

  • 0.2.10 - May 6, 2014
    • If name is only a title and one part, assume it’s a last name instead of a first name, with exceptions for some titles like ‘Sir’. (#7).

    • Add some judicial and other common titles. (#9)

  • 0.2.9 - Apr 1, 2014
    • Add a new nickname attribute containing anything in parenthesis or double quotes (Issue 33).

  • 0.2.8 - Oct 25, 2013
    • Add support for Python 3.3+. Thanks to @corbinbs.

  • 0.2.7 - Feb 13, 2013
    • Fix bug with multiple conjunctions in title

    • add legal and crown titles

  • 0.2.6 - Feb 12, 2013
    • Fix python 2.6 import error on logging.NullHandler

  • 0.2.5 - Feb 11, 2013
    • Set logging handler to NullHandler

    • Remove ‘ben’ from PREFIXES because it’s more common as a name than a prefix.

    • Deprecate BlankHumanNameError. Do not raise exceptions if full_name is empty string.

  • 0.2.4 - Feb 10, 2013
    • Adjust logging, don’t set basicConfig. Fix Issue 10 and Issue 26.

    • Fix handling of single lower case initials that are also conjunctions, e.g. “john e smith”. Re Issue 11.

    • Fix handling of initials with no space separation, e.g. “E.T. Jones”. Fix #11.

    • Do not remove period from first name, when present.

    • Remove ‘e’ from PREFIXES because it is handled as a conjunction.

    • Python 2.7+ required to run the tests. Mark known failures.

    • tests/test.py can now take an optional name argument that will return repr() for that name.

  • 0.2.3 - Fix overzealous “Mac” regex

  • 0.2.2 - Fix parsing error

  • 0.2.0
    • Significant refactor of parsing logic. Handle conjunctions and prefixes before parsing into attribute buckets.

    • Support attribute overriding by assignment.

    • Support multiple titles.

    • Lowercase titles constants to fix bug with comparison.

    • Move documentation to README.rst, add release log.

  • 0.1.4 - Use set() in constants for improved speed. setuptools compatibility - sketerpot

  • 0.1.3 - Add capitalization feature - twotwo

  • 0.1.2 - Add slice support