Release Log¶
2.4.0 - Unreleased
nameparser 2.4 is under development.
Behavior Changes
Fix a particle surname before a comma being split when a credential follows it.
HumanName("van der Berg, PhD")gives lastvan der Berg, suffixPhD, where 1.4.0 through 2.3.0 gave firstvan, lastder Berg. Where the parser counts the words before a comma, a particle and the word it attaches to now count as one, so a particle surname reads asBerg, PhDdoes, andAbu Bakar, PhDgives lastAbu Bakarwhere 1.4.0 through 2.3.0 gave firstAbu. The same count keeps a particle surname whole in front of a credential that is also a name:De La Cruz, Edgives firstEd, lastDe La Cruz, as 2.3.0 read it, andvan der Berg, MAgives lastvan der Berg, suffixMA, the readingSmith, MAgets below. Where the part is the surname alone, the comma also decides that a leading particle is not a first name:Van Johnson, Dr.gives lastVan Johnson, where 2.2 and 2.3 gave firstVan. A given name in front still makes two words, soJohn van Buren, Edreads suffixEdasJohn Smith, Eddoes. A connective surname is not counted as one:Ortega y Gasset, PhDstill gives firstOrtega, middley, asOrtega y Gassetreads on its own. See theC1entry ofdocs/design/decisions.md(closes #575)Fix a surname particle that is also a credential (vd, mc) being split away from the surname or the credentials around it.
HumanName("SMITH VD MA, JOHN")gives lastSMITH VD MA, where 2.2 and 2.3 gave lastSMITH MA, suffixVD, andHumanName("Doe, Jane PhD vd MA")gives suffixPhD vd MA, where 2.3 gave middleMA, suffixPhD vd. Avdwith nothing behind it still joins the last name:HumanName("Doe, Jane PhD vd")gives lastvd Doe. (closes #573)Fix a one-letter connective joining a name that gives no sign it is a connective.
HumanName("jose e maria santos")gives firstjose, middlee maria, lastsantos, where 1.4.0 through 2.3.0 gave firstjose e maria; andJUAN GARCIA Y LOPEZgives lastGARCIA Y LOPEZ, where every release since 1.4.0 read the bare capital as an initial and gave middleGARCIA Y. A single letter is an initial where the writing says so – a bare Latin capital in a name that is not written wholly in one case – and a name written wholly in one case says nothing either way, so the reading comes from the vocabulary there:ereads as an initial andyjoins. Mixed-case input is untouched in both directions:Jose e Maria Santosstill gives firstJose e MariaandJose E Maria Santosstill gives middleE Maria. Short names move in the derived views rather than the fields, P3’s three-word carve-out being unchanged:parse("john e smith").initials()isj. e. s.where 2.3.0 gavej. s., andHumanName("john e smith").capitalize()givesJohn E Smithwhere 2.3.0 gaveJohn e Smith;JUAN Y GARCIAmoves itscapitalize()the same way in reverse, givingJuan y Garcia, while its initials do not move at all:parse(...).initials()isJ. Y. G., what 2.3.0 gave and what 1.4.0’s own view gave, theYholding its part alone and contributing an initial again under the connective-initials fix further down this list (closes #461).HumanName.initials()agrees with the core on both – see the #528 bullet below, which closed a split this change opened and the same release closes. Seventeen names in the differential corpora are written in one case and carry a cased single-letter connective, and ten of them move something against 2.3.0. The Cyrillic reading is unchanged (Хосе И Мария Сантосstill gives firstХосе И Мария), and Arabicوnever enters the rule, having no case to be written against. ALexiconknob decides which letters are marked, so the reading is configurable rather than fixed. See theP3entry ofdocs/design/decisions.md(closes #383, closes #479)Fix HumanName.initials() reading a one-letter connective by vocabulary and written shape instead of by the parse.
HumanName("john e smith").initials()givesj. e. s., where every release from 1.4.0 through 2.3.0 gavej. s.;JUAN GARCIA Y LOPEZgivesJ. G. L.where 2.3.0 gaveJ. G. Y. L.;JUAN Y GARCIAdoes not move at all this cycle, givingJ. Y. G.on both surfaces as 2.3.0 and 1.4.0 did, since the connective-initials fix further down this list (closes #461) gives itsYan initial again. That first name read 1.4.0’s way at 2.3.0 and only there: 2.0.0 through 2.2.0 already gave today’s answer, by the unrelated bug the 2.3.0 note below records as fixed (the facade dropping a bare capital that is also a one-letter conjunction, #462), so against those three releases it does not move at all. The v1 facade decided whether a word was the connective by looking the word up and checking its shape, whileparse(...).initials()read the tag the parse recorded – so the change above, which reads a single letter in a one-case name from the vocabulary rather than from its case, moved one view and not the other. Both views of a parse now give the same answer. Mixed-case names are untouched on both, the writing having decided the letter:John E Smithis stillJ. E. S.andScott E. WernerstillS. E. W.. So is a one-case name whose letter is outside the marked set –maria y lopezis stillm. l.,yhaving joined before this release and after it. Two costs, and both match whatcapitalize()has always done: editingC.conjunctionsafter a name is parsed no longer changes its initials untilfull_nameis assigned again, and a name restored from a pickle, copied withcopy.copy/copy.deepcopy(the same state hooks), or built from keyword fields (HumanName(first=..., middle=..., last=...)) carries no tags, so its initials come from the vocabulary and can differ from a fresh parse of the same string. One private break, stated because a v1 subclass can hit it: an override of_process_initialwritten to v1’s(name_part, firstname=False)signature now raisesTypeErrorthe first timeinitials()runs, sinceinitials()passes the part’s tokens. Such an override has to accept atokenskeyword and pass it on –return super()._process_initial(name_part, firstname, tokens=tokens)– to receive this fix. Widening the signature without forwarding still works, but on the pre-#528 STRING path: the token call hands the override the group’s own text asname_partrather than an empty placeholder, sojohn e smithinitialsj. s.under such an override, not thej. e. s.above. A subclass overriding one of the publicfirst_list,middle_listorlast_listproperties keeps working too: that member takes the pre-2.4 vocabulary reading instead of the change above, while an un-overridden member still moves. See theR3entry ofdocs/design/decisions.md(closes #528)Fix a credential acronym that is also a surname being read by position alone.
HumanName("Jack MA")gives suffixMAwhere 2.0 through 2.3 gave lastMA, andJohn Smith Magives lastMawhere they gave suffixMa. In a name written in more than one case, an ambiguous acronym written in capitals is written the way a credential is written and is read as one even where removing it leaves no surname; one written in any other cased form that is not wholly lower is written the way a surname is written and stays one even where there are words to spare (John Smith maandJohn Smith ed– all lower, no contrast – give suffixma/edinstead). A name written wholly in one case says nothing either way and keeps the reading it had:JOHN SMITH MAis still a credential,ANH DOstill a surname,jack mastill a surname. The same reading reaches the comma forms, where the words-to-spare count is now a count of NAME words:Smith, MAgives lastSmith, suffixMA;Smith Jr., MAkeeps lastSmith; andJohn Smith, MA,John Smith, Ed,john smith, maandJOHN SMITH, MAall give a suffix again, which is what 1.4.0 read and 2.0 through 2.3 did not.Jack MaandAnh Doare unchanged. The LEAN is inert on a caseless script, but the comma count above is not – it asks name-word count, not case – so마틴 킹, MAand田中 太郎, MAalso give a suffix again (1.4.0 parity on the suffix, two pre-comma name words each) while the single-token毛泽东, MAdoes not move, having no case to write a contrast in either way. See theS2entry ofdocs/design/decisions.md(closes #289)Fix a bare trailing Meng or Lac being read as a credential and losing the family name: meng and lac are now acronyms that are also ordinary names.
HumanName("wang meng")gives firstwang, lastmeng, andparse()reports a suffix-or-name ambiguity, where every release from 2.0.0 through 2.3.0 gave suffixmengand no last name; 1.4.0 read lastmeng, so this is 1.4.0’s answer plus the flag.li mengandtran lacmove the same way,Wang, Menggives firstMeng, lastWangagain, andParser(policy=Policy(name_order=FAMILY_FIRST)).parse("Wang Meng")gives givenMengwhere 2.0.0 through 2.3.0 gave familyWang, suffixMengand no given name. With a full name in front the credential reading stays:john smith mengandnguyen van lackeep suffixmengandlac, now flagged. But a Title-caseNguyen Van Lacgives lastVan Lacwhere every release gave lastVan, suffixLac. The cost is the marking’s own, and it falls on the conventional spellings: aMEngorLAcwritten that way, in a name written in more than one case, is read the wayJohn Smith Mais (above), soJohn Smith MEnggives middleSmith, lastMEng, where every release gave suffixMEng. Where one of them LEADS a credential run nothing speaks for it:John Smith MEng PhDgives middleSmith, lastMEng, suffixPhD, where every release read suffixMEng PhD(MEng, PhDthrough 2.2); behind a credential, or written with its periods, it is read with the run (the next entry). After a commaSmith, MEngandSmith, menggive firstMEngandmeng(1.4.0’s reading, not the suffix 2.0 through 2.3 gave) andSmith, John MEnggives middleMEng(every release gave suffixMEng), and a bracketedJohn Smith (MEng)falls through to nickname, where every release gave suffixMEng. A lone credential after a comma behind a full name (John Smith, MEng) keeps the credential reading, and so does a run of them (the next entry). Meng is a common Chinese surname and given name, Lac a Vietnamese given name (Nguyen Van Lac) and a French surname; see thesuffix-acronym-collisionsentry ofdocs/design/decisions.md(closes #540)Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.
HumanName("John Smith, Ed Ma")gives firstJohn, lastSmith, suffixEd Ma, where 2.0 through 2.3 gave firstEd, middleMa, lastJohn Smith– 1.4.0’s reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (John Smith, MA), andparse()reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report when another credential opens it, the capitals having decided it (John Smith, PhD MA,John Smith, MS MA), while one opened by such an acronym reports on that first word, asJohn Smith, MAalone does (John Smith, MA PhD). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, sojohn smith, v phd magives firstjohn, lastsmith, suffixv phd ma, where 2.3 gave firstv, middlema, lastjohn smith.john smith, md maandJohn Smith, Ms Mamove the same way, where 2.0 through 2.3 gave titlemd/Ms– the second is the accepted cost,Msread as the suffix word it also is, asJohn Smith, Msalone already reads it – while one name word before the comma keeps the listing form (Smith, Ms Magives titleMs, firstMa). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential’s company whatever its case:John Smith PhD MEngandDoe, Jane PhD MEnggive suffixPhD MEng, the fields every release gave (PhD, MEngthrough 2.2), now reported, anddoe, jane v phd dogives suffixv phd dowhere 2.3.0 gave lastdo doe– a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks:Wang Ma PhDkeeps lastMa. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported:Smith, PhD Magives lastSmith, suffixPhD Ma, where 2.3 gave firstPhD, middleMa. A title that is also a credential (MD,Ms) opening that part stays a title and nothing in the part speaks, soSmith, MD PhD Makeeps titleMD, firstPhD, middleMaandSmith, Ms MD MatitleMs MD, firstMa, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, soWang M.Eng.gives suffixM.Eng., asWang M.A.does and as 2.0 through 2.3 did. See the #544 entry underS2indocs/design/decisions.md(closes #544)Fix a credential run after a comma losing the name in front of it when two surname particles stand side by side in the run.
HumanName("John Smith, PhD DO DO")gives firstJohn, lastSmith, suffixPhD DO DO, where 2.2 and 2.3 gave firstPhD, lastDO DO John Smithand 2.0 and 2.1 gave titlePhD, firstDO DO– 1.4.0’s reading, restored.DO,MCandVDare credentials and surname particles at once, and two particles next to each other had joined into one particle run that took the credentials apart; the part after the comma is now read as credentials before any particle can join it.John Smith, PhD vd DOgives suffixPhD vd DOthe same way, andJohn Smith, MD DO DOgives suffixMD DO DOwhere 2.3 gave titleMD, firstDO, middleDO. A singleDOis still left to its capitals (John Smith, PhD DOgives suffixPhD DO), and with one name word before the comma a part of credentials is all suffixes (Smith, PhD DO DOgives lastSmith, suffixPhD DO DO, andSmith, MA DO DOlastSmith, suffixMA DO DO, where 2.3 gave firstPhDand firstMA, lastDO DO Smith). See the #562 entry underC1indocs/design/decisions.md(closes #562)New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.
HumanName("John Smith X.Y.Z.")gives suffixX.Y.Z.where every release gave lastX.Y.Z., whileJack X.Y.Z.keeps its surname, the same words-to-spare rule a listed acronym takes – and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person’s initials are written, and two words before a comma may be one surname, soGarcía Márquez, G.J.keeps firstG.J.and lastGarcía Márquezand reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (John Smith, PhD X.Y.gives suffixPhD X.Y., whileGarcía Márquez, Ms G.J.keeps titleMs, firstG.J.). Three letters or more read by the count, soJohn Smith, X.Y.Z.gives suffixX.Y.Z.– and so doesGarcía Márquez, G.J.R., the accepted cost of the line, since initials are conventionally written apart (García Márquez, G. J. R.), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, sojohn smith x.y.z.reads the same way. Words the vocabulary does know are untouched (M.A.,Ph.D.,A.B.C.), a single trailing period is still not this shape (John Smith Xyz.keeps lastXyz.), and a dotted run at the FRONT of a name is untouched (J.R.R. Tolkien). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS – the roman numerals the suffix list holds, and the lone digit2– was reading as a generational suffix, soJack X.Y.I.gives lastX.Y.I.again, as 1.4.0 read it, whileMsc.Ed.,JD.CPAandLt.Gov.are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY:John Smith 1.4.2gives last1.4.2where 2.3 gave suffix1.4.2, andJohn Smith, 1.4.2gives first1.4.2, lastJohn Smith. Such a token reports nothing at any policy – it is no acronym either, the shape reading wanting every chunk alphabetic – and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way – setting it toFalsereads an unlisted dotted word as name material by position instead (John Smith X.Y.Z.keeps lastX.Y.Z.), the pre-2.4 reading for THAT half alone. See theS2andsuffix-acronym-collisionsentries ofdocs/design/decisions.md(closes #516)New Policy field unlisted_caps_suffixes: an unlisted all-caps word reads as a credential after a comma by default, and elsewhere on request. Its value is a
CapsSuffixes. The default,CapsSuffixes.AFTER_COMMA, reads such a word in the part right after a comma behind two or more name words, alone or in a run with other credentials:HumanName("John Smith, XYZ")gives firstJohn, lastSmith, suffixXYZ, where 1.4.0 through 2.3.0 gave firstXYZ, lastJohn Smith;John Smith, LEED APandJohn Smith, PhD XYZgive suffixLEED APandPhD XYZthe same way, andJohn Smith, RAIgives suffixRAIagain, as it did before 2.3. The all-caps surname convention writes the capitals at the end of a name or before a comma (Jean DUPONT,DUPONT, Jean) and never there. A word after a one-word surname stays the given name (Smith, XYZ), a two-letter word reads exactly as dotted initials do (García Márquez, MJandGarcía Márquez, MJ PhDkeep firstMJ), and the name has to contrast the capitals with a word of its own holding a capital whose last letter is lowercase (Smith,DiCaprio); a name typed with decomposed accents reads as its composed spelling. A surname written in capitals ends in a capital whatever is glued in front of it, so such a record keeps its given name beside its lowercase particles, titles and maiden clauses and beside a mixed-case credential (GISCARD d'ESTAING, VALÉRY,LLOYD FitzGERALD, RONALD,LLOYD WEBBER, ANDREW PhD), as does a name written wholly in lowercase.CapsSuffixes.EVERYWHEREalso reads the end of a name, the given part’s last word after a family comma and the word ending a maiden marker’s clause:.parse("John Smith XYZ")gives suffixXYZ, andJean Pierre DUPONTgives lastPierre, suffixDUPONT– why it is not the default.CapsSuffixes.OFFreads none of them and reports nothing; it is the way to keep a given name written in capitals after a two-word surname, which the default reads as a credential (García Márquez, GABRIELgives suffixGABRIEL). The field reaches the core parser only, throughParser(policy=Policy(unlisted_caps_suffixes=...)); aHumanNametracks the parser’s defaults, so the comma reading reaches it and the other two settings cannot be chosen from there. Neither this field norunlisted_dotted_suffixeshas a v1Constantsmanager. See theS2andC1entries ofdocs/design/decisions.md(closes #516, closes #564)The comma’s own decision about an ambiguous credential is now reported.
parse("Smith, MA").ambiguitiesnamessuffix-or-name, and so does every other decision at the ambiguous credential class – before or after a comma, in either direction, with no newAmbiguityKind(the family-comma attachment fork already reported this way, e.g.parse("Berg, Jan vd")). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence:John Smith, X.Y.Z.andJohn Smith, PhD X.Y.report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports:John Smith, X.Y. P.Q.gives lastSmith, suffixX.Y. P.Q.(#563). One report per decision:Smith, Mareports that the word was kept as the given name just asSmith, MAreports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did:John van der Berg Magives lastvan der Berg Maand namessuffix-or-name, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized:Steven Hardman, MD, DO, DDSno longer reportscomma-structure, on its written case. That is the whole of the losses over the differential corpora –John Smith, MD, R.A.I.is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release’s own development. The other movement an upgrader sees is a SWAP rather than a loss:Jack X.Y.I.reportedgiven-or-familyat 2.3.0 and reportssuffix-or-namehere, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker’s clause. See theS2andC1entries ofdocs/design/decisions.mdFix a credential ending the given part of a family-comma listing being read as a middle name in silence.
HumanName("Doe, John MA")gives firstJohn, lastDoe, suffixMA, where 2.0 through 2.3 gave middleMA– and 1.4.0 gave the suffix, so this restores v1’s reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone:Doe, John Makeeps middleMa, written the way a name is written, andDoe, John Edkeeps middleEd. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields – a declined name re-rendered without its comma,John Ma Doe, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent –Doe, John MA Smithgives middleMA Smithand reports nothing – while a credential run or a trailing title is transparent to it:Doe, John MA PhDgives suffixMA PhDandDoe, John MA Prof.gives titleProf.with suffixMA. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further:Doe, John Prof. MAnow gives titleProf.where it gave middleProf. MA, the trailing-title chain reaching a word the credential used to hide; andDoe, John van MAgives lastvan Doewith suffixMAwhere it gave middlevan MA, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read:DOE, MARY JO MA,doe, john ma,田中, 太郎 MAand김, 민준 MAall give a suffix. The unlisted dotted spelling moves with them without the parity claim –Doe, John X.Y.Z.gives suffixX.Y.Z.where 1.4.0 and 2.3.0 both gave a middle name – to match the comma-lessJohn Doe X.Y.Z.. One word is carved out:dois the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry).Doe, John DOgives suffixDO, whileDoe, John do,Doe, John Do,DOE, JOHN DOanddoe, john doare unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, soSMITH, JOHN DOkeeps lastDO SMITHasNASCIMENTO, EDSON ARANTES DOdoes – right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See theS2andP6entries ofdocs/design/decisions.md(closes #531)Change where a maiden marker takes a maiden name: only behind a surname, and only up to the credentials and titles the name ends with.
HumanName("Jane Doe nee Smith"),Mai Le née NguyenandDoe nee Smith, Janegive maidenSmithandNguyenas before, but a marker counts only in the name before any comma or in the surname part before a family comma, and only behind a name word. Anywhere else it is an ordinary word, as in 1.4.0:Doe, Jane nee Smithgives middlenee Smith,Smith, John, PhD née Jonesgives suffixPhD née Jones, andDr. nee Smithgives firstnee, lastSmith, where 2.0 through 2.3 gave maidenSmith,JonesandSmith– the last with no name at all. A credential in front of the marker ends the name (see the credential-run entry below), soJane Doe PhD nee Smithgives suffixPhD nee Smithwhere 2.0 through 2.3 gave suffixPhD, maidenSmith. A particle or a lone word is a surname here:Jane van nee Smithgives lastvan, maidenSmith, andJane Smith née Vgives lastSmith, maidenV, as 2.0 and 2.1 read it, where 2.2 and 2.3 gave lastnée, suffixV. The words the marker takes now end where the run of post-nominals and titles the name would end with if the clause were not written begins, and the words it gives up read as they would there:Jane Doe nee Smith MAgives maidenSmith, suffixMA, where 2.0 through 2.3 gave maidenSmith MAand said nothing, as do the one-caseJANE DOE NEE SMITH MAandjane doe nee smith ma;Jane Doe nee Smith MA PhDgives suffixMA PhDwhere 2.3.0 gave maidenSmith MA, suffixPhD, so the two orders of the credentials now agree;Jane Doe nee Smith DO DOgives suffixDO DO;Jane Doe nee Smith Prof.,Mary Smith née Jones Prof.andJane van der Berg nee Smith Prof.give titleProf.;Jane Doe nee Smith King.gives titleKing.;Jane Doe nee Smith MA Prof.andJane Doe nee Smith Prof. MAboth give titleProf., suffixMA; andJane Doe nee Smith V Prof.gives suffixV– each where 2.3.0 kept every word in the maiden name. A credential the clause gives up or keeps at that boundary is reported assuffix-or-name. The writing still decides, as it does at the end of a name with no clause:Jane Doe nee Smith Makeeps maidenSmith Ma, andJane Doe nee Yo-Yo Makeeps a two-word birth surname whole. The first word after the marker is always taken, the marker having announced a name:Jane Doe nee MAandJane Doe nee King.keep it, andJane Doe nee Prof. Dr.gives maidenProf., titleDr., where 2.3.0 gave maidenProf. Dr.. Before a family comma the clause keeps every word, as before:Doe nee Smith Prof., Janekeeps maidenSmith Prof.. Delimiters settle the question outright:Jane Doe (nee Smith MA)keeps the whole span, and brackets around a clause ending in a period are dropped as in 2.3.0, soJane Doe (nee Smith Prof.)gives titleProf.. The name #548 reported,Dr. nee Smith PhD Prof., gives titleDr. Prof., firstnee, lastSmith, suffixPhD, where 2.3.0 gave lastPhD, maidenSmith.John Smith nee Jones R.A.I.gives suffixR.A.I., as 2.3.0 read it. The rule replaces the clause rules this cycle’s #533 and #535 had added, and the readings those changes gave here are superseded. See theM2entry ofdocs/design/decisions.md(closes #601)Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.
parse("John Smith, Jr., Freiherr von Richthofen").ambiguitiesnamescomma-structurealone, where 2.0 through 2.3 also namedparticle-or-givenforvon– a word the same parse had put in the suffix, so the report described a reading it never made. This fix moves no field; the title change below (#603) movesFreiherrto the title. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain’s credential-acronym report this release adds (above) keeps the same bound:John Smith, PhD Do Mareports the comma’s decision once and not again forMa. See the 2026-09-28 bullet of theC1entry indocs/design/decisions.mdA credential after the name core starts a suffix run to the end of its part.
HumanName("John Smith PhD Jones")gives suffixPhD Jonesand reportssuffix-or-namefor the name word it took, where 2.3.0 gave middleSmith PhD, lastJones;Smith, John PhD Jonesgives suffixPhD Joneswhere 2.3.0 gave middleJones, suffixPhD. A title in the run is a title, soEric H. Holder Jr. Attorney Generalgives titleAttorney General, suffixJr., as the comma spellingEric H. Holder, Jr. Attorney Generalalready read, where 2.3.0 gave middleH. Holder Jr. Attorney, lastGeneral. Only an unambiguous credential or generational word starts the run, and only behind the name core – two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, soJohn Smith MA JonesandMohamed Ali Abd Allahread as before, and a bare title word starts nothing:Mary Jane King Smithkeeps middleJane King. See theS2entry ofdocs/design/decisions.md(closes #602)Fix a comma part that opens with a credential being read as a given name.
HumanName("John Smith, PhD Jones")gives firstJohn, lastSmith, suffixPhD Jonesand reportssuffix-or-namefor the name word it took, where 2.3.0 gave firstPhD, middleJones, lastJohn Smith(and 1.4.0 titlePhD, firstJones). With one word before the comma the whole part is suffixes:Doe, PhD Jonesgives lastDoe, suffixPhD Jones. A title word in the part stays a title. A word that is a title as well as a credential (MD,Ms) opens nothing, soSmith, Ms Janekeeps titleMs, firstJane. An undeclared delimiter becomes one more word of the part:Steven Hardman, RN - CRNAgives firstSteven, lastHardman, suffixRN - CRNA, where 1.4.0 through 2.3.0 gave firstRN, middle-, lastSteven Hardman. The split spelling now counts like the joined one, soSmith, John Ph. D. Jonesgives suffixPh. D. Jones, where 2.3.0 gave middleJones, suffixPh. D.. See theC1entry ofdocs/design/decisions.md(closes #603)Fix a comma part read as credentials being taken apart again by a later join.
HumanName("John Smith, Ph. D. and Mary Jones")gives firstJohn, lastSmith, suffixPh. D. and Mary Jones, where 2.3.0 gave firstPh. D. and Mary, middleJones, lastJohn Smith, andparse()reports each name word the part took: the parser now decides once, after every word is recognized, whether the part after a comma holds the credentials, and nothing joins across that decision afterwards. Titles joined byandstay one title (John Smith, Mr. and Mrs.keeps titleMr. and Mrs.). A surname joined byycounts as its words, as #575 already counted it in front of a credential:Ortega y Gasset, Dr.gives titleDr., firstOrtega, middley, lastGasset, where 2.3.0 gave lastOrtega y Gasset. And a surname with a suffix in it before the comma is the surname, asSmith Jr., Johnalways read it:Smith Jr., Esq.gives lastSmith, suffixJr., Esq., where 2.3.0 gave firstSmith, andSmith Jr., Vgives firstV, lastSmith, suffixJr., where 2.3.0 gave firstSmith, suffixJr., V. See the #613 entry underC1indocs/design/decisions.md(closes #613)Change a title word in a part after a second comma to read as a title.
HumanName("Eric H. Holder, Jr., Attorney General")gives titleAttorney General, suffixJr., where 1.4.0 through 2.3.0 gave suffixJr., Attorney General, and the part no longer reportscomma-structure. A word that is also a suffix stays a suffix (John Smith, MD, Ms). A title joined by a connective is read as a title, but its part keeps the report (Secretary of State). (#603)Add the Catalan and Polish surname link.
parse("Josep Carod i Rovira")gives familyCarod i Rovira, where every release from 1.4.0 through 2.3.0 gave middleCarod iwith familyRovira;Josep Lluis Carod i Roviragives middleLluiswith that same family; andCarod i Rovira, Josepgives it too, where they read familyCarod Roviraand took the link intosuffixas a generation marker.iis connective vocabulary now, the wayyalready was, and a connective counts as a name word wherever the three-word carve-out counts them – whatever else the vocabulary says the word is, which matters here becauseiis also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, soJohn Quincy Smith ikeeps suffixi,Josep Lluis Carod i IIIkeeps suffixi III, and the two-wordCarod ikeeps its generation reading. Written wholly in one case the letter reads as an initial and says so:JOSEP CAROD I ROVIRAandjosep carod i rovirakeep the fields they had and gain aconjunction-or-initialreport, which a one-case name gains wherever a bareiorIstands among the name’s own words – a letter inside a maiden clause is read by the clause’s rules and stays silent, asealready was – and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (parse("john smith i jr")gives middlesmith, familyiand suffixjrwhere it gave familysmithand suffixi jr). Case repair follows the reading: a lower-caseithe parse read as the generation is still title-cased bycapitalize(force=True)(Carod igivesCarod I, as every release did), while one standing among the name words keeps its lower case asyalways has (Carod i RoviragivesCarod i Rovira, whereCarod I Rovirawas the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way:HumanName("Jane Doe nee Puig i Soler")gives maidenPuig i Solerwith lastDoe, where 2.0 through 2.3 ended the birth name at the link and gave maidenPuigwith middleDoe i, lastSoler– and the same words would have joined into lastDoe i Solerunder the change above, carrying a word of the birth name into the current surname.Jane Doe née Kowalska i Nowakmoves with it, as does the all-lowerjane doe nee puig i soler; theyspelling always read this way and is untouched. The link still has to be joining:Jane Doe nee Puig ikeeps maidenPuigwith suffixi, andJane Doe nee Puig i IIIsuffixi III. A caller with Catalan or Polish data removes the entry fromconjunctions_ambiguousand gets the join in the one-case names too; a caller who wants none of this removesifromconjunctions, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it –parse("Dr. John i Smith").capitalized(force=True)keepsiwhere the off switch givesDr. John I Smith, andCarod i RoviraandJosep i Roviraare the same shape. Those two are the whole of what the switch does not undo, andtests/v2/test_properties.pystates them as its invariants’ only exemptions. A delimiter the caller declares throughPolicy(extra_suffix_delimiters=...)parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under(" - ",),Smith, John, PhD - i Solerkeeps suffixPhD, i Soler, as 2.3.0 read it. The default policy declares no such delimiter. See theP3andM2entries ofdocs/design/decisions.md(closes #397, closes #538)Fix a connective contributing no initial even where it is joining nothing.
parse("Juan de y").initials()givesJ. y., where every release gaveJ.whilefamily_basesaidy– two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, soJohn and Jane SmithgivesJ. J. S.where 2.0 through 2.3 gaveJ. a. J. S.and 1.4.0 the run-togetherJ a J. S.,Duke of EdinburghgivesD. E.where 2.0 through 2.3 gaveD. o. E.and 1.4.0D o E., andJohn & JanegivesJ. J.. The question is asked of the whole part and never of a word count, soJon Dough andhas baseDough andand keepsJ. D., andJuan Velasquez y GarciakeepsJ. V. G..HumanName.initials()moves with the core – over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it:JUAN Y GARCIAandمحمد و عليboth give the answer 1.4.0 gave. Parsing got cheaper by the same change – the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, soinitials()andcapitalize()disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See theR3entry ofdocs/design/decisions.md(closes #461)Change case repair’s exceptions map from replacement spellings to case masks, so md repairs to MD and phd to PhD.
HumanName("john smith phd").capitalize()givesJohn Smith PhDandjohn smith mdgivesJohn Smith MD, where every release from 1.4.0 through 2.3.0 gaveJohn Smith Ph.D.andJohn Smith M.D.. Acapitalization_exceptionsvalue is now the key’s own letters and digits in the case each should take, laid over the word as it was written, so the onephdentry repairsph.d.toPh.D.andJOHN SMITH PH.D.toJohn Smith Ph.D., and repair never adds or drops a character:john smith iii.givesJohn Smith III.where every release dropped the period. Punctuation in a value only marks which of its letters are joined and is never written into the word, which matters for a lone initial:john smith p.h.d.givesJohn Smith P.H.D., each letter the writer split off from the mask’s one runPhDbeing an initial, where every release gaveJohn Smith Ph.D.. The shipped map holdsphd→PhD,bsc→BScandmsc→MScand fifteen more of the listed post-nominals whose usual spelling is mixed case –DSc,PsyD,PharmD,MDiv,ThDandBtamong them – sojohn smith bscgivesJohn Smith BScandjohn smith psydgivesJohn Smith PsyD, where every release gaveJohn Smith BscandJohn Smith Psyd;mengandeddget no mask, since a mask applies wherever its word stands and both are also names (Meng Li,Edd Smith);md,ii,iiiandivleft it, a suffixmdnow repairing by the acronym repair listed under Additions and a suffix numeral by the numeral repair below. A mask still applies wherever its word stands (phd smithgivesPhD Smith), but a word that left the map and was parsed as anything but a suffix repairs as that reading:iv smithgivesIv Smithwhere every release gaveIV Smith, andMd Abdul KarimstaysMdunderforce=Truewhere every release gaveM.D.. A value that does not spell its key’s letters and digits –{"jr": "Junior"}– now raisesValueErrorwhen theLexiconis built, and at the first parse for a v1Constants, where the raise names the fix in v1’s own spelling (constants.capitalization_exceptions['jr'] = 'JR') rather than theLexicon()constructor call the same check offers a 2.0 caller; a value may still carry its own punctuation ({"md": "M.D."}is accepted, and repairsmdtoMD). Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 11 move on the defaultcapitalize()path and 78 underforce=True; no role field moves. The recipe is theR4entry’s 2026-09-23 MEASURED bullet indocs/design/decisions.md(closes #459)Fix case repair treating a suffix written in capitals as evidence that the whole name was cased on purpose.
HumanName("juan garcia III").capitalize()givesJuan Garcia III, where every release from 1.4.0 through 2.3.0 returned it untouched – v1’s test for it had been a known failure since 2012 (the Google Code tracker’s issue 22) – andjuan garcia PhDandJUAN GARCIA Jr.repair the same way. Repair still acts only on a name written wholly in one case, but the suffixes are left out of that test now: a credential or a generation written the way one is written says nothing about how the writer cased the name. A title still counts, soDr. juan garciais left alone, and so does every other word, nicknames and maiden names included (Juan garcia IIIandjane doe nee SMITH IIIare left alone too). A suffix written in more than one case is the writer’s spelling and is kept as written:john smith EdDgivesJohn Smith EdDandjuan garcia PsyDgivesJuan Garcia PsyD, and so does a garbled one,juan garcia IiigivingJuan Garcia Iii;force=Truerepairs it (John Smith EDD,Juan Garcia III), and every release returned those names untouched. The parser’s own reading of a name’s case is unchanged and still counts the suffix, so the two can differ:john e jones IIIgivesJohn e Jones III, the capitals making theea connective to the parser, wherejohn e jones iiigivesJohn E Jones III. Because the test follows the parser’s suffix reading,jack MAgivesJack MAandMD, PhDgivesMd PhD, where 2.3.0 left both untouched. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), the new test admits 17 and 12 of them move on the default path;force=Trueis unchanged. The recipe is theR5entry’s 2026-09-23 MEASURED bullet indocs/design/decisions.md(closes #492)Fix case repair capitalizing the connective inside a hyphenated compound surname.
HumanName("jose ortega-y-gasset").capitalize()givesJose Ortega-y-Gassetandmaria silva-e-sousagivesMaria Silva-e-Sousa, the lowercase connective 1.4.0, 2.0.0 and 2.1.0 gave, where 2.2.0 and 2.3.0 gaveJose Ortega-Y-GassetandMaria Silva-E-Sousa– so this restores 1.4.0’s answer after a 2.2 regression.JOSE ORTEGA-Y-GASSETgives the sameJose Ortega-y-Gasset, where every release gaveJose Ortega-Y-Gasset. A connective with a part on each side of it inside one hyphenated word keeps its lowercase, as the spaced spelling always has, in every role; at either end of the word it is ordinary name text, sojuan e-f smithstill givesJuan E-F Smithandjuan y-garciagivesJuan Y-Garcia. A single letter marked with a period is an initial there, never the connective, soj.-e.-p. dupontkeepsJ.-E.-P. Dupont. The hyphen is read as the writer’s join even in a name written wholly in one case, where the same letter spaced reads as an initial, so a bare hyphenated initial that spells a connective is lowered –J-E-P DUPONTgivesJ-e-P Dupontwhere every release gaveJ-E-P Dupont, a recorded boundary – while the spacedmaria silva e sousagivesMaria Silva E Sousa. Only connective vocabulary is read, so the MāoriTe Awanui-a-Rangi Blackstill repairs toTe Awanui-A-Rangi Blackunderforce=True. The shape is rare: none of the 1340 names in the differential corpora at the commit before this change (2026-09-23) carries it, the only movers being the rules document’s own example lines. See theR4entry ofdocs/design/decisions.md(closes #478)Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.
HumanName("john smith x.y.z.").capitalize()givesJohn Smith X.Y.Z.where every release gaveJohn Smith X.y.z., the dotted word being a suffix now (theunlisted_dotted_suffixeschange above) and repaired as a listed acronym is; andjohn smith vigivesJohn Smith VIwhere every release gaveJohn Smith Vi, withvii,viiiandixalike. Both are keyed on the suffix role:Jack X.Y.Z., which keeps its surname, still repairs as a name word (Jack X.y.z.underforce=True), andjohn smith xistill givesJohn Smith Xi, the parser readingxias the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (john smith B.Tech.givesJohn Smith B.Tech.), and reads all capitals underforce=True(John Smith B.TECH., where every release gaveJohn Smith B.tech.), which acapitalization_exceptionsmask such as{"btech": "BTech"}undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 underforce=True. The recipe is theR4entry’s 2026-09-23 MEASURED bullet indocs/design/decisions.md(#459)Fix case repair breaking a name typed with decomposed accents at each accent.
HumanName(unicodedata.normalize("NFD", "josé garcía")).capitalize()givesJosé García, where every release from 1.4.0 through 2.3.0 gaveJosé GarcíA: decomposed text (NFD, which macOS file names and some databases hand back) writesíasifollowed by a combining accent, and the letters after the accent were repaired as a separate word. A decomposed name now repairs the way its composed spelling does, Mac/Mc names and case-repair masks included, and the output keeps the form it was typed in. See theR4entry ofdocs/design/decisions.md(closes #542)Fix initials dropping the accent from a name typed with decomposed accents.
parse(unicodedata.normalize("NFD", "émile zola")).initials()givesé. z., where every release gavee. z.: the initial was a word’s first code point, which in decomposed text is the letter without its accent. Decomposed katakana lost its voicing mark the same way and now keeps it:マイケル ジャクソンtyped decomposed initialsマ. ジ., where every release gaveマ. シ..HumanName.initials()moves alike, a decomposed name’s initials stay decomposed, and a composed name’s initials do not change. See theR3entry ofdocs/design/decisions.md(closes #585)Change the parse pipeline to copy its state without dataclasses.replace. Every stage returns a copy of its frozen state, and several also copy tokens one at a time;
dataclasses.replacegoes throughfields()and__init__on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline’s own three dataclasses and checks them at import. One parse of the benchmark’s reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11,parse406 to 370 andHumanName443 to 407), and the call-count baselines move with them. Recomputable withuv run python tools/perf/call_count.py --against e0f1a2f; the counts for every interpreter are in theparse-costentry ofdocs/design/decisions.md. No user-visible behavior changes (#546)Fix a long given part after a family comma costing quadratic time. Since 2.3.0, parsing
"Doe, Jane " + "Smith " * ntook time growing with the square of the part’s length: going from 1,600 to 6,400 words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (Doe, Jane Smith Prof. Prof. ...) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553)Fix a name ending in alternating credentials and titles costing quadratic time. Since 2.3.0, parsing
"John Smith " + "MA Prof. " * ntook time growing with the square ofn: going from 400 to 1,600 pairs cost 12.9x the time, where 2.2.0 cost 4x. It costs 4x again (Python 3.11,HumanName, measured 2026-10-01). No field moves (closes #558)Fix many leading titles plus many surname particles costing quadratic time. Parsing
"Dr. " * n + "Jan " + "van Berg " * ntook time growing with the product of the two counts in every 2.x release: going from 400 to 1,600 of each cost 12.2x the time at 2.2.0 and 2.3.0, and 12.5x at 2.0.0. It costs 4x now (Python 3.11,HumanName, measured 2026-10-01). No field moves (closes #559)Fix a v1 ``Constants`` entry with stray whitespace raising or changing the parse. After
c.suffix_not_acronyms.add("ma "),HumanName("John Smith", c)raisedValueErrorin 2.0 through 2.3, and afterc.titles.add(" dean "),HumanName("dean john smith", c)gave titledean, where 1.4.0 read both as if the entry were absent. An entry that is empty or holds edge whitespace, a whitespace run or any whitespace character other than a single space, in any set or as acapitalization_exceptionskey, is now ignored with aUserWarningnaming it, as 1.4.0 ignored it. AConstantsrestored from a 1.4 pickle carries two such entries from 1.4.0’s own title list, and these are ignored without a warning, soHumanName("Actor John Smith", c)gives firstActoras on 1.4.0, where 2.0 through 2.3 gave titleActor. (closes #541)Fix a v1 ``Constants`` entry that is only a CJK full stop raising at the first parse. After
c.titles.add("。"),HumanName("john smith", c)raisedValueErrorin 2.3;。,.and。in any set or as acapitalization_exceptionskey are now ignored with the sameUserWarningas an entry with stray whitespace. 1.4.0 through 2.2.0 accepted such an entry and applied it to a name token that is nothing but that full stop, which 2.3 and later cannot match. An entry ending in such a full stop no longer raises either:c.suffix_not_acronyms.add("ma。")madeHumanName("jack ma", c)raiseValueErrorin 2.3, and now gives lastma, as without the entry. (closes #582)Fix a Japanese name written in halfwidth katakana being read given-first.
HumanName("山田 タロウ")gives last山田, firstタロウ, as山田 タロウdoes, where every release gave first山田, lastタロウ. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (高橋一郎 タロウ), the Chinese·divides between halfwidth kana (タロウ·ヤマダgives firstタロウ, lastヤマダwhere it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title:タナカ. Johngives firstタナカ.where it gave titleタナカ.. A name written wholly in katakana, halfwidth or not, still keeps the declared order, given-first by default (ヤマダ タロウgives firstヤマダ), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add(Script.KATAKANA, FAMILY_FIRST)toPolicy.script_orders(see East Asian names). See theW4entry ofdocs/design/decisions.md(closes #594)Fix a katakana name typed with a separate voicing mark being read family-first.
HumanName("ア゙イ タロウ")gives firstア゙イ, lastタロウ, asアイ タロウdoes, where 2.1 through 2.3 gave lastア゙イ, firstタロウ. A dakuten or handakuten after a kana that has no precomposed voiced form (ア゙,ン゙), or the spacing゛and゜, was read as hiragana, which made the name Japanese by script and turned it around; the mark now belongs to the kana it follows. It flipped the rest of the name too:マイケル ア゙イgives firstマイケルwhere it gave lastマイケル. Hiragana names and kanji-and-kana names with such a mark read as before. See theW4entry ofdocs/design/decisions.md(closes #596)Fix a declared suffix delimiter being taken into a joined name part instead of separating suffixes.
HumanName("Smith, John, PhD - and MD", suffix_delimiter=" - ").suffixisPhD, and MD, where 2.0 through 2.3 gavePhD - and MD;Smith, John, Puig - y SolergivesPuig, y Solerwhere they gavePuig - y Soler. Both are 1.4.0’s answers again. A delimiter declared throughsuffix_delimiterorPolicy(extra_suffix_delimiters=...)now separates a trailing suffix part as a comma typed in its place would, giving the same fields: a connective beside it never joins across it, a maiden clause ends at it, and the delimiter itself is dropped, where 2.0 through 2.3 dropped only a delimiter standing alone and kept one a connective had joined. A maiden marker in a part the delimiter separates takes no maiden name, as no marker after a suffix comma does (see the maiden-marker entry above):Smith, John, MD - née Jones Smithgives suffixMD, née Jones Smithand no maiden, where 2.0 through 2.3 gave maidenJones Smith. Only the parts after a suffix comma are affected – the credentials ofName, PhDor ofFamily, Given, PhD; in the name before the first comma, and in the given-name part ofFamily, Given, the delimiter is still a word, as in 1.4.0. The default policy declares no delimiter, so nothing changes without one. See theC1entry ofdocs/design/decisions.md(closes #549)Fix halfwidth corner brackets not being read as a nickname.
HumanName("山田 「タロー」 タロウ")gives nicknameタロー, last山田, firstタロウ, where every release gave middle「タロー」and first山田(the order moves with the halfwidth katakana change above; the bracket would otherwise have blocked it). The halfwidth「」are the corner brackets of legacy JIS X 0201 data, the same punctuation as「」, and are now a default nickname pair in both APIs:DEFAULT_NICKNAME_DELIMITERSgains("「", "」")and the 1.xnickname_delimitersgains the keyhalfwidth_corner_brackets. They are not limited to Japanese text:John 「Jack」 Smithgives nicknameJackwhere it gave middle「Jack」. AConstantsrestored from a pickle keeps the keys it was saved with, as it did when 2.0 added the other typographic pairs. See theN1entry ofdocs/design/decisions.md(closes #597)Fix the Irish particles Ó, Ní and Ua and the Malay binti being read as a middle name.
HumanName("Liam Ó Murchú")gives firstLiam, lastÓ Murchú, where 1.4.0 through 2.3.0 gave middleÓ, lastMurchú;Sinéad Ní Mhurchú,Seán Ua BuachallaandIna binti Navalamar(and the Singapore spellingbinte) move the same way.ÓandNíare never given names, soÓ Murchúalone is all last name, where every release gave firstÓ.Uaandbintican be, so a leading one stays the first name andparse()reportsparticle-or-given:Ua Buachallagives firstUa, lastBuachalla, as before.Ó.written with a period is still an initial:Juan Ó. Pérezkeeps middleÓ., whileJuan Ó Pérezgives lastÓ Pérez. Case repair writes the Irish particles capitalized,SEÁN Ó MURCHÚrepairing toSeán Ó Murchú, andbintiin lowercase:INA BINTI NAVALAMARrepairs toIna binti Navalamar, where 2.3.0 gaveIna Binti Navalamar. The Irish casing comes from newcapitalization_exceptionsentries, which now outrank the lowercase case repair gives a particle, and that holds for your own entries too: withconstants.capitalization_exceptions['van'] = 'Van',ludwig van beethovenrepairs toLudwig Van Beethoven, where 1.4.0 through 2.3.0 keptvan. SeeP7and the #604 entries undervocabulary-collisionsandR4indocs/design/decisions.md(closes #604)Fix salutations in Finnish, Estonian, Romanian, Croatian/Serbian, Icelandic, Czech, Lithuanian, Malay, Filipino and other languages being read as a first name.
HumanName("Herra Väinö Johansson")gives titleHerra, firstVäinö, lastJohansson, where 1.4.0 through 2.3.0 gave firstHerra, middleVäinö. The titles gain Mr/Mrs/Miss forms such asrouva,proua,doamna,gospođa,frú,meneer,paní,ponas,encikandginang; the word for Count in several of them (kreivi,krahv,hrabia,greve,graaf);familie(Familie Hansengives titleFamilie, lastHansen);knight; and the officescommissioner,counselandadministrator, which complete titles that half-worked:Police Commissioner James Gordongives titlePolice Commissioner, firstJames, where it gave titlePolice, firstCommissioner. Some of the new titles are also names, nearly always the last name (Greve,Knight), and those still read as the last name there:Gladys KnightandKnight, Gladysare unchanged (Gladys Knight., with a trailing period, gives titleKnight., asMary Jane King.already does). The cost falls on a name that begins with one of them, the costGrafalready pays:Greve Annagives titleGreve, lastAnna, and so doesHrabia AnnaunderFAMILY_FIRST; the few borne as first names (Herra,Batoni) lose them the same way. Words that commonly lead a real name stay out, among them VietnameseÔng, Polish and CzechPan, andMarshalandJustice.Knt(Knight) joins the post-nominals besideKt:Sir John Smith Kntgives suffixKnt, where it gave middleSmith, lastKnt. See thesalutation-titlesentry ofdocs/design/decisions.md(closes #606)
Additions
Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials. A subset of
conjunctionsholdingeandiby default; it is the knob for the change above rather than a switch. Portuguese data, whereelinks surnames the wayydoes in Spanish, takes it out:Lexicon.default().remove(conjunctions_ambiguous={"e"})restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one:Lexicon.default().add(conjunctions_ambiguous={"y"}). A v1Constantshas no manager of its own for it – deleting the word fromconjunctionsis what turns the marking off, the same rule the glued-honorific tails follow. Seedocs/customize.rst(#383, #479)Add AmbiguityKind.CONJUNCTION_OR_INITIAL, reported when a one-letter connective in a name written wholly in one case is read as an initial:
parse("jose e maria santos").ambiguitiesandparse("JOSE E MARIA SANTOS").ambiguitiesboth name it, anddetailnames the letter. That is the call the behavior change above had to make. A letter outside the marked set reports nothing, its reading not being in doubt, soJUAN GARCIA Y LOPEZis silent; so is every mixed-case name, where the writing decided it. See theP3entry ofdocs/design/decisions.md(#383, #479)Repair a credential acronym the case-repair exceptions map does not carry to all-caps instead of title-casing it.
HumanName("JOHN SMITH MBA").capitalize()givesJohn Smith MBAwhere every release since 1.4.0 gaveJohn Smith Mba;john smith jdgivesJohn Smith JD. The repair is keyed on the word having parsed in the suffix role from the acronym vocabulary, so a word that is an ordinary name merely sharing a spelling with an acronym is untouched, and the exceptions map is still asked first – since the case-mask change above it holds masks, sojohn smith bscgivesJohn Smith BScrather thanBSC, andmd, which left the map, is one of the acronyms this repair reaches (john smith mdgivesJohn Smith MD) – while the generationaljris unaffected (john smith jrgivesJohn Smith Jr). The given-name half of a mixed run is unchanged, soQC MPgivesQc MPwith theQC(given role) still title-cased and only theMP(suffix role) repaired. When this repair landed (PR #521, before the case-mask change), twenty-two names moved in the differential corpora on the defaultcapitalize()path and 119 underforce=True, every one a single-case name with an acronym suffix the map did not carry; no role field moves. See theR4entry ofdocs/design/decisions.md(#459)Remove ph from the default post-nominal acronyms. The fragment existed only so the merged
Ph. D.token could pass the acronym test on its first piece, and the repair above would have readjohn smith ph. d.asJohn Smith PH. D.; the parser merges the split spelling by its own rule, soHumanName("John Smith Ph. D.")still gives suffixPh. D.,john smith ph. d.capitalizes toJohn Smith Ph. D., andphd/Ph.D.parse as they did. The cost is a barephwith noD.behind it, dotted or not, alone or inside a credential run:HumanName("John Smith Ph.")gives middleSmith, lastPh., where every release since 1.4.0 gave suffixPh., andJohn Smith MD Ph.gives middleSmith MD, lastPh., theMDleaving the suffix with it. A caller who needs that back adds it:Lexicon.default().add(suffix_acronyms={"ph"}). See theExcluded (SUFFIX_ACRONYMS -- ph)entry ofdocs/design/decisions.md(#459)
2.3.0 - September 12, 2026
nameparser 2.3 is parsing fixes and new honorific vocabulary; nothing in the API is removed or renamed.
The fixes cluster around post-nominals and titles. A space-separated run of post-nominals keeps the spacing the writer typed, and the acronyms that are also surnames – Rai, Cha, Ba – no longer take a name’s family name. A run of titles addresses by its last, and a trailing abbreviated title reads as a title. CJK names and honorifics written with a full stop of any width now parse. The additions are renunciate and royal given-name titles, Devanagari and the first Bengali honorifics, and two
AmbiguityKindmembers that report a reading nothing in the name decided.One incompatibility: a
Lexiconpickled by 2.1.x or 2.2.x with a caller-added wide-stop or NFD entry no longer loads; the full-stop bullet below has the remedy.Behavior Changes
Fix HumanName.initials() dropping a middle- or family-group initial that is also a one-letter conjunction.
HumanName("Scott E. Werner").initials()givesS. E. W.again where 2.0.0 through 2.2.0 gaveS. W.;Juan Y. Garciaand a bare ASCII capitalJohn E Smithlikewise. v1 excluded initial-shaped words from its conjunction test and the 2.0 facade had not;parse(...).initials()was already right and is unchanged. A bare lowercasejohn e smithstill reads theeas the connective. See theR3entry ofdocs/design/decisions.md(closes #462)Record a 2.0.0 change to HumanName.initials() that no release note had classified: since 2.0.0 the facade initials each WORD of a name part, where 1.4.0 initialed a joined run as one group –
HumanName("Juan Velasquez y Garcia").initials()isJ. V. G.and wasJ. V G.;Abdul Salam HassanisA. S. H.and wasA S. H.. Nothing changes in 2.3.0; the differential gate now comparesinitials()(#484) and this is what it found. See thedifferential-ledger, the initials viewentry ofdocs/design/decisions.mdFix a space-separated run of post-nominals rendering with a comma the writer never typed.
HumanName("John Smith MD PhD").suffixisMD PhDand wasMD, PhDat every release since 1.4.0;Kenneth Clarke QC MPgivesQC MP, and the CJK honorific runs (김민준 박사 씨) follow the same rule. This is a deliberate deviation from 1.4.0, and the v1-parity suite is re-pinned to match. The comma forms are unchanged –HumanName("Smith, MD, PhD").suffixis stillMD, PhD– because the separator is now the comma the writer typed; a configured suffix delimiter still parts a run, as does a name word standing between two post-nominals. Round-tripping is fixed for these runs, which the 2.2.0 note below recorded as broken:str(HumanName("Smith, MD PhD"))isSmith MD PhDand re-parses to suffixMD PhD.str()is still a rendering rather than a canonical form. See theC1entry ofdocs/design/decisions.md(closes #436, closes #437)Fix Parser.revise() splitting a space-separated suffix value into comma-separated entries.
Parser().revise(n, suffix="MD PhD").suffixisMD PhDand wasMD, PhD. A suffix value’s entries are now derived from the value’s own commas by the rule a whole name uses – a comma parts two credentials and a space joins them – sorevise(n, suffix="MD, PhD")is still two entries, andrevise(n, suffix="Ph. D.")rendersPh. D.where the 2.2.0 note below acceptedPh., D.. One limit: a delimiter configured throughextra_suffix_delimitersparts a value only where the value’s own words read as a name with a tail segment, so in a run of post-nominals it stays a word; write a comma at the boundary instead.ParsedName.replace()is unchanged. See theC1entry ofdocs/design/decisions.md(closes #511)Remove rai and cha from the default post-nominal acronyms, so a trailing Rai or CHA keeps the family name.
HumanName("Aishwarya Rai")gives lastRai, where 2.0.0 through 2.2.0 gave suffixRaiand no last name at all – 1.4.0’s reading, restored. The cost is that a genuine credential written after a full name is no longer recognized:John Smith RAIgives lastRAI, andJohn Smith, RAIgives firstRAI, lastJohn Smith. A caller who needs either back adds it:Lexicon.default().add(suffix_acronyms={"cha"}). See thesuffix-acronym-collisionsentry ofdocs/design/decisions.md(closes #342)Mark ba as an acronym that is also an ordinary name, so a bare trailing Ba keeps the family name.
HumanName("Anna Ba")gives lastBaand reports a suffix-or-name ambiguity, where 2.0.0 through 2.2.0 gave suffixBaand no last name. The spaced full-name form keeps the credential reading –John Smith BAstill gives suffixBA, now flagged – and the dottedJohn Smith B.A.is an unflagged suffix. The comma forms move, which is the marking’s cost:John Smith, BAgives firstBA, lastJohn Smith, asSmith, Edalready did for the other ambiguous acronyms. WriteB.A.to keep the credential reading. Ba is a real surname in Vietnamese and Senegalese Fula, thema/Mashape exactly (#342)Fix a title run addressing by its first title rather than its last.
HumanName("Her Majesty Queen Elizabeth")gives firstElizabethwith an empty last name, where every release since 1.4.0 gave lastElizabeth. Several titles written together are one form of address and the one that does the addressing is the last, so the run is now matched whole or by its last word:Reverend Mother Teresa,Dr. Sir JohnandSir Sheikh abdul rahmanmove the same way. A run whose last word addresses by surname does not move:His Excellency Lord Duncanstill gives lastDuncan,lordnot being a given-name title. A caller’s multi-word entry still matches as a phrase. See theH1entry ofdocs/design/decisions.md(closes #489)Fix the leading title peel taking a name word and leaving a post-nominal to be the name.
HumanName("Dr King Jr")gives titleDr, lastKing, suffixJr, where every release since 1.4.0 gave titleDr King, lastJrand no suffix at all;Dr. King MDmoves the same way, and both now read as the comma spellingKing, Dr Jralways has. A name that is nothing but titles or nothing but post-nominals is untouched (Marquess of Bath,MD DDS), and so is a title written as one joined unit:Prince of Wales Jrkeeps titlePrince of Wales. Where the only word left is the title itself,Dr Jrgives firstDr, suffixJrand reports a title-or-name ambiguity. See theH3entry ofdocs/design/decisions.mdFix a trailing abbreviated title reading as a name word.
HumanName("John Smith Prof.")gives titleProf., firstJohn, lastSmith, where every release since 1.4.0 gave lastProf.and lost the surname;John Smith Dr.,John Smith Rev.andAndrew Perkins (Mgr.)move the same way, and the comma formSmith, John Prof.now agrees with the bare one where it gave middleProf.. A run chains from the end (John Smith Prof. Dr.gives titleProf. Dr.), a leading title keeps its place (Dr. John Smith Prof.gives titleDr. Prof.), and a post-nominal is read through it (John Smith Jr. Prof.gives suffixJr.). Only a listed title word wearing the abbreviation period is claimed: an unlisted abbreviation (John Smith Xyz.), a bare title word (John Smith Sir) and a post-nominal (John Smith Esq.) do not move. The reach is the whole title vocabulary, ordinary surnames in it included, soMary Jane King.gives titleKing.where the bareMary Jane Kingkeeps lastKing– accepted rather than prevented, the period being evidence the bare spelling never gives. The same argument holds in a native script:毛 泽东 Dr.gives titleDr., first泽东, last毛, where 2.2.0 gave lastDr.. See theH5entry ofdocs/design/decisions.md(closes #316)Remove esq from the default post-nominal acronyms, and assert the two post-nominal sets disjoint.
HumanName("John Smith E.S.Q.")gives middleSmith, lastE.S.Q., where every release since 1.4.0 gave suffixE.S.Q.;Esq,Esq.,ESQandesqare unchanged, the post-nominal word list carrying every single-token spelling. Esquire is a contraction rather than an initialism, and it was the one word in both post-nominal sets, which can now assert they do not overlap. A caller who needs the dotted spelling back adds it:Lexicon.default().add(suffix_acronyms={"esq"}). See thesuffix-acronym-collisionsentry ofdocs/design/decisions.mdFix the East Slavic and Turkic patronymic rotations overriding a declared family-first name order. With
patronymic_rulesopted in andPolicy(name_order=FAMILY_FIRST),Мицкевич Адам Юзефgave lastАдамthrough 2.2.0 and now gives lastМицкевич– the reading the declaration asks for – andoglu Ahmad Vali Aliwith Turkic handling gave lastAhmadand nowoglu. The rotations exist to restore the given-first reading a family-first listing hides, so under a declared family-first order the declaration decides. See theO1entry ofdocs/design/decisions.md(closes #384)Fix CJK honorifics and names written with a full stop of any width.
HumanName("김민준 씨.").suffixis씨.(the fullwidth stop a Japanese or Chinese IME produces by default), with last김and first민준, where every release since 2.1.0 gave first김, middle민준, last씨.and no suffix; the ideographic씨。and halfwidth씨。spellings move the same way. The vocabulary lookup now composes NFC, so a decomposedSeñorornéefrom macOS-origin data is recognized too. A period glued to a name word no longer breaks the name:양. 지훈gives last양., first지훈where 2.1.0 through 2.2.0 gave first양., middle지, last훈, and김민준씨.peels its honorific as the stop-less spelling does. The period stays on the word it was written with; nothing is rewritten. A Latin name written with ASCII periods is untouched (Smith. Johnstill reads titleSmith.), but a Latin or Cyrillic word wearing one of the three wider stops now reaches the vocabulary:Dr。 John Smithreads titleDr。where it read firstDr。. This retires the 2.2.0 note below that read period strictly. One incompatibility, by decision: aLexiconpickled by 2.1.x or 2.2.x that carries a caller-added entry the widened fold now changes – a non-ASCII entry written with a fullwidth or ideographic stop, or in NFD – no longer loads (ValueError: incompatible Lexicon pickle: entries are not normalized); the shipped vocabulary is unaffected, and the remedy is to rebuild theLexiconfrom its source rather than unpickle it. See thecjk-full-stopsentry ofdocs/design/decisions.md(closes #322, closes #323)
Additions
Add the renunciate titles to the given-name title list, so a renunciate’s one name is a given name.
HumanName("Swami Vivekananda")gives firstVivekanandawith an empty last name, where every release since 1.4.0 gave lastVivekananda;Guru Nanak,Baba RamdevandLama Zopamove the same way, and so do the Devanagari and Bengali spellings added below. Two name words behind the title are unchanged –Swami Vivekananda Saraswatikeeps lastSaraswati– and a surname-retaining title is untouched:Rabbi Cohenstill gives lastCohen.venerableis deliberately not in the list, the traditions using it splitting on whether the family name survives. See theindic-honorificsentry ofdocs/design/decisions.md(closes #346)Add prince and princess to the given-name title list, so a royal’s one name is a given name.
HumanName("Prince Harry")gives firstHarrywith an empty last name, where every release since 1.4.0 gave lastHarry;Princess AnneandHer Royal Highness Princess Annemove the same way. Two name words behind the title are unchanged –Prince Harry Windsorkeeps lastWindsor– and so isPrince of Wales Jr, the joined title addressing by its last word.lordandladyare deliberately NOT in the list: they address by given name only as a courtesy style (Lord Peter,Lady Diana) and by surname for every peer and every wife (Lord Byron,Lady Thatcher), which the list cannot express. The given-name collision is untouched:Prince Fieldernow gives firstFielder(#348). See theH1entry ofdocs/design/decisions.md(closes #519)Add trailing honorifics as post-nominal vocabulary: Latin
rinpoche, Devanagariजी,साहब,साहिब,साहेब,महाराज, and Bengaliসাহেব,বাবু,মহারাজ.HumanName("Lama Zopa Rinpoche")reads titleLama, firstZopa, suffixRinpoche;नरेन्द्र मोदी जीreads lastमोदी, suffixजी. They are recognized SPACED only and are deliberately absent fromLexicon.honorific_tails: Banerjee, Mukherjee and Chatterjee end in theजीsubstring (बनर्जी,मुखर्जी,चटर्जी), so a glued peel would cut a real family name in two, andगांधीजीstaying unpeeled is the accepted cost. Bengaliবাবুis trailing where Devanagariबाबूis a leading title (#344, #343)Add Devanagari honorifics (#344):
डॉक्टर,डा,प्रो,प्रोफेसर,प्राध्यापक,प्रा,पंडित,पं,सरदार,सुश्री,श्रीयुत,श्रीमान,सौ,बाबू,महात्मा,न्यायमूर्ति,मौलाना,जनाबandमहाराजाas titles, beside theश्री/श्रीमती/डॉthat shipped in 2.1.0, andस्वामी,गुरु,बाबा,संतas given-name titles.डॉक्टर शर्माreads titleडॉक्टर, lastशर्मा;स्वामी विवेकानंदreads firstविवेकानंदwith no last name. Dotted spellings (प्रो.,पं.) match the same entries. Excluded under the collision rule:कुमारी(Kumari is a given and a family name),बेगम,शेख,आचार्य,राजा/रानीandठाकुर, all borne as ordinary names (closes #344)Add Bengali honorifics – the first Bengali vocabulary in the default lexicon (#343):
ড,ডঃ,ডক্টর,ডাঃ,ডা,ডাক্তার,শ্রী,শ্রীমতী,জনাব,অধ্যাপক,প্রফেসর,বিচারপতি,মাওলানা,মুফতি,আলহাজ্ব,আলহাজ,মিঃ,মি,মিসেস,মোঃ,মো,মোসাঃ,মোসা,মোছাঃandমোছাas titles, andস্বামী,শ্রীল,গুরু,বাবাas given-name titles.ড. মুহাম্মদ ইউনূসreads titleড., firstমুহাম্মদ, lastইউনূস– the vocabulary beats the initial reading – while real initials are untouched:র. কে. নারায়ণis unchanged.মোঃ আবদুল করিমreads titleমোঃ, firstআবদুল, lastকরিম, the mirror of LatinMd; the visarga spelling and theমো.period spelling both match, and the women’sমোসাঃ/মোসা.rides the same pair of entries.ঠাকুরstays out, being Tagore. Latin transliterations (Sri,Pandit,Mst) are not added – they collide with real given names where the native scripts cannot – and belong to the opt-in packs of #345 (closes #343)Add AmbiguityKind.GIVEN_OR_FAMILY, reported when a name of one name word had nothing to decide which field it is:
parse("Andrew")still gives givenAndrewand now says that field was a convention rather than a reading – the library picks the given name under the default order and the family name under a declared family-first one, anddetailnames the field it picked.parse("Smith Jr.")reports it too, the suffix being peeled first. A name something DID decide stays silent –"Dr. Smith","Smith, Andrew","abdul"(bound given-name vocabulary),"J."(an initial’s shape) – and so does"毛泽东", where the writing system settles the order. No field moves anywhere. See theO5entry ofdocs/design/decisions.md(closes #449)Add AmbiguityKind.TITLE_OR_NAME, reported when an input that is nothing but honorifics had its last word read as the name:
parse("Lord Chancellor")still gives titleLord, familyChancellor, and now says so – this is a name parser, not a title parser, so handed a string with no name in it, it reads the last title word as one."His Holiness"and"Dr. King"move the same way,kingbeing title vocabulary. A title with an ordinary word behind it is silent ("Dr. Smith","King Charles"), and so is a lone title word:parse("Dr.")is a title with no name beside it. The same convention on the post-nominal vocabulary reports the existing suffix-or-name:parse("Rinpoche")gives givenRinpocheand flags it, as does"QC MP". No field moves for this change;Dr King Jr,Dr. King MDandDr Jrgain the report from the title-peel fix above. See theH4entry ofdocs/design/decisions.md(closes #491)
Documentation
Document the family-first and East Asian input shapes beside the three Latin ones. The input-shapes list in
usage.rstgrows from three forms to seven: forms 4 and 5 for a declaredFAMILY_FIRSTorFAMILY_FIRST_GIVEN_LASTorder, and forms 6 and 7 for the native East Asian arrangements the script carries on its own. A comma or a Latin wrapper around a CJK name is named as tolerated input – parsed best-effort, its handling changeable without notice – and thecustomize.rstcorrespondence between forms 2 and 4 is written out. Behavior is unchanged (#469)
2.2.0 - August 31, 2026
nameparser 2.2 is a rename plus about thirty parsing fixes.
The
nameparser.configword lists were still named for v1’s fields —PREFIXES,BOUND_FIRST_NAMES,FIRST_NAME_TITLES— while theLexiconthey feed has used particles and given names since 2.0. They now agree, and the lists are frozen, which retires editing one in place as a way to change a default. The rename itself changes no parse.The fixes cluster around surname particles, largely what a declared
name_ordermeans for Latin-script names, which this release settles; then maiden-name clauses, Arabic bound given names, and credentials after a comma. Most reach the default name order, and a bullet says so where its change is family-first only. Each names the shapes it moves, and the issue it closes carries the measurement.What breaks is code that writes to a default word list. Code that imports one by its 1.x name has until 3.0.
Breaking Changes
Add docs/design/ contributor documentation:
rules.md(the parser’s normative rules, with executable examples),decisions.md(the decision record) andmechanisms.md(the solution-pattern catalog). New tests execute every documented example and verify every code citationChange every vocabulary set in nameparser.config to a frozenset. Editing one in place –
TITLES.add("dean"), the old way of changing a global default – now raisesAttributeErrorat the line that writes it. To change the defaults forHumanName, build a privateConstantsand pass it (c = Constants(); c.titles.add("dean"); HumanName(name, constants=c)); for the 2.0 API, build a lexicon (Parser(lexicon=Lexicon.default().add(titles={"dean"}))). Mutating the sharedCONSTANTSstill works, but warns and goes away in 3.0.CAPITALIZATION_EXCEPTIONSis a mapping, not a set, and is unchanged. See Migrating from HumanName and Customizing the parser (#293)
Behavior Changes
Fix a title changing how the name behind it is read.
"Dr. Van Johnson"gave familyVan Johnsonwith no given name, and"Sir Van Johnson"gave givenVan Johnsonwith no family at all; both now read givenVan, familyJohnson– the reading the untitled"Van Johnson"has always had. A leading word that is both a title and a particle is unchanged:"St John Smith","Do John Smith"and"Freiherr von Richthofen"keep their readings. See theP2entry ofdocs/design/decisions.md(closes #367)Fix a given-name title keeping a bound given name from joining the word after it.
"Sheik abdul salam"read givenabdul, familysalam, and now reads givenabdul salamwith an empty family, as"Sir John"does;"الشيخ عبد الله"reads givenعبد الله. A title that addresses by family is unchanged ("Dr. abdul salam"). This also restores"Sheik Abu Bakar"to givenAbu Bakar, which the fix above had regressed, and drops thePARTICLE_OR_GIVENambiguity that name reported through 2.1 (closes #369)Fix a bound given name swallowing the family name before a single-letter generational suffix.
"abdul Smith V"read givenabdul Smithwith no family, where"abdul Smith II"and"abdul Smith Jr"read correctly; it now reads givenabdul, familySmith, suffixV, and so doIandX, for every bound given-name word. A suffix word before the numeral no longer hides it ("abdul Smith Jr V"reads familySmith). Shipped since 1.x: 1.4.0 read firstabdul Smith, lastV(closes #401)Fix a bound given name joining a suffix as “the word after it”.
"abdul Jr Smith Berg"read givenabdul Jrand now reads givenabdul, middleJr Smith, where"John Jr Smith Berg"puts it. Where the suffix was a split credential the bound word joined into it –"abdul Ph. D. Smith Berg"read suffixabdul Ph. D., a 2.0 regression – and now reads givenabdul, middleSmith, suffixPh. D.(closes #421)Fix a bound given name joining past a credential that the suffix rule then takes, leaving no family.
"abdul Smith Jr Ma"read givenabdul Smithwith no family and now reads familySmith, suffixJr, Ma, as"John Smith Jr Ma"does;"abdul Smith Ma"reads familySmith, suffixMa. Both as 1.4.0 read them."abdul Smith Berg Ma"keeps its join, and"Berg, abdul Sir"still reads givenabdul Sir(closes #425)Remove the Czech/Slovak abbreviation roz. from the default maiden markers. Marker matching is case-folded and period-insensitive, so
Roz– the diminutive of Rosalind – was the same string as the marker, and a marker takes every word after it:"Rosalind Roz Smith"read maidenSmithwith no family name at all. It and"Rosalind Roz Jones Smith"now read as 1.4.0 read them. The full participle is untouched –"Anna Nováková rozená Svobodová"still reads maidenSvobodová– and a caller who wants the abbreviation back adds it to their own lexicon:Parser(lexicon=Lexicon.default().add(maiden_markers={"roz"}))(found in #335’s review)Add the Polish maiden marker z domu to the default vocabulary, and let a maiden_markers entry be more than one word.
"Maria Kowalska z domu Nowak"now reads familyKowalska, maidenNowak, where every earlier version read the marker as part of the name (1.4.0: middleKowalska z domu, familyNowak). The bracketed spelling moves with it.maiden_markersandgiven_name_titlesare now the two fields exempt from the multi-word warning. See Customizing the parser (#434)Fix a bracketed maiden clause reading as a nickname because its brackets were not declared.
"Jane Smith nee Jones"gave maidenJoneswhile"Jane Smith (née Jones)"gave nicknamenée Jones; the bracketed spelling now reads familySmith, maidenJonestoo, and so does the Japanese"山田 花子(旧姓 佐藤)", which neededPolicy(maiden_delimiters=...)through 2.1. Every delimiter pair the parser ships moves the same way, quotes included. An interior clause no longer eats the name behind it ("Jane (née Jones) Smith"keeps familySmith), and two clauses beside each other each keep their own role ("Jane "Janey" Smith (née Jones)"reads nicknameJaney, maidenJones). A clause with no marker in it is still a nickname, which is whatPolicy(maiden_delimiters=...)remains for. This reachesHumanName(closes #335)Fix a particle chain and a maiden name taking a trailing generational numeral as a name word.
"John van der Berg V"read familyvan der Berg Vand"John née Jones Smith V"read maidenJones Smith V, where"John Smith V"reads suffixV; both now stop before the numeral, forIandXalike. A word before the numeral that is an initial keeps its reading ("John van der J. V"). The chain also stops before a bare credential with words to spare –"John van der Berg Ma"reads suffixMa, as 1.4.0 did – and no longer swallows the given name behind an unlisted abbreviation:"Xyz. van Johnson"and"Esq. van Gogh"read givenvan(closes #424)Fix a name losing its given/family split when a comma is followed only by an honorific.
"John Smith, Mr."returned the whole of"John Smith"as the family name and now gives givenJohn, familySmith, titleMr.: it is"Mr. John Smith"with the honorific moved to the end, and marks no surname boundary. A comma followed by an actual name still fixes the family ("John Smith, Jones"), and a single pre-comma piece has no split to keep ("Smith, Dr."is unchanged). The pre-comma name now also picks up the declared name order –"de Mesnil Jean, Dr."keeps familyde Mesnilunder a family-first order – and the particle-or-given ambiguity report ("Van Johnson, Mr.")Fix pure postnominals being claimed as titles:
jr,junior,phd,doandsehave left the defaulttitlesvocabulary, anddr/srahave left the suffix vocabulary they never belonged in, so"Smith, PhD"gives suffix rather than titlePhD. Twelve words are genuine duals and keep both memberships, with position deciding –"Lt. Smith"is a title,"Smith, LT"a postnominal, bareMdbefore a name the Bengali and South Asian abbreviation of Muhammad,MDafter it the degree. The cost is in leading position, where a dropped word now reads as a name:"PhD Smith"gives givenPhD, which is what makes"Do Nguyen"parse as the Vietnamese name it is.drandsraalso stop being recognized in trailing position, so"John Smith Dr."gives familyDr.. An ambiguous credential acronym (ma,ed,jd,do) counts as a suffix only when written with its periods, so"Jack Ma."keeps familyMa.as 1.4.0 read it. Routing a trailing title word totitleis a separate open question (#316)Fix a credential run after a one-word family comma reading as a title or a given name.
"Smith, Jr."and"Smith, PhD"now give suffixJr./PhDwhere they gave title, and"Smith, Ph. D. Jr."gives suffixPh. D. Jr.where the split credential landed in the given name – a regression from 1.4.0. The position right after a family comma is postnominal position. Vocabulary still decides which words qualify ("Smith, Dr."keeps titleDr.), the leading readings are untouched ("Sr. Garcia"is still titleSr.), and a name word in the run makes it the given-and-suffix reading it always had ("Smith, John Jr.") (closes #296, closes #325)Fix a space-separated credential run after a family comma rendering with a comma the name never had.
"Smith, MD PhD"gives suffixMD PhDwhere it gaveMD, PhD, and"Smith, CBE MC","Smith, BSc MBA"and"Smith, Dr. MD PhD"the same. The roles are unchanged; only the rendered string carried the extra comma. This reaches any family comma whose following segment holds no name word, not only a one-word family, so"John Smith, Jr. III"gives suffixJr. III– also what 1.4.0 gave. A run written with commas keeps them ("Smith, MD, PhD"), and a name written without a comma is unaffected and still renders its run comma-joined, so re-parsingstr()output does not reproduce the run (closes #429)Fix a one-character suffix word after a comma being read by the wrong neighbour.
"Smith, PSM I"gives suffixPSM Iwhere it gave givenPSMand suffixI, and"Smith, John V."gives middleV.where it gave suffixV.. Inside a comma part a suffix word short enough to be mistaken for an initial –I,Vand2in the shipped vocabulary – is read by what stands before it: behind a credential it describes that credential (PSM Iis Professional Scrum Master level I), and behind a name a period marks an abbreviation and so a middle initial. A numeral written bare after a name is still the generation it looks like ("Smith, John V"is suffixV), and a name with no comma is untouched (closes #430, closes #432)Fix a name opening with a particle that is never a given name being split at the particle under a family-first name order.
"de Mesnil"read as familyde, givenMesniland"de la Vega"as familyde, givenla Vega; each is now the whole surname, as it has always been in the default order, underFAMILY_FIRSTandFAMILY_FIRST_GIVEN_LASTalike. A word that can never be a given name leavesname_ordernothing to decide. Standing alone is the whole of it:"Juan de la Vega"underFAMILY_FIRSTstill reports givende la Vega. A leading particle that may be a given name is genuinely order-dependent and is untouched, so"van Gogh"still reads familyvan, givenGoghunder both family-first orders (closes #359)Fix a family name made only of particle words reporting no base, so the surname vanished from family_base and from the initials.
parse("Anh Do")gave familyDowithfamily_base''and initialsA., and underPolicy(name_order=FAMILY_FIRST)"Del Toro"gave familyDelthe same way. A particle standing alone in a name part is not doing a particle’s work there and now reads as an ordinary name word:"Anh Do"is baseDo, initialsA. D.;"Juan van der"is basevan der, initialsJ. v. d.;"Nguyen, Van Le"initialsV. L. N.where the middle name used to be dropped. The parse fields themselves do not move – only the derived views and the initials. See theR2entry ofdocs/design/decisions.md(closes #385, closes #402)Fix case repair lowercasing the words of a family name made only of particle words, where every other view already reads them as ordinary name words.
HumanName("ANH DO").capitalize()givesAnh Dowhere it gaveAnh do, and"anh van do"givesAnh Van Do. This DIFFERS FROM 1.4.0 deliberately and does not restore it: 1.4.0 returnedAnh do. The accepted cost is that a family which is nothing but particles capitalizes too, so"juan van der"givesJuan Van Der. A conjunction is untouched ("der, y van"givesy Van Der), and where the particles DO join a name word nothing changes ("juan de la vega"still givesJuan de la Vega) (closes #407)Change case repair to read the parser’s own conjunction tag instead of re-deciding, from the word’s spelling, whether a word is a conjunction or an initial. Two spellings of one name disagreed because of it:
"jose ortega-y-gasset"capitalized toJose Ortega-y-Gassetwhile"JOSE ORTEGA-Y-GASSET"gaveJose Ortega-Y-Gasset; both giveJose Ortega-Y-Gassetnow, a hyphenated token being one word to the parse whatever it contains. The spaced spelling is untouched and still repairs toJose Ortega y Gasset. A field assigned after the parse was never classified, so repair asks the vocabulary there – today’s vocabulary, which is narrower than 1.4.0 parity:h.last = "хосе и мария сантос"givesХосе И Мария Сантосon 1.4.0 andХосе и Мария Сантосhere. One reading changes for hand-builtTokens in the 2.0 API: an untagged token whose text is conjunction vocabulary now capitalizes as an ordinary name word. See theR4entry ofdocs/design/decisions.md(closes #458)Change the parse-cost benchmark to bound function calls per parse rather than wall-clock seconds. The old one-second bound failed four times on CI while the same code re-ran green on master; frame counts do not move under load. The bound is a per-interpreter band of ±2%, with a loose five-second backstop for what frame counts cannot see. Recomputable with
uv run python tools/perf/call_count.py --against v2.1.0; the counts and the per-PR attribution are in theparse-costentry ofdocs/design/decisions.md. No user-visible behavior changes (closes #475)Fix a name that opens with a spaced Ph. D. losing its surname.
parse("Ph. D. Van Johnson")read givenVan Johnsonwith an emptyfamilyand suffixPh. D.; it now reads titlePh., givenD., familyVan Johnson. A suffix never begins a name, and the split credential was the only shape that reached the defect. A family comma still opens a listing rather than a name ("John Smith Ph. D."and"Smith, Ph. D. Jr."keep their suffixes), and “the head” means the head of the string rather than of the name, so"Sir Ph. D. Van Johnson"is unchanged. This RESTORES 1.4.0. One accepted consequence:Parser.revise(suffix="Ph. D.")rendersPh., D.(closes #371)Fix a trailing surname particle being stranded as a standalone middle name under a family-first name order. Under
Policy(name_order=FAMILY_FIRST)the same listing written with a comma reads it as part of the surname:"Jong Anke de"gave familyJongwithdeleft as a middle name and now gives familyde Jong, givenAnke– the answerparse("Jong, Anke de")has always given.FAMILY_FIRSTis the only order that puts a trailing piece in a middle;FAMILY_FIRST_GIVEN_LASTputs it in the given slot, so"Nguyen Thi Van"under that order still reads givenVan."Beethoven Ludwig van"underFAMILY_FIRSTnow gives familyvan Beethoven. A particle standing alone in the given slot is no longer folded into the family either, so"Ménil de"reports givende. Nothing moves under the DEFAULT name order. See Customizing the parser for what declaring an order settles, and theP6entry ofdocs/design/decisions.mdfor the reasoning (closes #467)Fix initials() ordering a name differently from the fields of the same parse. Two rules fold words into the family and render them ahead of it –
Policy(middle_as_family=True)and the tussenvoegsel attachment after a family comma – and thefamilyfield honored the fold whereinitials()did not:parse("der, y van")gave familyvan derbut initialsy. d. v., and now givesy. v. d.. Undermiddle_as_familythis RESTORES v1, that option beingmiddle_name_as_last’s successor:"Doe, Dr. John A."givesJ. A. D.again where 2.0 through 2.2 gaveJ. D. A..HumanName.initials()was already right and is unchanged. See theR3entry ofdocs/design/decisions.md(closes #408)Fix a tussenvoegsel attached to the family name after a comma deciding a genuinely uncertain reading and reporting nothing.
"Van Johnson"reports aPARTICLE_OR_GIVENambiguity –Vanis a Dutch particle and a Vietnamese given name – while"Nguyen, Thi Van"picked the same word the same way, silently. The attachment now reports the fork it decides:"Nguyen, Thi Van","Berg, Jan van der"and"Vega, Juan de la"each gain aPARTICLE_OR_GIVEN, while a particle already read as a post-nominal reportsSUFFIX_OR_NAMEinstead ("Berg, Jan vd"). Worth knowing before you filter on this:"Beethoven, Ludwig van"– read exactly right – now carries a report too, nothing in the input separating it from"Nguyen, Thi Van".ambiguitiesis the only value that grows (closes #405)Fix a tussenvoegsel after a family comma being parsed as a middle name. Dutch and Belgian alphabetized listings move the particle behind the given name –
"Beethoven, Ludwig van"is how"Ludwig van Beethoven"is filed – and it was read as a middle name rather than as part of the surname:"Beethoven, Ludwig van"gave middlevan, lastBeethoven, and"Berg, Jan van der"gave middlevan der. Those now read familyvan Beethovenandvan der Berg. Two guards bound it: a name whose only given word is the particle keeps it ("Nguyen, Van"still reads givenVan), and where the word is BOTH particle and suffix vocabulary the attachment wins, so"Berg, Jan vd"reads familyvd Bergwhere 1.4.0 and 2.1 alike gave suffixvd– as doesmc. (closes #379, closes #380)Add abd to BOUND_GIVEN_NAMES, so the spellings that write the article as its own word join like the others do:
"abd Allah Smith"was givenabd, middleAllahand is now givenabd Allah.abdul,abdelandabdalwere already there, and the Arabic-scriptعبدhas covered the same word since 2.0, so only the Latin spelling was short. The word is also the postnominal ABD (“All But Dissertation”) and stays inSUFFIX_ACRONYMS: position tells the two readings apart, so"Jane Smith ABD","Jane Smith, ABD"and"Jane Smith A.B.D."all still read the credential as a suffix (#400)Change how far a leading never-given particle takes the surname when a family-first name_order is declared.
Policy(name_order=FAMILY_FIRST)read"de Mesnil Jean"as familyde Mesnil Jean– the whole name – and now reads familyde Mesnil, givenJean. The default order is unchanged, deliberately: with no order declared nothing marks where the surname ends, and a particle followed by several words really can be all surname (von Bergen Wessels); a caller who means familyde la Vegaplus givenJuanthere writes the comma. The stop cannot land inside a conjunction-joined run or a bound given-name pair:"de la Vega y Santos Juan"reads familyde la Vega y Santos,"ibn Awf abdul Rahman"givenabdul Rahman. Where two or more words are left over, the two family-first orders differ from each other for the first time:"de la Cruz Juan Carlos"reads givenJuan, middleCarlosunderFAMILY_FIRSTand the reverse underFAMILY_FIRST_GIVEN_LAST. See Customizing the parser, and theP1entry ofdocs/design/decisions.mdfor the reasoning (closes #395)Change the detail text of a PARTICLE_OR_GIVEN ambiguity to name the role the leading particle was actually given. It said “read as a given name” under every
name_order, which is false underPolicy(name_order=FAMILY_FIRST)– there"Van Johnson"reads familyVan, givenJohnson, and the report described the reading not taken. It now ends “read as a family name” in that case. Thekindis unchanged and staysPARTICLE_OR_GIVEN; only the human-readable text moved, and default-order output is identical (#355)Fix a maiden name being lost when a particle stood in front of the marker.
"Ursula Leyen geb. Albrecht"reported maidenAlbrechtcorrectly, but"Ursula von der Leyen geb. Albrecht"– the same words one particle chain apart – gave familyvon der Leyen geb. Albrechtand no maiden name at all, as did"Jane van der Berg née Jones". A suffix already stopped the particle chain; a marker now does too, so those read familyvon der LeyenmaidenAlbrechtand familyvan der BergmaidenJones. Under a family-first order"de la Cruz née Vega"now reads familyde la Cruz, maidenVega. Two limits remain, both recorded inrules.md#M2: a conjunction join and a bound given-name join each still absorb a marker first (closes #399)Move mc and ste into the never-given half of the particle vocabulary, and add los, las and das. The Spanish and Portuguese articles were absent from it entirely. A never-given particle opening a name folds into the family (
rules.md#P1) instead of being read as a given name, so"Mc Donald"was firstMc, lastDonaldand is now lastMc Donald;"Ste Marie","Los Santos","Las Casas"and"Das Silva"move the same way, and"Mc Donald Smith"becomes lastMc Donald Smith. ThePARTICLE_OR_GIVENambiguity goes with it formcandste.los,lasanddaswere not particles at all, so for those three the ordinary particle join fires from a non-leading position too:"Maria das Neves"is now lastdas Neves. (closes #360)Fix a bound given-name join leaving no family name when the name also carries a maiden clause, and stop the join absorbing the marker itself.
"abdul Berg née Jones"read givenabdul Bergwith an EMPTY family, where"abdul Berg"alone reads givenabdul, familyBerg; it now reads givenabdul, familyBerg, maidenJones. The join also declines when the piece it would absorb is a marker, so"van der Berg, abdul née Jones"reads givenabdul, familyvan der Berg, maidenJoneswhere it read givenabdul née. Where the bound word is ALSO suffix vocabulary, a declining join after a family comma leaves the post-nominal reading:"Berg, abd née Jones"reads familyBerg, suffixabd, maidenJones, as"Berg, abd"alone always has (closes #411)Fix a maiden clause changing how the rest of the name is read, and a connective join keeping the marker in the surname. The grouping rules that count a name’s words counted the maiden clause too:
"juan y garcia"reads givenjuan, middley, familygarcia, while"juan y garcia nee jones"read givenjuan y garciawith NO family name at all. A name of two or more name words now reads as it reads without its maiden clause, plus the maiden name. The connective join used to merge the marker into a multi-word piece, so"Jane van der Berg née y Jones"kept familyvan der Berg née y Jonesand now reads familyvan der Berg, maideny Jones. Two limits: a bound given-name word still never joins onto a marker standing as a word of its own ("Berg, abdul née PhD"), and a suffix-vocabulary word inside the maiden name stops the marker, so"Jane née Jr y Jones"now reads familyJr y Joneswith no maiden name (closes #412, closes #417, closes #418)Fix a title-plus-surname name losing its family name whenever anything stood beside it.
"Dr. Smith"reads familySmith, but"Dr. Smith née Jones"read givenSmithwith no family at all, and so did"Dr. Smith PhD"and"Dr. "Smitty" Smith". A suffix, a nickname and a maiden name each stand beside the name rather than in it, and the rule now counts name words alone: those read familySmithwith maidenJones, suffixPhDand nicknameSmittyrespectively, and"Freiherr von Richthofen geb. Albrecht"reads familyvon Richthofen, maidenAlbrecht. A given-name title still names no family:"Sir John née Jones"keeps givenJohn, exactly as"Sir John"does. One name moves where the nickname LEADS:"'Smitty' Dr. Jones"reads familyJoneswhere it read given (closes #410)Fix a name that is a surname and a maiden clause reporting no family at all.
"Smith née Jones"read givenSmithwith an emptyfamily; it now reads familySmith, maidenJones. A maiden marker announces a FORMER surname, which only means something beside a current one, so the lone name word left standing is the surname in use now. Every spelling of the shape moves, including a bracket pair you declared to mean maiden: underPolicy(maiden_delimiters=frozenset({("(", ")")})),"Smith (Jones)"reads familySmith, maidenJones."Smith (née Jones)"RESTORES 1.4.0, which read familySmithwith the clause as a nickname. Two shapes deliberately do NOT move: a word the vocabulary claims as a given name ("abd née Jones") and a word written as an initial ("J. née Jones Smith V") – this rule changes what POSITION decided and does not reach what a word already is. If you have code that reads the lone name word beside a maiden clause out of given, this is the release where it moves to family (closes #445)Change what a star import of the two 2.2 vocabulary modules binds.
nameparser.config.particlesandnameparser.config.bound_given_namesnow declare__all__, which their 1.x shims already did, sofrom ... import *binds their vocabulary alone – it also bound theassert_normalizedinvariant helper, and fromparticlestheBOUND_GIVEN_NAMESit imports only for a disjointness check. Importing a constant by name is unaffected, and no parse changes (#356)
Deprecations
Rename the four vocabularies whose 1.x names described the fields they feed in v1’s words, so the data layer matches the
Lexicon:1.x name
2.2 name
nameparser.config.prefixesprefixes.PREFIXESparticles.PARTICLESprefixes.NON_FIRST_NAME_PREFIXESparticles.NON_GIVEN_NAME_PARTICLESnameparser.config.bound_first_namesbound_first_names.BOUND_FIRST_NAMESbound_given_names.BOUND_GIVEN_NAMEStitles.FIRST_NAME_TITLEStitles.GIVEN_NAME_TITLESsuffixes.SUFFIX_NOT_ACRONYMSsuffixes.SUFFIX_WORDSEvery row above still resolves and is removed in 3.0. The two module rows are import paths: importing them still works and says nothing, since both modules are now empty shims. Reading a constant – by attribute access, by
from ... import, or byfrom ... import *– emits aDeprecationWarningnaming the module and constant to move to, once per line that reads it rather than once per process, so every place you have to edit is reported rather than only whichever one ran first.python -W error::DeprecationWarning -c "import yourapp"surfaces them; Python hidesDeprecationWarningoutside__main__. TheCONSTANTSattribute names (prefixes,non_first_name_prefixes,bound_first_names,first_name_titles,suffix_not_acronyms) are v1 facade surface and are unchanged. See Migrating from HumanName (#293)
2.1.0 - August 7, 2026
nameparser 2.1 makes East Asian names work without configuration. A name written wholly in Han or hangul, or in kanji with kana, is read family-first. An unspaced Korean name is split against the census surname list. CJK honorifics are recognized whether they are spaced or written against the name. The two conventions that need you to declare a language, Han segmentation and kana-aware division, ship as the opt-in
locales.ZHandlocales.JApacks. See East Asian names for how it fits together, and Customizing the parser for the switches that turn it off.Most of this is default-on, deliberately: wherever nameparser acts unasked, the script itself settles the convention and no language detection is involved. Latin-script names are unaffected. To restore 2.0’s reading of CJK text, use
Parser(policy=Policy( script_orders=(), segment_scripts=frozenset())).East Asian name support
Add the Chinese locale pack locales.ZH: opt-in Han segmentation for unspaced names like
毛泽东, with the surname vocabulary it needs. A pack rather than a default because a Chinese surname list corrupts Japanese names written in the same characters (高橋一郎would split高+橋一郎). Japanese data goes throughlocales.JAinstead. See Locale packs (#271)Add Lexicon.surnames, Policy.script_orders, Policy.segment_scripts, the Script enum and the DEFAULT_SCRIPT_ORDERS constant to the public API. This is the first behavior nameparser keys on the script a name is written in, allowed only where the script itself settles a convention and never as a proxy for guessing the language. See API reference (#271)
Add Lexicon.honorific_tails, the vocabulary the glued-honorific peel matches: entries that may be split off the end of a name token. It is a narrower set than the spaced honorific vocabulary, since a glued tail has no token boundary to lean on, and every entry must also be a
suffix_wordsentry. Extend both in one call, or adding tohonorific_tailsalone raisesValueError. See Customizing the parser (#308)Add AmbiguityKind.SEGMENTATION, reported when a surname split had a vocabulary-supported alternative:
"남궁민수"is 남궁 + 민수 by the compound surname but 남 + 궁민수 by the single-syllable one, and longest-match had to pick. A name with only one possible split reports nothing (#271)Add the Japanese locale pack locales.JA and the segmenter factory locales.ja_segmenter(), which together divide an unspaced Japanese name:
parser_for(locales.JA, segmenter=locales.ja_segmenter())reads山田太郎as family山田, given太郎. They are separate because no surname list can do this job, so the pack activates the stage and a third-party divider performs it.ja_segmenter()wraps namedivider-python, installed with the newnameparser[ja]extra; the core stays dependency-free.locales.available()is now('ja', 'ru', 'tr_az', 'zh'). See Locale packs (closes #272)Add Segmentation and the Segmenter type alias to the public API, plus the keyword-only Parser(segmenter=…) hook: any callable from a token’s text to a
Segmentationor toNoneto decline. It is consulted only for scripts inPolicy.segment_scripts, and only where the surname vocabulary declined first. Note what that ordering means when packs are stacked: a Japanese name opening on a listed Chinese surname never reaches the segmenter, so高橋一郎still splits高+橋一郎. The packs are alternatives, one per corpus. See Segmenters (#272)Add a construction-time UserWarning when a parser activates segmentation for scripts nothing can divide.
parser_for(locales.JA)withoutsegmenter=used to build a parser that behaved like a working one minus the feature, silently. It now names the dead scripts and the call to pass. Any configured segmenter or covering surname vocabulary silences it, so the default parser and thezhpack never warn. A from-scratch lexicon with no hangul surnames warns under the default policy, withPolicy(segment_scripts=frozenset())offered as the deactivation. See East Asian namesAdd the Script members HIRAGANA and KATAKANA. Two members rather than one
KANAbecause the parser treats them differently: hiragana never transcribes a foreign name, while a wholly-katakana name usually is one (#272)
Breaking Changes
Change the pickle compatibility of Policy and Lexicon: the new fields change the field layout their guarded
__setstate__checks, so a pickle written by 2.0.0 raisesValueErrornaming the missing fields instead of loading. Re-pickle after upgrading.HumanNamepickles are unaffected, since the facade pickles v1-shaped component state rather than these objectsChange the pickle compatibility of Parser the same way, for the new segmenter field. Re-pickle after upgrading. A
Parsercarrying a segmenter pickles only if that segmenter does, which a module-level function does and a closure or lambda does not (#272)Change one thing about parse totality:
parse()still never raises on any input, but a user-suppliedParser(segmenter=...)runs inside the parse and its own exceptions propagate rather than being absorbed. A failure there is a bug in your callable, not a fact about the name (#272)
Behavior Changes
Fix names written wholly in Han or hangul parsing given-first. Native-script CJK now reads family-first through the new
Policy.script_orderstable, so"毛 泽东"gives family毛where 1.x gave family泽东. A single unspaced token moves the same way, which is the easiest form of this to miss:"毛泽东"renders identically whichever field holds it, while the spaced form shows the change (str(HumanName("毛 泽东"))is now"泽东 毛"). No language detection is involved, and an explicit comma still wins. Default-on, and it reachesHumanName. See East Asian names (closes #271)Fix unspaced Korean names not splitting. The census surname list now ships as default vocabulary with hangul segmentation on by default, so
"김민준"parses family김, given민준where 1.x returned the whole string asfirst. Rendering follows the split. Nothing but Korean is written in hangul and its surnames are a closed set, which is what makes this safe as a default rather than a pack. Default-on. See Customizing the parser for the two switches (#271)Fix Japanese names carrying kana parsing given-first. The family-first rule extends to any name whose characters stay within kanji and kana while carrying at least one kana character, so
"高橋 みなみ"gives family高橋. The reasoning is the one hangul already uses: hiragana never transcribes a foreign name, and a transcription is kana alone, so kanji-plus-kana is a Japanese name in Japanese order. A name written wholly in katakana is excluded and stays positional, being predominantly a transcribed foreign name. Default-on. See East Asian names (#272)Fix names containing 〆 (U+3006, the shime mark opening Japanese surnames like 〆木) parsing given-first. The script classifier now counts it as Han, so these names take the family-first reading like any other wholly-Han name (#303)
Fix the katakana middle dot ・ (U+30FB, and its halfwidth twin U+FF65) being read as part of a name rather than as a divider. It now separates tokens exactly as a space does, so
"マイケル・ジャクソン"gives givenマイケル, familyジャクソンwhere 1.x left the whole string infirst. Native Japanese names never contain this character, so the separation is unconditional and the policy opt-outs do not cover it. Rendering returns the dot as a space (#272)Fix 间隔号-divided transcriptions parsing as one unsplit token. U+00B7, the interpunct Chinese text divides a transcribed foreign name with, is now a token separator between characters of a classified script, and a name it divides keeps its source order and is never segmented. It divides only between classified-script characters, so the Catalan punt volat in
Gal·lais untouched. The Japanese nakaguro is deliberately not a transcription marker, so高橋・一郎keeps its family-first reading (#298)Fix spaced CJK postnominal honorifics parsing as name parts. 씨, 박사, 선생님, 교수님, 군, 양, 先生, 女士, 小姐, 博士, 教授, 様 and 氏 now route to
suffix, so王小明 先生reads family王小明where the family-first default had made 先生 the given name. See East Asian names (closes #307)Fix glued CJK honorifics parsing as part of the name.
田中さん,김민준씨and王小明先生now split the honorific off the end of the name intosuffix. Previously it stayed in the name:田中さんwas entirely the family name, and김민준씨gave given 민준씨. The peel runs before the name is split or ordered, so김민준씨still divides into family 김, given 민준. Only entries that can never end a name peel; 양, 군, 氏, 博士 and 殿 are recognized in their spaced form only, since 김지양 is a given name and some ninety Japanese surnames end in 殿. Default-on, and a lone family name written with a glued honorific now divides where it did not. See East Asian names for the full set and Customizing the parser for the off-switch (closes #308)Fix a comma or a 间隔号 stopping the glued-honorific peel.
김, 민준씨now gives family 김, given 민준, suffix 씨, the same as the spaced김 민준씨, and likewise田中, 太郎さん. Previously each left the honorific inside the name. Both marks say where a name divides into surname and given, and an honorific is not part of the name in either reading. What a comma does instead is say which runs to look in: the two around a family comma. Anything past those is out of reach, so김, 민준 지훈씨peels while김, 민준, 지훈씨does not. Default-on. See East Asian names (closes #312)Fix a glued honorific staying inside the name when the whole post-comma remainder is a credential.
田中さん, V.and田中さん, Ph. D.now give up さん tosuffixthe way田中さん, PhDalready did. The peel had been taking the post-comma run for name text on the strength of the comma alone, walking into the credentials and abandoning the peel there. That run is now tested with the same rule that decides the comma structure, and declined where it is credentials, provided the part before the comma offers a peel site of its own. Where the credential itself lands is still the comma’s business and still differs by spelling. Default-on (closes #319)Fix an ASCII period after a CJK honorific stopping it being recognized.
씨.,様.,氏.,님.,군.,양.and殿.now route tosuffixlike their periodless spellings, where the trailing period had left them inside the name. The cause was v1’s initial regex, whose\wis Unicode-aware and matched a hangul syllable or Han ideograph as readily as a letter; a veto written for Latin was being asked of scripts it was never about. Alphabets keep their initials untouched, and"А. С. Пушкин"is unaffected. Read period strictly: only the ASCII full stop is covered, so"김민준 씨."written with the fullwidth stop still reads the honorific as the family name (superseded in 2.3.0, above). Default-on (#320)Fix NFD-decomposed input missing the East Asian defaults entirely. Script classification now normalizes to NFC before deciding, so a Korean or Japanese name typed on macOS, where decomposed text is routine, gets the same order rule as its composed twin. Segmentation matching deliberately stays raw, so an unspaced NFD hangul name is ordered correctly but not split, rather than being split in the wrong place. One gotcha: parse output preserves the encoding it was given, so for NFD input
name.family == "김"isFalseeven though it is the same name. Compare NFC-normalized text when comparing across encodings. See East Asian names (#272)Fix the Ukrainian conjunction й not joining the pieces around it. It is the euphonic alternate of
і, chosen by the surrounding sounds rather than by meaning, so real Ukrainian data carries both spellings."Олесь й Олена Коваленки"now gives given"Олесь й Олена"where theйpreviously landed inmiddle. Same treatment as theи/іentries added in 2.0.0: the conjunction joins only once the name has enough pieces, and a punctuated initial still wins, so"Й. Сліпий"is unaffected. Raised in a comment on #267Add the Japanese maiden-name marker 旧姓 to the default vocabulary.
"山田花子 旧姓 佐藤"now gives family山田花子and maiden佐藤, where 1.4.0 left the marker in the name. It sits beside the Cyrillicурожд.and Germangeb.entries rather than inlocales.JA, on the rule that admitted those: a native-script marker cannot collide with a Latin-script name, so it is safe as a default. Matching is whole-token, so the marker has to be a token, which for Japanese means a space or a configured delimiter must divide it from the name. The fullwidth colon does not, so"山田(旧姓:佐藤)"still returns maiden"旧姓:佐藤"; that one wants the head-peel #317 tracks. Default-on. See Customizing the parser (#309)Fix a maiden marker inside bracketed content staying in the maiden value. Where a delimiter pair is routed to
maidenbyPolicy(maiden_delimiters=...), a marker at the head of the clause is now dropped the way it always has been in the bare form, so"Jane Smith (née Jones)"gives maidenJones, the same answer as the unbracketed spelling. Each clause loses its own leading marker, and only where the clause holds more than one token, so"Jane Smith (Nee) (Jones)"still gives maidenNee Jones(Neeis a real surname). Scope before you count on it:Policy.maiden_delimitersis empty by default, so under the default policy brackets route tonicknameand none of this applies. See Customizing the parser (closes #329)
Documentation
Correct the documented scope of the period-abbreviation title rule. It was described as applying to “a leading word”, which was never true of any comma form: the rule runs at the front of the part that carries the given name, which after a family comma is the part after the comma, so
"Morse, Det. Insp. Jane"gives titleDet. Insp.. Behavior is unchanged and matches 1.4.0; only the description was wrong
2.0.0 - July 27, 2026
Two release candidates preceded this release (rc1 on 2026-07-23, rc2 on 2026-07-26). Please report anything the migration missed on issue #284. The notes below describe 2.0.0 as a whole.
nameparser 2.0 adds a new parsing API alongside
HumanName.parse()returns an immutableParsedNamewhose seven fields are named for what they are (given,family) rather than where they sit in a Western name, configured by two frozen value objects – aLexiconof vocabulary and aPolicyof behavior – instead of a mutable global.HumanNamekeeps working: it is now a compatibility facade over the same pipeline, and it stays through 2.x. The removals below are the deprecations announced in 1.3.0 and 1.4.0 coming due; if your code runs warning-free on 1.4.0, most of them will not affect you. See Migrating from HumanName for the field-by-field map.The 2.0 API
Add parse(text), returning an immutable ParsedName with seven fields –
title,given,middle,family,suffix,nickname,maiden– plus the derived viewsgiven_names,surnames,family_baseandfamily_particles. Parsing is a pure function of the text, aLexiconand aPolicy: nothing global is consulted and nothing is mutatedAdd Lexicon, the frozen vocabulary object.
Lexicon.default()is the shipped vocabulary andLexicon.empty()is a blank one;add(**entries),remove(**entries)and field-wise|all return new instances. Its fields aretitles,given_name_titles,suffix_acronyms,suffix_words,suffix_acronyms_ambiguous,particles,particles_ambiguous,conjunctions,bound_given_names,maiden_markersandcapitalization_exceptionsAdd
Policy, the frozen behavior object:name_order,patronymic_rules,middle_as_family,nickname_delimiters,maiden_delimiters,extra_suffix_delimiters,lenient_comma_suffixes,strip_emojiandstrip_bidiAdd
Parser, a reusable parser bound to a lexicon and policy (Parser(lexicon=..., policy=...).parse(text)), andparser_for(*locales, base=None), which folds locale packs onto a base parserAdd
name_orderand the constantsGIVEN_FIRST,FAMILY_FIRSTandFAMILY_FIRST_GIVEN_LAST, so a family-first name can be parsed as written rather than reordered by hand (the configuration half of #270; the locale packs below complete it)Add
PatronymicRulewith the membersEAST_SLAVICandTURKIC. v1’s singlepatronymic_name_orderflag enabled both detectors at once;Policy(patronymic_rules=...)lets you enable either one aloneAdd PolicyPatch and the UNSET sentinel for partial policy deltas that compose – set-valued fields union, scalar fields override with later winning. This is the mechanism locale packs are built from, and
UNSETis only needed when you must distinguish “not set” from a realFalseorNoneAdd Token, Span and Role: every field is backed by tokens carrying exact
(start, end)offsets into the original string, reachable withtokens_for(Role.GIVEN). This replaces v1’s*_listattributes and makes it possible to highlight or re-slice the input the parse came from.Roleis aStrEnum, so members compare and stringify as their field names (Role.GIVEN == "given"), matchingAmbiguityKind;tokens_for()accepts a role’s string name too, and raisesValueErrorfor an unknown roleAdd
STABLE_TAGSto the public API: the four documentedToken.tagsvalues (particle,conjunction,initial,joined)Add
Policy.patched(patch), applying aPolicyPatchdirectly without wrapping it in a locale packAdd Parser.matches(a, b) and Parser.capitalized(name): the
ParsedNamemethods of the same names fall back to the default configuration for str/omitted arguments, which is silently wrong for names parsed with a customParserAdd
Parser.revise(name, **fields):ParsedName.replace()with the replacement text classified by the parser’s vocabulary, so particle/initial/suffix-join behavior survives the editAdd Ambiguity and the AmbiguityKind enum, so a parse reports what it had to guess at instead of guessing silently. The kinds emitted today are
particle-or-given,suffix-or-name,suffix-or-nickname,unbalanced-delimiterandcomma-structure;orderis reserved and not yet emitted. The two suffix kinds cover the post-nominals that are also ordinary words:"John Smith MA"reports thatMAwas read as a credential rather than a surname, and"JEFFREY (JD) BRICKEN"that the delimitedJDwas read as a nickname rather than a suffix. A reading the vocabulary settles on its own —"John Smith M.A.","Andrew Perkins (MBA)"— is not a guess and reports nothingAdd ParsedName output and comparison methods:
render(spec),initials(),capitalized()(which returns a new value rather than mutating in place),as_dict()(whoseinclude_emptyflag is keyword-only, unlikeHumanName.as_dict()’s),replace(**fields),matches()andcomparison_key()Add
Locale, the public pack type. Writing your own needs no registration – construct aLocaleand pass it toparser_for()Ship a fully typed public API (PEP 561): the core modules are checked under strict mypy settings, and nameparser 2.0 has no runtime dependencies
Breaking Changes
Raise the minimum Python to 3.11 and drop the last runtime dependency,
typing_extensions(#257). Python 3.10 is no longer supportedRemove HumanName.__eq__ and __hash__ (deprecated in 1.3.0, #223): instances now compare and hash by identity, so
HumanName("John Smith") == "John Smith"isFalsewhere 1.x returnedTrue. This changes result silently rather than raising – it is the one removal that can pass unnoticed into production. Usematches()to ask whether two names are the same, andcomparison_key()as a dict key or sort keyRemove every v1 parsing hook and subclass extension point:
pre_process,post_process,parse_pieces,parse_nicknames,join_on_conjunctions, theis_*predicates,cap_word/cap_piece,handle_firstnames,fix_phdand the rest. A subclass that overrides one gets aDeprecationWarningat construction naming the hooks, because the facade delegates to the coreParserand never calls them (closes #280). Customize throughLexicon/PolicyinsteadRemove regex configuration:
CONSTANTS.regexesis now a read-only proxy. Reads still work, butCONSTANTS.regexes.bidi = False, item assignment andConstants(regexes=...)all raiseTypeError. If you followed 1.3.1’s advice to keep bidi marks withCONSTANTS.regexes.bidi = False, that opt-out is nowPolicy(strip_bidi=False)on the 2.0 API; the same applies toregexes.emojiandPolicy(strip_emoji=False)Remove bytes input and the encoding argument (#245): passing
bytestoHumanNameor to a set manager raisesTypeErrorwith a decode hint, andSetManager.add_with_encoding()andDEFAULT_ENCODINGare gone. Decode first, then useadd()Remove SetManager.__call__ (#243); iterate the manager or call set(manager).
remove()of a missing member now raisesKeyErrorlikeset.remove–discard()is the ignore-missing form – and the set operators|,&,-and^return a plainsetRemove HumanName slice access and item assignment (#258):
name[1:-3]andname['first'] = valueraiseTypeError. String-key reads (name['first']) and iteration are unchanged; assign fields as plain attributesRemove
Constants.empty_attribute_default(#255): empty fields are always''. Assigning it raisesAttributeError; a pickle carrying the key still loads, with the value ignoredRemove
constants=None(#261): bothHumanName(..., constants=None)andhn.C = NoneraiseTypeError. UseConstants()for library defaults orCONSTANTS.copy()for a private snapshotRemove silent unknown-key access on the mapping managers (#256):
CONSTANTS.capitalization_exceptions.typoandCONSTANTS.regexes.typoraiseAttributeErrornaming the miss, instead of returningNone/EMPTY_REGEX(capitalization_exceptionsalso lists the known keys)..get()remains available on both for intentional soft accessRemove support for
Constantspickles written by nameparser 1.2.x or earlier (#279): loading one raisesValueErrortelling you to re-pickle under 1.3/1.4 firstRemove the dead
regexes.no_vowelspattern (#268) and theConstants.suffixes_prefixes_titlescached union; neither was read by the parserChange
HumanName’s*_listattributes to read-only properties: they remain readable snapshots, butname.first_list = [...]now raisesAttributeErrorChange the
Constantsconstructor to keyword-only; positional construction no longer works
Behavior Changes
Recognize maiden-name markers –
née,nee,geb.,roz.and the Scandinavian participle forms – and route the following name to the newmaidenfield. 1.x folded them intomiddle/last, so"Jane Smith née Jones"parsed asmiddle="Smith née"(closes #274)Recognize typographic nickname delimiters by default in both APIs (closes #273): smart quotes (
“Jack”), German and Polish low-high quotes („Hansi“), Swedish right-right quotes (”Ann”), guillemets in either direction («Petit»,»Hansi«), CJK corner brackets (「タロ」,『ハナ』) and fullwidth parentheses. In 1.x these leaked intomiddleas literal text. Curly single quotes stay excluded, because U+2019 is the apostrophe in “O’Connor”Fix the pre-comma piece being routed to first when everything after the comma is a suffix or title:
"Andrews, M.D."now reads familyAndrews/ suffixM.D.where 1.x read givenM.D./ familyAndrews, and"Smith, Dr."movesSmithfromfirsttofamily(the title was already correct in 1.x). The piece before a comma is definitionally the family nameFix a lone recognized trailing suffix with no comma being routed to first/last:
"Johnson PhD"and"Mr. Johnson PhD"now keep the suffix insuffixFix a split “Ph. D.” credential being read as two tokens; it now classifies as one suffix, replacing v1’s fix_phd hook. This now holds wherever the credential sits: 1.x healed the pair only when it trailed, so
"Ph. D. John Smith"parsed as titlePh./ givenD.with the real given name pushed tomiddle; it now reads givenJohn, familySmith, suffixPh. D.Fix chargé d’affaires: shipped as one unmatchable
TITLESentry since it was added, it is now two chainable entries (chargé,d'affaires), so the title is recognized; the unaccented spellingchargeships too, likeattaché/attacheRemove seven multi-word
SUFFIX_ACRONYMSentries (leed ap,nicet i–nicet iv,psm i,psm ii) that could never match in any release; splitting them would swallow real names (“John Leed”, “Smith, A.P.”), so they are dropped insteadAdd a UserWarning when a multi-word entry is stored in a per-word Lexicon field or as a capitalization_exceptions key – such entries can never match
Parse an input with no alphanumeric character to an empty name in both APIs. 1.x kept pure punctuation as a name part, so
"."gavefirst="."andbool()wasTrue; 2.0 empties it, keepingbool(parse(x))an honest “did I get a name?” test. The check is Unicode-aware, so names in any script are unaffected; only inputs that are entirely punctuation or symbols (".","- -") change. Junk embedded in a name with real content – the stray dot in"John . Smith"– is still kept, since that parse is already truthyFold a leading never-given particle into the family name. Note that
Lexicon.particles_ambiguousis the complement of v1’snon_first_name_prefixes, not a rename – it lists the particles that may double as a given name, where v1 listed the ones that may not. Copying a v1 customization across without inverting it silently reverses the behavior; see Migrating from HumanNameAdd ma and do to suffix_acronyms_ambiguous, the set of post-nominals that are also ordinary surnames. An entry there is read as a credential when the name can spare it — written with periods (
"John Smith M.A."), or when removing it still leaves a given and a family name ("John Smith MA"→ suffixMA). With only two pieces to go around, the surname reading wins instead:"Jack Ma"and"Anh Do"keep their family names. As a side effect, a parenthesized or quoted"(MA)"/"(DO)"now falls through to nickname parsing rather than escaping tosuffix, since inside delimiters the nickname reading is the plausible oneChange comparison_key() and matches() in both APIs to fold with str.casefold() where 1.4 used str.lower(), so Unicode case pairs compare equal –
"STRASSE"matches"Straße", and a Greek final sigma matches its regular form. This is strictly more permissive: anything 1.4 matched still matches. Vocabulary normalization deliberately still useslower(), for v1 parityChange the 2.0 API’s default render()/str() spec to show every non-empty field:
'{title} {given} "{nickname}" {middle} {family} ({maiden}) {suffix}'. The quoted nickname round-trips exactly; the parenthesized maiden re-parses as a nickname, a deliberate choice of presentation over lossless round-trip – usenée {maiden}in a custom spec if you need it to survive a reparse.HumanNamekeeps v1’s ownstring_formatdefault, unchangedChange delimiter-overlap precedence in the 2.0 API: a pair listed in
Policy.maiden_delimitersis dropped from the effective nickname set, soPolicy(maiden_delimiters={("(", ")")})alone routes parenthesized content tomaiden. The default nickname set is exported asDEFAULT_NICKNAME_DELIMITERS.HumanNamekeeps v1’s nickname-wins precedence, so no existing behavior changesChange suffix-delimiter rendering when a custom suffix delimiter is configured: for
suffix_delimiter="/"and"John Smith, RN/CRNA", 1.x split the token and renderedsuffix="RN, CRNA"where 2.0 keeps it whole as"RN/CRNA". Role assignment is unchanged; only rendering differsCorrect a long-standing typo in the shipped vocabulary that 1.x carried:
actorandtelevisionhad trailing spaces inTITLES, so"actor" in TITLESwasFalse. Parse output is unchanged – the parser normalizes entries on ingest, which is exactly what let the typo go unnoticed – but code that tests membership against the exported constants directly will see corrected results. The data modules now assert their invariants at import time, alongside the onesprefixes.pyalready checked; entries must be stored lowercase and whitespace-free, so this class of typo now fails the build
International name support
Add nameparser.locales with the first two packs:
locales.RU(East Slavic patronymic order) andlocales.TR_AZ(Turkic patronymic markers). Packs are pure data folded in atparser_for(locales.RU), they compose (parser_for(locales.RU, locales.TR_AZ)unions the rules), and they are never auto-detected – there is no reliable way to detect a name’s language from the name alone.locales.available()andlocales.get(code)look packs up by code; loading is lazy, so importing the package imports no pack (completes #270; #271/#272/#146 stay staged for 2.x)Add non-Latin vocabulary to the default lexicon (#269): Cyrillic, Greek, Arabic and Hebrew titles, conjunctions and name particles. Native-script entries cannot collide with Latin-script names, which is what makes them safe to enable by default. Deferred pending vetting: the Cyrillic
мл/стsuffixes and the bare Greekκtitle, which collides with the initial-plus-surname shape. Behavior note:محمد بن سلمانnow chainsبنonto the family name where 1.x read it as a middle nameAdd Arabic-script bound given names (#269):
عبد, the kunya pairأبو/ابو, andأم/امjoin the following word into the given name exactly as their transliterations (abdul,abu,umm) always did, soعبد الرحمن محمدparses givenعبد الرحمن, familyمحمدwhere 1.x split it into given plus middleAdd Arabic honorific titles and the conjunction و (#269): the doctor, professor, hajj, sheikha and engineer forms (
الدكتور/الدكتورة/دكتور/دكتورة,الأستاذ/الأستاذة/أستاذ/أستاذة,الحاج/الحاجة,الشيخة,مهندس) as given-name titles, since Arabic honorifics precede the given name likeالشيخ. Deferred under the collision rule: bareسيد/شيخ/أمير/سلطان(all common given names), theد.abbreviation (bareدwould swallow initials), and the Ottoman post-nominalsباشا/بك/أفندي(which survive as family names)Add Hebrew honorifics and post-nominals, and Devanagari titles (#269): the Israeli honorifics
גברת,פרופ'/פרופ׳,פרופסור,עו"ד/עו״דandהרבas titles;ז"ל/ז״לandשליט"א/שליט״אas suffixes, in both gershayim spellings; and Devanagariश्री,श्रीमतीandडॉ. Latinsri/shriwere deliberately not added, because they collide with real given names where the native script cannot. Deferred: bareרב(an ordinary word meaning “many”) andברas a particle (Bar is a common modern given name)
Compatibility layer
Reimplement HumanName as a facade over the 2.0 pipeline, and Constants as a shim resolving to a (Lexicon, Policy) snapshot with a shared parser cache. Fields, aggregates, mutation through
name.C.titles.add(...), rendering defaults,capitalize(),matches(),comparison_key(), iteration,as_dict()and pickling are all preserved, andnameparser.parserandnameparser.configremain importable. The compatibility layer ships through 2.x and is removed in 3.0Note that CONSTANTS.capitalize_name and force_mixed_case_capitalization are still honored through the facade, but the 2.0 API never capitalizes during parse() – call
capitalized()when you want itAdd
Rolemembers (and their string valuesgiven/family) as validHumanNamesubscript keys:hn[Role.GIVEN]returnshn.firstAdd a UserWarning when assigning HumanName.given or .family: the facade spells those attributes
first/last, so the assignment creates an inert stray attribute while the parse keeps the old value. The assignment still happens (v1-legal ad-hoc attributes keep working); the warning names the v1 spelling to use
Command line
Rewrite python -m nameparser over the 2.0 API with a real argument parser. It prints the parse plus its capitalized form and initials by default, takes
--jsonto emit the fields as JSON (python -m nameparser --json "Doe, John"), takes--locale CODEto apply a pack, and exits with a usage message on bad arguments
Documentation
Rewrite the documentation new-API-first: a new front page and README,
usage.rstas a tour of the 2.0 API, a principle-firstcustomize.rst, a reference split into the 2.0 API and the compatibility layer, and new pages for How the parser works, Locale packs and Migrating from HumanName – the last carrying full attribute and configuration maps from v1 names to 2.0 names (#262). The 1.x documentation remains online as the readthedocsstablebuild
Parsing changes were checked against a 652-name differential corpus – names harvested from the v1 test banks, plus names reported in the issue tracker – and everything not listed above parses identically between 1.4.0 and 2.0. The harness lives in
tools/differential/in the source repository (it is development tooling, not part of the installed package) – see its README to reproduce the comparison against your own names.Changed since 2.0.0rc1 (for anyone who tested the release candidate)
Rolebecame aStrEnum;str(Role.GIVEN)is now"given"ParsedName.tokens_for()raisesValueErrorfor unknown roles instead of returning no tokens; it also accepts role-name stringsParsedName.as_dict()’sinclude_emptyis keyword-onlyHumanNamesubscripting acceptsRolemembersAssigning
HumanName.given/.familywarns (the facade spells themfirst/last; the assignment was and remains an inert stray attribute)The eight multi-word vocabulary entries that could never match were repaired (
chargé d'affairessplit; seven credential acronyms removed), and storing a new multi-word entry now warns; those eight are dropped silently, not warned about, when a restoredConstantspickle carries all eight of them (the pre-2.0 signature)PolicyPatch’s repr shows only the fields a patch sets, instead of all nine with UNSET sentinelsNew since rc1:
STABLE_TAGS,Policy.patched(),Parser.matches(),Parser.capitalized(), andParser.revise()– see the API section above
1.4.0 - July 12, 2026
Add Constants.copy(), a detached deep copy that preserves the source instance’s current customizations (unlike Constants(), which always starts from library defaults) – useful as
CONSTANTS.copy()for a private snapshot of the shared config (#260)Deprecate passing constants=None to HumanName (or assigning hn.C = None): it silently builds a fresh
Constants(), discarding any customizations the caller may have expected to carry over from the sharedCONSTANTS. EmitsDeprecationWarning; will raiseTypeErrorin 2.0. Useconstants=Constants()for fresh library defaults orconstants=CONSTANTS.copy()for a private snapshot instead (closes #260)Deprecate assigning Constants.empty_attribute_default for removal in 2.0 (#255): once
Nonesupport goes, the only legal value left is the default'', so a dial with one position isn’t configuration. EmitsDeprecationWarning; reading the attribute is unaffectedDeprecate unknown-key attribute access on TupleManager/RegexTupleManager (CONSTANTS.regexes.typo, CONSTANTS.capitalization_exceptions.typo, etc.) for removal in 2.0 (#256): a misspelled or omitted key currently degrades silently (
None/EMPTY_REGEX) with no traceback pointing at the typo. EmitsDeprecationWarningnaming the miss and the known keys; will raiseAttributeErrorin 2.0..get()remains available for intentional soft accessDeprecate HumanName slice access (name[1:-3]) and item assignment (name[‘first’] = value) for removal in 2.0 (#258): field access by position has no real use case, and item assignment duplicates plain attribute assignment. Both emit
DeprecationWarning; string-key access (name['first']) is unaffectedDeprecate SetManager.add_with_encoding() itself for removal in 2.0 (#245), regardless of argument type: use
add()instead (decoding bytes first). Previously only thebytespath warned; thestrpath was silent even though the whole method goes awayDeprecate loading a legacy-format Constants pickle (written by nameparser <= 1.2.x, before the 1.3.0 pickle fix) for removal in 2.0 (#279):
__setstate__’s migration shim currently skips the stale computed-property key silently. EmitsDeprecationWarningonce per call telling users to re-pickle; will raiseValueErrorin 2.0Fix the
"Lastname, Firstname"comma format not being recognized when the input uses the Arabic comma،(U+060C, the standard comma in Arabic/Persian/Urdu text) or the fullwidth CJK comma,(U+FF0C) instead of the ASCII comma: both variants now also split the format and no longer leak into the parsed output (closes #265)
1.3.1 - July 11, 2026
Fix invisible Unicode bidirectional control characters (LRM/RLM/ALM, the embedding/override marks, and the isolates U+2066–U+2069) surviving parsing and sticking to
first/last/etc., so a copy-pasted right-to-left name silently failed equality and dedup. They are now stripped in preprocessing like emoji; disable viaCONSTANTS.regexes.bidi = False(closes #266)Fix str() corrupting name text containing the substring “None” when empty_attribute_default is None (e.g. “Nonez Smith” rendered as “z Smith”): empty attributes are now substituted as
''before the format string is applied, instead of scrubbing the interpolated"None"from the output afterward (closes #254)
1.3.0 - July 5, 2026
Breaking Changes & Deprecations
Deprecate HumanName.__eq__ and __hash__ for removal in 2.0 (#223): the current design’s three promises — case-insensitive equality, equality with plain strings, and hashability — are mutually inconsistent (equal objects can hash differently), equality depends on
string_format, andmaidenis invisible to it. Both now emitDeprecationWarningnaming the replacement; behavior is otherwise unchanged until 2.0 (closes #224)Deprecate bytes input for removal in 2.0 (#245): passing
bytestoHumanName/full_nameor toSetManager.add()/add_with_encoding()now emitsDeprecationWarning— decode first, e.g.value.decode('utf-8'). Theencodingconstructor argument is deprecated with itDeprecate SetManager.__call__ for removal in 2.0 (#243): calling a manager returns the raw underlying set, so mutating the result bypasses normalization and cache invalidation; iterate the manager or copy with
set(manager)insteadAdd SetManager.discard(), and deprecate remove() of a missing member (#243): it currently does nothing but will raise
KeyErrorin 2.0, matchingset.remove; usediscard()for intentional ignore-missing removal. Removing present members is unchanged and does not warnFix HumanName acting as its own iterator with a stored cursor: breaking out of a loop, iterating in nested loops, or calling
len(name)mid-loop corrupted subsequent iteration;iter(name)now returns a fresh independent iterator each time.next(name)on the instance itself (undocumented) now raisesTypeError— callnext(iter(name))instead (closes #225)Remove the vestigial unparsable attribute: the guard that was meant to set it has been unreachable since 2013 (v0.2.9), so it has reported
Falsefor every parsed name for over a decade; checklen(name) == 0to detect an empty parseRemove
__ne__; Python 3 derives!=from__eq__automaticallyChange internal initials helper
__process_initial__to_process_initial: double-underscore-both-sides names are reserved for Python special methods; subclasses overriding the old name must rename their overrideChange
REGEXESfrom asetof(name, pattern)tuples to adict, so a duplicate name is a visible overwrite in the source instead of a nondeterministic winner at import time; code iteratingREGEXESdirectly now gets keys instead of pairs — use.items()(#227)Change
CAPITALIZATION_EXCEPTIONSfrom a tuple of(key, value)tuples to adict; code iterating it directly now gets keys instead of pairs — use.items()(#233)
Behavior Changes (affect existing parse output)
Add
bound_first_namesset toConstants; bound Arabic given-name prefixes (abdul,abu, etc.) now join forward to form a single first name (e.g."abdul salam ahmed salem"→first="abdul salam",middle="ahmed",last="salem"). Disable viaCONSTANTS.bound_first_names.clear(). Default-on: changes parsing output for names with these prefixes. (#150)Treat an unrecognized, multi-letter token ending in a period in the leading title run (before the first name is set), e.g.
"Major.", as atitleinstead of afirstname; internal-period abbreviations ("E.T.") and single-letter initials ("J.") are unaffected. Default-on: changes parsing of names with a leading unknown period-abbreviation (closes #109)Fix parsing writing back into the Constants it reads (usually the shared module-level CONSTANTS): pieces derived while parsing a name — period-joined titles/suffixes like
"Lt.Gov."and conjunction-joined pieces like"Mr. and Mrs."or"von und zu"— are now tracked per parse instead of being permanentlyadd()-ed to the config, so parse results no longer depend on which names were parsed earlier in the process and parsing no longer mutates shared state across threadsFix
__hash__to lowercase the name like__eq__does, so equalHumanNameinstances hash equal and behave correctly in sets and dicts
New comparison methods
Add matches() and comparison_key() for explicit name comparison:
matches()compares parsed components case-insensitively (parsingstrarguments first, soname.matches("Smith, John")andname.matches("John Smith")both match) andcomparison_key()returns a hashable tuple of the seven components for dedup, dict keys, and sorting (#224)
New name fields
Add a first-class
maidenfield andmaiden_delimiterstoConstants, so a delimiter (e.g. parenthesis) can be routed tomaideninstead ofnicknamefor alternate/maiden surnames, e.g."Baker (Johnson), Jenny"(closes #22)Add
given_names(andgiven_names_list) attribute as aggregate of first and middle names, mirroringsurnames(closes #157)Add
last_base,last_prefixes(and_listvariants) for splitting last-name prefix particles (tussenvoegsels) from the core surname (#130, #132)
New customization options
Add
initials_separatortoConstantsandHumanNameto control spacing between consecutive initials within a name group (#171)Add
suffix_delimitertoConstantsandHumanNamefor parsing suffixes separated by arbitrary delimiters, e.g."RN - CRNA"(#156)Add
nickname_delimiterstoConstantsfor registering additional nickname-delimiter regex patterns at runtime, without subclassing (closes #110, #112)Add
suffix_acronyms_ambiguoustoConstantsfor acronym suffixes that also read as given-name nicknames (e.g."JD","Ed"), used when disambiguating parenthesized/quoted content (#111)
International name support
Add
patronymic_name_orderflag toConstantsandHumanNamefor opt-in detection and reordering of Russian formal-order names (Surname GivenName Patronymic) (#85)Add Turkic (Azerbaijani/Central-Asian) patronymic detection to
patronymic_name_order, rotating the reversed 4-token formal shape (Surname GivenName PatronymicRoot Marker, e.g.oglu/qizi) into Western order (#185)Add
middle_name_as_lastflag toConstantsandHumanNamefor opt-in folding of middle names into the last name, for naming systems with no middle-name concept (e.g. Arabic patronymic chaining) (#133)Add non_first_name_prefixes to Constants: a leading particle that is never a first name (e.g.
"de Mesnil","dos Santos") now parses as a surname with an empty first name, instead of treating the particle as the first name (closes #121)Add international honorifics to
TITLES(#187)Add German/Austrian nobility and ecclesiastical titles to
TITLES(closes #101)Add German/Dutch last-name prefixes and title/degree suffixes; fix
join_on_conjunctions()to register multi-word prefix chains (e.g."von und zu") as prefixes, mirroring existing title handling (closes #18)
Parsing fixes
Fix suffix boundary lookup for prefixed last names with a title before and after (e.g.
"dr Vincent van Gogh dr"producing a corrupted middle name) (closes #100)Fix a repeated prefix word in a prefix chain (e.g. “Juan de la de la Vega”) silently dropping the earlier occurrence in join_on_conjunctions(): value-based
pieces.index(prefix)lookups re-found the wrong occurrence once the list had already been mutated by prior joins; prefix positions are now tracked positionally instead of re-derived by value (closes #208)Fix a trailing suffix being silently dropped after an empty comma segment, e.g.
"Doe, John,, Jr."losing the"Jr."Fix degenerate comma input (a bare
","or an empty comma segment, e.g."Doe,, Jr.","John Doe, Jr.,,") leaving an empty-string member infirst_list,last_list, orsuffix_list; whitespace-only tokens assigned via the setters are dropped the same wayFix suffix-shaped parenthesized/quoted content (e.g.
"(Ret)","(MBA)") being misclassified as a nickname instead of a suffix (closes #111)Fix single-character symbol conjunctions (e.g.
"&","/") being ignored in short names (#173)Fix recognition of single-letter roman numeral suffixes (e.g.
"I","V") in suffix-comma format (closes #136)Fix recognition of trailing
suffix_not_acronyms(e.g."Jr.") in lastname-comma format (closes #144)Fix missing comma between
'msc'and'mscmsm'insuffix_acronyms, which silently concatenated them into a bogus'mscmscmsm'entry (#111)Fix
'apn aprn'split into separatesuffix_acronymsentries so each is recognized independently (closes #155)
Formatting and output fixes
Fix
IndexErrorininitials()/initials_list()when a*_listattribute was assigned directly with an element containing unnormalized whitespace (e.g.name.middle_list = ['Q R']), bypassing the parser’s whitespace normalization (closes #232)Fix initials() emitting a stray empty initial (e.g. “J. . V.”) – or raising
TypeErrorwhenempty_attribute_defaultisNone– for name parts with no initialable words, e.g. a prefix-only middle name like"de la"Fix capitalization of suffix acronyms written with dots, e.g.
"M.D."(closes #141)Fix extra whitespace before punctuation in
str()output when astring_formatfield is empty (closes #139)Fix spurious leading space in surnames and empty token in suffix list after
capitalize()with an empty middle or suffix (#164)
API correctness and cleanup
Fix the five non-cached-union
SetManager-backedConstantsattributes (first_name_titles,conjunctions,bound_first_names,non_first_name_prefixes,suffix_acronyms_ambiguous) accepting non-SetManagerassignment silently (e.g.constants.conjunctions = 'and'), degrading membership checks into substring tests with no error; assignment now raisesTypeErrorlike the four cached-union attributes already did (closes #241)Fix
HumanName.Caccepting an invalidconstantsvalue on post-construction assignment (e.g.hn.C = 'garbage'), bypassing the constructor’s validation and failing later with an unrelatedAttributeError;Cis now a property that validates on assignment too (closes #239)Fix
TupleManager(andRegexTupleManager) accepting a bare string/bytes argument (raising a crypticdict-internalsValueError) or an iterable of 2-character strings (silently shredding each into a key/value pair, e.g.Constants(capitalization_exceptions=['ii'])becoming{'i': 'i'}); both now raiseTypeErrorwith a clear message (closes #242)Fix
SetManager.__contains__being the one operation that didn’t normalize (lowercase, strip leading/trailing periods) its operand, so e.g.'Dr.' in constants.titlescould returnFalseeven though the title was correctly configured; membership checks now normalize likeadd()/remove()/the constructor/the set operators (closes #244)Fix a bare string passed to a set-backed
Constantsargument (e.g.Constants(titles='dr')), toSetManager, or as aSetManagerset-operator operand (e.g.constants.titles |= 'esq') being silently split into single characters, replacing or polluting the set and producing wrong parses with no error; it now raisesTypeErrorwith the suggested fix — wrap strings in a list, decodebytesfirst (closes #238)Fix SetManager set operators and the constructor skipping the lowercase/strip-edge-periods normalization that add() applies:
constants.titles |= ['Esq.']kept a raw'Esq.'the parser’s lookups could never match,titles & ['Dr.']missed'dr', andConstants(titles=[...])stored raw elements that silently never matched; elements and operands are now normalized everywhere, and non-strelements (bytes,None, numbers) raiseTypeErrorinstead of crashing cryptically or being coercedFix the constants constructor argument silently discarding Constants subclass instances: the exact-type check replaced them with fresh defaults, throwing away the caller’s configuration. Subclass instances are now used as given; anything that is neither
Nonenor aConstantsinstance now raisesTypeErrorinstead of being silently swapped for defaults (closes #226)Fix
Constantscustomizations, singleton identity, andTupleManagersubclass being lost acrosspickle/deepcopyround-trips (#167, #168, #169)Fix
is_rootname()returning stale results afteradd()/remove()ontitles,prefixes,suffix_acronyms, orsuffix_not_acronyms(#166)Fix the library logger calling
setLevel(logging.ERROR)on import, which silently discarded log records regardless of an application’s own logging configuration; the logger now leaves its level atNOTSETand lets the application control verbosity (closes #228)Minor internal cleanups: drop a dead length check in the initials helper, simplify double-wrapped
len(list(...))calls, and other small parser tidy-ups with no behavior change (closes #229)Change
Constants.__repr__to report collection sizes and non-default scalar config, replacing the uninformative<Constants() instance>(#221)
- 1.2.1 - June 19, 2026
Fix
initials()interpolating the literalNonefor empty name parts whenempty_attribute_default = None(e.g."J. None D."); empty parts now render as an empty string and a fully-empty result returnsempty_attribute_defaultAdd
python -m nameparser "Name String"command-line helper that prints a parsed nameReorganize the test suite from a single
tests.pyinto atests/pytest package
- 1.2.0 - June 11, 2026
Drop Python 2 and Python < 3.10 support; Python 3.10–3.14 now required
Add type hints and type declarations (PEP 561
py.typedmarker)Migrate build tooling to
pyproject.toml, dropsetup.pyRemove dead Python 2 compatibility shims (
ENCODINGconstant,next()aliases)Modernize CI: uv-based workflow, trusted publishing to PyPI, Dependabot
- 1.1.3 - September 20, 2023
Fix case when we have two same prefixes in the name ()#147)
- 1.1.2 - November 13, 2022
Add support for attributes in constructor (#140)
Make HumanName instances hashable (#138)
Update repr for names with single quotes (#137)
- 1.1.1 - January 28, 2022
Fix bug in is_suffix handling of lists (#129)
- 1.1.0 - January 3, 2022
Add initials support (#128)
Add more titles and prefixes (#120, #127, #128, #119)
- 1.0.6 - February 8, 2020
Fix Python 3.8 syntax error (#104)
- 1.0.5 - Dec 12, 2019
Fix suffix parsing bug in comma parts (#98)
Fix deprecation warning on Python 3.7 (#94)
Improved capitalization support of mixed case names (#90)
Remove “elder” from titles (#96)
Add post-nominal list from Wikipedia to suffixes (#93)
- 1.0.4 - June 26, 2019
Better nickname handling of multiple single quotes (#86)
full_name attribute now returns formatted string output instead of original string (#87)
- 1.0.3 - April 18, 2019
fix sys.stdin usage when stdin doesn’t exist (#82)
support for escaping log entry arguments (#84)
- 1.0.2 - Oct 26, 2018
Fix handling of only nickname and last name (#78)
- 1.0.1 - August 30, 2018
Fix overzealous regex for “Ph. D.” (#43)
Add surnames attribute as aggregate of middle and last names
- 1.0.0 - August 30, 2018
Fix support for nicknames in single quotes (#74)
Change prefix handling to support prefixes on first names (#60)
Fix prefix capitalization when not part of lastname (#70)
Handle erroneous space in “Ph. D.” (#43)
- 0.5.8 - August 19, 2018
Add “Junior” to suffixes (#76)
Add “dra” and “srta” to titles (#77)
- 0.5.7 - June 16, 2018
Fix doc link (#73)
Fix handling of “do” and “dos” Portuguese prefixes (#71, #72)
- 0.5.6 - January 15, 2018
Fix python version check (#64)
- 0.5.5 - January 10, 2018
Support J.D. as suffix and Wm. as title
- 0.5.4 - December 10, 2017
Add Dr to suffixes (#62)
Add the full set of Italian derivatives from “di” (#59)
Add parameter to specify the encoding of strings added to constants, use ‘UTF-8’ as fallback (#67)
Fix handling of names composed entirely of conjunctions (#66)
- 0.5.3 - June 27, 2017
Remove emojis from initial string by default with option to include emojis (#58)
- 0.5.2 - March 19, 2017
Added names scrapped from VIAF data, thanks daryanypl (#57)
- 0.5.1 - August 12, 2016
Fix error for names that end with conjunction (#54)
- 0.5.0 - August 4, 2016
Refactor join_on_conjunctions(), fix #53
- 0.4.1 - July 25, 2016
Remove “bishop” from titles because it also could be a first name
Fix handling of lastname prefixes with periods, e.g. “Jane St. John” (#50)
- 0.4.0 - June 2, 2016
Remove “CONSTANTS.suffixes”, replaced by “suffix_acronyms” and “suffix_not_acronyms” (#49)
Add “du” to prefixes
Add “sheikh” variations to titles
Add parameter to force capitalization of mixed case strings
- 0.3.16 - March 24, 2016
Clarify LGPL licence version (#47)
Skip pickle tests if pickle not installed (#48)
- 0.3.15 - March 21, 2016
Fix string format when empty_attribute_default = None (#45)
Include tests in release source tarball (#46)
- 0.3.14 - March 18, 2016
Add CONSTANTS.empty_attribute_default to customize value returned for empty attributes (#44)
- 0.3.13 - March 14, 2016
Improve string format handling (#41)
- 0.3.12 - March 13, 2016
Fix first name clash with suffixes (#42)
Fix encoding of constants added via the python shell
Add “MSC” to suffixes, fix #41
- 0.3.11 - October 17, 2015
Fix bug capitalization exceptions (#39)
- 0.3.10 - September 19, 2015
Fix encoding of byte strings on python 2.x (#37)
- 0.3.9 - September 5, 2015
Separate suffixes that are acronyms to handle periods differently, fixes #29, #21
Don’t find titles after first name is filled, fixes (#27)
Add “chair” titles (#37)
- 0.3.8 - September 2, 2015
Use regex to check for roman numerals at end of name (#36)
Add DVM to suffixes
- 0.3.7 - August 30, 2015
Speed improvement, 3x faster
Make HumanName instances pickleable
- 0.3.6 - August 6, 2015
Fix strings that start with conjunctions (#20)
handle assigning lists of names to a name attribute
support dictionary-like assignment of name attributes
- 0.3.5 - August 4, 2015
Fix handling of string encoding in python 2.x (#34)
Add support for dictionary key access, e.g. name[‘first’]
add ‘santa’ to prefixes, add ‘cpa’, ‘csm’, ‘phr’, ‘pmp’ to suffixes (#35)
Fix prefixes before multi-part last names (#23)
Fix capitalization bug (#30)
- 0.3.4 - March 1, 2015
Fix #24, handle first name also a prefix
Fix #26, last name comma format when lastname is also a title
- 0.3.3 - Aug 4, 2014
Allow suffixes to be chained (#8)
Handle trailing suffix in last name comma format (#3). Removes support for titles with periods but no spaces in them, e.g. “Lt.Gen.”. (#21)
- 0.3.2 - July 16, 2014
Retain original string in “original” attribute.
Collapse white space when using custom string format.
Fix #19, single comma name format may have trailing suffix
- 0.3.1 - July 5, 2014
Fix Pypi package, include new config module.
- 0.3.0 - July 4, 2014
Refactor configuration to simplify modifications to constants (backwards incompatible)
use unicode_literals to simplify Python 2 & 3 support.
Generate documentation using sphinx and host on readthedocs.
- 0.2.10 - May 6, 2014
If name is only a title and one part, assume it’s a last name instead of a first name, with exceptions for some titles like ‘Sir’. (#7).
Add some judicial and other common titles. (#9)
- 0.2.9 - Apr 1, 2014
Add a new nickname attribute containing anything in parenthesis or double quotes (Issue 33).
- 0.2.8 - Oct 25, 2013
Add support for Python 3.3+. Thanks to @corbinbs.
- 0.2.7 - Feb 13, 2013
Fix bug with multiple conjunctions in title
add legal and crown titles
- 0.2.6 - Feb 12, 2013
Fix python 2.6 import error on logging.NullHandler
- 0.2.5 - Feb 11, 2013
Set logging handler to NullHandler
Remove ‘ben’ from PREFIXES because it’s more common as a name than a prefix.
Deprecate BlankHumanNameError. Do not raise exceptions if full_name is empty string.
- 0.2.4 - Feb 10, 2013
Adjust logging, don’t set basicConfig. Fix Issue 10 and Issue 26.
Fix handling of single lower case initials that are also conjunctions, e.g. “john e smith”. Re Issue 11.
Fix handling of initials with no space separation, e.g. “E.T. Jones”. Fix #11.
Do not remove period from first name, when present.
Remove ‘e’ from PREFIXES because it is handled as a conjunction.
Python 2.7+ required to run the tests. Mark known failures.
tests/test.py can now take an optional name argument that will return repr() for that name.
0.2.3 - Fix overzealous “Mac” regex
0.2.2 - Fix parsing error
- 0.2.0
Significant refactor of parsing logic. Handle conjunctions and prefixes before parsing into attribute buckets.
Support attribute overriding by assignment.
Support multiple titles.
Lowercase titles constants to fix bug with comparison.
Move documentation to README.rst, add release log.
0.1.4 - Use set() in constants for improved speed. setuptools compatibility - sketerpot
0.1.3 - Add capitalization feature - twotwo
0.1.2 - Add slice support