Wikiwordbōc angwiktionary https://ang.wiktionary.org/wiki/H%C4%93afodtramet MediaWiki 1.47.0-wmf.21 case-sensitive Media Syndrig Mōtung Brūcend Brūcendmōtung Wikiwordbōc Wikiwordbōcmōtung Ymele Ymelmōtung MediaWiki MediaWikimōtung Bysen Bysenmōtung Help Helpmōtung Flocc Floccmōtung Ætēaca Ætēacmōtung TimedText TimedText talk Module Module talk Event Event talk geong 0 950 54838 52298 2026-09-26T21:54:26Z Deadend0914 7211 /* {{Wordcynn|Tōgeīecendlic}} */ 54838 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== *[[IPA]] /junɡ/, [juŋɡ] *[[SAMPA]]: /jung/, [juNg] ===Wordstǣr=== Ofcymþ þæt orgermanisce word jung. === {{Wordcynn|Tōgeīecendlic}} === {{Tógeíecendlic-Tabule| 1. Stæpe=geong |2. Stæpe=giengra |3. Stæpe=giengest |Othera_declinunga=geong (declīnung) }} #níwan geboren #: Ā sċyle '''ġeong''' mon wesan ġeōmormōd, heard heortan ġeþōht, swylċe habban sċeal blīþe ġebǣro, ēac þon brēostċeare, sinsorgna ġedreag sȳ æt him sylfum ġelong eal his worulde wyn, sȳ ful wīde fāh feorres folclondes, þæt mīn frēond siteð under stānhliþe, storme behrīmed, wine wēriġmōd, wætre beflōwen on drēorsele. #nā eald, nīwe [[Flocc:Englisc tōgeīecendlic]] [[Flocc:Unefenlic Tógeíecendlic]] ax7m63j2rlqewkg3121cm2p3vp3jq5t junio 0 1195 54850 51319 2026-09-26T23:20:34Z Deadend0914 7211 /* Nama */ 54850 wikitext text/x-wiki ==Spēonisc== ===Rihtstefn=== * IPA: [ˈxunjo] ===Nama=== {{es-noun|m}} # [[Sēremōnaþ]], se 6a mōnaþ þæs gēares ===Sēo ēac=== *[[mes]] [[Flocc:Spēonisce mōnþas]] 3rc7t17lvncwv10xib53mf1oj1us3nm declīnung 0 1249 54810 52203 2026-09-26T18:25:42Z Deadend0914 7211 /* Wendunga */ 54810 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== === {{f}} {{Wordcynn|Nama}} === {{Naman-Tabule| Hwā oþþe hwæt? (Ānfeald)=sēo declīnung |Hwā oþþe hwæt? (Manigfeald)=þā declīnunga |Hwæs? (Ānfeald)=þǣre declīnunge |Hwæs? (Manigfeald)=þāra declīnunga |Hwǣm? (Ānfeald)=þǣre declīnunge |Hwǣm? (Manigfeald)=þǣm declīnungum |Hwȳ? (Ānfeald)=þǣre declīnunge |Hwȳ? (Manigfeald)=þǣm declīnungum |Hwone? (Ānfeald)=þā declīnunge |Hwone? (Manigfeald)=þā declīnunga }} '''Declīnung''' ''wīflic'' #[[Gebīgednes]], sēo ansīen oþþe þā ansīene, þe word habbaþ tō tācnienne hira brūcunge in ferse. ==== Fruma ==== *Ofgangen of þǣm Lǣdenan worde '''[[declinare]]'''. ==== Wendunga ==== *[[Nīwenglisc]]: [[declension]], [[conjugation]] [[Category:Englisc nama]] [[Category:Englisc wīflic nama]] [[Category:Grammaticcræft]] riovy1j57eqapih4ol8ewgzy0l119pu 54811 54810 2026-09-26T18:26:26Z Deadend0914 7211 /* {{sprǣc|ang}} */ 54811 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== === {{f}} {{Wordcynn|Nama}} === {{Naman-Tabule| Hwā oþþe hwæt? (Ānfeald)=sēo declīnung |Hwā oþþe hwæt? (Manigfeald)=þā declīnunga |Hwæs? (Ānfeald)=þǣre declīnunge |Hwæs? (Manigfeald)=þāra declīnunga |Hwǣm? (Ānfeald)=þǣre declīnunge |Hwǣm? (Manigfeald)=þǣm declīnungum |Hwȳ? (Ānfeald)=þǣre declīnunge |Hwȳ? (Manigfeald)=þǣm declīnungum |Hwone? (Ānfeald)=þā declīnunge |Hwone? (Manigfeald)=þā declīnunga }} '''Declīnung''' ''wīflic'' #[[Gebīgednes]], sēo ansīen oþþe þā ansīene, þe word habbaþ tō tācnienne hira brūcunge in ferse. ==== Fruma ==== *Ofgangen of þǣm Lǣdenan worde '''[[declinare]]'''. ==== Wendunga ==== *[[Nīwenglisc]]: [[declension]], [[conjugation]] [[Category:Englisc nama]] [[Category:Englisc wīflic nama]] [[Category:Stæfcræft]] lvyneawroaxzyynlwo1pf6yjtt9vwmp dōn (Norþhymbrisc) 0 2039 54808 13596 2026-09-26T18:20:17Z Deadend0914 7211 /* Englisc */ 54808 wikitext text/x-wiki #redirect [[dōn]] 0t0e2lsaq5efudzeh3y49gof0stlrh2 dōn (Anglisc) 0 2040 54809 42919 2026-09-26T18:20:39Z Deadend0914 7211 Redirected page to [[dōn]] 54809 wikitext text/x-wiki #redirect [[dōn]] 0t0e2lsaq5efudzeh3y49gof0stlrh2 seal 0 2615 54864 52812 2026-09-26T23:37:21Z Deadend0914 7211 /* Nama */ 54864 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:// ===Nama=== {{en-noun}} #[[seolh]] (dēor) ====Fruma==== ====Wendunga==== [[Category:Nīwe Englisc nama]] [[Category:Nīwe Englisc dēor]] bkun7eu35oglx1fa6hnggrmbu67c26c child 0 2663 54849 52414 2026-09-26T23:15:20Z Deadend0914 7211 54849 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA: /tʃaɪld/ ===Wordstǣr=== * Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kiltham|kiltham]] ====Sweostorword==== * Niðerlendisc: [[kind]] * Þēodisc: [[Kind]] ===Nama=== # [[cild]], [[bearn]] #: In one case, a mother was deported and took her 2-year-old '''child''' with her, while the other involves another mother deported and her 4- and 7-year-old '''children''' went with her, the American Civil Liberties Union and the National Immigration Project, among other organizations, said in a news release Friday. #*:On anum gelimpe, mōdor wæs ādyden and tōc hire twīwintre '''bearn''' mid hire, hwíl þæt óðer belimpþ an ōþer mōdor ādyden and hire fēowerwintre and seofonwintre '''bearn''' ēodon mid hire, þæt American Civil Liberties Union and þæt National Immigration Project, hérongemong, sæġdon inn an gewrit Fríandæg # [[cnafa]] #: [[cnapa]] #: [[lȳtling]] #: [[magutimber]] jyltr28sjnkmqnmcscs5euecb5esx7n race 0 2892 54839 52777 2026-09-26T22:03:43Z Deadend0914 7211 /* {{sprǣc|en}} */ 54839 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:/reɪs/ ====Wordstǣr 1==== *Ofgangen of Middelengliscum [[race]], of Ealdum Frenciscum [[race]], of Ealdum Ītaliscum [[razza]] ===Nama=== '''race''' (mnf. races) #[[cynn]], [[cnōsl]] ====Word mid gelīcum Sweotolungum==== #[[progeny]] ====Wendunga==== Englisc: [[cynn]], [[cnōsl]] [[Category:Nīwe Englisc nama]] ====Wordstǣr 2=== *Ofgangen of Middelenglisce [[race]], of Englisce [[ræs]] ===Nama=== '''race''' (mnf. races) # ræs ====Wendunga==== Englisc: [[ræs]] q4spljvga4wvk3fx8v7ed0syd4ck6f2 54840 54839 2026-09-26T22:03:52Z Deadend0914 7211 /* =Wordstǣr 2 */ 54840 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:/reɪs/ ====Wordstǣr 1==== *Ofgangen of Middelengliscum [[race]], of Ealdum Frenciscum [[race]], of Ealdum Ītaliscum [[razza]] ===Nama=== '''race''' (mnf. races) #[[cynn]], [[cnōsl]] ====Word mid gelīcum Sweotolungum==== #[[progeny]] ====Wendunga==== Englisc: [[cynn]], [[cnōsl]] [[Category:Nīwe Englisc nama]] ====Wordstǣr 2==== *Ofgangen of Middelenglisce [[race]], of Englisce [[ræs]] ===Nama=== '''race''' (mnf. races) # ræs ====Wendunga==== Englisc: [[ræs]] mtm3enluuejd3rnsrzcgvgsr6inbr4q 54841 54840 2026-09-26T22:04:05Z Deadend0914 7211 /* Wordstǣr 2 */ 54841 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:/reɪs/ ====Wordstǣr 1==== *Ofgangen of Middelengliscum [[race]], of Ealdum Frenciscum [[race]], of Ealdum Ītaliscum [[razza]] ===Nama=== '''race''' (mnf. races) #[[cynn]], [[cnōsl]] ====Word mid gelīcum Sweotolungum==== #[[progeny]] ====Wendunga==== Englisc: [[cynn]], [[cnōsl]] [[Category:Nīwe Englisc nama]] ===Wordstǣr 2=== *Ofgangen of Middelenglisce [[race]], of Englisce [[ræs]] ===Nama=== '''race''' (mnf. races) # ræs ====Wendunga==== Englisc: [[ræs]] a6a3beb5jbpcwc0s8a81gattk03qfdy 54842 54841 2026-09-26T22:04:32Z Deadend0914 7211 /* Nama */ 54842 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:/reɪs/ ====Wordstǣr 1==== *Ofgangen of Middelengliscum [[race]], of Ealdum Frenciscum [[race]], of Ealdum Ītaliscum [[razza]] ===Nama=== '''race''' (mnf. races) #[[cynn]], [[cnōsl]] ====Word mid gelīcum Sweotolungum==== #[[progeny]] ====Wendunga==== Englisc: [[cynn]], [[cnōsl]] [[Category:Nīwe Englisc nama]] ===Wordstǣr 2=== *Ofgangen of Middelenglisce [[race]], of Englisce [[ræs]] ===Nama=== '''race''' (mnf. races) # [[ræs]] ====Wendunga==== Englisc: [[ræs]] qnsut74qzjei4wh4n1y0jo3xq43ijzl inheritance 0 5810 54812 52653 2026-09-26T18:27:52Z Deadend0914 7211 /* {{sprǣc|en}} */ 54812 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA:// ===Wordstǣr=== {{inh+|en|enm|enheritaunce}}, {{m|enm|inheritaunce}} ===Nama=== '''inheritance''' (mnf. inheritances) #[[ierfe]] [[Flocc:Nīwe Englisc nama]] qbik9efnauetsef3z7cc76lo01fantu use 0 5895 54815 52915 2026-09-26T18:37:59Z Deadend0914 7211 /* {{sprǣc|en}} */ 54815 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== IPA:/juːs/ ===Nama=== '''use''', (mnf. uses) #[[notu]], [[nytt]] ===Wendunga=== * Englisc: [[notu]], [[nytt]] [[Category:Nīwenglisc nama]] ===Word=== '''use''', # tō [[notienne]] ===Wendunga=== * Englisc: [[notian]] 2dhpro7dcrbxth8qfz2edzzr186ogpc sprǣc 0 6017 54843 53117 2026-09-26T22:41:21Z Deadend0914 7211 /* {{sprǣc|ang}} */ 54843 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /spræːtʃ/ === {{f}} {{Wordcynn|Nama}} === {{Naman-Tabule| Hwā oþþe hwæt? (Ānfeald)=sēo sprǣc |Hwā oþþe hwæt? (Manigfeald)=þā sprǣca |Hwæs? (Ānfeald)=þǣre sprǣce |Hwæs? (Manigfeald)=þāra sprǣca |Hwǣm? (Ānfeald)=þǣre sprǣce |Hwǣm? (Manigfeald)=þǣm sprǣcum |Hwȳ? (Ānfeald)=þǣre sprǣce |Hwȳ? (Manigfeald)=þǣm sprǣcum |Hwone? (Ānfeald)=þā sprǣce |Hwone? (Manigfeald)=þā sprǣca }} {{ang-noun|f|head=sprǣċ}} # Sēo gesprecene oþþe gewritene [[endebyrdnes]] tō ferienne þancas and blisse betwēonum mannum. ====Fruma==== ====Word mid gelīcum Sweotolungum==== #[[gereord]], [[reord]] ====Wendunga==== * Nīwenglisc: [[language]] [[Flocc:Englisc nama]] [[Flocc:Englisc wīflic nama]] mluxuodyod386v435x3p05onbdkv7re 54844 54843 2026-09-26T22:41:31Z Deadend0914 7211 /* {{f}} {{Wordcynn|Nama}} */ 54844 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /spræːtʃ/ === {{f}} {{Wordcynn|Nama}} === {{Naman-Tabule| Hwā oþþe hwæt? (Ānfeald)=sēo sprǣc |Hwā oþþe hwæt? (Manigfeald)=þā sprǣca |Hwæs? (Ānfeald)=þǣre sprǣce |Hwæs? (Manigfeald)=þāra sprǣca |Hwǣm? (Ānfeald)=þǣre sprǣce |Hwǣm? (Manigfeald)=þǣm sprǣcum |Hwȳ? (Ānfeald)=þǣre sprǣce |Hwȳ? (Manigfeald)=þǣm sprǣcum |Hwone? (Ānfeald)=þā sprǣce |Hwone? (Manigfeald)=þā sprǣca }} {{ang-noun|wif|head=sprǣċ}} # Sēo gesprecene oþþe gewritene [[endebyrdnes]] tō ferienne þancas and blisse betwēonum mannum. ====Fruma==== ====Word mid gelīcum Sweotolungum==== #[[gereord]], [[reord]] ====Wendunga==== * Nīwenglisc: [[language]] [[Flocc:Englisc nama]] [[Flocc:Englisc wīflic nama]] oybqzhbb509ug68ugvi4h6cucx7xsns 54845 54844 2026-09-26T22:41:40Z Deadend0914 7211 /* {{f}} {{Wordcynn|Nama}} */ 54845 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /spræːtʃ/ === {{f}} {{Wordcynn|Nama}} === {{Naman-Tabule| Hwā oþþe hwæt? (Ānfeald)=sēo sprǣc |Hwā oþþe hwæt? (Manigfeald)=þā sprǣca |Hwæs? (Ānfeald)=þǣre sprǣce |Hwæs? (Manigfeald)=þāra sprǣca |Hwǣm? (Ānfeald)=þǣre sprǣce |Hwǣm? (Manigfeald)=þǣm sprǣcum |Hwȳ? (Ānfeald)=þǣre sprǣce |Hwȳ? (Manigfeald)=þǣm sprǣcum |Hwone? (Ānfeald)=þā sprǣce |Hwone? (Manigfeald)=þā sprǣca }} {{ang-noun|wif|head=sprǣc}} # Sēo gesprecene oþþe gewritene [[endebyrdnes]] tō ferienne þancas and blisse betwēonum mannum. ====Fruma==== ====Word mid gelīcum Sweotolungum==== #[[gereord]], [[reord]] ====Wendunga==== * Nīwenglisc: [[language]] [[Flocc:Englisc nama]] [[Flocc:Englisc wīflic nama]] czfbuloey751rznhbjbg4z1wk8jw0gs hlāf 0 6386 54847 53627 2026-09-26T22:43:12Z Deadend0914 7211 /* Declinung */ 54847 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednesa=== * hlǣf ===Rihtstefn=== * /hlɑːf/ ===Wordstǣr=== Ealdorlic Germanisc: ===Werlic Nama=== ====Nīwenglisc Andgiet==== 1: {{Wendung|en|loaf}}, [[bread]], [[cake]], [[food]] ====Declinung==== {{ang-decl-noun-a-m|hlāf}} ====Gesibbword==== [[hlāford]] "wr" = lord, master, ruler <br> hlāfgang "wr" = meal, meeting,<br> hlāfweard "wr" = steward <br> ====Gesetnyssum==== hlāf-ǣta "wr" = loaf-eater, dependent <br> hlāf-ofn "wr" = oven, baker's oven <br> hlāf-hwǣte "wr" = bread-wheat, wheat <br> hlāf-gebroc "wf", hlāf-gebrecu "nw" = bit of bread <br> hlāfsēnung "wf" = blessing of bread <br> ====Bōccræft Fruma==== ====Wendunga==== *[[Ealdhēahþēodisc]]: [[hlaib]] *[[Ealdnoren]]: {{Wendung|non|hleifr}} *[[Hebrēisc]]: {{Wendung|he|כיכר לחם}} *[[Gotisc]]: {{Wendung|got|𐌷𐌻𐌰𐌹𐍆𐍃}} (hlaifs) *[[Īslendisc]]: {{Wendung|is|hleifur}} *[[Þēodisc]]: {{Wendung|de|Laib}} [[Flocc:Englisc nama]] [[Flocc:Englisc werlic nama]] 3v1ljamus836ttsw5agyon8hcnxprk4 dēaþ 0 6423 54833 54727 2026-09-26T21:39:32Z Deadend0914 7211 /* {{sprǣc|ang}} */ 54833 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednessa=== * deað * dǣþ ===Rihtstefn=== * /dæːɑθ/ ===Wordstǣr=== ====Sweostorword==== *Denisc: {{Wendung|da|død}} *[[Fresisc]]: {{Wendung|fy|dea}} ''wr'' *Īslendisc: {{Wendung|is|dauði}} ''wr'' *Niðerlendisc: {{Wendung|nl|dood}} ''wr'' *Norþwegisc: {{Wendung|no|død}} ''wr'' *[[Seaxisc]]: {{Wendung|nds|Dood}} ''wr'' *Swēonisc: {{Wendung|sv|död}} ''gm'' *Þēodisc: {{Wendung|de|Tod}} ''wr'' ===Werlīc Nama=== # Þæs līfes ende. #: '''''Dēaþ''' nis nān belimp on līfe. '''Dēaðes''' man ne gebītt. Gif man ēcnesse understent nā tō endelēasre tīdlengðe, ac tō tīdlēasnesse, þonne leofaþ sē on ēcnesse þe leofaþ on andweardnesse. Ūre līf is endelēas ealswā ūre gesihtfeld gemǣrelēas biþ.'' ====Nīwenglisc Andgiet==== # [[death]] ====Declīnung==== {{ang-decl-nama-a-m|dēaþ}} ====Gesetnyssum==== [[ær]]-dēaþ - early death, death before one's time, premature death <br> dēaþ-[[bēam]] "wr" - death tree <br> dēaþ-[[bedd]] "nw" - death bed, grave, bed of death, <br> dēaþ-[[cwalu]] "wf" - plague death, violent death, <br> dēaþ-[[dæg]] "wr" - death day, day of death, <br> dēaþ-denu "wf" - valley of death, <br> fær-dēaþ, dēaþ-ræs "wr" - sudden death, <br> dēaþ-scūa "wr", dēaþ-scufa "wr" - death shadow, shadow of death, <br> dēaþ-[[scyld]] "wf" - crime worthy of death <br> dēaþ-stede "wr" - place of death <br> dēaþ-[[wang]] "wr" - death plain, <br> dēaþ-wīc "nw" - death's place, dwelling of the dead, <br> dēaþ-[[wyrd]] "wf" - death fate, fate of death <br> [[ende]]-dēað -death as the end of life <br> [[gúð]]-dēað - battle death, death in battle <br> [[mere]]-dēaþ - sea death, death at sea, <br> [[swylt]]-dēaþ - death (die-death) <br> [[wæl]]-dēaþ - death in battle, <br> [[wundor]]-dēaþ - wondrous death <br> ====Wendunga==== *Frencisc: {{Wendung|fr|mort}} ''wf'' *Grēcisc: {{Wendung|el|θάνατος}} [ˈθa.na.to̞s] ''wr'', {{Wendung|el|θανατάς}} [θa.na.ˈtas] ''wr'', {{Wendung|el|πεθαμός}} [pe̞.θa.ˈmo̞s] ''wr'', {{Wendung|el|αποθαμός}} [a.po̞.θa.ˈmo̞s] {{m}}, {{Wendung|el|χάρος}} [ˈxa.ro̞s] ''wr'' *Hebrēisc: {{Wendung|he|מוות}} (mavet) ''wf'' *Italisc: {{Wendung|it|morte}} ''wf'' *Lǣden: {{Wendung|la|mors}} ''wf'', {{Wendung|la|exitium}} ''nā'', {{Wendung|la|quietus}} ''wr'' *Persisc: {{Wendung|fa|مرگ}} *Scyttisc Gǣlisc: {{Wendung|gd|bàs}} ''wr'' [[Flocc:Englisc nama]] [[Flocc: Englisc werlīc nama]] 9qcv5kskcc5u1l3hnw3ukbj4yat0ldb 54870 54833 2026-09-26T23:51:34Z Deadend0914 7211 /* Declīnung */ 54870 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednessa=== * deað * dǣþ ===Rihtstefn=== * /dæːɑθ/ ===Wordstǣr=== ====Sweostorword==== *Denisc: {{Wendung|da|død}} *[[Fresisc]]: {{Wendung|fy|dea}} ''wr'' *Īslendisc: {{Wendung|is|dauði}} ''wr'' *Niðerlendisc: {{Wendung|nl|dood}} ''wr'' *Norþwegisc: {{Wendung|no|død}} ''wr'' *[[Seaxisc]]: {{Wendung|nds|Dood}} ''wr'' *Swēonisc: {{Wendung|sv|död}} ''gm'' *Þēodisc: {{Wendung|de|Tod}} ''wr'' ===Werlīc Nama=== # Þæs līfes ende. #: '''''Dēaþ''' nis nān belimp on līfe. '''Dēaðes''' man ne gebītt. Gif man ēcnesse understent nā tō endelēasre tīdlengðe, ac tō tīdlēasnesse, þonne leofaþ sē on ēcnesse þe leofaþ on andweardnesse. Ūre līf is endelēas ealswā ūre gesihtfeld gemǣrelēas biþ.'' ====Nīwenglisc Andgiet==== # [[death]] ====Declīnung==== {{ang-decl-noun-a-m|dēaþ}} ====Gesetnyssum==== [[ær]]-dēaþ - early death, death before one's time, premature death <br> dēaþ-[[bēam]] "wr" - death tree <br> dēaþ-[[bedd]] "nw" - death bed, grave, bed of death, <br> dēaþ-[[cwalu]] "wf" - plague death, violent death, <br> dēaþ-[[dæg]] "wr" - death day, day of death, <br> dēaþ-denu "wf" - valley of death, <br> fær-dēaþ, dēaþ-ræs "wr" - sudden death, <br> dēaþ-scūa "wr", dēaþ-scufa "wr" - death shadow, shadow of death, <br> dēaþ-[[scyld]] "wf" - crime worthy of death <br> dēaþ-stede "wr" - place of death <br> dēaþ-[[wang]] "wr" - death plain, <br> dēaþ-wīc "nw" - death's place, dwelling of the dead, <br> dēaþ-[[wyrd]] "wf" - death fate, fate of death <br> [[ende]]-dēað -death as the end of life <br> [[gúð]]-dēað - battle death, death in battle <br> [[mere]]-dēaþ - sea death, death at sea, <br> [[swylt]]-dēaþ - death (die-death) <br> [[wæl]]-dēaþ - death in battle, <br> [[wundor]]-dēaþ - wondrous death <br> ====Wendunga==== *Frencisc: {{Wendung|fr|mort}} ''wf'' *Grēcisc: {{Wendung|el|θάνατος}} [ˈθa.na.to̞s] ''wr'', {{Wendung|el|θανατάς}} [θa.na.ˈtas] ''wr'', {{Wendung|el|πεθαμός}} [pe̞.θa.ˈmo̞s] ''wr'', {{Wendung|el|αποθαμός}} [a.po̞.θa.ˈmo̞s] {{m}}, {{Wendung|el|χάρος}} [ˈxa.ro̞s] ''wr'' *Hebrēisc: {{Wendung|he|מוות}} (mavet) ''wf'' *Italisc: {{Wendung|it|morte}} ''wf'' *Lǣden: {{Wendung|la|mors}} ''wf'', {{Wendung|la|exitium}} ''nā'', {{Wendung|la|quietus}} ''wr'' *Persisc: {{Wendung|fa|مرگ}} *Scyttisc Gǣlisc: {{Wendung|gd|bàs}} ''wr'' [[Flocc:Englisc nama]] [[Flocc: Englisc werlīc nama]] 3lc6edwjeeqomm6y0o987yegaf0di2m cild 0 6428 54824 54784 2026-09-26T21:01:10Z Deadend0914 7211 /* Nāhwæðer Nama */ 54824 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong mann þe ne weorþaþ ġīet mann oþþe wif #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-noun-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] 3wgndh3wumkaojh7vhnn78qmdx6s7pt 54825 54824 2026-09-26T21:01:34Z Deadend0914 7211 /* Nāhwæðer Nama */ 54825 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong mann þe ne weorþaþ ġīet [[mann]] oþþe [[wif]] #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-noun-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] r1xr7vhc8iv71drhtzv4mrjmg35dyj4 54835 54825 2026-09-26T21:46:17Z Deadend0914 7211 /* Declinung */ 54835 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong mann þe ne weorþaþ ġīet [[mann]] oþþe [[wif]] #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-nama-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] gstk1tg62bwp34uqb0zugykrp9r77a3 54836 54835 2026-09-26T21:47:16Z Deadend0914 7211 /* Nāhwæðer Nama */ 54836 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong cnafa þe ne weorþaþ ġīet [[mann]] oþþe [[wif]] #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-nama-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] 5kgwdlqweus3d943r1hfa0jxfhzf8g7 54837 54836 2026-09-26T21:48:56Z Deadend0914 7211 /* Nāhwæðer Nama */ 54837 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong [[cnafa]] oþþe [[mæġden]] þe ne weorþaþ ġīet [[mann]] oþþe [[wif]] #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-nama-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] 0311ml51g3610swr11h5vciq0skmxbg 54846 54837 2026-09-26T22:42:32Z Deadend0914 7211 /* Declinung */ 54846 wikitext text/x-wiki =={{sprǣc|ang}}== [[File:ENFANTS-BLONDS.jpg|thumb|'''cild''']] ===Rihtstefn=== * /tʃild/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kelþą|kelþą]] ===Nāhwæðer Nama=== # ġeong [[cnafa]] oþþe [[mæġden]] þe ne weorþaþ ġīet [[mann]] oþþe [[wif]] #: Sē ēadega Cūðbeorht, þā þā hē wæs eahtawintre ċild, rann swā swā him his nytenlīċe ield tyhte plegende mid his efnealdum. ====Nīwenglisc Andgiet==== 1: [[child]] : infant : youth ====Declinung==== {{ang-decl-noun-z-n|cild}} ====Gesibbword==== cildlic - youth <br> cildisc - childish <br> cildhād - childhood <br> cildfaru - child baring <br> cildsung - childishness <br> ====Gesetnyssum==== [[cniht]]-cild - youth, male child, <br> cradol-cild - cradle child, child in the cradle <br> hyse-cild, wæpned-cild - male child <br> [[mæden]]-cild, wīf-cild - female child <br> [[mōdor]]-cild - child of one's mother <br> munuc-cild - "monk-child" monastic child <br> cild-[[geong]] - "adj" child-youth, young <br> ====Wendunga==== * Denisc: [[barn]] * Ealdnoren: [[barn]] * Frencisc: [[enfant]] * Hebrēisc: [[ילד]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Noren: [[barn]] * Niðerlendisc: [[kind]] * Persisc: [[کودک]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] 5gri6wilft7uzb903m17wqzek0mxjm8 bearn 0 6429 54805 52189 2026-09-26T18:05:17Z Deadend0914 7211 /* Declīnung */ 54805 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednesa=== * barn * beorn ===Rihtstefn=== * /bæɑrn/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kiltham|kiltham]] ====Sweostorword==== * Denisc: [[barn]] * Noren: [[barn]] ====Nāhwæðer Nama==== 1: [[cild]] "nw" <br> 2: [[cnafa]] "wr" - youth :[[cnapa]] "wr" - youth :[[lȳtling]] "wr" :[[magutimber]] "nw" =====Bisen===== Se Cyning nǣfde næfre '''bearn'''. ====Declīnung==== {{ang-decl-noun-a-n|bearn}} ====Gesetnyssum==== frum-bearn -first child <br> dryht-bearn - princely child, royal child <br> cyne-bearn - Christ, royal child <br> fréo-bearn - child of gentile birth <br> god-bearn - divine child, Christ, son of God, <br> wæpned-bearn - male child <br> sige-bearn - victorious child, victor child, <br> wūsc-bearn - little child <br> [[folc]]-bearn - folk child, child of man <br> bearn-cennicge "wf" - mother <br> bearn-cennung "wf" - childbirth <br> bearn-ēacen, bearn-ēacnigende, bearn-ēaca, bearn-ēacnung, baarn-eacnigen - "adj" pregnant, with child, <br> bearn-gebyrda "wf" - child bearing <br> bearn-lēas - "adj" childless <br> bearn-lēast "wf" - childlessness <br> bearn-myrðra "wr", bearnmyrðre "wr" - child murder, child murderer <br> sweostor-bearn - sister's child, nephew, niece, <br> ====Wendunga==== * Frencisc: [[enfant]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Niðerlendisc: [[kind]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] o3au004n3cb0hxc2ctgbigp24gcc5cx 54821 54805 2026-09-26T20:46:23Z Deadend0914 7211 /* Nāhwæðer Nama */ 54821 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednesa=== * barn * beorn ===Rihtstefn=== * /bæɑrn/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kiltham|kiltham]] ====Sweostorword==== * Denisc: [[barn]] * Noren: [[barn]] ====Nāhwæðer Nama==== # Hwæs cild =====Bisen===== Se Cyning nǣfde næfre '''bearn'''. ====Declīnung==== {{ang-decl-noun-a-n|bearn}} ====Gesetnyssum==== frum-bearn -first child <br> dryht-bearn - princely child, royal child <br> cyne-bearn - Christ, royal child <br> fréo-bearn - child of gentile birth <br> god-bearn - divine child, Christ, son of God, <br> wæpned-bearn - male child <br> sige-bearn - victorious child, victor child, <br> wūsc-bearn - little child <br> [[folc]]-bearn - folk child, child of man <br> bearn-cennicge "wf" - mother <br> bearn-cennung "wf" - childbirth <br> bearn-ēacen, bearn-ēacnigende, bearn-ēaca, bearn-ēacnung, baarn-eacnigen - "adj" pregnant, with child, <br> bearn-gebyrda "wf" - child bearing <br> bearn-lēas - "adj" childless <br> bearn-lēast "wf" - childlessness <br> bearn-myrðra "wr", bearnmyrðre "wr" - child murder, child murderer <br> sweostor-bearn - sister's child, nephew, niece, <br> ====Wendunga==== * Frencisc: [[enfant]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Niðerlendisc: [[kind]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] guexzdm1qveiehwnojfgw176z1x50m3 54822 54821 2026-09-26T20:47:01Z Deadend0914 7211 /* Nāhwæðer Nama */ 54822 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednesa=== * barn * beorn ===Rihtstefn=== * /bæɑrn/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kiltham|kiltham]] ====Sweostorword==== * Denisc: [[barn]] * Noren: [[barn]] ====Nāhwæðer Nama==== # Hwæs cild #: Se Cyning nǣfde næfre '''bearn'''. ====Declīnung==== {{ang-decl-noun-a-n|bearn}} ====Gesetnyssum==== frum-bearn -first child <br> dryht-bearn - princely child, royal child <br> cyne-bearn - Christ, royal child <br> fréo-bearn - child of gentile birth <br> god-bearn - divine child, Christ, son of God, <br> wæpned-bearn - male child <br> sige-bearn - victorious child, victor child, <br> wūsc-bearn - little child <br> [[folc]]-bearn - folk child, child of man <br> bearn-cennicge "wf" - mother <br> bearn-cennung "wf" - childbirth <br> bearn-ēacen, bearn-ēacnigende, bearn-ēaca, bearn-ēacnung, baarn-eacnigen - "adj" pregnant, with child, <br> bearn-gebyrda "wf" - child bearing <br> bearn-lēas - "adj" childless <br> bearn-lēast "wf" - childlessness <br> bearn-myrðra "wr", bearnmyrðre "wr" - child murder, child murderer <br> sweostor-bearn - sister's child, nephew, niece, <br> ====Wendunga==== * Frencisc: [[enfant]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Niðerlendisc: [[kind]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] lps79a6030qe3je6cqmq86f1tl80zem 54823 54822 2026-09-26T20:47:40Z Deadend0914 7211 /* Nāhwæðer Nama */ 54823 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednesa=== * barn * beorn ===Rihtstefn=== * /bæɑrn/ ===Wordstǣr=== Ealdorlic Germanisc: *[[Ætēaca:Ealdorlic Germanisc/kiltham|kiltham]] ====Sweostorword==== * Denisc: [[barn]] * Noren: [[barn]] ====Nāhwæðer Nama==== # Hwæs [[cild]] #: Se Cyning nǣfde næfre '''bearn'''. ====Declīnung==== {{ang-decl-noun-a-n|bearn}} ====Gesetnyssum==== frum-bearn -first child <br> dryht-bearn - princely child, royal child <br> cyne-bearn - Christ, royal child <br> fréo-bearn - child of gentile birth <br> god-bearn - divine child, Christ, son of God, <br> wæpned-bearn - male child <br> sige-bearn - victorious child, victor child, <br> wūsc-bearn - little child <br> [[folc]]-bearn - folk child, child of man <br> bearn-cennicge "wf" - mother <br> bearn-cennung "wf" - childbirth <br> bearn-ēacen, bearn-ēacnigende, bearn-ēaca, bearn-ēacnung, baarn-eacnigen - "adj" pregnant, with child, <br> bearn-gebyrda "wf" - child bearing <br> bearn-lēas - "adj" childless <br> bearn-lēast "wf" - childlessness <br> bearn-myrðra "wr", bearnmyrðre "wr" - child murder, child murderer <br> sweostor-bearn - sister's child, nephew, niece, <br> ====Wendunga==== * Frencisc: [[enfant]] * Italisc: [[bambino]] * Lǣden: [[infans]] * Niðerlendisc: [[kind]] * Spēonisc: [[niño]] * Þēodisc: [[Kind]] [[Flocc:Englisc nama]] [[Flocc:Englisce lēode]] [[Flocc:Englisc mǣgþ]] 6iyhjekdnlbr3j21plyazex9qxgouk9 Bysen:ang-decl-noun-a-m 10 7954 54871 54685 2026-09-26T23:52:29Z Deadend0914 7211 54871 wikitext text/x-wiki {{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strang ''a''-stefn<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}{{#if:{{{vowel|}}}||e}}s<!-- -->|7={{{datsg|{{{1}}}{{#if:{{{vowel|}}}||e}}}}}<!-- -->|2={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|4={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|6={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}|n}}a<!-- -->|8={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}m{{#if:{{{vowel|}}}|,{{{2|{{{1}}}}}}um}}<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|masculine a-stem nouns}}<!-- --><noinclude>{{documentation}}</noinclude> snz6kg6eh33novzz9k8wee8bxp0p8iv Module:gender and number/data 828 8023 54863 54779 2026-09-26T23:35:20Z Deadend0914 7211 54863 Scribunto text/plain local data = {} local insert = table.insert -- A list of all possible "parts" that a specification can be made out of. For each part, we list the class it's in -- (gender, animacy, etc.), the associated category (if any) and the display form. In a given gender/number spec, only -- one part of each class is allowed. `display` is how the code is diplayed to the user and should normally be wrapped -- in <abbr title="tooltip">...</abbr> with an explanatory tooltip. If not, it will automatically be wrapped in this -- fashion. If `req` is true, a category "Requests for TYPE in LANG entries" will be generated, except for the code "?", -- which is special-cased; TYPE is "gender" unless the POS is "verb", in which case it is "aspect". data.codes = { ["?"] = {type = "other", req = true, display = '<abbr title="gender incomplete">?</abbr>'}, -- FIXME: The following should be either eliminated in favor of g! or converted to a general "gender/number unattested". ["?!"] = {type = "other", display = "gender unattested"}, -- Genders ["wer"] = {type = "gender", cat = "masculine POS", display = '<abbr title="werlic cynn">wer</abbr>'}, ["wif"] = {type = "gender", cat = "feminine POS", display = '<abbr title="wiflic cynn">wif</abbr>'}, ["m"] = {type = "gender", cat = "masculine POS", display = '<abbr title="werlic cynn">wer</abbr>'}, ["f"] = {type = "gender", cat = "feminine POS", display = '<abbr title="wiflic cynn">wif</abbr>'}, ["n"] = {type = "gender", cat = "neuter POS", display = '<abbr title="nāhwæðer cynn">n</abbr>'}, ["c"] = {type = "gender", cat = "common-gender POS", display = '<abbr title="common gender">c</abbr>'}, ["gneut"] = {type = "gender", cat = "gender-neutral POS", display = "gender-neutral"}, ["g!"] = {type = "gender", display = "gender unattested"}, ["g?"] = {type = "gender", req = true, display = "gender unspecified"}, -- Animacy -- Animate = either animal or personal (for Russian, etc.) ["an"] = {type = "animacy", cat = "animate POS", display = '<abbr title="animate">anim</abbr>'}, ["in"] = {type = "animacy", cat = "inanimate POS", display = '<abbr title="inanimate">inan</abbr>'}, -- Animal (for Ukrainian, Belarusian, Polish, etc.) ["anml"] = {type = "animacy", cat = "animal POS", display = "animal"}, -- Personal (for Ukrainian, Belarusian, Polish, etc.) ["pr"] = {type = "animacy", cat = "personal POS", display = '<abbr title="personal">pers</abbr>'}, ["np"] = {type = "animacy", cat = "nonpersonal POS", display = '<abbr title="nonpersonal">npers</abbr>'}, ["an!"] = {type = "animacy", display = "animacy unattested"}, ["an?"] = {type = "animacy", req = true, display = "animacy unspecified"}, -- Definiteness ["def"] = {type = "definiteness", cat = "definite POS", display = '<abbr title="definite">def</abbr>'}, ["indef"] = {type = "definiteness", cat = "indefinite POS", display = '<abbr title="indefinite">indef</abbr>'}, -- Virility (for Polish) ["vr"] = {type = "virility", cat = "virile POS", display = '<abbr title="virile (= masculine personal)">vir</abbr>'}, ["nv"] = {type = "virility", cat = "nonvirile POS", display = '<abbr title="nonvirile (= other than masculine personal)">nvir</abbr>'}, -- Numbers ["s"] = {type = "number", display = '<abbr title="ānfeald rime">sg</abbr>'}, ["d"] = {type = "number", cat = "dualia tantum", display = '<abbr title="twifeald rime">du</abbr>'}, ["p"] = {type = "number", cat = "pluralia tantum", display = '<abbr title="manigfeald rime">pl</abbr>'}, ["num!"] = {type = "number", display = "number unattested"}, ["num?"] = {type = "number", req = true, display = "number unspecified"}, -- Verb qualifiers ["impf"] = {type = "aspect", cat = "imperfective POS", display = '<abbr title="imperfective aspect">impf</abbr>'}, ["pf"] = {type = "aspect", cat = "perfective POS", display = '<abbr title="perfective aspect">pf</abbr>'}, ["asp!"] = {type = "aspect", display = "aspect unattested"}, ["asp?"] = {type = "aspect", req = true, display = "aspect unspecified"}, } -- Combined codes that are equivalent to giving multiple specs. `mf` is the same as specifying two separate specs, -- one with `m` in it and the other with `f`. `mfbysense` is similar but is used for nouns that can be either masculine -- or feminine according as to whether they refer to masculine or feminine beings. local combinations = { ["biasp"] = {codes = {"impf", "pf"}}, ["anin"] = {codes = {"an", "in"}}, -- "bianimate" doesn't exist as a linguistic term } for _, comb in ipairs{"mf", "mn", "fm", "fn", "cn", "nm", "nf", "nc", "mfn", "mnf", "fmn", "fnm", "nmf", "nfm"} do local codes = {} for ch in comb:gmatch(".") do insert(codes, ch) end combinations[comb] = {codes = codes} combinations[comb .. "equiv"] = {codes = codes, display = '<abbr title="different genders do not affect the meaning">same meaning</abbr>'} if comb == "mf" or comb == "fm" then combinations[comb .. "bysense"] = {codes = codes, cat = "masculine and feminine POS by sense", display = '<abbr title="according to the gender of the referent">by sense</abbr>'} end end data.combinations = combinations -- Categories when multiple gender/number codes of a given type occur in different specs (two or more of the same type -- cannot occur in a single spec). data.multicode_cats = { ["gender"] = "POS with multiple genders", ["animacy"] = "POS with multiple animacies", ["aspect"] = "biaspectual POS", } return data 4yjum8e1bggbftozkecoe546shr606q Bysen:ang-decl-noun-z-n 10 8024 54848 54783 2026-09-26T22:49:33Z Deadend0914 7211 54848 wikitext text/x-wiki {{ang-decl-noun<!-- -->|type=strang ''z''-stefn<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}es<!-- -->|7={{{1}}}e<!-- -->|2={{{1}}}ru<!-- -->|4={{{1}}}ru<!-- -->|6={{{1}}}ra<!-- -->|8={{{1}}}rum<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|z-stem nouns}}<!-- --><noinclude>{{documentation}}</noinclude> l5hr7jakaaz03vdj2932ldaici195hs Module:en-hēafodword 828 8036 54829 54804 2026-09-26T21:21:48Z Deadend0914 7211 54829 Scribunto text/plain local export = {} local pos_functions = {} --[==[ Bócere from 2020 on: mostly Benwing2, with significant contributions from Theknightwho. Based on a prior version by Rua (by now mostly rewritten), with contributions from Erutuon and others (see history for full attribution). ]==] local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local require = require local thurf_thonne_node = require("Module:thurf thonne node") local en_netwerdlicnesa_module = "Module:en-netwerdlicnesa" local heafodword_netwerdlicnesa_module = "Module:heafodword netwerdlicnesa" local heafodword_module = "Module:heafodword" local declinung_netwerdlicnesa_module = "Module:declinung netwerdlicnesa" local parse_netwerdlicnesa_module = "Module:parse netwerdlicnesa" local JSON_module = "Module:JSON" local labels_module = "Module:labels" local links_module = "Module:links" local parameters_module = "Module:parameters" local string_netwerdlicnesa_module = "Module:string netwerdlicnesa" local table_module = "Module:table" local netwerdlicnesa_module = "Module:netwerdlicnesa" local yesno_module = "Module:yesno" local iut = thurf_thonne_node(declinung_netwerdlicnesa_module) local put = thurf_thonne_node(parse_netwerdlicnesa_module) local m_heafodword_netwerdlicnesa = thurf_thonne_node(heafodword_netwerdlicnesa_module) local add_links_to_multiword_term = thurf_thonne_node(heafodword_netwerdlicnesa_module, "add_links_to_multiword_term") local add_suffix = thurf_thonne_node(en_netwerdlicnesa_module, "add_suffix") local apply_link_modifiers = thurf_thonne_node(heafodword_netwerdlicnesa_module, "apply_link_modifiers") local concat = table.concat local deepEquals = thurf_thonne_node(table_module, "deepEquals") local dump = mw.dumpObject local format_categories = thurf_thonne_node(netwerdlicnesa_module, "format_categories") local full_heafodword = thurf_thonne_node(heafodword_module, "full_heafodword") local get_label_info = thurf_thonne_node(labels_module, "get_label_info") local get_link_page = thurf_thonne_node(links_module, "get_link_page") local glossary_link = thurf_thonne_node(heafodword_netwerdlicnesa_module, "glossary_link") local insert = table.insert local insertIfNot = thurf_thonne_node(table_module, "insertIfNot") local ipairs = ipairs local is_regular_plural = thurf_thonne_node(en_netwerdlicnesa_module, "is_regular_plural") local list_to_set = thurf_thonne_node(table_module, "listToSet") local pairs = pairs local process_params = thurf_thonne_node(parameters_module, "process") local remove = table.remove local remove_links = thurf_thonne_node(links_module, "remove_links") local replacement_escape = thurf_thonne_node(string_netwerdlicnesa_module, "replacement_escape") local shallowCopy = thurf_thonne_node(table_module, "shallowCopy") local singularize = thurf_thonne_node(en_netwerdlicnesa_module, "singularize") local split = thurf_thonne_node(string_netwerdlicnesa_module, "split") local toJSON = thurf_thonne_node(JSON_module, "toJSON") local toNFD = mw.ustring.toNFD local type = type local ulen = thurf_thonne_node(string_netwerdlicnesa_module, "len") local ulower = thurf_thonne_node(string_netwerdlicnesa_module, "lower") local umatch = thurf_thonne_node(string_netwerdlicnesa_module, "match") local u = thurf_thonne_node(string_netwerdlicnesa_module, "char") local ugsub = thurf_thonne_node(string_netwerdlicnesa_module, "gsub") local lang = require("Module:spraeca").getByCode("en") local langname = lang:getCanonicalName() local list_param = {list = true, disallow_holes = true} local list_allow_holes = {list = true, allow_holes = true} local boolean_param = {type = "boolean"} local function ine(val) if val == "" then return nil else return val end end local function track(page) require("Module:debug/track")("en-heafodword/" .. page) return true end ------------------------------------------- UTILITY FUNCTIONS ------------------------------------------ -- Parse and return an declinung not requiring additional processing. The raw arguments come from `args[field]`, which -- is parsed for inline modifiers. local function parse_declinung(args, field, is_head) local argfield = field if type(argfield) == "table" then argfield = argfield[1] end return m_heafodword_netwerdlicnesa.parse_term_list_with_modifiers { paramname = field, forms = args[argfield], splitchar = ",", is_head = is_head, } end -- Insert the parsed declinunga in `terms` (as parsed by `parse_declinung`) into `data.declinunga`, with label -- `label` and optional accelerator spec `accel`. local function insert_declinung(data, terms, label, accel, no_label) for _, termobj in ipairs(terms) do m_heafodword_netwerdlicnesa.remove_termobj_field_modifiers(termobj) end m_heafodword_netwerdlicnesa.insert_declinung { headdata = data, terms = terms, label = label, no_label = no_label, accel = accel and {form = accel} or nil, } end -- Insert a fixed label `label` into the declinunga for `data`. If `originating_term` is supplied, copy the decorations -- from it into the fixed label. local function insert_fixed_declinung(data, label, originating_term) m_heafodword_netwerdlicnesa.insert_fixed_declinung { headdata = data, originating_term = originating_term, label = label, } end -- Parse and insert an declinung not requiring additional processing into `data.declinunga`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the declinunga are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_declinung(data, args, field, label, accel) m_heafodword_netwerdlicnesa.parse_and_insert_declinung { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- These functions are used directly in the <> format as well as in the utility functions #2 below. local function compute_double_last_cons_stem(term) local last_cons = term:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("Verb stem '" .. term .. "' must end in a consonant to use ++") end return term .. last_cons end local function compute_plusplus_s_form(term, default_s_form) if term:find("[szx]$") then -- regas -> regasses, derez -> derezzes return compute_double_last_cons_stem(term) .. "es" else return default_s_form end end -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local iparams = { [1] = true, } local iargs = require("Module:parameters").process(frame.args, iparams) local parargs = frame:getParent().args local poscat = iargs[1] local pos_in_1 = not poscat if pos_in_1 then poscat = ine(parargs[1]) or mw.title.getCurrentTitle().fullText == "Template:en-head" and "interjection" or error("Part of speech must be specified in 1=") poscat = require(heafodword_module).canonicalize_pos(poscat) end local indexing_poscat = pos_in_1 and "head" or poscat local params = { ["head"] = list_param, ["id"] = true, ["json"] = boolean_param, ["sort"] = true, ["splithyph"] = boolean_param, ["nosplithyph"] = boolean_param, ["hyphspace"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {type = "boolean_param", alias_of = "nolink"}, ["suffix"] = boolean_param, ["nosuffix"] = boolean_param, ["nomultiwordcat"] = boolean_param, ["abbr"] = list_param, ["the"] = true, ["def"] = {alias_of = "the"}, ["pagename"] = true, -- for testing } if pos_in_1 then params[1] = {required = true} -- required but ignored as already processed above end local pos_data = pos_functions[indexing_poscat] local pos_func if pos_data then local pos_params = pos_data.params if pos_params then for key, val in pairs(pos_params) do params[key] = val end end pos_func = pos_data.func end local args = process_params(parargs, params) -- Account for unsupported titles, e.g. 'C|N>K' instead of 'Unsupported titles/C through N to K'. local pagename = args.pagename or mw.loadData("Module:heafodword/data").pagename local user_specified_heads = parse_declinung(args, "head", "is_head") local heads = user_specified_heads local autohead if args.nolink or not pagename:find("[ '%-]") then autohead = pagename else local en_no_split_apostrophe_words = list_to_set { "one's", "someone's", "he's", "she's", "it's", } local en_include_hyphen_prefixes = list_to_set { -- We don't include things that are also words even though they are often (perhaps mostly) prefixes, e.g. -- "be", "counter", "cross", "extra", "half", "mid", "over", "pan", "under". "acro", "acousto", "Afro", "agro", "anarcho", "angio", "Anglo", "ante", "anti", "arch", "auto", "bi", "bio", "cis", "co", "cryo", "crypto", "de", "demi", "eco", "electro", "Euro", "ex", "Greco", "hemi", "hydro", "hyper", "hypo", "infra", "Indo", "inter", "intra", "Judeo", "macro", "meta", "micro", "mini", "multi", "neo", "neuro", "non", "para", "peri", "post", "pre", "pro", "proto", "pseudo", "re", "semi", "sub", "super", "trans", "un", "vice", } local function is_english(term) local title = mw.title.new(term) if title and title.exists then local content = title:getContent() if content and content:find("==English==\n") then return true end end return false end local function en_split_hyphen_when_space(word) if not word:find("-", nil, true) then return nil end if args.hyphspace then return "[[" .. word:gsub("%-+", " ") .. "|" .. word .. "]]" end if args.nosplithyph then return "[[" .. word .. "]]" end if not args.splithyph then local space_word = word:gsub("%-+", " ") if is_english(space_word) then return "[[" .. space_word .. "|" .. word .. "]]" end if is_english(word) then return "[[" .. word .. "]]" end end return nil end local function en_split_apostrophe(word) local base = word:match("^(.*)'s$") if base then return "[[" .. base .. "]][[-'s|'s]]" end -- Only treat final apostrophe as possessive if preceded by something that looks like a plural ending in /z/. -- In particular we don't want to do it for words like [[truckin']]. base = word:match("^(.*[sxz])'$") if base then if base:find("s$") then local sg = singularize(base) if is_english(sg) then return "[[" .. sg .. "|" .. base .. "]][[-'|']]" end end return "[[" .. base .. "]][[-'|']]" end return "[[" .. word .. "]]" end autohead = add_links_to_multiword_term(pagename, { split_hyphen_when_space = en_split_hyphen_when_space, split_apostrophe = en_split_apostrophe, no_split_apostrophe_words = en_no_split_apostrophe_words, include_hyphen_prefixes = en_include_hyphen_prefixes, }) end if not heads[1] then heads = {{term = autohead}} else for _, headobj in ipairs(heads) do local head = headobj.term if head:find("^~") then head = apply_link_modifiers(autohead, head:sub(2), lang) headobj.term = head elseif head:find("^[!?]$") then -- If explicit head= just consists of ! or ?, add it to the end of the default head. headobj.term = autohead .. head end if head == autohead then track("redundant-head") end end end -- handle the=/def= if args.the == "~" then local newheads = {} for _, headobj in ipairs(heads) do local barehead = shallowCopy(headobj) insert(newheads, barehead) headobj.term = "the " .. headobj.term insert(newheads, headobj) end heads = newheads elseif args.the then local the = require(yesno_module)(args.the) if the then for _, headobj in ipairs(heads) do headobj.term = "the " .. headobj.term end end end local data = { lang = lang, pos_category = poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, -- We use our own splitting algorithm so the redundant head cat will be inaccurate. no_redundant_head_cat = true, declinunga = {}, nomultiwordcat = args.nomultiwordcat, sort_key = args.sort, pagename = pagename, id = args.id, force_cat_output = force_cat, } local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local is_suffix = false if args.suffix or not args.nosuffix and pagename:find("^%-") and not pagename:find("^%-%-") and poscat ~= "suffix forms" then is_suffix = true data.pos_category = "suffixes" local singular_poscat = singularize(poscat) inscat(singular_poscat .. "-forming suffixes") insert(data.declinunga, {label = singular_poscat .. "-forming suffix"}) end if pos_func then pos_func(args, data, is_suffix) end local extra_categories = {} if pagename:find("[Qq]") then -- Check for q not followed by u. We want to exclude things like [[13q deletion syndrome]] and [[BFOQ]] that -- don't have a lowercase letter on either side, as well as things like [[& seq.]] and [[acq.]] that are -- abbreviations for words containing a following u. -- -- Approximate range of combining diacritics; we want to remove them so the checks below for -- a lowercase letter next to the q aren't tripped up by diacritics on the letter. local u300 = u(0x0300) local u36F = u(0x036F) local pagename_no_diacritics = ugsub(toNFD(pagename), "[" .. u300 .. "-" .. u36F .. "]", "") if pagename_no_diacritics:find("[Qq][a-tv-z]") or pagename_no_diacritics:find("[a-z]q[^u.]") or pagename_no_diacritics:find("[a-z]q$") then inscat("words containing Q not followed by U") end end -- toNFD performs decomposition, so letters that decompose to an ASCII -- vowel and a diacritic, such as é, are counted as vowels and do not do not -- need to be included in the pattern. if not umatch(ulower(toNFD(pagename)), "[aeiouyæœøəªºαεηιουω]") then inscat("words spelled without vowels") end if pagename:find("yre$") then inscat('words ending in "-yre"') end if not pagename:find(" ") and ulen(pagename) >= 25 then insert(extra_categories, "Long " .. langname .. " words") end if pagename:find("^[^aeiou ]*a[^aeiou ]*e[^aeiou ]*i[^aeiou ]*o[^aeiou ]*u[^aeiou ]*$") then inscat("words that use all vowels in alphabetical order") end parse_and_insert_declinung(data, args, "abbr", "abbreviation") if args.json then return toJSON(data) end return full_heafodword(data) .. (extra_categories[1] and format_categories(extra_categories, lang, args.sort) or "") end local function make_default_comparative(word) if word == "good" or word == "well" then return {"better"} elseif word == "bad" or word == "badly" then return {"worse"} elseif word == "far" then return {"further", "farther"} else return {add_suffix(word, "r")} end end local function make_default_superlative(word) if word == "good" or word == "well" then return {"best"} elseif word == "bad" or word == "badly" then return {"worst"} elseif word == "far" then return {"furthest", "farthest"} else return {add_suffix(word, "st.superlative")} end end -- This function does the common work between adjectives and adverbs. local function process_comparative_args(data, args, plpos) local pagename = data.pagename local comps = parse_declinung(args, 1) local sups = parse_declinung(args, "sup") local outcomps, outsups if args.componly then if comps[1] then error("Can't specify comparatives of comparative-only " .. plpos) end insert(data.declinunga, {label = glossary_link("comparative") .. " form only"}) insert(data.categories, langname .. " comparative-only " .. plpos) -- Set to empty list so we don't get any comparatives output, but process superlatives if specified. outcomps = {} if not sups[1] then -- Set to empty list so we don't get any superlatives output unless explicitly given. outsups = {} end elseif args.suponly then if comps[1] or sups[1] then error("Can't specify comparatives or superlatives of or superlative-only " .. plpos) end insert(data.declinunga, {label = glossary_link("superlative") .. " form only"}) insert(data.categories, langname .. " superlative-only " .. plpos) return end -- If the first parameter is ?, then don't show anything, just return. if comps[1] and comps[1].term == "?" then if comps[2] then error("Can't specify additional comparatives along with '?'") end if sups[1] then error("Can't specify superlatives along with '?' for the comparative") end return end if comps[1] and comps[1].term == "-" then local hyphencomp = remove(comps, 1) -- Remove the "-" but retain for decorations. -- Not (generally) comparable; may occasionally have a comparative if comps[1] then insert_fixed_declinung(data, "not generally <<comparable>>", hyphencomp) elseif not sups[1] then insert_fixed_declinung(data, "not <<comparable>>", hyphencomp) insert(data.categories, langname .. " uncomparable " .. plpos) return else -- No comparative, but a superlative. insert_declinung() will correctly generate 'no comparative' if we -- pass in "-" as the value. outcomps = {hyphencomp} end elseif not comps[1] then comps = {{term = "more"}} end if not outcomps then -- not if we set `outcomps` to "-" above or processed a comparative-only term outcomps = {} -- Go over each parameter given and create a comparative and superlative form. for _, compobj in ipairs(comps) do local comp = compobj.term if comp == "-" then error("Comparative of '-' only allowed as first comparative") end if comp == "+" then comp = "+more" elseif comp == "more" and pagename ~= "many" and pagename ~= "much" then comp = "+more" elseif comp == "further" and pagename ~= "far" then comp = "+further" elseif comp == "better" and pagename ~= "good" and pagename ~= "well" then comp = "+better" elseif comp:find("~") then comp = comp:gsub("~", replacement_escape(pagename)) end compobj.origterm = comp if comp == "+more" then comp = "more [[" .. pagename .. "]]" elseif comp == "+further" then comp = {"further [[" .. pagename .. "]]", "farther [[" .. pagename .. "]]"} elseif comp == "+better" then comp = "better [[" .. pagename .. "]]" elseif comp == "er" then -- Add -er. comp = add_suffix(pagename, "r") elseif comp == "ier" then if pagename:sub(-1) ~= "y" then error("Can't specify 'ier' comparative unless the term ends with 'y': " .. pagename) end comp = pagename:gsub("e?y$", "ier") elseif comp:find("^%+") then local special = m_heafodword_netwerdlicnesa.get_special_indicator(comp, "noerror") if special then comp = m_heafodword_netwerdlicnesa.handle_multiword(pagename, special, make_default_comparative) end end if type(comp) == "table" and not comp[2] then comp = comp[1] end if type(comp) == "table" then for i = 1, #comp - 1 do local outobj = shallowCopy(compobj) outobj.term = comp[i] insert(outcomps, outobj) end compobj.term = comp[#comp] insert(outcomps, compobj) else compobj.term = comp insert(outcomps, compobj) end end end if sups[1] and sups[1].term == "-" then if sups[2] then error("Can't specify '-' as superlative followed by further values") end -- No superlative. insert_declinung() will correctly generate 'no superlative' if we pass in "-" as the value. outsups = sups else if not sups[1] then sups = {{term = "+"}} end end -- `outsups` will be set if we set `outsups` to "-" above or processed a comparative-only term without superlatives. if not outsups then outsups = {} local function process_sup(sup, special, supobj, compobj) if special then sup = m_heafodword_netwerdlicnesa.handle_multiword(pagename, special, make_default_superlative) elseif sup == "-" or sup == "+" then error(("Internal error: Superlative value of '%s' should have been handled earlier"):format(sup)) elseif sup == "+most" then sup = "most [[" .. pagename .. "]]" elseif sup == "+furthest" then sup = {"furthest [[" .. pagename .. "]]", "farthest [[" .. pagename .. "]]"} elseif sup == "+best" then sup = "best [[" .. pagename .. "]]" elseif sup == "est" then -- Add -est. sup = add_suffix(pagename, "st.superlative") elseif sup == "iest" then if pagename:sub(-1) ~= "y" then error("Can't specify 'iest' superlative unless the term ends with 'y': " .. pagename) end sup = pagename:gsub("e?y$", "iest") end if type(sup) == "table" and not sup[2] then sup = sup[1] end if compobj then supobj = shallowCopy(supobj) supobj = m_heafodword_netwerdlicnesa.combine_termobj_decorations(supobj, compobj) end if type(sup) == "table" then for i = 1, #sup - 1 do local outobj = shallowCopy(supobj) outobj.term = sup[i] insert(outsups, outobj) end supobj.term = sup[#sup] insert(outsups, supobj) else supobj.term = sup insert(outsups, supobj) end end for _, supobj in ipairs(sups) do local sup = supobj.term if sup == "-" then error("Superlative of '-' only allowed as first superlative") end if sup == "+" then if not comps[1] then error("Superlative of '+' can't be specified when there are no comparatives") end for _, compobj in ipairs(comps) do local comp = compobj.origterm local special if comp == "+more" then sup = "+most" elseif comp == "+further" then sup = "+furthest" elseif comp == "+better" then sup = "+best" elseif comp == "er" then sup = "est" elseif comp == "ier" then sup = "iest" else if comp:find("^%+") then special = m_heafodword_netwerdlicnesa.get_special_indicator(comp, "noerror") end if not special then -- If the full comparative was given, then derive the superlative by replacing -er with -- -est. if comp:sub(-2) == "er" then sup = comp:sub(1, -3) .. "est" else error(("The superlative cannot be derived automatically from comparative '%s' because it doesn't end in -er"):format(comp)) end end end process_sup(sup, special, supobj, compobj) end else local special = m_heafodword_netwerdlicnesa.get_special_indicator(sup, "noerror") -- Do some work here rather than in process_sup() so we don't end up double-processing a term with a '~' -- in it or a term that happens to be 'most' or similar after substitution of ~ in the comparative. if not special then if sup == "most" and pagename ~= "many" and pagename ~= "much" then sup = "+most" elseif sup == "furthest" and pagename ~= "far" then sup = "+furthest" elseif sup == "best" and pagename ~= "good" and pagename ~= "well" then sup = "+best" elseif sup:find("~") then sup = sup:gsub("~", replacement_escape(pagename)) end end process_sup(sup, special, supobj) end end end insert_declinung(data, outcomps, "<<comparative>>", "comparative") insert_declinung(data, outsups, "<<superlative>>", "superlative") end pos_functions["adjectives"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adjectives") end, } pos_functions["adverbs"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adverbs") end, } local function escape(str) return (str:gsub("\\([:#])", "\\\\%1") :gsub("[:#]", "\\%0")) end local function canonicalize_plural(pl, pagename, pos) if pl == "+" then return escape(add_suffix(pagename, "s.plural", pos)) elseif pl == "++" then return escape(compute_plusplus_s_form(pagename, add_suffix(pagename, "s.plural", pos))) elseif pl == "*" then return escape(pagename) elseif pl == "ies" then if pagename:sub(-1) == "y" then return escape(pagename:gsub("e?y$", pl)) end error("Can't specify 'ies' plural unless the term ends with 'y'.") elseif pl == "s" or pl == "es" or pl == "'s" then return escape(pagename .. pl) end end local function do_nouns(args, data, pos) local pagename = data.pagename pos = pos or "noun" local plurals = parse_declinung(args, 1) local function insert_plurale_tantum_declinunga(is_plural_only, originating_label) if args.sg[1] then insert_fixed_declinung(data, "normally plural", originating_label) parse_and_insert_declinung(data, args, "sg", "singular") elseif is_plural_only then insert_fixed_declinung(data, "plural only", originating_label) end if args.attr[1] then parse_and_insert_declinung(data, args, "attr", "attributive") end end local function first_pl_term() return plurals[1] and plurals[1].term or nil end if first_pl_term() == "p" then -- plurale tantum if plurals[2] then error("With plurale tantum noun, can't specify more than one plural") end data.genders = {"p"} -- this should auto-insert the correct 'pluralia tantum' category insert_plurale_tantum_declinunga("plural only", plurals[1]) return end local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local need_default_plural = pos == "noun" if first_pl_term() == "sp" then -- construed as singular or plural sp = remove(plurals, 1) -- Remove the "sp" but retain it for its decorations. inscat("nouns construed as singular or plural") data.genders = {"s", "p"} -- this should auto-insert the correct 'pluralia tantum' category insert_plurale_tantum_declinunga(nil, sp) need_default_plural = false elseif first_pl_term() == "-" then -- Uncountable noun; may occasionally have a plural local hyphpl = remove(plurals, 1) -- Remove the "-" but retain for decorations. inscat("uncountable nouns") -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert_fixed_declinung(data, "usually <<uncountable>>", hyphpl) else insert_fixed_declinung(data, "<<uncountable>>", hyphpl) end need_default_plural = false elseif first_pl_term() == "#" then -- Usually countable (e.g., "grilled cheese") local hashpl = remove(plurals, 1) -- Remove the "#" but retain for decorations. insert_fixed_declinung(data, "usually <<countable>>", hashpl) inscat("uncountable nouns") inscat("countable nouns") -- If no plural was given, add a default one now if not plurals[1] then plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end elseif first_pl_term() == "~" then -- Mixed countable/uncountable noun, always has a plural local tildepl = remove(plurals, 1) -- Remove the "~" but retain for decorations. insert_fixed_declinung(data, "<<countable>> and <<uncountable>>", tildepl) inscat("uncountable nouns") inscat("countable nouns") -- If no plural was given, add a default one now if not plurals[1] then plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end end -- Plural is unknown if first_pl_term() == "?" then local questionpl = remove(plurals, 1) -- Remove the "?" but retain for decorations. -- Not desired; see [[Wiktionary:Tea_room/2021/August#"Plural unknown or uncertain"]] -- insert_fixed_declinung(data, "plural unknown or uncertain", questionpl) inscat("nouns with unknown or uncertain plurals") if plurals[1] then error("Can't specify explicit plurals along with '?' for unknown/uncertain plural") end return end -- Plural is not attested if first_pl_term() == "!" then local exclampl = remove(plurals, 1) -- Remove the "!" but retain for decorations. insert_fixed_declinung(data, "plural not attested", exclampl) inscat("nouns with unattested plurals") if plurals[1] then error("Can't specify explicit plurals along with '!' for unattested plural") end return end -- If no plural was given, maybe add a default one, otherwise (when "-" was given or proper noun) return. if not plurals[1] then if not need_default_plural then inscat("uncountable nouns") return end plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end -- There are plural forms to show, so show them. inscat("countable nouns") local irregular, indeclinable for i, pl in ipairs(plurals) do local canon_pl = canonicalize_plural(pl.term, pagename, pos) if canon_pl then pl.term = canon_pl end local pl_term = get_link_page(pl.term, lang) if not (pagename:find(" ") or is_regular_plural(pl_term, pagename)) then irregular = true if pl_term == pagename then indeclinable = true end end end if irregular then inscat("nouns with irregular plurals") end if indeclinable then inscat("indeclinable nouns") end insert_declinung(data, plurals, "plural", "p") end -- Return the parameters to be used for nouns and proper nouns. Currently the same. local noun_params = { [1] = list_param, ["pl\1qual"] = {list = true, allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the plural", }, -- The following four only used for pluralia tantum (1=p) ["sg"] = list_param, ["attr"] = list_param, } pos_functions["nouns"] = { params = noun_params, func = do_nouns, } pos_functions["proper nouns"] = { params = noun_params, func = function(args, data) return do_nouns(args, data, "proper noun") end, } local function base_default_verb_forms(verb) return escape(add_suffix(verb, "s.verb")), escape(add_suffix(verb, "ing")), escape(add_suffix(verb, "d")) end local function default_verb_forms(verb) local full_s_form, full_ing_form, full_ed_form = base_default_verb_forms(verb) if verb:find(" ") then local first, rest = verb:match("^(.-)( .*)$") local first_s_form, first_ing_form, first_ed_form = base_default_verb_forms(first) return full_s_form, full_ing_form, full_ed_form, first_s_form .. rest, first_ing_form .. rest, first_ed_form .. rest, first, rest else return full_s_form, full_ing_form, full_ed_form, nil, nil, nil, nil, nil end end local function compute_double_last_cons_stem_of_split_verb(verb, ending) local first, rest = verb:match("^(.-)( .*)$") if not first then error("Verb '" .. verb .. "' must have a space in it to use **") end local last_cons = first:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("First word '" .. first .. "' must end in a consonant to use **") end return first .. last_cons .. ending .. rest end local function check_non_nil_star_form(form, pagename) if form == nil then error("Verb '" .. pagename .. "' must have a space in it to use *, **, *l, *! or *'") end return form end local function sub_tilde(form, pagename) if not form then return nil end if form:find("~") then form = form:gsub("~", replacement_escape(pagename)) end return form end local deprecated_qual_replaced_by_inline_modifier = { list = true, allow_holes = true, replaced_by = false, instead = "use an inline modifier <q:...> or <l:...> on the value" } pos_functions["verbs"] = { params = { [1] = {list = "pres_3sg", disallow_holes = true}, ["pres_3sg\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [2] = {list = "pres_ptc", disallow_holes = true}, ["pres_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [3] = {list = "past", disallow_holes = true}, ["past\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [4] = {list = "past_ptc", allow_holes = true}, ["past_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, ["noautolinkverb"] = boolean_param, ["angle_bracket"] = boolean_param, }, func = function(args, data) -- Get parameters local par1s local par2s = parse_declinung(args, {2, "pres_ptc"}) local par3s = parse_declinung(args, {3, "past"}) local par4s = parse_declinung(args, {4, "past_ptc"}) local pres_3sgs, pres_ptcs, pasts, past_ptcs local pagename = data.pagename ------------------------------------------- UTILITY FUNCTIONS #2 ------------------------------------------ -- These functions are used in both in the separate-parameter format and in the override params such as past_ptc2=. local full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed local lemma local function set_lemma_and_default_forms(the_lemma) lemma = the_lemma full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed, lemma_first, lemma_rest = default_verb_forms(the_lemma) end local function canonicalize_s_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_s elseif form == "*" then return check_non_nil_star_form(split_default_s, lemma) elseif form == "++" then return compute_plusplus_s_form(lemma, full_default_s) elseif form == "**" then if lemma:find("^[^ ]*[szx] ") then return compute_double_last_cons_stem_of_split_verb(lemma, "es") else return check_non_nil_star_form(split_default_s, lemma) end elseif form == "+!" then return lemma .. "s" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "s" .. lemma_rest elseif form == "+'" then return lemma .. "'s" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'s" .. lemma_rest elseif form == "+l" then if lemma:find("[szx]$") then return {{term = full_default_s, l = {"US"}}, {term = compute_plusplus_s_form(lemma, full_default_s), l = {"UK"}}} else return compute_plusplus_s_form(lemma, full_default_s) end elseif form == "*l" then if lemma:find("^[^ ]*[szx] ") then return {{term = check_non_nil_star_form(split_default_s, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "es"), l = {"UK"}}} else return check_non_nil_star_form(split_default_s, lemma) end else return sub_tilde(form, lemma) end end local function canonicalize_ing_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ing elseif form == "*" then return check_non_nil_star_form(split_default_ing, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ing" elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ing") elseif form == "+!" then return lemma .. "ing" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ing" .. lemma_rest elseif form == "+'" then return lemma .. "'ing" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'ing" .. lemma_rest elseif form == "+l" then return {{term = full_default_ing, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ing", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ing, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ing"), l = {"UK"}}} else return sub_tilde(form, lemma) end end local function canonicalize_ed_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ed elseif form == "*" then return check_non_nil_star_form(split_default_ed, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ed" elseif form == "+!" then return lemma .. "ed" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ed" .. lemma_rest elseif form == "+'" then return {{term = lemma .. "'d"}, {term = lemma .. "'ed"}} elseif form == "*'" then return {{term = check_non_nil_star_form(lemma_first) .. "'d" .. lemma_rest}, {term = check_non_nil_star_form(lemma_first) .. "'ed" .. lemma_rest}} elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ed") elseif form == "+l" then return {{term = full_default_ed, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ed", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ed, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ed"), l = {"UK"}}} else return sub_tilde(form, lemma) end end -- FIXME: options should be "+", "*", "++", "**", "+n", "*n", "++n" and "**n", but not "n" local function canonicalize_en_form(form) if form == "n" then track("n4") return add_suffix(lemma, "n") end return canonicalize_ed_form(form) end --------------------------------- MAIN PARSING/CONJUGATING CODE -------------------------------- local is_angle_bracket = args.angle_bracket if is_angle_bracket then if par2s[1] or par3s[1] or par4s[1] then error("Can't specify explicit values for 2=, 3= or 4= along with the angle-bracket format") end elseif is_angle_bracket == nil and not par2s[1] and not par3s[1] and not par4s[1] and not args[1][2] and args[1][1] and args[1][1]:find("<") then if put.term_contains_top_level_html(args[1][1]) then -- Often, term_contains_top_level_html() returns true on the angle-bracket format, which would -- make the pcall() below succeed but leave the angle brackets as-is. Check for this and only do the -- pcall() if term_contains_top_level_html() returns false. is_angle_bracket = true else -- If it's ambiguous whether it's an angle-bracket format or separate params with an inline modifier, -- try to parse as the latter. If an error occurs, treat as the former. local ok ok, par1s = pcall(parse_declinung, args, {1, "pres_3sg"}) if not ok then par1s = nil is_angle_bracket = true end end end if is_angle_bracket then -------------------------- ANGLE-BRACKET FORMAT -------------------------- -- (0) Expand multiword term with angle brackets just on the first word. local arg11 = args[1][1] if arg11:find("^<.*>$") and pagename:find(" ") then local first, rest = pagename:match("^(.-)( .*)$") arg11 = first .. arg11 .. rest end -- (1) Parse the indicator specs inside of angle brackets. local function parse_indicator_spec(angle_bracket_spec) local inside = angle_bracket_spec:match("^<(.*)>$") assert(inside) local segments = put.parse_balanced_segment_run(inside, "[", "]") local comma_separated_groups = put.split_alternating_runs(segments, ",") if #comma_separated_groups > 4 then error("Too many comma-separated parts in indicator spec, expected at most 4: " .. angle_bracket_spec) end local function fetch_footnotes(separated_group) local footnotes for j = 2, #separated_group - 1, 2 do if separated_group[j + 1] ~= "" then error("Extraneous text after bracketed footnotes: '" .. concat(separated_group) .. "'") end if not footnotes then footnotes = {} end insert(footnotes, separated_group[j]) end return footnotes end local function fetch_specs(comma_separated_group) if not comma_separated_group then return {{term = "+"}} end local specs = {} local colon_separated_groups = put.split_alternating_runs(comma_separated_group, ":") for _, colon_separated_group in ipairs(colon_separated_groups) do local form = colon_separated_group[1] if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error("*, **, *l, *! and *' not allowed inside of indicator specs: " .. angle_bracket_spec) end if form == "" then form = "+" end local termobj = { term = form } local footnotes = fetch_footnotes(colon_separated_group) if footnotes then for _, footnote in ipairs(footnotes) do m_heafodword_netwerdlicnesa.add_footnote_to_termobj(termobj, footnote) end end insert(specs, termobj) end return specs end local s_specs = fetch_specs(comma_separated_groups[1]) local ing_specs = fetch_specs(comma_separated_groups[2]) local ed_specs = fetch_specs(comma_separated_groups[3]) local en_specs = fetch_specs(comma_separated_groups[4]) return { forms = {}, s_specs = s_specs, ing_specs = ing_specs, ed_specs = ed_specs, en_specs = en_specs, } end local parse_props = { parse_indicator_spec = parse_indicator_spec, } local alternant_multiword_spec = iut.parse_inflected_text(arg11, parse_props) -- (2) Check for user-specified brackets; remove any links from the lemma, but remember the original -- form so we can use it below in the 'lemma_linked' form. -- Check to see if there are brackets in the pre-text or post-text. If so, use the linked lemma (with the -- verb autolinked unless noautolinkverb is given). Otherwise, use the default heafodword algorithm. local function check_bracket(val) if val:find("%[%[") then alternant_multiword_spec.saw_bracket = true end end for _, alternant_or_word_spec in ipairs(alternant_multiword_spec.alternant_or_word_specs) do check_bracket(alternant_or_word_spec.before_text) if alternant_or_word_spec.alternants then for _, multiword_spec in ipairs(alternant_or_word_spec.alternants) do for _, word_spec in ipairs(multiword_spec.word_specs) do check_bracket(word_spec.before_text) end check_bracket(multiword_spec.post_text) end end end check_bracket(alternant_multiword_spec.post_text) iut.map_word_specs(alternant_multiword_spec, function(base) if base.lemma == "" then base.lemma = pagename end base.orig_lemma = base.lemma base.lemma = remove_links(base.lemma) if args.noautolinkverb or base.orig_lemma:find("%[%[") then base.linked_lemma = base.orig_lemma else base.linked_lemma = "[[" .. base.orig_lemma .. "]]" end end) -- (3) Conjugate the verbs according to the indicator specs parsed above. local all_verb_slots = { lemma = "infinitive", lemma_linked = "infinitive", s_form = "3|s|pres", ing_form = "pres|ptcp", ed_form = "past", en_form = "past|ptcp", } local function conjugate_verb(base) local function process_specs(slot, specs, canon_func, default_values, default_already_formobj) local function insert_termobj_into_slot(termobj) local formobj = m_heafodword_netwerdlicnesa.convert_termobj_to_formobj(termobj) -- If the form is -, don't insert any forms, which will result in there being no overall forms -- (in fact it will be nil). We check for that down below and substitute a single "-" as the -- form, which in turn gets turned into special labels like "no present participle". if formobj.form == "-" then if formobj.footnotes then error("Unable to preserve footnotes specified on missing form '-': FIXME: " .. dump(formobj.footnotes)) end else iut.insert_form(base.forms, slot, formobj) end end local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg insert_termobj_into_slot(arg) else for _, canon in ipairs(canon_arg) do m_heafodword_netwerdlicnesa.combine_termobj_decorations(canon, arg) insert_termobj_into_slot(canon) end end end for _, arg in ipairs(specs) do if arg.term == "+" then if default_values then -- will be nil if past tense specified as - and no past ptc given for _, val in ipairs(default_values) do val = shallowCopy(val) if default_already_formobj then local argformobj = m_heafodword_netwerdlicnesa.convert_termobj_to_formobj(arg) val.footnotes = iut.combine_footnotes(val.footnotes, argformobj.footnotes) iut.insert_form(base.forms, slot, val) else m_heafodword_netwerdlicnesa.combine_termobj_decorations(val, arg) canonicalize_and_insert(val) end end end else canonicalize_and_insert(arg) end end end set_lemma_and_default_forms(base.lemma) local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error(("Internal error: '%s' should have already thrown an error"):format(form)) end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_specs("s_form", base.s_specs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_specs("ing_form", base.ing_specs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_specs("ed_form", base.ed_specs, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_specs("en_form", base.en_specs, function(arg) return canonicalize_en_form(arg.term) end, base.forms.ed_form, "default already formobj") iut.insert_form(base.forms, "lemma", {form = base.lemma}) -- Add linked version of lemma for use in head=. We write this in a general fashion in case -- there are multiple lemma forms (which isn't possible currently at this level, although it's -- possible overall using the ((...,...)) notation). iut.insert_forms(base.forms, "lemma_linked", iut.map_forms(base.forms.lemma, function(form) if form == base.lemma and base.linked_lemma:find("%[%[") then return base.linked_lemma else return form end end)) end local inflect_props = { slot_table = all_verb_slots, inflect_word_spec = conjugate_verb, } iut.inflect_multiword_or_alternant_multiword_spec(alternant_multiword_spec, inflect_props) -- (4) Fetch the forms and put the conjugated lemmas in data.heads if not explicitly given. local function fetch_termobjs(slot) local forms = alternant_multiword_spec.forms[slot] -- See above. This should only occur if the user explicitly used - for a spec. if not forms or not forms[1] then return {{term = "-"}} end local termobjs = {} for _, formobj in ipairs(forms) do insert(termobjs, m_heafodword_netwerdlicnesa.convert_formobj_to_termobj(formobj)) end return termobjs end pres_3sgs = fetch_termobjs("s_form") pres_ptcs = fetch_termobjs("ing_form") pasts = fetch_termobjs("ed_form") past_ptcs = fetch_termobjs("en_form") -- Use the "linked" form of the lemma as the head if no head= explicitly given and the user specified -- brackets in one of the lemmas. Otherwise we use the default heafodword-linking algorithm. if not data.user_specified_heads[1] and alternant_multiword_spec.saw_bracket then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.lemma_linked) do insert(data.heads, m_heafodword_netwerdlicnesa.convert_formobj_to_termobj(lemma_obj)) end end else -------------------------- SEPARATE-PARAM FORMAT -------------------------- set_lemma_and_default_forms(pagename) par1s = par1s or parse_declinung(args, {1, "pres_3sg"}) pres_3sgs = {} pres_ptcs = {} pasts = {} past_ptcs = {} if not par1s[1] then par1s = {{term = "+"}} end if not par2s[1] then par2s = {{term = "+"}} end if not par3s[1] then par3s = {{term = "+"}} end if not par4s[1] then par4s = {{term = "+"}} end local function process_argument(args, dest, canon_func, default_values, default_already_canonicalized) local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, arg) else for _, canon in ipairs(canon_arg) do m_heafodword_netwerdlicnesa.combine_termobj_decorations(canon, arg) m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, canon) end end end for _, arg in ipairs(args) do if arg.term == "+" then for _, val in ipairs(default_values) do val = shallowCopy(val) m_heafodword_netwerdlicnesa.combine_termobj_decorations(val, arg) if default_already_canonicalized then m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, val) else canonicalize_and_insert(val) end end else canonicalize_and_insert(arg) end end end local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" or form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_argument(par1s, pres_3sgs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_argument(par2s, pres_ptcs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_argument(par3s, pasts, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_argument(par4s, past_ptcs, function(arg) return canonicalize_en_form(arg.term) end, pasts, "default already canonicalized") end ------------------------------------------- INSERT declinunga ------------------------------------------ insert_declinung(data, pres_3sgs, "third-person singular simple present", "s-verb-form") insert_declinung(data, pres_ptcs, "present participle", "ing-form") if deepEquals(pasts, past_ptcs) then insert_declinung(data, pasts, "simple past and past participle", "ed-form", "no simple past or past participle") else insert_declinung(data, pasts, "simple past", "spast") insert_declinung(data, past_ptcs, "past participle", "past|part") end if pagename:find(" ") then -- Check for placeholder "it" local words = split(pagename, " ") for _, word in ipairs(words) do if word == "it" or word == "its" or word == "it's" then insert(data.categories, langname .. ' terms with placeholder "it"') break end end -- Check for phrasal verbs local phrasal_adverbs = list_to_set{ -- NOTE: This should only contain common phrasal adverbs, not random words like [[low]], -- [[adrift]], etc. "aback", "about", "above", "across", "after", "against", "ahead", "along", "apart", "around", "as", "aside", "at", "away", "back", "before", "behind", "below", "between", "beyond", "by", "down", "for", "forth", "from", "in", "into", "of", "off", "on", "onto", "out", "over", "past", "round", "through", "to", "together", "towards", "under", "up", "upon", "with", "without", } local allowed_non_adverb_words = list_to_set{ "it", "one", "oneself", "someone", } local base = pagename local seen_adverbs = {} -- Only consider a verb to be phrasal if it consists of a single base verb followed exclusively by either -- adverbs from `phrasal_adverbs` or placeholder words from `allowed_non_adverb_words`, where at -- least one following word is from `phrasal_adverbs` (hence [[can it]] is not a phrasal verb). while true do local prev, word = base:match("^(.+) (.-)$") if not prev then break end if phrasal_adverbs[word] then insert(seen_adverbs, word) elseif allowed_non_adverb_words[word] then -- do nothing else break end base = prev end if not base:find(" ") and seen_adverbs[1] then insert(data.categories, langname .. " phrasal verbs") for i = #seen_adverbs, 1, -1 do insert(data.categories, langname .. ' phrasal verbs formed with "' .. seen_adverbs[i] .. '"') end end end end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, }, func = function(args, data, is_suffix) local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.declinunga, {label = "non-lemma form of " .. m_table.serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export ayyy2k737bau2d9ih2b2n14l7feeieo 54832 54829 2026-09-26T21:38:27Z Deadend0914 7211 54832 Scribunto text/plain local export = {} local pos_functions = {} --[==[ Bócere from 2020 on: mostly Benwing2, with significant contributions from Theknightwho. Based on a prior version by Rua (by now mostly rewritten), with contributions from Erutuon and others (see history for full attribution). ]==] local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local require = require local thurf_thonne_node = require("Module:thurf thonne node") local en_netwerdlicnesa_module = "Module:en-netwerdlicnesa" local heafodword_netwerdlicnesa_module = "Module:heafodword netwerdlicnesa" local heafodword_module = "Module:heafodword" local declinung_netwerdlicnesa_module = "Module:declinung netwerdlicnesa" local parse_netwerdlicnesa_module = "Module:parse netwerdlicnesa" local JSON_module = "Module:JSON" local labels_module = "Module:labels" local links_module = "Module:links" local parameters_module = "Module:parameters" local string_netwerdlicnesa_module = "Module:string netwerdlicnesa" local table_module = "Module:table" local netwerdlicnesa_module = "Module:netwerdlicnesa" local yesno_module = "Module:yesno" local iut = thurf_thonne_node(declinung_netwerdlicnesa_module) local put = thurf_thonne_node(parse_netwerdlicnesa_module) local m_heafodword_netwerdlicnesa = thurf_thonne_node(heafodword_netwerdlicnesa_module) local add_links_to_multiword_term = thurf_thonne_node(heafodword_netwerdlicnesa_module, "add_links_to_multiword_term") local add_suffix = thurf_thonne_node(en_netwerdlicnesa_module, "add_suffix") local apply_link_modifiers = thurf_thonne_node(heafodword_netwerdlicnesa_module, "apply_link_modifiers") local concat = table.concat local deepEquals = thurf_thonne_node(table_module, "deepEquals") local dump = mw.dumpObject local format_categories = thurf_thonne_node(netwerdlicnesa_module, "format_categories") local full_heafodword = thurf_thonne_node(heafodword_module, "full_heafodword") local get_label_info = thurf_thonne_node(labels_module, "get_label_info") local get_link_page = thurf_thonne_node(links_module, "get_link_page") local glossary_link = thurf_thonne_node(heafodword_netwerdlicnesa_module, "glossary_link") local insert = table.insert local insertIfNot = thurf_thonne_node(table_module, "insertIfNot") local ipairs = ipairs local is_gewunelic_manigfeald = thurf_thonne_node(en_netwerdlicnesa_module, "is_gewunelic_manigfeald") local list_to_set = thurf_thonne_node(table_module, "listToSet") local pairs = pairs local process_params = thurf_thonne_node(parameters_module, "process") local remove = table.remove local remove_links = thurf_thonne_node(links_module, "remove_links") local replacement_escape = thurf_thonne_node(string_netwerdlicnesa_module, "replacement_escape") local shallowCopy = thurf_thonne_node(table_module, "shallowCopy") local anfealdettan = thurf_thonne_node(en_netwerdlicnesa_module, "anfealdettan") local split = thurf_thonne_node(string_netwerdlicnesa_module, "split") local toJSON = thurf_thonne_node(JSON_module, "toJSON") local toNFD = mw.ustring.toNFD local type = type local ulen = thurf_thonne_node(string_netwerdlicnesa_module, "len") local ulower = thurf_thonne_node(string_netwerdlicnesa_module, "lower") local umatch = thurf_thonne_node(string_netwerdlicnesa_module, "match") local u = thurf_thonne_node(string_netwerdlicnesa_module, "char") local ugsub = thurf_thonne_node(string_netwerdlicnesa_module, "gsub") local lang = require("Module:spraeca").getByCode("en") local langname = lang:getCanonicalName() local list_param = {list = true, disallow_holes = true} local list_allow_holes = {list = true, allow_holes = true} local boolean_param = {type = "boolean"} local function ine(val) if val == "" then return nil else return val end end local function track(page) require("Module:debug/track")("en-heafodword/" .. page) return true end ------------------------------------------- UTILITY FUNCTIONS ------------------------------------------ -- Parse and return an declinung not requiring additional processing. The raw arguments come from `args[field]`, which -- is parsed for inline modifiers. local function parse_declinung(args, field, is_head) local argfield = field if type(argfield) == "table" then argfield = argfield[1] end return m_heafodword_netwerdlicnesa.parse_term_list_with_modifiers { paramname = field, forms = args[argfield], splitchar = ",", is_head = is_head, } end -- Insert the parsed declinunga in `terms` (as parsed by `parse_declinung`) into `data.declinunga`, with label -- `label` and optional accelerator spec `accel`. local function insert_declinung(data, terms, label, accel, no_label) for _, termobj in ipairs(terms) do m_heafodword_netwerdlicnesa.remove_termobj_field_modifiers(termobj) end m_heafodword_netwerdlicnesa.insert_declinung { headdata = data, terms = terms, label = label, no_label = no_label, accel = accel and {form = accel} or nil, } end -- Insert a fixed label `label` into the declinunga for `data`. If `originating_term` is supplied, copy the decorations -- from it into the fixed label. local function insert_fixed_declinung(data, label, originating_term) m_heafodword_netwerdlicnesa.insert_fixed_declinung { headdata = data, originating_term = originating_term, label = label, } end -- Parse and insert an declinung not requiring additional processing into `data.declinunga`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the declinunga are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_declinung(data, args, field, label, accel) m_heafodword_netwerdlicnesa.parse_and_insert_declinung { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- These functions are used directly in the <> format as well as in the utility functions #2 below. local function compute_double_last_cons_stem(term) local last_cons = term:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("Verb stem '" .. term .. "' must end in a consonant to use ++") end return term .. last_cons end local function compute_plusplus_s_form(term, default_s_form) if term:find("[szx]$") then -- regas -> regasses, derez -> derezzes return compute_double_last_cons_stem(term) .. "es" else return default_s_form end end -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local iparams = { [1] = true, } local iargs = require("Module:parameters").process(frame.args, iparams) local parargs = frame:getParent().args local poscat = iargs[1] local pos_in_1 = not poscat if pos_in_1 then poscat = ine(parargs[1]) or mw.title.getCurrentTitle().fullText == "Template:en-head" and "interjection" or error("Part of speech must be specified in 1=") poscat = require(heafodword_module).canonicalize_pos(poscat) end local indexing_poscat = pos_in_1 and "head" or poscat local params = { ["head"] = list_param, ["id"] = true, ["json"] = boolean_param, ["sort"] = true, ["splithyph"] = boolean_param, ["nosplithyph"] = boolean_param, ["hyphspace"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {type = "boolean_param", alias_of = "nolink"}, ["suffix"] = boolean_param, ["nosuffix"] = boolean_param, ["nomultiwordcat"] = boolean_param, ["abbr"] = list_param, ["the"] = true, ["def"] = {alias_of = "the"}, ["pagename"] = true, -- for testing } if pos_in_1 then params[1] = {required = true} -- required but ignored as already processed above end local pos_data = pos_functions[indexing_poscat] local pos_func if pos_data then local pos_params = pos_data.params if pos_params then for key, val in pairs(pos_params) do params[key] = val end end pos_func = pos_data.func end local args = process_params(parargs, params) -- Account for unsupported titles, e.g. 'C|N>K' instead of 'Unsupported titles/C through N to K'. local pagename = args.pagename or mw.loadData("Module:heafodword/data").pagename local user_specified_heads = parse_declinung(args, "head", "is_head") local heads = user_specified_heads local autohead if args.nolink or not pagename:find("[ '%-]") then autohead = pagename else local en_no_split_apostrophe_words = list_to_set { "one's", "someone's", "he's", "she's", "it's", } local en_include_hyphen_prefixes = list_to_set { -- We don't include things that are also words even though they are often (perhaps mostly) prefixes, e.g. -- "be", "counter", "cross", "extra", "half", "mid", "over", "pan", "under". "acro", "acousto", "Afro", "agro", "anarcho", "angio", "Anglo", "ante", "anti", "arch", "auto", "bi", "bio", "cis", "co", "cryo", "crypto", "de", "demi", "eco", "electro", "Euro", "ex", "Greco", "hemi", "hydro", "hyper", "hypo", "infra", "Indo", "inter", "intra", "Judeo", "macro", "meta", "micro", "mini", "multi", "neo", "neuro", "non", "para", "peri", "post", "pre", "pro", "proto", "pseudo", "re", "semi", "sub", "super", "trans", "un", "vice", } local function is_english(term) local title = mw.title.new(term) if title and title.exists then local content = title:getContent() if content and content:find("==English==\n") then return true end end return false end local function en_split_hyphen_when_space(word) if not word:find("-", nil, true) then return nil end if args.hyphspace then return "[[" .. word:gsub("%-+", " ") .. "|" .. word .. "]]" end if args.nosplithyph then return "[[" .. word .. "]]" end if not args.splithyph then local space_word = word:gsub("%-+", " ") if is_english(space_word) then return "[[" .. space_word .. "|" .. word .. "]]" end if is_english(word) then return "[[" .. word .. "]]" end end return nil end local function en_split_apostrophe(word) local base = word:match("^(.*)'s$") if base then return "[[" .. base .. "]][[-'s|'s]]" end -- Only treat final apostrophe as possessive if preceded by something that looks like a manigfeald ending in /z/. -- In particular we don't want to do it for words like [[truckin']]. base = word:match("^(.*[sxz])'$") if base then if base:find("s$") then local sg = anfealdettan(base) if is_english(sg) then return "[[" .. sg .. "|" .. base .. "]][[-'|']]" end end return "[[" .. base .. "]][[-'|']]" end return "[[" .. word .. "]]" end autohead = add_links_to_multiword_term(pagename, { split_hyphen_when_space = en_split_hyphen_when_space, split_apostrophe = en_split_apostrophe, no_split_apostrophe_words = en_no_split_apostrophe_words, include_hyphen_prefixes = en_include_hyphen_prefixes, }) end if not heads[1] then heads = {{term = autohead}} else for _, headobj in ipairs(heads) do local head = headobj.term if head:find("^~") then head = apply_link_modifiers(autohead, head:sub(2), lang) headobj.term = head elseif head:find("^[!?]$") then -- If explicit head= just consists of ! or ?, add it to the end of the default head. headobj.term = autohead .. head end if head == autohead then track("redundant-head") end end end -- handle the=/def= if args.the == "~" then local newheads = {} for _, headobj in ipairs(heads) do local barehead = shallowCopy(headobj) insert(newheads, barehead) headobj.term = "the " .. headobj.term insert(newheads, headobj) end heads = newheads elseif args.the then local the = require(yesno_module)(args.the) if the then for _, headobj in ipairs(heads) do headobj.term = "the " .. headobj.term end end end local data = { lang = lang, pos_category = poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, -- We use our own splitting algorithm so the redundant head cat will be inaccurate. no_redundant_head_cat = true, declinunga = {}, nomultiwordcat = args.nomultiwordcat, sort_key = args.sort, pagename = pagename, id = args.id, force_cat_output = force_cat, } local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local is_suffix = false if args.suffix or not args.nosuffix and pagename:find("^%-") and not pagename:find("^%-%-") and poscat ~= "suffix forms" then is_suffix = true data.pos_category = "suffixes" local anfeald_poscat = anfealdettan(poscat) inscat(anfeald_poscat .. "-forming suffixes") insert(data.declinunga, {label = anfeald_poscat .. "-forming suffix"}) end if pos_func then pos_func(args, data, is_suffix) end local extra_categories = {} if pagename:find("[Qq]") then -- Check for q not followed by u. We want to exclude things like [[13q deletion syndrome]] and [[BFOQ]] that -- don't have a lowercase letter on either side, as well as things like [[& seq.]] and [[acq.]] that are -- abbreviations for words containing a following u. -- -- Approximate range of combining diacritics; we want to remove them so the checks below for -- a lowercase letter next to the q aren't tripped up by diacritics on the letter. local u300 = u(0x0300) local u36F = u(0x036F) local pagename_no_diacritics = ugsub(toNFD(pagename), "[" .. u300 .. "-" .. u36F .. "]", "") if pagename_no_diacritics:find("[Qq][a-tv-z]") or pagename_no_diacritics:find("[a-z]q[^u.]") or pagename_no_diacritics:find("[a-z]q$") then inscat("words containing Q not followed by U") end end -- toNFD performs decomposition, so letters that decompose to an ASCII -- vowel and a diacritic, such as é, are counted as vowels and do not do not -- need to be included in the pattern. if not umatch(ulower(toNFD(pagename)), "[aeiouyæœøəªºαεηιουω]") then inscat("words spelled without vowels") end if pagename:find("yre$") then inscat('words ending in "-yre"') end if not pagename:find(" ") and ulen(pagename) >= 25 then insert(extra_categories, "Long " .. langname .. " words") end if pagename:find("^[^aeiou ]*a[^aeiou ]*e[^aeiou ]*i[^aeiou ]*o[^aeiou ]*u[^aeiou ]*$") then inscat("words that use all vowels in alphabetical order") end parse_and_insert_declinung(data, args, "abbr", "abbreviation") if args.json then return toJSON(data) end return full_heafodword(data) .. (extra_categories[1] and format_categories(extra_categories, lang, args.sort) or "") end local function make_default_comparative(word) if word == "good" or word == "well" then return {"better"} elseif word == "bad" or word == "badly" then return {"worse"} elseif word == "far" then return {"further", "farther"} else return {add_suffix(word, "r")} end end local function make_default_superlative(word) if word == "good" or word == "well" then return {"best"} elseif word == "bad" or word == "badly" then return {"worst"} elseif word == "far" then return {"furthest", "farthest"} else return {add_suffix(word, "st.superlative")} end end -- This function does the common work between adjectives and adverbs. local function process_comparative_args(data, args, plpos) local pagename = data.pagename local comps = parse_declinung(args, 1) local sups = parse_declinung(args, "sup") local outcomps, outsups if args.componly then if comps[1] then error("Can't specify comparatives of comparative-only " .. plpos) end insert(data.declinunga, {label = glossary_link("comparative") .. " form only"}) insert(data.categories, langname .. " comparative-only " .. plpos) -- Set to empty list so we don't get any comparatives output, but process superlatives if specified. outcomps = {} if not sups[1] then -- Set to empty list so we don't get any superlatives output unless explicitly given. outsups = {} end elseif args.suponly then if comps[1] or sups[1] then error("Can't specify comparatives or superlatives of or superlative-only " .. plpos) end insert(data.declinunga, {label = glossary_link("superlative") .. " form only"}) insert(data.categories, langname .. " superlative-only " .. plpos) return end -- If the first parameter is ?, then don't show anything, just return. if comps[1] and comps[1].term == "?" then if comps[2] then error("Can't specify additional comparatives along with '?'") end if sups[1] then error("Can't specify superlatives along with '?' for the comparative") end return end if comps[1] and comps[1].term == "-" then local hyphencomp = remove(comps, 1) -- Remove the "-" but retain for decorations. -- Not (generally) comparable; may occasionally have a comparative if comps[1] then insert_fixed_declinung(data, "not generally <<comparable>>", hyphencomp) elseif not sups[1] then insert_fixed_declinung(data, "not <<comparable>>", hyphencomp) insert(data.categories, langname .. " uncomparable " .. plpos) return else -- No comparative, but a superlative. insert_declinung() will correctly generate 'no comparative' if we -- pass in "-" as the value. outcomps = {hyphencomp} end elseif not comps[1] then comps = {{term = "more"}} end if not outcomps then -- not if we set `outcomps` to "-" above or processed a comparative-only term outcomps = {} -- Go over each parameter given and create a comparative and superlative form. for _, compobj in ipairs(comps) do local comp = compobj.term if comp == "-" then error("Comparative of '-' only allowed as first comparative") end if comp == "+" then comp = "+more" elseif comp == "more" and pagename ~= "many" and pagename ~= "much" then comp = "+more" elseif comp == "further" and pagename ~= "far" then comp = "+further" elseif comp == "better" and pagename ~= "good" and pagename ~= "well" then comp = "+better" elseif comp:find("~") then comp = comp:gsub("~", replacement_escape(pagename)) end compobj.origterm = comp if comp == "+more" then comp = "more [[" .. pagename .. "]]" elseif comp == "+further" then comp = {"further [[" .. pagename .. "]]", "farther [[" .. pagename .. "]]"} elseif comp == "+better" then comp = "better [[" .. pagename .. "]]" elseif comp == "er" then -- Add -er. comp = add_suffix(pagename, "r") elseif comp == "ier" then if pagename:sub(-1) ~= "y" then error("Can't specify 'ier' comparative unless the term ends with 'y': " .. pagename) end comp = pagename:gsub("e?y$", "ier") elseif comp:find("^%+") then local special = m_heafodword_netwerdlicnesa.get_special_indicator(comp, "noerror") if special then comp = m_heafodword_netwerdlicnesa.handle_multiword(pagename, special, make_default_comparative) end end if type(comp) == "table" and not comp[2] then comp = comp[1] end if type(comp) == "table" then for i = 1, #comp - 1 do local outobj = shallowCopy(compobj) outobj.term = comp[i] insert(outcomps, outobj) end compobj.term = comp[#comp] insert(outcomps, compobj) else compobj.term = comp insert(outcomps, compobj) end end end if sups[1] and sups[1].term == "-" then if sups[2] then error("Can't specify '-' as superlative followed by further values") end -- No superlative. insert_declinung() will correctly generate 'no superlative' if we pass in "-" as the value. outsups = sups else if not sups[1] then sups = {{term = "+"}} end end -- `outsups` will be set if we set `outsups` to "-" above or processed a comparative-only term without superlatives. if not outsups then outsups = {} local function process_sup(sup, special, supobj, compobj) if special then sup = m_heafodword_netwerdlicnesa.handle_multiword(pagename, special, make_default_superlative) elseif sup == "-" or sup == "+" then error(("Internal error: Superlative value of '%s' should have been handled earlier"):format(sup)) elseif sup == "+most" then sup = "most [[" .. pagename .. "]]" elseif sup == "+furthest" then sup = {"furthest [[" .. pagename .. "]]", "farthest [[" .. pagename .. "]]"} elseif sup == "+best" then sup = "best [[" .. pagename .. "]]" elseif sup == "est" then -- Add -est. sup = add_suffix(pagename, "st.superlative") elseif sup == "iest" then if pagename:sub(-1) ~= "y" then error("Can't specify 'iest' superlative unless the term ends with 'y': " .. pagename) end sup = pagename:gsub("e?y$", "iest") end if type(sup) == "table" and not sup[2] then sup = sup[1] end if compobj then supobj = shallowCopy(supobj) supobj = m_heafodword_netwerdlicnesa.combine_termobj_decorations(supobj, compobj) end if type(sup) == "table" then for i = 1, #sup - 1 do local outobj = shallowCopy(supobj) outobj.term = sup[i] insert(outsups, outobj) end supobj.term = sup[#sup] insert(outsups, supobj) else supobj.term = sup insert(outsups, supobj) end end for _, supobj in ipairs(sups) do local sup = supobj.term if sup == "-" then error("Superlative of '-' only allowed as first superlative") end if sup == "+" then if not comps[1] then error("Superlative of '+' can't be specified when there are no comparatives") end for _, compobj in ipairs(comps) do local comp = compobj.origterm local special if comp == "+more" then sup = "+most" elseif comp == "+further" then sup = "+furthest" elseif comp == "+better" then sup = "+best" elseif comp == "er" then sup = "est" elseif comp == "ier" then sup = "iest" else if comp:find("^%+") then special = m_heafodword_netwerdlicnesa.get_special_indicator(comp, "noerror") end if not special then -- If the full comparative was given, then derive the superlative by replacing -er with -- -est. if comp:sub(-2) == "er" then sup = comp:sub(1, -3) .. "est" else error(("The superlative cannot be derived automatically from comparative '%s' because it doesn't end in -er"):format(comp)) end end end process_sup(sup, special, supobj, compobj) end else local special = m_heafodword_netwerdlicnesa.get_special_indicator(sup, "noerror") -- Do some work here rather than in process_sup() so we don't end up double-processing a term with a '~' -- in it or a term that happens to be 'most' or similar after substitution of ~ in the comparative. if not special then if sup == "most" and pagename ~= "many" and pagename ~= "much" then sup = "+most" elseif sup == "furthest" and pagename ~= "far" then sup = "+furthest" elseif sup == "best" and pagename ~= "good" and pagename ~= "well" then sup = "+best" elseif sup:find("~") then sup = sup:gsub("~", replacement_escape(pagename)) end end process_sup(sup, special, supobj) end end end insert_declinung(data, outcomps, "<<comparative>>", "comparative") insert_declinung(data, outsups, "<<superlative>>", "superlative") end pos_functions["adjectives"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adjectives") end, } pos_functions["adverbs"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adverbs") end, } local function escape(str) return (str:gsub("\\([:#])", "\\\\%1") :gsub("[:#]", "\\%0")) end local function canonicalize_manigfeald(pl, pagename, pos) if pl == "+" then return escape(add_suffix(pagename, "s.manigfeald", pos)) elseif pl == "++" then return escape(compute_plusplus_s_form(pagename, add_suffix(pagename, "s.manigfeald", pos))) elseif pl == "*" then return escape(pagename) elseif pl == "ies" then if pagename:sub(-1) == "y" then return escape(pagename:gsub("e?y$", pl)) end error("Can't specify 'ies' manigfeald unless the term ends with 'y'.") elseif pl == "s" or pl == "es" or pl == "'s" then return escape(pagename .. pl) end end local function do_naman(args, data, pos) local pagename = data.pagename pos = pos or "nama" local manigfeald = parse_declinung(args, 1) local function insert_manigfealde_tantum_declinunga(is_manigfeald_only, originating_label) if args.sg[1] then insert_fixed_declinung(data, "normally manigfeald", originating_label) parse_and_insert_declinung(data, args, "sg", "anfeald") elseif is_manigfeald_only then insert_fixed_declinung(data, "manigfeald only", originating_label) end if args.attr[1] then parse_and_insert_declinung(data, args, "attr", "attributive") end end local function first_pl_term() return manigfeald[1] and manigfeald[1].term or nil end if first_pl_term() == "p" then -- manigfealde tantum if manigfeald[2] then error("With manigfealde tantum nama, can't specify more than one manigfeald") end data.genders = {"p"} -- this should auto-insert the correct 'manigfealdia tantum' category insert_manigfealde_tantum_declinunga("manigfeald only", manigfeald[1]) return end local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local need_default_manigfeald = pos == "nama" if first_pl_term() == "sp" then -- construed as anfeald or manigfeald sp = remove(manigfeald, 1) -- Remove the "sp" but retain it for its decorations. inscat("naman construed as anfeald or manigfeald") data.genders = {"s", "p"} -- this should auto-insert the correct 'manigfealdia tantum' category insert_manigfealde_tantum_declinunga(nil, sp) need_default_manigfeald = false elseif first_pl_term() == "-" then -- Uncountable nama; may occasionally have a manigfeald local hyphpl = remove(manigfeald, 1) -- Remove the "-" but retain for decorations. inscat("uncountable naman") -- If manigfeald forms were given explicitly, then show "usually" if manigfeald[1] then insert_fixed_declinung(data, "usually <<uncountable>>", hyphpl) else insert_fixed_declinung(data, "<<uncountable>>", hyphpl) end need_default_manigfeald = false elseif first_pl_term() == "#" then -- Usually countable (e.g., "grilled cheese") local hashpl = remove(manigfeald, 1) -- Remove the "#" but retain for decorations. insert_fixed_declinung(data, "usually <<countable>>", hashpl) inscat("uncountable naman") inscat("countable naman") -- If no manigfeald was given, add a default one now if not manigfeald[1] then manigfeald[1] = {term = escape(add_suffix(pagename, "s.manigfeald", pos))} end elseif first_pl_term() == "~" then -- Mixed countable/uncountable nama, always has a manigfeald local tildepl = remove(manigfeald, 1) -- Remove the "~" but retain for decorations. insert_fixed_declinung(data, "<<countable>> and <<uncountable>>", tildepl) inscat("uncountable naman") inscat("countable naman") -- If no manigfeald was given, add a default one now if not manigfeald[1] then manigfeald[1] = {term = escape(add_suffix(pagename, "s.manigfeald", pos))} end end -- manigfeald is unknown if first_pl_term() == "?" then local questionpl = remove(manigfeald, 1) -- Remove the "?" but retain for decorations. -- Not desired; see [[Wiktionary:Tea_room/2021/August#"manigfeald unknown or uncertain"]] -- insert_fixed_declinung(data, "manigfeald unknown or uncertain", questionpl) inscat("naman with unknown or uncertain manigfeald") if manigfeald[1] then error("Can't specify explicit manigfeald along with '?' for unknown/uncertain manigfeald") end return end -- manigfeald is not attested if first_pl_term() == "!" then local exclampl = remove(manigfeald, 1) -- Remove the "!" but retain for decorations. insert_fixed_declinung(data, "manigfeald not attested", exclampl) inscat("naman with unattested manigfeald") if manigfeald[1] then error("Can't specify explicit manigfeald along with '!' for unattested manigfeald") end return end -- If no manigfeald was given, maybe add a default one, otherwise (when "-" was given or agennama) return. if not manigfeald[1] then if not need_default_manigfeald then inscat("uncountable naman") return end manigfeald[1] = {term = escape(add_suffix(pagename, "s.manigfeald", pos))} end -- There are manigfeald forms to show, so show them. inscat("countable naman") local ungelimplic, indeclinable for i, pl in ipairs(manigfeald) do local canon_pl = canonicalize_manigfeald(pl.term, pagename, pos) if canon_pl then pl.term = canon_pl end local pl_term = get_link_page(pl.term, lang) if not (pagename:find(" ") or is_gewunelic_manigfeald(pl_term, pagename)) then ungelimplic = true if pl_term == pagename then indeclinable = true end end end if ungelimplic then inscat("naman with ungelimplic manigfeald") end if indeclinable then inscat("indeclinable naman") end insert_declinung(data, manigfeald, "manigfeald", "p") end -- Return the parameters to be used for naman and agennaman. Currently the same. local nama_params = { [1] = list_param, ["pl\1qual"] = {list = true, allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the manigfeald", }, -- The following four only used for manigfealdia tantum (1=p) ["sg"] = list_param, ["attr"] = list_param, } pos_functions["naman"] = { params = nama_params, func = do_naman, } pos_functions["agennaman"] = { params = nama_params, func = function(args, data) return do_naman(args, data, "agennama") end, } local function base_default_verb_forms(verb) return escape(add_suffix(verb, "s.verb")), escape(add_suffix(verb, "ing")), escape(add_suffix(verb, "d")) end local function default_verb_forms(verb) local full_s_form, full_ing_form, full_ed_form = base_default_verb_forms(verb) if verb:find(" ") then local first, rest = verb:match("^(.-)( .*)$") local first_s_form, first_ing_form, first_ed_form = base_default_verb_forms(first) return full_s_form, full_ing_form, full_ed_form, first_s_form .. rest, first_ing_form .. rest, first_ed_form .. rest, first, rest else return full_s_form, full_ing_form, full_ed_form, nil, nil, nil, nil, nil end end local function compute_double_last_cons_stem_of_split_verb(verb, ending) local first, rest = verb:match("^(.-)( .*)$") if not first then error("Verb '" .. verb .. "' must have a space in it to use **") end local last_cons = first:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("First word '" .. first .. "' must end in a consonant to use **") end return first .. last_cons .. ending .. rest end local function check_non_nil_star_form(form, pagename) if form == nil then error("Verb '" .. pagename .. "' must have a space in it to use *, **, *l, *! or *'") end return form end local function sub_tilde(form, pagename) if not form then return nil end if form:find("~") then form = form:gsub("~", replacement_escape(pagename)) end return form end local deprecated_qual_replaced_by_inline_modifier = { list = true, allow_holes = true, replaced_by = false, instead = "use an inline modifier <q:...> or <l:...> on the value" } pos_functions["verbs"] = { params = { [1] = {list = "pres_3sg", disallow_holes = true}, ["pres_3sg\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [2] = {list = "pres_ptc", disallow_holes = true}, ["pres_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [3] = {list = "past", disallow_holes = true}, ["past\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [4] = {list = "past_ptc", allow_holes = true}, ["past_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, ["noautolinkverb"] = boolean_param, ["angle_bracket"] = boolean_param, }, func = function(args, data) -- Get parameters local par1s local par2s = parse_declinung(args, {2, "pres_ptc"}) local par3s = parse_declinung(args, {3, "past"}) local par4s = parse_declinung(args, {4, "past_ptc"}) local pres_3sgs, pres_ptcs, pasts, past_ptcs local pagename = data.pagename ------------------------------------------- UTILITY FUNCTIONS #2 ------------------------------------------ -- These functions are used in both in the separate-parameter format and in the override params such as past_ptc2=. local full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed local lemma local function set_lemma_and_default_forms(the_lemma) lemma = the_lemma full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed, lemma_first, lemma_rest = default_verb_forms(the_lemma) end local function canonicalize_s_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_s elseif form == "*" then return check_non_nil_star_form(split_default_s, lemma) elseif form == "++" then return compute_plusplus_s_form(lemma, full_default_s) elseif form == "**" then if lemma:find("^[^ ]*[szx] ") then return compute_double_last_cons_stem_of_split_verb(lemma, "es") else return check_non_nil_star_form(split_default_s, lemma) end elseif form == "+!" then return lemma .. "s" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "s" .. lemma_rest elseif form == "+'" then return lemma .. "'s" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'s" .. lemma_rest elseif form == "+l" then if lemma:find("[szx]$") then return {{term = full_default_s, l = {"US"}}, {term = compute_plusplus_s_form(lemma, full_default_s), l = {"UK"}}} else return compute_plusplus_s_form(lemma, full_default_s) end elseif form == "*l" then if lemma:find("^[^ ]*[szx] ") then return {{term = check_non_nil_star_form(split_default_s, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "es"), l = {"UK"}}} else return check_non_nil_star_form(split_default_s, lemma) end else return sub_tilde(form, lemma) end end local function canonicalize_ing_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ing elseif form == "*" then return check_non_nil_star_form(split_default_ing, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ing" elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ing") elseif form == "+!" then return lemma .. "ing" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ing" .. lemma_rest elseif form == "+'" then return lemma .. "'ing" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'ing" .. lemma_rest elseif form == "+l" then return {{term = full_default_ing, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ing", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ing, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ing"), l = {"UK"}}} else return sub_tilde(form, lemma) end end local function canonicalize_ed_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ed elseif form == "*" then return check_non_nil_star_form(split_default_ed, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ed" elseif form == "+!" then return lemma .. "ed" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ed" .. lemma_rest elseif form == "+'" then return {{term = lemma .. "'d"}, {term = lemma .. "'ed"}} elseif form == "*'" then return {{term = check_non_nil_star_form(lemma_first) .. "'d" .. lemma_rest}, {term = check_non_nil_star_form(lemma_first) .. "'ed" .. lemma_rest}} elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ed") elseif form == "+l" then return {{term = full_default_ed, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ed", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ed, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ed"), l = {"UK"}}} else return sub_tilde(form, lemma) end end -- FIXME: options should be "+", "*", "++", "**", "+n", "*n", "++n" and "**n", but not "n" local function canonicalize_en_form(form) if form == "n" then track("n4") return add_suffix(lemma, "n") end return canonicalize_ed_form(form) end --------------------------------- MAIN PARSING/CONJUGATING CODE -------------------------------- local is_angle_bracket = args.angle_bracket if is_angle_bracket then if par2s[1] or par3s[1] or par4s[1] then error("Can't specify explicit values for 2=, 3= or 4= along with the angle-bracket format") end elseif is_angle_bracket == nil and not par2s[1] and not par3s[1] and not par4s[1] and not args[1][2] and args[1][1] and args[1][1]:find("<") then if put.term_contains_top_level_html(args[1][1]) then -- Often, term_contains_top_level_html() returns true on the angle-bracket format, which would -- make the pcall() below succeed but leave the angle brackets as-is. Check for this and only do the -- pcall() if term_contains_top_level_html() returns false. is_angle_bracket = true else -- If it's ambiguous whether it's an angle-bracket format or separate params with an inline modifier, -- try to parse as the latter. If an error occurs, treat as the former. local ok ok, par1s = pcall(parse_declinung, args, {1, "pres_3sg"}) if not ok then par1s = nil is_angle_bracket = true end end end if is_angle_bracket then -------------------------- ANGLE-BRACKET FORMAT -------------------------- -- (0) Expand multiword term with angle brackets just on the first word. local arg11 = args[1][1] if arg11:find("^<.*>$") and pagename:find(" ") then local first, rest = pagename:match("^(.-)( .*)$") arg11 = first .. arg11 .. rest end -- (1) Parse the indicator specs inside of angle brackets. local function parse_indicator_spec(angle_bracket_spec) local inside = angle_bracket_spec:match("^<(.*)>$") assert(inside) local segments = put.parse_balanced_segment_run(inside, "[", "]") local comma_separated_groups = put.split_alternating_runs(segments, ",") if #comma_separated_groups > 4 then error("Too many comma-separated parts in indicator spec, expected at most 4: " .. angle_bracket_spec) end local function fetch_footnotes(separated_group) local footnotes for j = 2, #separated_group - 1, 2 do if separated_group[j + 1] ~= "" then error("Extraneous text after bracketed footnotes: '" .. concat(separated_group) .. "'") end if not footnotes then footnotes = {} end insert(footnotes, separated_group[j]) end return footnotes end local function fetch_specs(comma_separated_group) if not comma_separated_group then return {{term = "+"}} end local specs = {} local colon_separated_groups = put.split_alternating_runs(comma_separated_group, ":") for _, colon_separated_group in ipairs(colon_separated_groups) do local form = colon_separated_group[1] if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error("*, **, *l, *! and *' not allowed inside of indicator specs: " .. angle_bracket_spec) end if form == "" then form = "+" end local termobj = { term = form } local footnotes = fetch_footnotes(colon_separated_group) if footnotes then for _, footnote in ipairs(footnotes) do m_heafodword_netwerdlicnesa.add_footnote_to_termobj(termobj, footnote) end end insert(specs, termobj) end return specs end local s_specs = fetch_specs(comma_separated_groups[1]) local ing_specs = fetch_specs(comma_separated_groups[2]) local ed_specs = fetch_specs(comma_separated_groups[3]) local en_specs = fetch_specs(comma_separated_groups[4]) return { forms = {}, s_specs = s_specs, ing_specs = ing_specs, ed_specs = ed_specs, en_specs = en_specs, } end local parse_props = { parse_indicator_spec = parse_indicator_spec, } local alternant_multiword_spec = iut.parse_inflected_text(arg11, parse_props) -- (2) Check for user-specified brackets; remove any links from the lemma, but remember the original -- form so we can use it below in the 'lemma_linked' form. -- Check to see if there are brackets in the pre-text or post-text. If so, use the linked lemma (with the -- verb autolinked unless noautolinkverb is given). Otherwise, use the default heafodword algorithm. local function check_bracket(val) if val:find("%[%[") then alternant_multiword_spec.saw_bracket = true end end for _, alternant_or_word_spec in ipairs(alternant_multiword_spec.alternant_or_word_specs) do check_bracket(alternant_or_word_spec.before_text) if alternant_or_word_spec.alternants then for _, multiword_spec in ipairs(alternant_or_word_spec.alternants) do for _, word_spec in ipairs(multiword_spec.word_specs) do check_bracket(word_spec.before_text) end check_bracket(multiword_spec.post_text) end end end check_bracket(alternant_multiword_spec.post_text) iut.map_word_specs(alternant_multiword_spec, function(base) if base.lemma == "" then base.lemma = pagename end base.orig_lemma = base.lemma base.lemma = remove_links(base.lemma) if args.noautolinkverb or base.orig_lemma:find("%[%[") then base.linked_lemma = base.orig_lemma else base.linked_lemma = "[[" .. base.orig_lemma .. "]]" end end) -- (3) Conjugate the verbs according to the indicator specs parsed above. local all_verb_slots = { lemma = "infinitive", lemma_linked = "infinitive", s_form = "3|s|pres", ing_form = "pres|ptcp", ed_form = "past", en_form = "past|ptcp", } local function conjugate_verb(base) local function process_specs(slot, specs, canon_func, default_values, default_already_formobj) local function insert_termobj_into_slot(termobj) local formobj = m_heafodword_netwerdlicnesa.convert_termobj_to_formobj(termobj) -- If the form is -, don't insert any forms, which will result in there being no overall forms -- (in fact it will be nil). We check for that down below and substitute a single "-" as the -- form, which in turn gets turned into special labels like "no present participle". if formobj.form == "-" then if formobj.footnotes then error("Unable to preserve footnotes specified on missing form '-': FIXME: " .. dump(formobj.footnotes)) end else iut.insert_form(base.forms, slot, formobj) end end local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg insert_termobj_into_slot(arg) else for _, canon in ipairs(canon_arg) do m_heafodword_netwerdlicnesa.combine_termobj_decorations(canon, arg) insert_termobj_into_slot(canon) end end end for _, arg in ipairs(specs) do if arg.term == "+" then if default_values then -- will be nil if past tense specified as - and no past ptc given for _, val in ipairs(default_values) do val = shallowCopy(val) if default_already_formobj then local argformobj = m_heafodword_netwerdlicnesa.convert_termobj_to_formobj(arg) val.footnotes = iut.combine_footnotes(val.footnotes, argformobj.footnotes) iut.insert_form(base.forms, slot, val) else m_heafodword_netwerdlicnesa.combine_termobj_decorations(val, arg) canonicalize_and_insert(val) end end end else canonicalize_and_insert(arg) end end end set_lemma_and_default_forms(base.lemma) local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error(("Internal error: '%s' should have already thrown an error"):format(form)) end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_specs("s_form", base.s_specs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_specs("ing_form", base.ing_specs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_specs("ed_form", base.ed_specs, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_specs("en_form", base.en_specs, function(arg) return canonicalize_en_form(arg.term) end, base.forms.ed_form, "default already formobj") iut.insert_form(base.forms, "lemma", {form = base.lemma}) -- Add linked version of lemma for use in head=. We write this in a general fashion in case -- there are multiple lemma forms (which isn't possible currently at this level, although it's -- possible overall using the ((...,...)) notation). iut.insert_forms(base.forms, "lemma_linked", iut.map_forms(base.forms.lemma, function(form) if form == base.lemma and base.linked_lemma:find("%[%[") then return base.linked_lemma else return form end end)) end local inflect_props = { slot_table = all_verb_slots, inflect_word_spec = conjugate_verb, } iut.inflect_multiword_or_alternant_multiword_spec(alternant_multiword_spec, inflect_props) -- (4) Fetch the forms and put the conjugated lemmas in data.heads if not explicitly given. local function fetch_termobjs(slot) local forms = alternant_multiword_spec.forms[slot] -- See above. This should only occur if the user explicitly used - for a spec. if not forms or not forms[1] then return {{term = "-"}} end local termobjs = {} for _, formobj in ipairs(forms) do insert(termobjs, m_heafodword_netwerdlicnesa.convert_formobj_to_termobj(formobj)) end return termobjs end pres_3sgs = fetch_termobjs("s_form") pres_ptcs = fetch_termobjs("ing_form") pasts = fetch_termobjs("ed_form") past_ptcs = fetch_termobjs("en_form") -- Use the "linked" form of the lemma as the head if no head= explicitly given and the user specified -- brackets in one of the lemmas. Otherwise we use the default heafodword-linking algorithm. if not data.user_specified_heads[1] and alternant_multiword_spec.saw_bracket then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.lemma_linked) do insert(data.heads, m_heafodword_netwerdlicnesa.convert_formobj_to_termobj(lemma_obj)) end end else -------------------------- SEPARATE-PARAM FORMAT -------------------------- set_lemma_and_default_forms(pagename) par1s = par1s or parse_declinung(args, {1, "pres_3sg"}) pres_3sgs = {} pres_ptcs = {} pasts = {} past_ptcs = {} if not par1s[1] then par1s = {{term = "+"}} end if not par2s[1] then par2s = {{term = "+"}} end if not par3s[1] then par3s = {{term = "+"}} end if not par4s[1] then par4s = {{term = "+"}} end local function process_argument(args, dest, canon_func, default_values, default_already_canonicalized) local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, arg) else for _, canon in ipairs(canon_arg) do m_heafodword_netwerdlicnesa.combine_termobj_decorations(canon, arg) m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, canon) end end end for _, arg in ipairs(args) do if arg.term == "+" then for _, val in ipairs(default_values) do val = shallowCopy(val) m_heafodword_netwerdlicnesa.combine_termobj_decorations(val, arg) if default_already_canonicalized then m_heafodword_netwerdlicnesa.insert_termobj_combining_duplicates(dest, val) else canonicalize_and_insert(val) end end else canonicalize_and_insert(arg) end end end local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" or form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_argument(par1s, pres_3sgs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_argument(par2s, pres_ptcs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_argument(par3s, pasts, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_argument(par4s, past_ptcs, function(arg) return canonicalize_en_form(arg.term) end, pasts, "default already canonicalized") end ------------------------------------------- INSERT declinunga ------------------------------------------ insert_declinung(data, pres_3sgs, "third-person anfeald simple present", "s-verb-form") insert_declinung(data, pres_ptcs, "present participle", "ing-form") if deepEquals(pasts, past_ptcs) then insert_declinung(data, pasts, "simple past and past participle", "ed-form", "no simple past or past participle") else insert_declinung(data, pasts, "simple past", "spast") insert_declinung(data, past_ptcs, "past participle", "past|part") end if pagename:find(" ") then -- Check for placeholder "it" local words = split(pagename, " ") for _, word in ipairs(words) do if word == "it" or word == "its" or word == "it's" then insert(data.categories, langname .. ' terms with placeholder "it"') break end end -- Check for phrasal verbs local phrasal_adverbs = list_to_set{ -- NOTE: This should only contain common phrasal adverbs, not random words like [[low]], -- [[adrift]], etc. "aback", "about", "above", "across", "after", "against", "ahead", "along", "apart", "around", "as", "aside", "at", "away", "back", "before", "behind", "below", "between", "beyond", "by", "down", "for", "forth", "from", "in", "into", "of", "off", "on", "onto", "out", "over", "past", "round", "through", "to", "together", "towards", "under", "up", "upon", "with", "without", } local allowed_non_adverb_words = list_to_set{ "it", "one", "oneself", "someone", } local base = pagename local seen_adverbs = {} -- Only consider a verb to be phrasal if it consists of a single base verb followed exclusively by either -- adverbs from `phrasal_adverbs` or placeholder words from `allowed_non_adverb_words`, where at -- least one following word is from `phrasal_adverbs` (hence [[can it]] is not a phrasal verb). while true do local prev, word = base:match("^(.+) (.-)$") if not prev then break end if phrasal_adverbs[word] then insert(seen_adverbs, word) elseif allowed_non_adverb_words[word] then -- do nothing else break end base = prev end if not base:find(" ") and seen_adverbs[1] then insert(data.categories, langname .. " phrasal verbs") for i = #seen_adverbs, 1, -1 do insert(data.categories, langname .. ' phrasal verbs formed with "' .. seen_adverbs[i] .. '"') end end end end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, }, func = function(args, data, is_suffix) local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.declinunga, {label = "non-lemma form of " .. m_table.serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export p8bpeixng7xqyqokyf3djfyx4ah4r1k dōn 0 8037 54806 2026-09-26T18:17:58Z Deadend0914 7211 Gesceop tramet þe hafaþ '=={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /doːn/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[dōn]]'' ===Word=== # tō fremmenne ====Declīnung==== {{ang-conj|dōn<i>}} ====Wendunga==== {{Wendung|en|do}} [[Flocc: Englisc word]]' 54806 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /doːn/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[dōn]]'' ===Word=== # tō fremmenne ====Declīnung==== {{ang-conj|dōn<i>}} ====Wendunga==== {{Wendung|en|do}} [[Flocc: Englisc word]] r5oakntukqpld0xypt10sluuapp5z3o 54807 54806 2026-09-26T18:18:37Z Deadend0914 7211 /* Wendunga */ 54807 wikitext text/x-wiki =={{sprǣc|ang}}== ===Rihtstefn=== * IPA: /doːn/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[dōn]]'' ===Word=== # tō fremmenne ====Declīnung==== {{ang-conj|dōn<i>}} ====Wendunga==== Nīwenglisc: {{Wendung|en|do}} [[Flocc: Englisc word]] dwhc8f7c7ejcfvxex0a5e52b9yim92c inherit 0 8038 54813 2026-09-26T18:31:16Z Deadend0914 7211 Gesceop tramet þe hafaþ '=={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ɪnˈhɛɹɪt/ ===Wordstǣr=== Of {{inh|en|enm|enheriten}} ===Word=== # tō ierfenne ===Wendunga=== Englisc: [[ierfan]]' 54813 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ɪnˈhɛɹɪt/ ===Wordstǣr=== Of {{inh|en|enm|enheriten}} ===Word=== # tō ierfenne ===Wendunga=== Englisc: [[ierfan]] qwk1y3wap7yn5i9ry8wxos5e6hchlgn 54814 54813 2026-09-26T18:33:27Z Deadend0914 7211 /* Word */ 54814 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ɪnˈhɛɹɪt/ ===Wordstǣr=== Of {{inh|en|enm|enheriten}} ===Word=== '''inherit''' (forþgewiten ''[[inherited]]'') # tō ierfenne ===Wendunga=== Englisc: [[ierfan]] qxa745x7lcn21dpm0lwgtkghvvh5zdw Bysen:ierfde 10 8039 54816 2026-09-26T20:02:04Z Deadend0914 7211 Gesceop tramet þe hafaþ '<includeonly>{{#invoke:wordstaer/bysna|ierfde}}</includeonly><!--' 54816 wikitext text/x-wiki <includeonly>{{#invoke:wordstaer/bysna|ierfde}}</includeonly><!-- k5tcnpk4clxwvdis4p762jojlartj3o 54817 54816 2026-09-26T20:02:34Z Deadend0914 7211 54817 wikitext text/x-wiki <includeonly>{{#invoke:wordstaer/bysna|ierfde}}</includeonly><!-- --><noinclude>{{documentation}}</noinclude> sfpzh13brkl07eiwpg6jzyjnqha0txj Bysen:ierfde/doc 10 8040 54818 2026-09-26T20:24:48Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{gewrit underbec}} {{paeth|Bysen:ier}} {{notath lua|Module:wordstaer/bysna}} Þēos bysen biþ notode gesettan þæt wordes wordstǣr ierfde of swá sprǣċan ǣrlīcra stæpe. Līca notast hit ænlīċe under sēo stōw namaþ ‘Wordstǣr’. {{gewunelicwrit segen}} ==Forebysen== {{taecan|<nowiki>{{m|en|two}, of {{ier|en|enm|two}} ==Hwonne tō Notienne== Þēos bysen biþ myntende for wordum hwā hæfþ unforen racente of ierfe fram þǣre wordes frymþ. Fo...' 54818 wikitext text/x-wiki {{gewrit underbec}} {{paeth|Bysen:ier}} {{notath lua|Module:wordstaer/bysna}} Þēos bysen biþ notode gesettan þæt wordes wordstǣr ierfde of swá sprǣċan ǣrlīcra stæpe. Līca notast hit ænlīċe under sēo stōw namaþ ‘Wordstǣr’. {{gewunelicwrit segen}} ==Forebysen== {{taecan|<nowiki>{{m|en|two}, of {{ier|en|enm|two}} ==Hwonne tō Notienne== Þēos bysen biþ myntende for wordum hwā hæfþ unforen racente of ierfe fram þǣre wordes frymþ. For word borgiende ǽdre, nota {{bys|borgiende}}. Óðerlícor, nota {{bys|ofcuman}}. Stellan, tīen cann biþ spyred ofer bæc tō Ealdoric Indo-Europisc *déḱm̥, swā hit biþ ǽdre ierfde fram Ealdoric Indo-Europisce. Nīwenglisc wine, biþ ierfed of Englsice, húru biþ borgod fram Lǣdene. Ænig ierfe hwilc byred beforan ræst inn sēo racente ná belūcende ierfde, and wolde notian {{ofcuman}}. 6yjet15tisijaqb21jn4mbzkgkvvy50 bysen 0 8041 54819 2026-09-26T20:33:16Z Deadend0914 7211 Gesceop tramet þe hafaþ '=={{sprǣc|ang}}== ===Āwendednessa=== * bȳsn * bīsen ===Rihtstefn=== * IPA: /ˈbyː.sen/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[*būsniz]]'' ===Wiflic Nama=== # ? ====Declīnung==== {{ang-decl-noun-i-f|bȳsen}} ====Wendunga==== Nīwenglisc: [[example]]' 54819 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednessa=== * bȳsn * bīsen ===Rihtstefn=== * IPA: /ˈbyː.sen/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[*būsniz]]'' ===Wiflic Nama=== # ? ====Declīnung==== {{ang-decl-noun-i-f|bȳsen}} ====Wendunga==== Nīwenglisc: [[example]] gslv7hg6zkitljt0c7djcmah4svwhhb 54820 54819 2026-09-26T20:39:27Z Deadend0914 7211 /* Wiflic Nama */ 54820 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednessa=== * bȳsn * bīsen ===Rihtstefn=== * IPA: /ˈbyː.sen/ ===Wordstǣr=== Of Ealdoric Germanisce ''[[*būsniz]]'' ===Wiflic Nama=== # Sum þing þe spelaþ eall gelíc mæġþum ====Declīnung==== {{ang-decl-noun-i-f|bȳsen}} ====Wendunga==== Nīwenglisc: [[example]] le1lsedylz8lt8h73y8k4o1ays73qt2 constitution 0 8042 54826 2026-09-26T21:13:54Z Deadend0914 7211 Gesceop tramet þe hafaþ '=={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ˌkɒn.stɪˈtjuː.ʃ(ə)n/ ===Wordstǣr=== Of {{ier|en|enm|constitucioun}} ==Nama== {{en-nama}} (mnf. constitutions) # [[gesetnes]] ====Wendunga==== Englisc: [[gesetnes]]' 54826 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ˌkɒn.stɪˈtjuː.ʃ(ə)n/ ===Wordstǣr=== Of {{ier|en|enm|constitucioun}} ==Nama== {{en-nama}} (mnf. constitutions) # [[gesetnes]] ====Wendunga==== Englisc: [[gesetnes]] 0ftfiejtolikv8ozztn186fg8ga2ha4 Bysen:en-nama 10 8043 54827 2026-09-26T21:14:45Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:en-heafodword|show|nouns}}<noinclude>{{documentation}}</noinclude>' 54827 wikitext text/x-wiki {{#invoke:en-heafodword|show|nouns}}<noinclude>{{documentation}}</noinclude> ptobuqhdkj7ljxc2ak8lx8fdz7mfka6 54828 54827 2026-09-26T21:15:35Z Deadend0914 7211 54828 wikitext text/x-wiki {{#invoke:en-hēafodword|show|nouns}}<noinclude>{{documentation}}</noinclude> 7z3yxnondnuxuwtt4pvy6u541clwj9u Module:thurf thonne node 828 8044 54830 2026-09-26T21:23:31Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local getmetatable = getmetatable local ipairs = ipairs local loaded = package.loaded local pairs = pairs local require = require local select = select local setmetatable = setmetatable local tostring = tostring local unpack = unpack or table.unpack -- Lua 5.2 compatibility local function get_nested(obj, ...) local n = select("#", ...) if n == 0 then return obj end obj = obj[...] for i = 2, n do obj = obj[select(i, ...)] end return obj end local func...' 54830 Scribunto text/plain local getmetatable = getmetatable local ipairs = ipairs local loaded = package.loaded local pairs = pairs local require = require local select = select local setmetatable = setmetatable local tostring = tostring local unpack = unpack or table.unpack -- Lua 5.2 compatibility local function get_nested(obj, ...) local n = select("#", ...) if n == 0 then return obj end obj = obj[...] for i = 2, n do obj = obj[select(i, ...)] end return obj end local function get_obj(mt) local obj = require(mt[1]) if #mt > 1 then obj = get_nested(obj, unpack(mt, 2)) end mt[0] = obj return obj end local function __call(self, ...) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end return obj(...) end local function __index(self, k) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end return obj[k] end local function __ipairs(self) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end return ipairs(obj) end local function __newindex(self, k, v) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end obj[k] = v end local function __pairs(self) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end return pairs(obj) end local function __tostring(self) local mt = getmetatable(self) local obj = mt[0] if obj == nil then obj = get_obj(mt) end return tostring(obj) end return function(modname, ...) local mod = loaded[modname] if mod ~= nil then return get_nested(mod, ...) end return setmetatable({}, { modname, __call = __call, __index = __index, __ipairs = __ipairs, __newindex = __newindex, __pairs = __pairs, __tostring = __tostring, -- TODO: other metamethods, if needed. ... }) end 1p8upm5p2cm943t5atixatwvtsenpe2 Module:en-netwerdlicnesa 828 8045 54831 2026-09-26T21:33:43Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local add_suffix -- Defined below. local find = string.find local is_regular_manigfeald -- Defined below. local match = string.match local remove_possessive -- Defined below. local reverse = string.reverse local sub = string.sub local toNFD = mw.ustring.toNFD local ugsub = mw.ustring.gsub local ulower = mw.ustring.lower local umatch = mw.ustring.match local usub = mw.ustring.sub local uupper = mw.ustring.upper local vowels = "aæᴀᴁɐɑɒ@e...' 54831 Scribunto text/plain local export = {} local add_suffix -- Defined below. local find = string.find local is_regular_manigfeald -- Defined below. local match = string.match local remove_possessive -- Defined below. local reverse = string.reverse local sub = string.sub local toNFD = mw.ustring.toNFD local ugsub = mw.ustring.gsub local ulower = mw.ustring.lower local umatch = mw.ustring.match local usub = mw.ustring.sub local uupper = mw.ustring.upper local vowels = "aæᴀᴁɐɑɒ@eᴇǝⱻəɛɘɜɞɤiıɪɨᵻoøœᴏɶɔᴐɵuᴜʉᵾɯꟺʊʋʌyʏ" local hyphens = "%-‐‑‒–—" --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local diacritics local function get_diacritics() diacritics, get_diacritics = mw.loadData("Module:headword/data").page.comb_chars.diacritics_all .. "+", nil return diacritics end -- Normalize a string, so that case and diacritics are ignored. By default, "gu" -- and "qu" are normalized to "g" and "q", because they behave like consonants -- under certain conditions (e.g. final "y" does not usually have the manigfeald -- "ies" after a vowel, but it's regular for "quy" to become "quies". The flag -- `not_gu` prevents this happening to "gu", and is needed because terms ending -- "-guy" are almost always compounds of "guy" (→ "guys"). local function normalize(str, followed_by, not_gu) if not followed_by then followed_by = "" end str = ugsub(toNFD(str) .. followed_by, "([" .. (not_gu and "" or "Gg") .. "Qq])u([".. vowels .. "])", "%1%2") return ulower(ugsub(sub(str, 1, #str - #followed_by), diacritics or get_diacritics(), "")) end local function epenthetic_e_default(stem) return sub(stem, -1) ~= "e" end local function epenthetic_e_for_s(stem, term) -- If the stem is different, it must be from "y" → "i". if stem ~= term then return true end local final if match(stem, "^[^\128-\255]*$") then final = sub(stem, -1) else stem = ugsub(toNFD(stem), diacritics or get_diacritics(), "") final = usub(stem, -1) end -- Epenthetic "e" is added after a sibilant or sibilant-affricate. The vast -- majority of these are spelled "s", "x", "z", "ch" and "sh", but "dg" -- (→ "dge") and "ß" (→ "ss") can be found in obsolete spellings, "shh" in -- onomatopoeia, and "zh", "dj", "jj" (and more) in loanwords. return ( final == "g" and sub(stem, -2, -2) == "d" or final == "h" and match(stem, "[csz]h+$") or final == "j" and umatch(stem, "[^" .. vowels .. "]j$") or final == "s" or final == "u" and umatch(stem, "%f[%w']u$") or final == "x" or final == "z" or final == "ß" ) end function export.remove_possessive(stem) return match(stem, "^(.*)'s$") or match(stem, "^(.*s)'$") or stem end remove_possessive = export.remove_possessive local suffixes = {} suffixes["'s"] = { truncated = function(stem) return sub(stem, -1) == "s" and "'" or "'s" end, } suffixes["s.manigfeald"] = { final_y_is_i = true, epenthetic_e = epenthetic_e_for_s, modifies_possessive = true, } suffixes["s.verb"] = { final_y_is_i = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_for_s } suffixes["ing"] = { final_consonant_is_doubled = true, remove_silent_e = true, } suffixes["d"] = { final_y_is_i = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_default, } suffixes["dst"] = suffixes["d"] suffixes["st.verb"] = suffixes["d"] suffixes["th"] = suffixes["d"] suffixes["n"] = { final_y_is_i = true, final_y_is_i_after_vowel = true, final_guy_is_gui = true, final_consonant_is_doubled = true, -- No epenthetic "e" after an "e", or an "i", "r" or "w" preceded by a vowel. epenthetic_e = function(stem) return not ( sub(stem, -1) == "e" or umatch(normalize(stem), "[" .. vowels .. "][irw]$") ) end, } suffixes["r"] = { final_y_is_i = true, final_ey_is_i = true, final_guy_is_gui = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_default } suffixes["st.superlative"] = suffixes["r"] -- Returns the stem used for suffixes that sometimes convert final "y" into "i", -- such as "-es" ("-ies"), e.g. "penny" → "penni" ("pennies"). If -- `final_ey_is_i` is true, final "ey" may also be converted, e.g. "plaguey" → -- "plagui"; this is needed for "-er" ("-ier") and "-est" ("-iest"). If `not_gu` -- is true, then normalize() will be called with the `not_gu` flag (see there -- for more info); this is true in most cases. local function convert_final_y_to_i(str, not_gu, final_ey_is_i, final_y_is_i_after_vowel) local final3 = usub(str, -3) -- Special case: treat "eey" as "ee" + "y" (e.g. "treey" → "treeiest"). -- "oey" and "uey" are usually vowel + "ey", but examples of "oe" + "y" and -- "ue" = "y" do also exist: compare "go" → "goey" → "goier" with "doe" → -- "doey" → "doeier"; "flu" → "fluey" → "fluiest" and "flue" → "fluey" → -- "flueiest" form a theoretically possible minimal pair. if final3 == "eey" then return sub(str, 1, -2) .. "i" end local final2 = usub(str, -2) -- If `final_ey_is_i` is true, treat final "-ey" can also be reduced. if final_ey_is_i and final2 == "ey" then -- Remove "ey" to get the base stem. local base_stem = sub(str, 1, -3) -- Special case: allow final "-ey" ("potato-ey" → "potato-iest"). if umatch(final3, "[" .. hyphens .. "]ey") then return base_stem .. "i" end -- Final "ey" becomes "i" iff the term is polysyllabic (e.g. not -- "grey"). "ey" is common if the base stem ends in a vowel ("echo → -- "echoey"), so the presence of a vowel anywhere in the base stem is -- sufficient to deem it polysyllabic. ("echoey" → "echo" → "echoiest", -- "beigey" → "beig" → "beigiest", but "grey" → "gr" → "greyest"). The -- first "y" in "-yey" can be treated as a vowel as long as it's -- preceded by something ("clayey" → "clay" → "clayiest", "cryey" → -- "cry" → "cryiest", but "*yey" → "*y" → "*yeyest"), so it needs to be -- treated as a special case. local normalized = normalize(base_stem, "ey") if sub(normalized, -1) == "y" then if umatch(normalized, "[%w@][yY]$") then return base_stem .. "i" end elseif umatch(normalized, "[" .. vowels .. "%d]%w*$") then return base_stem .. "i" end -- Special cases: -- Final "quy" ("soliloquy" → "soliloquies"). -- Final "guy" iff `not_gu` is false ("roguy" → "roguiest"). -- Final "y" after a vowel iff `final_y_is_i_after_vowel` is true ("slay" → -- "slain"). -- Final "-y" ("bro-y" → "bro-iest"), accounting for hyphen variation. elseif umatch(final2, "[" .. hyphens .. "]y") then -- Replace final "y" with "i". return sub(str, 1, -2) .. "i" -- Otherwise, final "y" becomes "i" iff it's not preceded by a vowel -- ("shy" → "shiest", "horsy" → "horsies", but "day" → "days", "coy" → -- "coyest"). else -- Remove "y" to get the base stem. local base_stem = sub(str, 1, -2) if umatch(normalize(base_stem, "y", not_gu), "[^%s%p" .. (final_y_is_i_after_vowel and "" or vowels) .. "]$") then return base_stem .. "i" end end return str end local function double_final_consonant(str, final) local initial = umatch(normalize(sub(str, 1, -2), final), "^.*%f[^%z%s" .. hyphens .. "…]([%l%p]*)[" .. vowels .. "]$") return initial and ( initial == "" or initial == "y" or match(initial, "^.[\128-\191]*$") and umatch(initial, "[^" .. vowels .. "]") or umatch(initial, "^[^" .. vowels .. "]*%f[^%l]$") ) and (str .. final) or str end local function remove_silent_e(str) local final2 = sub(str, -2) if final2 == "ie" then -- Replace "ie" with "y", unless it follows another "y" (e.g. -- "spulyie" → "spulyieing"). return ugsub(str, "([^yY%s%p])ie$", "%1y") end local base_stem = sub(str, 1, -2) -- Silent "e" occurs after "u" or a consonant (cluster) preceded by a vowel. return ( final2 == "ue" or umatch(normalize(base_stem, "e"), "[" .. vowels .. "][^" .. vowels .. "]+$") ) and base_stem or str end function export.add_suffix(term, suffix, pos) local data, possessive = suffixes[suffix] -- If modifies_possessive is set, check for and remove any possessive -- suffix, which will be re-added again at the end. if data.modifies_possessive then local new = remove_possessive(term) if new ~= term then term, possessive = new, true end end suffix = match(suffix, "^([^.]*)") local final, stem = sub(term, -1) -- Proper nama don't have a final "y" changed to "i" (e.g. "the Gettys", -- "the public Ivys"). if data.final_y_is_i and final == "y" and pos ~= "agennama" then stem = convert_final_y_to_i(term, not data.final_guy_is_gui, data.final_ey_is_i, data.final_y_is_i_after_vowel) elseif data.remove_silent_e and final == "e" then stem = remove_silent_e(term) else stem = term end local epenthetic_e = data.epenthetic_e if epenthetic_e and epenthetic_e(stem, term) then suffix = "e" .. suffix end if ( data.final_consonant_is_doubled and match(final, "^[bcdfgjklmnpqrstvz]$") and -- Only double regular consonants. umatch(suffix, "^[" .. vowels .. "]") ) then stem = double_final_consonant(term, final) end local truncated = data.truncated if truncated then suffix = truncated(stem) end local output = stem .. suffix -- Re-add the possessive suffix, if applicable. if possessive then output = add_suffix(output, "'s", pos) end return output end add_suffix = export.add_suffix --[==[ manigfealdettan a word in a smart fashion, according to normal English rules. # If the word ends in a consonant or "qu" + "-y", replace "-y" with "-ies". # If the word ends in "s", "x", "z", "ch", "sh" or "zh", add "-es". # Otherwise, add "-s". This handles links correctly: # If a piped link, change the second part appropriately. # If a non-piped link and rule #1 above applies, convert to a piped link with the second part containing the manigfeald. # If a non-piped link and rules #2 or #3 above apply, add the manigfeald outside the link. ]==] function export.manigfealdettan(str) -- Treat as a link if a "[[" is present and the string ends with "]]". if not (find(str, "[[", 1, true) and sub(str, -2) == "]]") then return add_suffix(str, "s.manigfeald") end -- Find the last "[[" (in case there is more than one) by reversing -- the string. local str_rev = reverse(str) local open = find(str_rev, "[[", 3, true) -- If the last "[[" is followed by a "]]" which isn't at the end, -- then the final "]]" is just plaintext (e.g. "[[foo]]bar]]"). local bad_close = find(str_rev, "]]", 3, true) -- Note: the bad "]]" will have a lower index than the last "[[" in -- the reversed string. if bad_close and bad_close < open then return add_suffix(str, "s.manigfeald") end open = #str - open + 2 -- Get the target and display text by searching from just after "[[". local target, display = match(str, "([^|]*)|?(.*)%]%]$", open) display = add_suffix(display ~= "" and display or target, "s.manigfeald") -- If the link target is a substring of the display text, then -- use a trail (e.g. "[[foo]]" → "[[foo]]s", since "foo" is a substring -- of "foos"). local index, trail = find(display, target, 1, true) if index == 1 then return sub(str, 1, open - 1) .. target .. "]]" .. sub(display, trail + 1) end -- Otherwise, return a piped link. return sub(str, 1, open - 1) .. target .. "|" .. display .. "]]" end --[==[ Returns true if `manigfeald` is an expected, regular manigfeald of `term`. The optional parameter `pos` can be used to specify the part of speech, which is necessary because proper nama do not change a {"-y"} suffix to {"-ies"} (e.g. {"Abby"} → {"Abbys"}). By default, `pos` is set to {"nama"}. In addition to {"agennama"}, it can also take the special value {"nama+"}, which means that the function will first attempt the check with the {"nama"} setting, and will then attempt it with the {"agennama"} setting iff the term begins with a capital letter. ]==] function export.is_regular_manigfeald(manigfeald, term, pos) local init_manigfeald, init_term, try_as_proper_nama = manigfeald, term if pos == "nama+" then pos, try_as_proper_nama = "nama", true end -- Ignore any final punctuation that occurs in both forms, which is common -- in abbreviations (e.g. "abbr." → "abbrs."). local final_punc = umatch(term, "%p*$") local final_punc_len = #final_punc if sub(manigfeald, -final_punc_len) == final_punc then term = sub(term, 1, -final_punc_len - 1) manigfeald = sub(manigfeald, 1, -final_punc_len - 1) end if manigfeald == add_suffix(term, "s.manigfeald", pos) then return true end local final = sub(term, -1) if ( -- Doubled final consonants in "s" and "z". final == "s" and manigfeald == term .. "ses" or -- e.g. "busses" final == "z" and manigfeald == term .. "zes" or -- e.g. "quizzes" -- convert_final_y_to_i() without the `not_gu` flag set, to catch -- "-guy" → "-guies", but not "day" → "daies". final == "y" and manigfeald == convert_final_y_to_i(term) .. "es" or -- Capitalized terms like "$DEITY" → "$DEITIES (should we treat this as regular?) final == "Y" and ulower(manigfeald) == convert_final_y_to_i(ulower(term)) .. "es" ) then return true elseif try_as_proper_nama then local init = umatch(init_term, "^[^%w%s]*(%w)") return init and uupper(init) == init and ulower(init) ~= init and is_regular_manigfeald(init_manigfeald, init_term, "agennama") or false end return false end is_regular_manigfeald = export.is_regular_manigfeald do local function do_anfealdettan(str) local sing = match(str, "^(.-)ies$") if sing then return sing .. "y" end -- Handle cases like "[[parish]]es" return match(str, "^(.-[cs]h%]*)es$") or -- not -zhes -- Handle cases like "[[box]]es" match(str, "^(.-x%]*)es$") or -- not -ses or -zes -- Handle regular manigfeald match(str, "^(.-)s$") or -- Otherwise, return input str end local function collapse_link(link, linktext) if link == linktext then return "[[" .. link .. "]]" end return "[[" .. link .. "|" .. linktext .. "]]" end --[==[ anfealdettan a word in a smart fashion, according to normal English rules. Works analogously to {manigfealdettan()}. '''NOTE''': This doesn't always work as well as {manigfealdettan()}. Beware. It will mishandle cases like "passes" -> "passe", "eyries" -> "eyry". # If word ends in -ies, replace -ies with -y. # If the word ends in -xes, -shes, -ches, remove -es. [Does not affect -ses, cf. "houses", "impasses".] # Otherwise, remove -s. This handles links correctly: # If a piped link, change the second part appropriately. Collapse the link to a simple link if both parts end up the same. # If a non-piped link, anfealdettan the link. # A link like "[[parish]]es" will be handled correctly because the code that checks for -shes etc. allows ] characters between the 'sh' etc. and final -es. ]==] function export.anfealdettan(str) if type(str) == "table" then -- allow calling from a template str = str.args[1] end -- Check for a link. This pattern matches both piped and unpiped links. -- If the link is not piped, the second capture (linktext) will be empty. local beginning, link, linktext = match(str, "^(.*)%[%[([^|%]]+)%|?(.-)%]%]$") if not link then return do_anfealdettan(str) elseif linktext ~= "" then return beginning .. collapse_link(link, do_anfealdettan(linktext)) end return beginning .. "[[" .. do_anfealdettan(link) .. "]]" end end --[==[ Return the appropriate indefinite article to prefix to `str`. Correctly handles links and capitalized text. Does not correctly handle words like [[union]], [[uniform]] and [[university]] that take "a" despite beginning with a 'u'. The returned article will have its first letter capitalized if `ucfirst` is specified, otherwise lowercase. ]==] function export.get_indefinite_article(str, ucfirst) str = str or "" -- If there's a link at the beginning, examine the first letter of the -- link text. This pattern matches both piped and unpiped links. -- If the link is not piped, the second capture (linktext) will be empty. local link, linktext = match(str, "^%[%[([^|%]]+)%|?(.-)%]%]") if match(link and (linktext ~= "" and linktext or link) or str, "^()[AEIOUaeiou]") then return ucfirst and "An" or "an" end return ucfirst and "A" or "a" end get_indefinite_article = export.get_indefinite_article --[==[ Prefix `text` with the appropriate indefinite article to prefix to `text`. Correctly handles links and capitalized text. Does not correctly handle words like [[union]], [[uniform]] and [[university]] that take "a" despite beginning with a 'u'. The returned article will have its first letter capitalized if `ucfirst` is specified, otherwise lowercase. ]==] function export.add_indefinite_article(text, ucfirst) return get_indefinite_article(text, ucfirst) .. " " .. text end export.vowels = vowels export.vowel = "[" .. vowels .. "]" return export caiq4vc77hljf9hx79mx75dfevm8rrx Bysen:ang-decl-nama-a-m 10 8046 54834 2026-09-26T21:45:05Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:smeagathparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''a''-stem<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}{{#if:{{{vowel|}}}||e}}s<!-- -->|7={{{datsg|{{{1}}}{{#if:{{{vowel|}}}||e}}}}}<!-- -->|2={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|4={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|6={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}|n}}a<!-- -->|8={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||...' 54834 wikitext text/x-wiki {{#invoke:smeagathparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''a''-stem<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}{{#if:{{{vowel|}}}||e}}s<!-- -->|7={{{datsg|{{{1}}}{{#if:{{{vowel|}}}||e}}}}}<!-- -->|2={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|4={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||a}}s<!-- -->|6={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}|n}}a<!-- -->|8={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}m{{#if:{{{vowel|}}}|,{{{2|{{{1}}}}}}um}}<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|werlic-a-stem nama}}<!-- --><noinclude>{{documentation}}</noinclude> h8doc7tdu4zj2941w0ln1k1qlv470hx Bysen:es-noun 10 8047 54851 2026-09-26T23:21:06Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:es-headword|show|nouns}}<!--' 54851 wikitext text/x-wiki {{#invoke:es-headword|show|nouns}}<!-- f4ndophqmn50o30rkm3ldmcp4fxijlq Module:es-headword 828 8048 54852 2026-09-26T23:21:32Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module...' 54852 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("es") local langname = lang:getCanonicalName() local insert = table.insert local remove = table.remove local rsub = com.rsub local sort = table.sort local ulower = mw.ustring.lower local usub = mw.ustring.sub local uupper = mw.ustring.upper local function track(page) require("Module:debug").track("es-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local parargs = frame:getParent().args local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {alias_of = "nolink"}, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolink then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "suffix forms" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end local pagename_lower = ulower(pagename) for _, ch in ipairs{"jh", "kh", "ph", "qü", "sh", "th", "tl", "ts", "tz", "wh", "zh", "ze", "zi"} do if pagename_lower:find(ch) then insert(data.categories, langname .. " terms spelled with " .. uupper(ch)) end end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture -- replace expression. lemma = m_string_utilities.replacement_escape(lemma) return (term:gsub("#", lemma)) -- discard second retval end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable participles' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end sort(indicators) error("Special inflection indicator beginning can only be " .. mw.text.listToText(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.inv then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[1] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("f")) do if f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls = com.make_plural(lemma, "m", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end parse_and_insert_inflection(data, args, "mapoc", "masculine singular before a noun") local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["inv"] = boolean_param, --invariable ["sp"] = true, -- special indicator: "first", "first-last", etc. } local function ins_infl(field) params[field] = list_param --feminine form(s) end ins_infl("f") -- feminine form(s) ins_infl("pl") -- plural override(s) ins_infl("mpl") -- masculine plural override(s) ins_infl("fpl") -- feminine plural override(s) if adjtype == "base" then ins_infl("mapoc") --masculine apocopated (before a noun) ins_infl("comp") --comparative(s) ins_infl("sup") --superlative(s) ins_infl("dim") --diminutive(s) ins_infl("aug") --augmentative(s) params["fonly"] = boolean_param -- feminine only params["gneut"] = boolean_param -- gender-neutral adjective e.g. [[latine]] params["hascomp"] = true -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } pos_functions["comparative adjectives"] = { params = get_adjective_params("comp"), func = do_adjective, pos_category = "adjectives", } pos_functions["superlative adjectives"] = { params = get_adjective_params("sup"), func = function(args, data) do_adjective(args, data, true) end, pos_category = "adjectives", } ----------------------------------------------------------------------------------------- -- Adverbs -- ----------------------------------------------------------------------------------------- pos_functions["adverbs"] = { params = { ["sup"] = list_param, --superlative(s) }, func = function(args, data) parse_and_insert_inflection(data, args, "sup", "superlative") end, } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- pos_functions["cardinal numbers"] = { params = { ["f"] = list_param, --feminine(s) ["mapoc"] = list_param, --masculine apocopated form(s) }, func = function(args, data) insert(data.categories, 1, langname .. " cardinal numbers") if args.f[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, "f", "feminine") end parse_and_insert_inflection(data, args, "mapoc", "masculine before a noun") end, pos_category = "numerals", } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"m", "f", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end -- Display additional inflection information for a noun local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_make_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_make_plural = saw_gneut and "gneut" or gender_for_irreg_ending local lemma = data.pagename local plurals = {} if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_make_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_make_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end end if #plurals > 1 then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end -- Gather masculines/feminines. For each one, generate the corresponding plural(s). `field` is the name of the -- field containing the masculine or feminine forms (normally "m" or "f"); `gender` is "m" or "f" for the gender -- of the forms; `inflect` is a function of one or two arguments to generate the default masculine or feminine from -- the lemma (the arguments are the lemma and optionally a "special" flag to indicate how to handle multiword -- lemmas, and the function is normally make_feminine or make_masculine from [[Module:es-common]]); and -- `default_plurals` is a list into which the corresponding default plurals of the gathered or generated masculine -- or feminine forms are stored. Note that there may be more default plurals than masculines or feminines, because -- some terms have multiple possible plurals. local function handle_mf(field, gender, inflect, default_plurals) local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = inflect(lemma) else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = inflect(lemma, special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|es|MFSING}} plobj.accel = {form = "p", lemma = mf.term} insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("f", "f", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("m", "m", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, gender, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, -- assume each plural goes with the corresponding singular -- and use each corresponding singular as the lemma in the accelerator. -- The generated entry will have # {{plural of|es|SINGULAR}} as the -- definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", "f", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", "m", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end insert_noun_inflection(plurals, "plural", "p") insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Maybe add category 'Spanish nouns with irregular gender' (or similar) local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (rfind(irreg_gender_lemma, "o$") and (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf")) or (irreg_gender_lemma:find("a$") and (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf")) then insert(data.categories, langname .. " nouns with irregular gender") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["nouns"] = { params = get_noun_params(), func = do_noun, } pos_functions["proper nouns"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["verbs"] = { params = { [1] = {}, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["pagename"] = {}, -- for testing ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, }, func = function(args, data) local preses, prets, parts if args.attn then insert(data.categories, "Requests for attention concerning " .. langname) return end local es_verb = require(es_verb_module) local alternant_multiword_spec = es_verb.do_generate_forms(args, "es-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and specforms[slot][1] end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), } elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), } else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), } end end preses = do_finite("pres", "present") prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } if args.pres[1] or args.pret[1] or args.part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:es-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use {{es-verb form of}} instead. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:es-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Spanish equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["phrases"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "m", "masculine") parse_and_insert_inflection(data, args, "f", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. require("Module:table").serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export ksx716atyb6uz6c4j2jriej6azirkk3 54858 54852 2026-09-26T23:27:26Z Deadend0914 7211 54858 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("es") local langname = lang:getCanonicalName() local insert = table.insert local remove = table.remove local rsub = com.rsub local sort = table.sort local ulower = mw.ustring.lower local usub = mw.ustring.sub local uupper = mw.ustring.upper local function track(page) require("Module:debug").track("es-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local parargs = frame:getParent().args local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {alias_of = "nolink"}, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolink then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "suffix forms" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end local pagename_lower = ulower(pagename) for _, ch in ipairs{"jh", "kh", "ph", "qü", "sh", "th", "tl", "ts", "tz", "wh", "zh", "ze", "zi"} do if pagename_lower:find(ch) then insert(data.categories, langname .. " terms spelled with " .. uupper(ch)) end end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture -- replace expression. lemma = m_string_utilities.replacement_escape(lemma) return (term:gsub("#", lemma)) -- discard second retval end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable participles' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end sort(indicators) error("Special inflection indicator beginning can only be " .. mw.text.listToText(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.inv then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[1] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("wif")) do if f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls = com.make_plural(lemma, "wer", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end parse_and_insert_inflection(data, args, "mapoc", "masculine singular before a noun") local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["inv"] = boolean_param, --invariable ["sp"] = true, -- special indicator: "first", "first-last", etc. } local function ins_infl(field) params[field] = list_param --feminine form(s) end ins_infl("wif") -- feminine form(s) ins_infl("pl") -- plural override(s) ins_infl("mpl") -- masculine plural override(s) ins_infl("fpl") -- feminine plural override(s) if adjtype == "base" then ins_infl("mapoc") --masculine apocopated (before a noun) ins_infl("comp") --comparative(s) ins_infl("sup") --superlative(s) ins_infl("dim") --diminutive(s) ins_infl("aug") --augmentative(s) params["fonly"] = boolean_param -- feminine only params["gneut"] = boolean_param -- gender-neutral adjective e.g. [[latine]] params["hascomp"] = true -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } pos_functions["comparative adjectives"] = { params = get_adjective_params("comp"), func = do_adjective, pos_category = "adjectives", } pos_functions["superlative adjectives"] = { params = get_adjective_params("sup"), func = function(args, data) do_adjective(args, data, true) end, pos_category = "adjectives", } ----------------------------------------------------------------------------------------- -- Adverbs -- ----------------------------------------------------------------------------------------- pos_functions["adverbs"] = { params = { ["sup"] = list_param, --superlative(s) }, func = function(args, data) parse_and_insert_inflection(data, args, "sup", "superlative") end, } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- pos_functions["cardinal numbers"] = { params = { ["f"] = list_param, --feminine(s) ["mapoc"] = list_param, --masculine apocopated form(s) }, func = function(args, data) insert(data.categories, 1, langname .. " cardinal numbers") if args.f[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, "f", "feminine") end parse_and_insert_inflection(data, args, "mapoc", "masculine before a noun") end, pos_category = "numerals", } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"m", "f", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end -- Display additional inflection information for a noun local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_make_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_make_plural = saw_gneut and "gneut" or gender_for_irreg_ending local lemma = data.pagename local plurals = {} if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_make_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_make_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end end if #plurals > 1 then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end -- Gather masculines/feminines. For each one, generate the corresponding plural(s). `field` is the name of the -- field containing the masculine or feminine forms (normally "m" or "f"); `gender` is "m" or "f" for the gender -- of the forms; `inflect` is a function of one or two arguments to generate the default masculine or feminine from -- the lemma (the arguments are the lemma and optionally a "special" flag to indicate how to handle multiword -- lemmas, and the function is normally make_feminine or make_masculine from [[Module:es-common]]); and -- `default_plurals` is a list into which the corresponding default plurals of the gathered or generated masculine -- or feminine forms are stored. Note that there may be more default plurals than masculines or feminines, because -- some terms have multiple possible plurals. local function handle_mf(field, gender, inflect, default_plurals) local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = inflect(lemma) else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = inflect(lemma, special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|es|MFSING}} plobj.accel = {form = "p", lemma = mf.term} insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("f", "f", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("wer", "wer", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, gender, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, -- assume each plural goes with the corresponding singular -- and use each corresponding singular as the lemma in the accelerator. -- The generated entry will have # {{plural of|es|SINGULAR}} as the -- definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", "f", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", "wer", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end insert_noun_inflection(plurals, "plural", "p") insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Maybe add category 'Spanish nouns with irregular gender' (or similar) local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (rfind(irreg_gender_lemma, "o$") and (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf")) or (irreg_gender_lemma:find("a$") and (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf")) then insert(data.categories, langname .. " nouns with irregular gender") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["nouns"] = { params = get_noun_params(), func = do_noun, } pos_functions["proper nouns"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["verbs"] = { params = { [1] = {}, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["pagename"] = {}, -- for testing ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, }, func = function(args, data) local preses, prets, parts if args.attn then insert(data.categories, "Requests for attention concerning " .. langname) return end local es_verb = require(es_verb_module) local alternant_multiword_spec = es_verb.do_generate_forms(args, "es-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and specforms[slot][1] end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), } elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), } else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), } end end preses = do_finite("pres", "present") prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } if args.pres[1] or args.pret[1] or args.part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:es-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use {{es-verb form of}} instead. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:es-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Spanish equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["phrases"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "wer", "masculine") parse_and_insert_inflection(data, args, "wif", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. require("Module:table").serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export mb8dsqzel25txtt220z9ilhrndw0aj0 54859 54858 2026-09-26T23:28:26Z Deadend0914 7211 54859 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("es") local langname = lang:getCanonicalName() local insert = table.insert local remove = table.remove local rsub = com.rsub local sort = table.sort local ulower = mw.ustring.lower local usub = mw.ustring.sub local uupper = mw.ustring.upper local function track(page) require("Module:debug").track("es-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local parargs = frame:getParent().args local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {alias_of = "nolink"}, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolink then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "suffix forms" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end local pagename_lower = ulower(pagename) for _, ch in ipairs{"jh", "kh", "ph", "qü", "sh", "th", "tl", "ts", "tz", "wh", "zh", "ze", "zi"} do if pagename_lower:find(ch) then insert(data.categories, langname .. " terms spelled with " .. uupper(ch)) end end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture -- replace expression. lemma = m_string_utilities.replacement_escape(lemma) return (term:gsub("#", lemma)) -- discard second retval end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable participles' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end sort(indicators) error("Special inflection indicator beginning can only be " .. mw.text.listToText(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.inv then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[1] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("wif")) do if f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls = com.make_plural(lemma, "wer", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end parse_and_insert_inflection(data, args, "mapoc", "masculine singular before a noun") local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["inv"] = boolean_param, --invariable ["sp"] = true, -- special indicator: "first", "first-last", etc. } local function ins_infl(field) params[field] = list_param --feminine form(s) end ins_infl("wif") -- feminine form(s) ins_infl("pl") -- plural override(s) ins_infl("mpl") -- masculine plural override(s) ins_infl("fpl") -- feminine plural override(s) if adjtype == "base" then ins_infl("mapoc") --masculine apocopated (before a noun) ins_infl("comp") --comparative(s) ins_infl("sup") --superlative(s) ins_infl("dim") --diminutive(s) ins_infl("aug") --augmentative(s) params["fonly"] = boolean_param -- feminine only params["gneut"] = boolean_param -- gender-neutral adjective e.g. [[latine]] params["hascomp"] = true -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } pos_functions["comparative adjectives"] = { params = get_adjective_params("comp"), func = do_adjective, pos_category = "adjectives", } pos_functions["superlative adjectives"] = { params = get_adjective_params("sup"), func = function(args, data) do_adjective(args, data, true) end, pos_category = "adjectives", } ----------------------------------------------------------------------------------------- -- Adverbs -- ----------------------------------------------------------------------------------------- pos_functions["adverbs"] = { params = { ["sup"] = list_param, --superlative(s) }, func = function(args, data) parse_and_insert_inflection(data, args, "sup", "superlative") end, } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- pos_functions["cardinal numbers"] = { params = { ["f"] = list_param, --feminine(s) ["mapoc"] = list_param, --masculine apocopated form(s) }, func = function(args, data) insert(data.categories, 1, langname .. " cardinal numbers") if args.f[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, "f", "feminine") end parse_and_insert_inflection(data, args, "mapoc", "masculine before a noun") end, pos_category = "numerals", } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"wer", "wif", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end -- Display additional inflection information for a noun local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_make_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_make_plural = saw_gneut and "gneut" or gender_for_irreg_ending local lemma = data.pagename local plurals = {} if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_make_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_make_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end end if #plurals > 1 then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end -- Gather masculines/feminines. For each one, generate the corresponding plural(s). `field` is the name of the -- field containing the masculine or feminine forms (normally "m" or "f"); `gender` is "m" or "f" for the gender -- of the forms; `inflect` is a function of one or two arguments to generate the default masculine or feminine from -- the lemma (the arguments are the lemma and optionally a "special" flag to indicate how to handle multiword -- lemmas, and the function is normally make_feminine or make_masculine from [[Module:es-common]]); and -- `default_plurals` is a list into which the corresponding default plurals of the gathered or generated masculine -- or feminine forms are stored. Note that there may be more default plurals than masculines or feminines, because -- some terms have multiple possible plurals. local function handle_mf(field, gender, inflect, default_plurals) local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = inflect(lemma) else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = inflect(lemma, special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|es|MFSING}} plobj.accel = {form = "p", lemma = mf.term} insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("f", "f", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("wer", "wer", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, gender, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, -- assume each plural goes with the corresponding singular -- and use each corresponding singular as the lemma in the accelerator. -- The generated entry will have # {{plural of|es|SINGULAR}} as the -- definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", "f", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", "wer", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end insert_noun_inflection(plurals, "plural", "p") insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Maybe add category 'Spanish nouns with irregular gender' (or similar) local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (rfind(irreg_gender_lemma, "o$") and (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf")) or (irreg_gender_lemma:find("a$") and (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf")) then insert(data.categories, langname .. " nouns with irregular gender") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["nouns"] = { params = get_noun_params(), func = do_noun, } pos_functions["proper nouns"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["verbs"] = { params = { [1] = {}, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["pagename"] = {}, -- for testing ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, }, func = function(args, data) local preses, prets, parts if args.attn then insert(data.categories, "Requests for attention concerning " .. langname) return end local es_verb = require(es_verb_module) local alternant_multiword_spec = es_verb.do_generate_forms(args, "es-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and specforms[slot][1] end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), } elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), } else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), } end end preses = do_finite("pres", "present") prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } if args.pres[1] or args.pret[1] or args.part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:es-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use {{es-verb form of}} instead. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:es-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Spanish equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["phrases"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "wer", "masculine") parse_and_insert_inflection(data, args, "wif", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. require("Module:table").serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export bzi3a4ukph5wam75olzfb7woeus8x45 54861 54859 2026-09-26T23:32:37Z Deadend0914 7211 54861 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("es") local langname = lang:getCanonicalName() local insert = table.insert local remove = table.remove local rsub = com.rsub local sort = table.sort local ulower = mw.ustring.lower local usub = mw.ustring.sub local uupper = mw.ustring.upper local function track(page) require("Module:debug").track("es-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local parargs = frame:getParent().args local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {alias_of = "nolink"}, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolink then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "suffix forms" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end local pagename_lower = ulower(pagename) for _, ch in ipairs{"jh", "kh", "ph", "qü", "sh", "th", "tl", "ts", "tz", "wh", "zh", "ze", "zi"} do if pagename_lower:find(ch) then insert(data.categories, langname .. " terms spelled with " .. uupper(ch)) end end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture -- replace expression. lemma = m_string_utilities.replacement_escape(lemma) return (term:gsub("#", lemma)) -- discard second retval end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable participles' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end sort(indicators) error("Special inflection indicator beginning can only be " .. mw.text.listToText(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.inv then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[1] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("wif")) do if f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls = com.make_plural(lemma, "wer", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "wif", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end parse_and_insert_inflection(data, args, "mapoc", "masculine singular before a noun") local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["inv"] = boolean_param, --invariable ["sp"] = true, -- special indicator: "first", "first-last", etc. } local function ins_infl(field) params[field] = list_param --feminine form(s) end ins_infl("wif") -- feminine form(s) ins_infl("pl") -- plural override(s) ins_infl("mpl") -- masculine plural override(s) ins_infl("fpl") -- feminine plural override(s) if adjtype == "base" then ins_infl("mapoc") --masculine apocopated (before a noun) ins_infl("comp") --comparative(s) ins_infl("sup") --superlative(s) ins_infl("dim") --diminutive(s) ins_infl("aug") --augmentative(s) params["fonly"] = boolean_param -- feminine only params["gneut"] = boolean_param -- gender-neutral adjective e.g. [[latine]] params["hascomp"] = true -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } pos_functions["comparative adjectives"] = { params = get_adjective_params("comp"), func = do_adjective, pos_category = "adjectives", } pos_functions["superlative adjectives"] = { params = get_adjective_params("sup"), func = function(args, data) do_adjective(args, data, true) end, pos_category = "adjectives", } ----------------------------------------------------------------------------------------- -- Adverbs -- ----------------------------------------------------------------------------------------- pos_functions["adverbs"] = { params = { ["sup"] = list_param, --superlative(s) }, func = function(args, data) parse_and_insert_inflection(data, args, "sup", "superlative") end, } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- pos_functions["cardinal numbers"] = { params = { ["f"] = list_param, --feminine(s) ["mapoc"] = list_param, --masculine apocopated form(s) }, func = function(args, data) insert(data.categories, 1, langname .. " cardinal numbers") if args.f[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, "f", "feminine") end parse_and_insert_inflection(data, args, "mapoc", "masculine before a noun") end, pos_category = "numerals", } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"wer", "wif", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end -- Display additional inflection information for a noun local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_make_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_make_plural = saw_gneut and "gneut" or gender_for_irreg_ending local lemma = data.pagename local plurals = {} if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_make_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_make_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end end if #plurals > 1 then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end -- Gather masculines/feminines. For each one, generate the corresponding plural(s). `field` is the name of the -- field containing the masculine or feminine forms (normally "m" or "f"); `gender` is "m" or "f" for the gender -- of the forms; `inflect` is a function of one or two arguments to generate the default masculine or feminine from -- the lemma (the arguments are the lemma and optionally a "special" flag to indicate how to handle multiword -- lemmas, and the function is normally make_feminine or make_masculine from [[Module:es-common]]); and -- `default_plurals` is a list into which the corresponding default plurals of the gathered or generated masculine -- or feminine forms are stored. Note that there may be more default plurals than masculines or feminines, because -- some terms have multiple possible plurals. local function handle_mf(field, gender, inflect, default_plurals) local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = inflect(lemma) else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = inflect(lemma, special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|es|MFSING}} plobj.accel = {form = "p", lemma = mf.term} insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("wif", "wif", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("wer", "wer", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, gender, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, -- assume each plural goes with the corresponding singular -- and use each corresponding singular as the lemma in the accelerator. -- The generated entry will have # {{plural of|es|SINGULAR}} as the -- definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", "f", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", "wer", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end insert_noun_inflection(plurals, "plural", "p") insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Maybe add category 'Spanish nouns with irregular gender' (or similar) local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (rfind(irreg_gender_lemma, "o$") and (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf")) or (irreg_gender_lemma:find("a$") and (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf")) then insert(data.categories, langname .. " nouns with irregular gender") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["nouns"] = { params = get_noun_params(), func = do_noun, } pos_functions["proper nouns"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["verbs"] = { params = { [1] = {}, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["pagename"] = {}, -- for testing ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, }, func = function(args, data) local preses, prets, parts if args.attn then insert(data.categories, "Requests for attention concerning " .. langname) return end local es_verb = require(es_verb_module) local alternant_multiword_spec = es_verb.do_generate_forms(args, "es-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and specforms[slot][1] end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), } elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), } else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), } end end preses = do_finite("pres", "present") prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } if args.pres[1] or args.pret[1] or args.part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:es-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use {{es-verb form of}} instead. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:es-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Spanish equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["phrases"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "wer", "masculine") parse_and_insert_inflection(data, args, "wif", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. require("Module:table").serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export arjkxwxiu38u6mpqmw0y7tr0cyuzevk 54862 54861 2026-09-26T23:33:17Z Deadend0914 7211 54862 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local rfind = mw.ustring.find local rmatch = mw.ustring.match local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:es-common") local en_utilities_module = "Module:en-utilities" local es_verb_module = "Module:es-verb" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("es") local langname = lang:getCanonicalName() local insert = table.insert local remove = table.remove local rsub = com.rsub local sort = table.sort local ulower = mw.ustring.lower local usub = mw.ustring.sub local uupper = mw.ustring.upper local function track(page) require("Module:debug").track("es-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local parargs = frame:getParent().args local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {alias_of = "nolink"}, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolink then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "suffix forms" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end local pagename_lower = ulower(pagename) for _, ch in ipairs{"jh", "kh", "ph", "qü", "sh", "th", "tl", "ts", "tz", "wh", "zh", "ze", "zi"} do if pagename_lower:find(ch) then insert(data.categories, langname .. " terms spelled with " .. uupper(ch)) end end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture -- replace expression. lemma = m_string_utilities.replacement_escape(lemma) return (term:gsub("#", lemma)) -- discard second retval end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable participles' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end sort(indicators) error("Special inflection indicator beginning can only be " .. mw.text.listToText(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.inv then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[1] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("f")) do if f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls = com.make_plural(lemma, "m", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end parse_and_insert_inflection(data, args, "mapoc", "masculine singular before a noun") local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["inv"] = boolean_param, --invariable ["sp"] = true, -- special indicator: "first", "first-last", etc. } local function ins_infl(field) params[field] = list_param --feminine form(s) end ins_infl("f") -- feminine form(s) ins_infl("pl") -- plural override(s) ins_infl("mpl") -- masculine plural override(s) ins_infl("fpl") -- feminine plural override(s) if adjtype == "base" then ins_infl("mapoc") --masculine apocopated (before a noun) ins_infl("comp") --comparative(s) ins_infl("sup") --superlative(s) ins_infl("dim") --diminutive(s) ins_infl("aug") --augmentative(s) params["fonly"] = boolean_param -- feminine only params["gneut"] = boolean_param -- gender-neutral adjective e.g. [[latine]] params["hascomp"] = true -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } pos_functions["comparative adjectives"] = { params = get_adjective_params("comp"), func = do_adjective, pos_category = "adjectives", } pos_functions["superlative adjectives"] = { params = get_adjective_params("sup"), func = function(args, data) do_adjective(args, data, true) end, pos_category = "adjectives", } ----------------------------------------------------------------------------------------- -- Adverbs -- ----------------------------------------------------------------------------------------- pos_functions["adverbs"] = { params = { ["sup"] = list_param, --superlative(s) }, func = function(args, data) parse_and_insert_inflection(data, args, "sup", "superlative") end, } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- pos_functions["cardinal numbers"] = { params = { ["f"] = list_param, --feminine(s) ["mapoc"] = list_param, --masculine apocopated form(s) }, func = function(args, data) insert(data.categories, 1, langname .. " cardinal numbers") if args.f[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, "f", "feminine") end parse_and_insert_inflection(data, args, "mapoc", "masculine before a noun") end, pos_category = "numerals", } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"m", "f", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end -- Display additional inflection information for a noun local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_make_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_make_plural = saw_gneut and "gneut" or gender_for_irreg_ending local lemma = data.pagename local plurals = {} if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_make_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_make_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end end if #plurals > 1 then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end -- Gather masculines/feminines. For each one, generate the corresponding plural(s). `field` is the name of the -- field containing the masculine or feminine forms (normally "m" or "f"); `gender` is "m" or "f" for the gender -- of the forms; `inflect` is a function of one or two arguments to generate the default masculine or feminine from -- the lemma (the arguments are the lemma and optionally a "special" flag to indicate how to handle multiword -- lemmas, and the function is normally make_feminine or make_masculine from [[Module:es-common]]); and -- `default_plurals` is a list into which the corresponding default plurals of the gathered or generated masculine -- or feminine forms are stored. Note that there may be more default plurals than masculines or feminines, because -- some terms have multiple possible plurals. local function handle_mf(field, gender, inflect, default_plurals) local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = inflect(lemma) else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = inflect(lemma, special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|es|MFSING}} plobj.accel = {form = "p", lemma = mf.term} insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("f", "f", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("m", "m", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, gender, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, -- assume each plural goes with the corresponding singular -- and use each corresponding singular as the lemma in the accelerator. -- The generated entry will have # {{plural of|es|SINGULAR}} as the -- definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", "f", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", "m", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end insert_noun_inflection(plurals, "plural", "p") insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Maybe add category 'Spanish nouns with irregular gender' (or similar) local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (rfind(irreg_gender_lemma, "o$") and (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf")) or (irreg_gender_lemma:find("a$") and (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf")) then insert(data.categories, langname .. " nouns with irregular gender") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["nouns"] = { params = get_noun_params(), func = do_noun, } pos_functions["proper nouns"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["verbs"] = { params = { [1] = {}, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["pagename"] = {}, -- for testing ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, }, func = function(args, data) local preses, prets, parts if args.attn then insert(data.categories, "Requests for attention concerning " .. langname) return end local es_verb = require(es_verb_module) local alternant_multiword_spec = es_verb.do_generate_forms(args, "es-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and specforms[slot][1] end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), } elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), } else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), } end end preses = do_finite("pres", "present") prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } if args.pres[1] or args.pret[1] or args.part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:es-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use {{es-verb form of}} instead. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:es-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Spanish equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["phrases"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "m", "masculine") parse_and_insert_inflection(data, args, "f", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. require("Module:table").serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export ksx716atyb6uz6c4j2jriej6azirkk3 Module:es-common 828 8049 54853 2026-09-26T23:22:04Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local romut_module = "Module:romance utilities" local u = require("Module:string/char") local rsplit = mw.text.split local rfind = mw.ustring.find local rmatch = mw.ustring.match local rsubn = mw.ustring.gsub local toNFD = mw.ustring.toNFD local TILDE = u(0x0303) -- tilde = ̃ local DIA = u(0x0308) -- diaeresis = ̈ local CEDILLA = u(0x0327) -- cedilla = ̧ local TEMPC1 = u(0xFFF1) local TEMPC2 = u(0xFFF2) local TEMPV1 = u(0xFFF3) local DIV...' 54853 Scribunto text/plain local export = {} local romut_module = "Module:romance utilities" local u = require("Module:string/char") local rsplit = mw.text.split local rfind = mw.ustring.find local rmatch = mw.ustring.match local rsubn = mw.ustring.gsub local toNFD = mw.ustring.toNFD local TILDE = u(0x0303) -- tilde = ̃ local DIA = u(0x0308) -- diaeresis = ̈ local CEDILLA = u(0x0327) -- cedilla = ̧ local TEMPC1 = u(0xFFF1) local TEMPC2 = u(0xFFF2) local TEMPV1 = u(0xFFF3) local DIV = u(0xFFF4) local vowel = "aeiouáéíóúý" .. TEMPV1 local V = "[" .. vowel .. "]" local AV = "[áéíóúý]" -- accented vowel local W = "[iyuw]" -- glide local C = "[^" .. vowel .. ".]" export.vowel = vowel export.V = V export.AV = AV export.W = W export.C = C local remove_accent = { ["á"] = "a", ["é"] = "e", ["í"] = "i", ["ó"] = "o", ["ú"] = "u", ["ý"] = "y" } local add_accent = { ["a"] = "á", ["e"] = "é", ["i"] = "í", ["o"] = "ó", ["u"] = "ú", ["y"] = "ý" } export.remove_accent = remove_accent export.add_accent = add_accent local prepositions = { "al? ", "del? ", "como ", "con ", "en ", "para ", "por ", } -- version of rsubn() that discards all but the first return value local function rsub(term, foo, bar) local retval = rsubn(term, foo, bar) return retval end export.rsub = rsub -- apply rsub() repeatedly until no change local function rsub_repeatedly(term, foo, bar) while true do local new_term = rsub(term, foo, bar) if new_term == term then return term end term = new_term end end export.rsub_repeatedly = rsub_repeatedly function export.decompose(text) -- decompose everything but ç, ñ and ü text = toNFD(text) text = rsub(text, ".[" .. TILDE .. DIA .. CEDILLA .. "]", { ["c" .. CEDILLA] = "ç", ["C" .. CEDILLA] = "Ç", ["n" .. TILDE] = "ñ", ["N" .. TILDE] = "Ñ", ["u" .. DIA] = "ü", ["U" .. DIA] = "Ü", }) return text end -- Apply vowel alternation to stem. function export.apply_vowel_alternation(stem, alternation) local ret, err -- Treat final -gu, -qu as a consonant, so the previous vowel can alternate (e.g. conseguir -> consigo). -- This means a verb in -guar can't have a u-ú alternation but I don't think there are any verbs like that. stem = rsub(stem, "([gq])u$", "%1" .. TEMPC1) local before_last_vowel, last_vowel, after_last_vowel = rmatch(stem, "^(.*)(" .. V .. ")(.-)$") if alternation == "ie" then if last_vowel == "e" or last_vowel == "i" then -- allow i for adquirir -> adquiero, inquirir -> inquiero, etc. ret = before_last_vowel .. "ie" .. after_last_vowel else err = "should have -e- or -i- as the last vowel" end elseif alternation == "ye" then if last_vowel == "e" then ret = before_last_vowel .. "ye" .. after_last_vowel else err = "should have -e- as the last vowel" end elseif alternation == "ue" then if last_vowel == "o" or last_vowel == "u" then -- allow u for jugar -> juego; correctly handle avergonzar -> avergüenzo ret = ( last_vowel == "o" and before_last_vowel:find("g$") and before_last_vowel .. "üe" .. after_last_vowel or before_last_vowel .. "ue" .. after_last_vowel ) else err = "should have -o- or -u- as the last vowel" end elseif alternation == "hue" then if last_vowel == "o" then ret = before_last_vowel .. "hue" .. after_last_vowel else err = "should have -o- as the last vowel" end elseif alternation == "i" then if last_vowel == "e" then ret = before_last_vowel .. "i" .. after_last_vowel else err = "should have -i- as the last vowel" end elseif alternation == "í" then if last_vowel == "e" or last_vowel == "i" then -- allow e for reír -> río, sonreír -> sonrío ret = before_last_vowel .. "í" .. after_last_vowel else err = "should have -e- or -i- as the last vowel" end elseif alternation == "ú" then if last_vowel == "u" then ret = before_last_vowel .. "ú" .. after_last_vowel else err = "should have -u- as the last vowel" end else error("Unrecognized vowel alternation '" .. alternation .. "'") end ret = ret and ret:gsub(TEMPC1, "u") or nil return {ret = ret, err = err} end -- Syllabify a word. This implements the full syllabification algorithm, based on the corresponding code -- in [[Module:es-pronunc]]. This is more than is needed for the purpose of this module, which doesn't -- care so much about syllable boundaries, but won't hurt. function export.syllabify(word) word = DIV .. word .. DIV -- gu/qu + front vowel; make sure we treat the u as a consonant; a following -- i should not be treated as a consonant ([[alguien]] would become ''álguienes'' -- if pluralized) word = rsub(word, "([gq])u([eiéí])", "%1" .. TEMPC2 .. "%2") local vowel_to_glide = { ["i"] = TEMPC1, ["u"] = TEMPC2 } -- i and u between vowels should behave like consonants ([[paranoia]], [[baiano]], [[abreuense]], -- [[alauita]], [[Malaui]], etc.) word = rsub_repeatedly(word, "(" .. V .. ")([iu])(" .. V .. ")", function(v1, iu, v2) return v1 .. vowel_to_glide[iu] .. v2 end ) -- y between consonants or after a consonant at the end of the word should behave like a vowel -- ([[ankylosaurio]], [[cryptomeria]], [[brandy]], [[cherry]], etc.) word = rsub_repeatedly(word, "(" .. C .. ")y(" .. C .. ")", function(c1, c2) return c1 .. TEMPV1 .. c2 end ) word = rsub_repeatedly(word, "(" .. V .. ")(" .. C .. W .. "?" .. V .. ")", "%1.%2") word = rsub_repeatedly(word, "(" .. V .. C .. ")(" .. C .. V .. ")", "%1.%2") word = rsub_repeatedly(word, "(" .. V .. C .. "+)(" .. C .. C .. V .. ")", "%1.%2") word = rsub(word, "([pbcktdg])%.([lr])", ".%1%2") word = rsub_repeatedly(word, "(" .. C .. ")%.s(" .. C .. ")", "%1s.%2") -- Any aeo, or stressed iu, should be syllabically divided from a following aeo or stressed iu. word = rsub_repeatedly(word, "([aeoáéíóúý])([aeoáéíóúý])", "%1.%2") word = rsub_repeatedly(word, "([ií])([ií])", "%1.%2") word = rsub_repeatedly(word, "([uú])([uú])", "%1.%2") word = rsub(word, "([" .. DIV .. TEMPC1 .. TEMPC2 .. TEMPV1 .. "])", { [DIV] = "", [TEMPC1] = "i", [TEMPC2] = "u", [TEMPV1] = "y", }) return rsplit(word, "%.") end -- Return the index of the (last) stressed syllable. function export.stressed_syllable(syllables) -- If a syllable is stressed, return it. for i = #syllables, 1, -1 do if rfind(syllables[i], AV) then return i end end -- Monosyllabic words are stressed on that syllable. if #syllables == 1 then return 1 end local i = #syllables -- Unaccented words ending in a vowel or a vowel + s/n are stressed on the preceding syllable. if rfind(syllables[i], V .. "[sn]?$") then return i - 1 end -- Remaining words are stressed on the last syllable. return i end -- Add an accent to the appropriate vowel in a syllable, if not already accented. function export.add_accent_to_syllable(syllable) -- Don't do anything if syllable already stressed. if rfind(syllable, AV) then return syllable end -- Prefer to accent an a/e/o in case of a diphthong or triphthong (the first one if for some reason -- there are multiple, which should not occur with the standard syllabification algorithm); -- otherwise, do the last i or u in case of a diphthong ui or iu. if rfind(syllable, "[aeo]") then return rsub(syllable, "^(.-)([aeo])", function(prev, v) return prev .. add_accent[v] end) end return rsub(syllable, "^(.*)([iu])", function(prev, v) return prev .. add_accent[v] end) end -- Remove any accent from a syllable. function export.remove_accent_from_syllable(syllable) return rsub(syllable, AV, remove_accent) end -- Return true if an accent is needed on syllable number `sylno` if that syllable were to receive the stress, -- given the syllables of a word. The current accent may be on any syllable. function export.accent_needed(syllables, sylno) -- Diphthongs iu and ui are normally stressed on the second vowel, so if the accent is on the first vowel, -- it's needed. if rfind(syllables[sylno], "íu") or rfind(syllables[sylno], "úi") then return true end -- If the default-stressed syllable is different from `sylno`, accent is needed. local unaccented_syllables = {} for _, syl in ipairs(syllables) do table.insert(unaccented_syllables, export.remove_accent_from_syllable(syl)) end local would_be_stressed_syl = export.stressed_syllable(unaccented_syllables) if would_be_stressed_syl ~= sylno then return true end -- At this point, we know that the stress would by default go on `sylno`, given the syllabification in -- `syllables`. Now we have to check for situations where removing the accent mark would result in a -- different syllabification. For example, países -> `pa.i.ses` but removing the accent mark would lead -- to `pai.ses`. Similarly, río -> `ri.o` but removing the accent mark would lead to single-syllable `rio`. -- We need to check whether (a) the stress falls on an i or u; (b) in the absence of an accent mark, the -- i or u would form a diphthong with a preceding or following vowel and the stress would be on that vowel. -- The conditions are slightly different when dealing with preceding or following vowels because ui and ui -- diphthongs are by default stressed on the second vowel. We also have to ignore h between the vowels. local accented_syllable = export.add_accent_to_syllable(unaccented_syllables[sylno]) if sylno > 1 and rfind(unaccented_syllables[sylno - 1], "[aeo]$") and rfind(accented_syllable, "^h?[íú]") then return true end if sylno < #syllables then if rfind(accented_syllable, "í$") and rfind(unaccented_syllables[sylno + 1], "^h?[aeou]") or rfind(accented_syllable, "ú$") and rfind(unaccented_syllables[sylno + 1], "^h?[aeio]") then return true end end return false end function export.make_plural(form, gender, special) local retval = require(romut_module).handle_multiword(form, special, function(term) return export.make_plural(term, gender) end, prepositions) if retval then return retval end if gender == "gneut" and rfind(form, "[x@]$") then return {form .. "s"} end -- ends in unstressed vowel or á, é, ó if rfind(form, "[aeiouáéó]$") then return {form .. "s"} end -- ends in í or ú if rfind(form, "[íú]$") then return {form .. "es", form .. "s"} end -- ends in a vowel + z if rfind(form, V .. "z$") then return {rsub(form, "z$", "ces")} end -- ends in cons + s/z if rfind(form, C.."[sz]$") then return {form} end -- ends in s/z + cons if rfind(form, "[sz]"..C.."$") then return {form} end local syllables = export.syllabify(form) -- ends in s or x with more than 1 syllable, last syllable unstressed if syllables[2] and rfind(form, "[sx]$") and not rfind(syllables[#syllables], AV) then return {form} end -- ends in l, r, n, d, z, or j with 3 or more syllables, stressed on third to last syllable if syllables[3] and rfind(form, "[lrndzj]$") and rfind(syllables[#syllables - 2], AV) then return {form} end -- ends in an accented vowel + consonant if rfind(form, AV .. C .. "$") then return {rsub(form, "(.)(.)$", function(vowel, consonant) return export.remove_accent[vowel] .. consonant .. "es" end)} end -- ends in a vowel + y, l, r, n, d, j, s, x if rfind(form, "[aeiou][ylrndjsx]$") then -- two or more syllables: add stress mark to plural; e.g. joven -> jóvenes if syllables[2] and rfind(form, "n$") then syllables[#syllables - 1] = export.add_accent_to_syllable(syllables[#syllables - 1]) return {table.concat(syllables, "") .. "es"} end return {form .. "es"} end -- ends in a vowel + ch if rfind(form, "[aeiou]ch$") then return {form .. "es"} end -- ends in two consonants if rfind(form, C .. C .. "$") then return {form .. "s"} end -- ends in a vowel + consonant other than l, r, n, d, z, j, s, or x if rfind(form, "[aeiou][^aeioulrndzjsx]$") then return {form .. "s"} end return nil end function export.make_feminine(form, special) local retval = require(romut_module).handle_multiword(form, special, export.make_feminine, prepositions) if retval then if #retval ~= 1 then error("Internal error: Should have one return value for make_feminine: " .. table.concat(retval, ",")) end return retval[1] end if form:find("o$") then local retval = form:gsub("o$", "a") -- discard second retval return retval end local function make_stem(form) return rsub( form, "^(.+)(.)(.)$", function (before_stress, stressed_vowel, after_stress) return before_stress .. (export.remove_accent[stressed_vowel] or stressed_vowel) .. after_stress end) end if rfind(form, "[áíó]n$") or rfind(form, "[éí]s$") or rfind(form, "[dtszxñ]or$") or rfind(form, "ol$") then -- holgazán, comodín, bretón (not común); francés, kirguís (not mandamás); -- volador, agricultor, defensor, avizor, flexor, señor (not posterior, bicolor, mayor, mejor, menor, peor); -- español, mongol return make_stem(form) .. "a" end return form end function export.make_masculine(form, special) local retval = require(romut_module).handle_multiword(form, special, export.make_masculine, prepositions) if retval then if #retval ~= 1 then error("Internal error: Should have one return value for make_masculine: " .. table.concat(retval, ",")) end return retval[1] end if form:find("dora$") then local retval = form:gsub("a$", "") -- discard second retval return retval end if form:find("a$") then local retval = form:gsub("a$", "o") -- discard second retval return retval end return form end return export qq46tgo6gtotmf220p5pp0dpz3e6kus Module:languages/data/2 828 8050 54854 2026-09-26T23:22:36Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared -- Ideally, we want to move these into [[Module:languages/data]], but because (a) it's necessary to use require on that module, and (b) they're only used in this data modu...' 54854 Scribunto text/plain local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared -- Ideally, we want to move these into [[Module:languages/data]], but because (a) it's necessary to use require on that module, and (b) they're only used in this data module, it's less memory-efficient to do that at the moment. If it becomes possible to use mw.loadData, then these should be moved there. s["de-Latn-sortkey"] = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer .. c.ringabove, from = {"æ", "œ", "ß"}, to = {"ae", "oe", "ss"} } s["de-Latn-standardchars"] = "AaÄäBbCcDdEeFfGgHhIiJjKkLlMmNnOoÖöPpQqRrSsẞßTtUuÜüVvWwXxYyZz" s["ka-stripdiacritics"] = {remove_diacritics = c.circ} s["no-sortkey"] = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.dacute .. c.caron .. c.cedilla, remove_exceptions = {"å"}, from = {"æ", "ø", "å"}, to = {"z" .. p[1], "z" .. p[2], "z" .. p[3]} } s["no-standardchars"] = "AaBbDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsTtUuVvYyÆæØøÅå" .. c.punc s["sa-Deva-stripdiacritics"] = { -- Don't use remove_diacritics for accent marks, as १ and ३ should also be removed if (and only if) they carry any. from = {"ॐ", "[१३]?[" .. c.anudatta .. c.udatta .. c.dsvarita .. c.tsvarita .. "]+"}, to = {"ओँ"}, } s["tg-stripdiacritics"] = {remove_diacritics = c.grave .. c.acute} s["tk-stripdiacritics"] = {remove_diacritics = c.macron} local m = {} m["aa"] = { "Afar", 27811, "cus-eas", "Latn, Ethi", strip_diacritics = { Latn = {remove_diacritics = c.acute}, }, } m["ab"] = { "Abkhaz", 5111, "cau-abz", "Cyrl, Geor, Latn", translit = { Cyrl = "ab-translit", -- Geor translit in [[Module:scripts/data]] }, override_translit = true, display_text = { Cyrl = s["cau-Cyrl-displaytext"] }, strip_diacritics = { Cyrl = { remove_diacritics = c.acute, from = {"^а%-"}, to = {"а"}, }, Latn = s["cau-Latn-stripdiacritics"], }, sort_key = { Cyrl = { from = { "х'ә", -- 3 chars "гь", "гә", "ӷь", "ҕь", "ӷә", "ҕә", "дә", "ё", "жь", "жә", "ҙә", "ӡә", "ӡ'", "кь", "кә", "қь", "қә", "ҟь", "ҟә", "ҫә", "тә", "ҭә", "ф'", "хь", "хә", "х'", "ҳә", "ць", "цә", "ц'", "ҵә", "ҵ'", "шь", "шә", "џь", -- 2 chars "ӷ", "ҕ", "ҙ", "ӡ", "қ", "ҟ", "ԥ", "ҧ", "ҫ", "ҭ", "ҳ", "ҵ", "ҷ", "ҽ", "ҿ", "ҩ", "џ", "ә", -- 1 char "^а", }, to = { "х" .. p[4], "г" .. p[1], "г" .. p[2], "г" .. p[5], "г" .. p[6], "г" .. p[7], "г" .. p[8], "д" .. p[1], "е" .. p[1], "ж" .. p[1], "ж" .. p[2], "з" .. p[2], "з" .. p[4], "з" .. p[5], "к" .. p[1], "к" .. p[2], "к" .. p[4], "к" .. p[5], "к" .. p[7], "к" .. p[8], "с" .. p[2], "т" .. p[1], "т" .. p[3], "ф" .. p[1], "х" .. p[1], "х" .. p[2], "х" .. p[3], "х" .. p[6], "ц" .. p[1], "ц" .. p[2], "ц" .. p[3], "ц" .. p[5], "ц" .. p[6], "ш" .. p[1], "ш" .. p[2], "ы" .. p[3], "г" .. p[3], "г" .. p[4], "з" .. p[1], "з" .. p[3], "к" .. p[3], "к" .. p[6], "п" .. p[1], "п" .. p[2], "с" .. p[1], "т" .. p[2], "х" .. p[5], "ц" .. p[4], "ч" .. p[1], "ч" .. p[2], "ч" .. p[3], "ы" .. p[1], "ы" .. p[2], "ь" .. p[1], "", } }, }, } m["ae"] = { "Avestan", 29572, "ira-cen", "Avst, Gujr, Deva", translit = { Avst = "Avst-translit" }, } m["af"] = { "Afrikaans", 14196, "gmw-frk", "Latn, Arab", ancestors = "nl", sort_key = { Latn = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.diaer .. c.ringabove .. c.cedilla .. "'", from = {"['ʼ]n"}, to = {"n" .. p[1]} } }, } m["ak"] = { "Akan", 28026, "alv-ctn", "Latn", } m["am"] = { "Amharic", 28244, "sem-eth", "Ethi", translit = "Ethi-translit", } m["an"] = { "Aragonese", 8765, "roa-nar", "Latn", } m["ar"] = { "Arabic", 13955, "sem-arb", "Arab, Hebr, Syrc, Brai, Nbat", translit = { Arab = "ar-translit" }, strip_diacritics = { Arab = "ar-stripdiacritics", }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["as"] = { "Assamese", 29401, "inc-bas", "as-Beng", ancestors = "inc-mas", translit = "as-translit", } m["av"] = { "Avar", 29561, "cau-ava", "Cyrl, Latn, Arab", ancestors = "oav", translit = { Cyrl = "cau-nec-translit", Arab = "ar-translit", }, override_translit = true, display_text = { Cyrl = s["cau-Cyrl-displaytext"], }, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, sort_key = { Cyrl = { from = {"гъ", "гь", "гӏ", "ё", "кк", "къ", "кь", "кӏ", "лъ", "лӏ", "тӏ", "хх", "хъ", "хь", "хӏ", "цӏ", "чӏ"}, to = {"г" .. p[1], "г" .. p[2], "г" .. p[3], "е" .. p[1], "к" .. p[1], "к" .. p[2], "к" .. p[3], "к" .. p[4], "л" .. p[1], "л" .. p[2], "т" .. p[1], "х" .. p[1], "х" .. p[2], "х" .. p[3], "х" .. p[4], "ц" .. p[1], "ч" .. p[1]} }, }, } m["ay"] = { "Aymara", 4627, "sai-aym", "Latn", } m["az"] = { "Azerbaijani", 9292, "trk-ogz", "Latn, Cyrl, Arab", ancestors = "trk-oat", dotted_dotless_i = true, strip_diacritics = { Latn = { from = {"ʼ"}, to = {"'"}, }, Arab = { module = "ar-stripdiacritics", ["from"] = { "ۆ", "ۇ", "وْ", "ڲ", "ؽ", }, ["to"] = { "و", "و", "و", "گ", "ی", }, }, }, display_text = { Latn = { from = {"'"}, to = {"ʼ"} } }, sort_key = { Latn = { from = { "i", -- Ensure "i" comes after "ı". "ç", "ə", "ğ", "x", "ı", "q", "ö", "ş", "ü", "w" }, to = { "i" .. p[1], "c" .. p[1], "e" .. p[1], "g" .. p[1], "h" .. p[1], "i", "k" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1], "z" .. p[1] } }, Cyrl = { from = {"ғ", "ә", "ы", "ј", "ҝ", "ө", "ү", "һ", "ҹ"}, to = {"г" .. p[1], "е" .. p[1], "и" .. p[1], "и" .. p[2], "к" .. p[1], "о" .. p[1], "у" .. p[1], "х" .. p[1], "ч" .. p[1]} }, }, } m["ba"] = { "Bashkir", 13389, "trk-kbu", "Cyrl", translit = "ba-translit", override_translit = true, sort_key = { from = {"ғ", "ҙ", "ё", "ҡ", "ң", "ө", "ҫ", "ү", "һ", "ә"}, to = {"г" .. p[1], "д" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "с" .. p[1], "у" .. p[1], "х" .. p[1], "э" .. p[1]} }, } m["be"] = { "Belarusian", 9091, "zle", "Cyrl, Latn", ancestors = "zle-mbe", translit = { Cyrl = "be-translit", }, strip_diacritics = { Cyrl = { remove_diacritics = c.grave .. c.acute, }, Latn = { remove_diacritics = c.grave .. c.acute, remove_exceptions = {"Ć", "ć", "Ń", "ń", "Ś", "ś", "Ź", "ź"}, }, }, sort_key = { Cyrl = { remove_diacritics = c.grave .. c.acute, from = {"ґ", "ё", "і", "ў"}, to = {"г" .. p[1], "е" .. p[1], "и" .. p[1], "у" .. p[1]} }, Latn = { remove_diacritics = c.grave .. c.acute, remove_exceptions = {"ć", "ń", "ś", "ź"}, from = {"ć", "č", "dz", "dź", "dž", "ch", "ł", "ń", "ś", "š", "ŭ", "ź", "ž"}, to = {"c" .. p[1], "c" .. p[2], "d" .. p[1], "d" .. p[2], "d" .. p[3], "h" .. p[1], "l" .. p[1], "n" .. p[1], "s" .. p[1], "s" .. p[2], "u" .. p[1], "z" .. p[1], "z" .. p[2]} }, }, standard_chars = { Cyrl = "АаБбВвГгДдЕеЁёЖжЗзІіЙйКкЛлМмНнОоПпРрСсТтУуЎўФфХхЦцЧчШшЫыЬьЭэЮюЯя", Latn = "AaBbCcĆćČčDdEeFfGgHhIiJjKkLlŁłMmNnŃńOoPpRrSsŚśŠšTtUuŬŭVvYyZzŹźŽž", (c.punc:gsub("'", "")) -- Exclude apostrophe. }, } m["bg"] = { "Bulgarian", 7918, "zls", "Cyrl", ancestors = "cu-bgm", translit = "bg-translit", strip_diacritics = { remove_diacritics = c.grave .. c.acute, remove_exceptions = {"%f[^%z%s]ѝ%f[%z%s]"}, }, sort_key = { remove_diacritics = c.grave .. c.acute, remove_exceptions = {"%f[^%z%s]ѝ%f[%z%s]"}, }, standard_chars = "АаБбВвГгДдЕеЖжЗзИиЙйКкЛлМмНнОоПпРрСсТтУуФфХхЦцЧчШшЩщЪъЬьЮюЯя" .. c.punc, } m["bh"] = { "Bihari", 135305, "inc-eas", "Deva", } m["bi"] = { "Bislama", 35452, "crp", "Latn", ancestors = "en", } m["bm"] = { "Bambara", 33243, "dmn-emn", "Latn, Nkoo", sort_key = { Latn = { from = {"ɛ", "ɲ", "ŋ", "ɔ"}, to = {"e" .. p[1], "n" .. p[1], "n" .. p[2], "o" .. p[1]} }, }, } m["bn"] = { "Bengali", 9610, "inc-bas", "Beng, Newa", ancestors = "inc-mbn", translit = { Beng = "bn-translit" }, } m["bo"] = { "Tibetan", 34271, "sit-tib", "Tibt", -- sometimes Deva? ancestors = "xct", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["br"] = { "Breton", 12107, "cel-brs", "Latn", ancestors = "xbm", sort_key = { from = {"ch", "c['ʼ’]h"}, to = {"c" .. p[1], "c" .. p[2]} }, } m["ca"] = { "Catalan", 7026, "roa-ocr", "Latn", ancestors = "roa-oca", sort_key = {remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla .. "·"}, standard_chars = "AaÀàBbCcÇçDdEeÉéÈèFfGgHhIiÍíÏïJjLlMmNnOoÓóÒòPpQqRrSsTtUuÚúÜüVvXxYyZz·" .. c.punc, } m["ce"] = { "Chechen", 33350, "cau-vay", "Cyrl, Latn, Arab", translit = { Cyrl = "cau-nec-translit", Arab = "ar-translit", }, override_translit = true, display_text = { Cyrl = s["cau-Cyrl-displaytext"] }, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, sort_key = { Cyrl = { from = {"аь", "гӏ", "ё", "кх", "къ", "кӏ", "оь", "пӏ", "тӏ", "уь", "хь", "хӏ", "цӏ", "чӏ", "юь", "яь"}, to = {"а" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "к" .. p[2], "к" .. p[3], "о" .. p[1], "п" .. p[1], "т" .. p[1], "у" .. p[1], "х" .. p[1], "х" .. p[2], "ц" .. p[1], "ч" .. p[1], "ю" .. p[1], "я" .. p[1]} }, }, } m["ch"] = { "Chamorro", 33262, "poz", "Latn", sort_key = { remove_diacritics = "'", from = {"å", "ch", "ñ", "ng"}, to = {"a" .. p[1], "c" .. p[1], "n" .. p[1], "n" .. p[2]} }, } m["co"] = { "Corsican", 33111, "roa-itr", "Latn", sort_key = { from = {"chj", "ghj", "sc", "sg"}, to = {"c" .. p[1], "g" .. p[1], "s" .. p[1], "s" .. p[2]} }, standard_chars = "AaÀàBbCcDdEeÈèFfGgHhIiÌìÏïJjLlMmNnOoÒòPpQqRrSsTtUuÙùÜüVvZz" .. c.punc, } m["cr"] = { "Cree", 33390, "alg", "Latn, Cans", translit = { Cans = "cr-translit" }, } m["cs"] = { "Czech", 9056, "zlw", "Latn", ancestors = "cs-ear", sort_key = { from = {"á", "č", "ď", "é", "ě", "ch", "í", "ň", "ó", "ř", "š", "ť", "ú", "ů", "ý", "ž"}, to = {"a" .. p[1], "c" .. p[1], "d" .. p[1], "e" .. p[1], "e" .. p[2], "h" .. p[1], "i" .. p[1], "n" .. p[1], "o" .. p[1], "r" .. p[1], "s" .. p[1], "t" .. p[1], "u" .. p[1], "u" .. p[2], "y" .. p[1], "z" .. p[1]} }, standard_chars = "AaÁáBbCcČčDdĎďEeÉéĚěFfGgHhIiÍíJjKkLlMmNnŇňOoÓóPpRrŘřSsŠšTtŤťUuÚúŮůVvYyÝýZzŽž" .. c.punc, } m["cu"] = { "Old Church Slavonic", 35499, "zls", "Cyrs, Glag, Zname", translit = { Cyrs = "Cyrs-translit", Glag = "Glag-translit" }, -- Cyrs strip_diacritics, sort_key in [[Module:scripts/data]] } m["cv"] = { "Chuvash", 33348, "trk-ogr", "Cyrl", ancestors = "cv-mid", translit = "cv-translit", override_translit = true, sort_key = { from = {"ӑ", "ё", "ӗ", "ҫ", "ӳ"}, to = {"а" .. p[1], "е" .. p[1], "е" .. p[2], "с" .. p[1], "у" .. p[1]} }, } m["cy"] = { "Welsh", 9309, "cel-brw", "Latn", ancestors = "wlm", sort_key = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer .. "'", from = {"ch", "dd", "ff", "ng", "ll", "ph", "rh", "th"}, to = {"c" .. p[1], "d" .. p[1], "f" .. p[1], "g" .. p[1], "l" .. p[1], "p" .. p[1], "r" .. p[1], "t" .. p[1]} }, standard_chars = "ÂâAaBbCcDdEeÊêFfGgHhIiÎîLlMmNnOoÔôPpRrSsTtUuÛûWwŴŵYyŶŷ" .. c.punc, } m["da"] = { "Danish", 9035, "gmq-eas", "Latn", ancestors = "gmq-oda", sort_key = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.dacute .. c.caron .. c.cedilla, remove_exceptions = {"å"}, from = {"æ", "ø", "å"}, to = {"z" .. p[1], "z" .. p[2], "z" .. p[3]} }, standard_chars = "AaBbDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsTtUuVvYyÆæØøÅå" .. c.punc, } m["de"] = { "German", 188, "gmw-hgm", "Latn, Latf, Brai", ancestors = "de-ear", sort_key = { Latn = s["de-Latn-sortkey"], Latf = s["de-Latn-sortkey"], }, standard_chars = { Latn = s["de-Latn-standardchars"], Latf = s["de-Latn-standardchars"], Brai = c.braille, c.punc } } m["dv"] = { "Dhivehi", 32656, "inc-ins", "Thaa, Diak", translit = { Thaa = "dv-translit", Diak = "Diak-translit", }, ancestors = "dv-old", override_translit = true, } m["dz"] = { "Dzongkha", 33081, "sit-tib", "Tibt", ancestors = "xct", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["ee"] = { "Ewe", 30005, "alv-gbe", "Latn", sort_key = { remove_diacritics = c.tilde, from = {"ɖ", "dz", "ɛ", "ƒ", "gb", "ɣ", "kp", "ny", "ŋ", "ɔ", "ts", "ʋ"}, to = {"d" .. p[1], "d" .. p[2], "e" .. p[1], "f" .. p[1], "g" .. p[1], "g" .. p[2], "k" .. p[1], "n" .. p[1], "n" .. p[2], "o" .. p[1], "t" .. p[1], "v" .. p[1]} }, } m["el"] = { "Greek", 9129, "grk", "Grek, Polyt, Brai", ancestors = "el-kth", translit = "el-translit", override_translit = true, -- Grek and Polyt display_text, strip_diacritics, sort_key in [[Module:scripts/data]] standard_chars = { Grek = "΅·ͺ΄ΑαΆάΒβΓγΔδΕεέΈΖζΗηΉήΘθΙιΊίΪϊΐΚκΛλΜμΝνΞξΟοΌόΠπΡρΣσςΤτΥυΎύΫϋΰΦφΧχΨψΩωΏώ", Brai = c.braille, c.punc }, } m["en"] = { "English", 1860, "gmw-ang", "Latn, Brai, Shaw, Dsrt", -- entries in Shaw or Dsrt might require prior discussion wikimedia_codes = "en, simple", ancestors = "en-ear", sort_key = { Latn = { -- Many of these are needed for sorting language names. remove_diacritics = "'\"%-%.,%s·ʻʼ" .. c.diacritics, -- These are found in pagenames. from = {"[ɒæ🅱¢©ᴄðđəǝɜɡħʜıɨłŋɲøɔœꝑꝓꝕßʋ]"}, to = {{ ["ɒ"] = "a", ["æ"] = "ae", ["🅱"] = "b", ["¢"] = "c", ["©"] = "c", ["ᴄ"] = "c", ["ð"] = "d", ["đ"] = "d", ["ə"] = "e", ["ǝ"] = "e", ["ɜ"] = "e", ["ɡ"] = "g", ["ħ"] = "h", ["ʜ"] = "h", ["ı"] = "i", ["ɨ"] = "i", ["ł"] = "l", ["ŋ"] = "n", ["ɲ"] = "n", ["ø"] = "o", ["ɔ"] = "o", ["œ"] = "oe", ["ꝑ"] = "p", ["ꝓ"] = "p", ["ꝕ"] = "p", ["ß"] = "ss", ["ʋ"] = "v", }}, }, }, standard_chars = { Latn = "AaBbCcDdEeFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvWwXxYyZz", Brai = c.braille, c.punc }, } m["eo"] = { "Esperanto", 143, "art", "Latn", sort_key = { remove_diacritics = c.grave .. c.acute, from = {"ĉ", "ĝ", "ĥ", "ĵ", "ŝ", "ŭ"}, to = {"c" .. p[1], "g" .. p[1], "h" .. p[1], "j" .. p[1], "s" .. p[1], "u" .. p[1]} }, standard_chars = "AaBbCcĈĉDdEeFfGgĜĝHhĤĥIiJjĴĵKkLlMmNnOoPpRrSsŜŝTtUuŬŭVvZz" .. c.punc, } m["es"] = { "Spanish", 1321, "roa-cas", "Latn, Brai", ancestors = "es-ear", sort_key = { Latn = { remove_exceptions = {"ñ"}, remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.diaer .. c.cedilla, from = {"ª", "æ", "ñ", "º", "œ"}, to = {"a", "ae", "n" .. p[1], "o", "oe"} }, }, standard_chars = { Latn = "AaÁáBbCcDdEeÉéFfGgHhIiÍíJjLlMmNnÑñOoÓóPpQqRrSsTtUuÚúÜüVvXxYyZz", Brai = c.braille, c.punc }, } m["et"] = { "Estonian", 9072, "urj-fin", "Latn", sort_key = { from = { "š", "ž", "õ", "ä", "ö", "ü", -- 2 chars "z" -- 1 char }, to = { "s" .. p[1], "s" .. p[3], "w" .. p[1], "w" .. p[2], "w" .. p[3], "w" .. p[4], "s" .. p[2] } }, standard_chars = "AaBbDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsTtUuVvÕõÄäÖöÜü" .. c.punc, } m["eu"] = { "Basque", 8752, "euq", "Latn", sort_key = { from = {"ç", "ñ"}, to = {"c" .. p[1], "n" .. p[1]} }, standard_chars = "AaBbDdEeFfGgHhIiJjKkLlMmNnÑñOoPpRrSsTtUuXxZz" .. c.punc, } m["fa"] = { "Persian", 9168, "ira-swi", "Arab, Hebr", ancestors = "fa-cls", strip_diacritics = { Arab = { -- character "ۂ" code U+06C2 to "ه" and "هٔ" (U+0647 + U+0654) to "ه"; hamzatu l-waṣli to a regular alif from = {"هٔ", "ٱ"}, -- character "ۂ" code U+06C2 to "ه"; hamzatu l-waṣli to a regular alif to = {"ه", "ا"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.superalef, }, }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["ff"] = { "Fula", 33454, "alv-fwo", "Latn, Adlm", } m["fi"] = { "Finnish", 1412, "urj-fin", "Latn", display_text = { from = {"'"}, to = {"’"} }, strip_diacritics = { -- used to indicate gemination of the next consonant remove_diacritics = "ˣ", from = {"’"}, to = {"'"}, }, sort_key = { -- [[Appendix:Finnish alphabet#Collation]] + "aͤ" and "oͤ" as historical variants of "ä" and "ö". remove_diacritics = "'’:" .. c.diacritics, remove_exceptions = { "a[" .. c.ringabove .. c.diaer .. c.small_e .. "]", -- åäaͤ "o[" .. c.diaer .. c.tilde .. c.dacute .. c.small_e .. "]", -- öõőoͤ "u[" .. c.diaer .. c.dacute .. "]" -- üű }, from = {"æ", "[ðđ]", "ł", "ŋ", "œ", "ß", "þ", "u[" .. c.diaer .. c.dacute .. "]", "å", "aͤ", "o[" .. c.tilde .. c.dacute .. c.small_e .. "]", "ø", "(.)['%-]"}, to = {"ae", "d", "l", "n", "oe", "ss", "th", "y", "z" .. p[1], "ä", "ö", "ö", "%1"} }, standard_chars = "AaBbDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsTtUuVvYyÄäÖö" .. c.punc, } m["fj"] = { "Fijian", 33295, "poz-pcc", "Latn", } m["fo"] = { "Faroese", 25258, "gmq-ins", "Latn", sort_key = { from = {"á", "ð", "í", "ó", "ú", "ý", "æ", "ø"}, to = {"a" .. p[1], "d" .. p[1], "i" .. p[1], "o" .. p[1], "u" .. p[1], "y" .. p[1], "z" .. p[1], "z" .. p[2]} }, standard_chars = "AaÁáBbDdÐðEeFfGgHhIiÍíJjKkLlMmNnOoÓóPpRrSsTtUuÚúVvYyÝýÆæØø" .. c.punc, } m["fr"] = { "French", 150, "roa-oil", "Latn, Brai", ancestors = "frm", sort_key = { Latn = s["roa-oil-sortkey"] }, standard_chars = { Latn = "AaÀàÂâBbCcÇçDdEeÉéÈèÊêËëFfGgHhIiÎîÏïJjLlMmNnOoÔôŒœPpQqRrSsTtUuÙùÛûÜüVvXxYyZz", Brai = c.braille, c.punc }, } m["fy"] = { "West Frisian", 27175, "gmw-fri", "Latn", sort_key = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer, from = {"y"}, to = {"i"} }, standard_chars = "AaâäàÆæBbCcDdEeéêëèFfGgHhIiïìYyỳJjKkLlMmNnOoôöòPpRrSsTtUuúûüùVvWwZz" .. c.punc, } m["ga"] = { "Irish", 9142, "cel-gae", "Latn, Latg", ancestors = "mga", sort_key = { remove_diacritics = c.acute, from = {"ḃ", "ċ", "ḋ", "ḟ", "ġ", "ṁ", "ṗ", "ṡ", "ṫ"}, to = {"bh", "ch", "dh", "fh", "gh", "mh", "ph", "sh", "th"} }, standard_chars = "AaÁáBbCcDdEeÉéFfGgHhIiÍíLlMmNnOoÓóPpRrSsTtUuÚúVv" .. c.punc, } m["gd"] = { "Scottish Gaelic", 9314, "cel-gae", "Latn, Latg", ancestors = "mga", sort_key = {remove_diacritics = c.grave .. c.acute}, standard_chars = "AaÀàBbCcDdEeÈèFfGgHhIiÌìLlMmNnOoÒòPpRrSsTtUuÙù" .. c.punc, } m["gl"] = { "Galician", 9307, "roa-gap", "Latn", sort_key = { remove_diacritics = c.acute, from = {"ñ"}, to = {"n" .. p[1]} }, standard_chars = "AaÁáBbCcDdEeÉéFfGgHhIiÍíÏïLlMmNnÑñOoÓóPpQqRrSsTtUuÚúÜüVvXxZz" .. c.punc, } m["gu"] = { "Gujarati", 5137, "inc-wes", "Arab, Gujr", ancestors = "inc-mgu", translit = { Gujr = "gu-translit", }, strip_diacritics = { Arab = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.kasra .. c.shadda .. c.sukun}, Gujr = {remove_diacritics = "઼"}, }, } m["gv"] = { "Manx", 12175, "cel-gae", "Latn", ancestors = "mga", sort_key = {remove_diacritics = c.cedilla .. "-"}, standard_chars = "AaBbCcÇçDdEeFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvWwYy" .. c.punc, } m["ha"] = { "Hausa", 56475, "cdc-wst", "Latn, Arab", strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron} }, sort_key = { Latn = { from = {"ɓ", "b'", "ɗ", "d'", "ƙ", "k'", "sh", "ƴ", "'y"}, to = {"b" .. p[1], "b" .. p[2], "d" .. p[1], "d" .. p[2], "k" .. p[1], "k" .. p[2], "s" .. p[1], "y" .. p[1], "y" .. p[2]} }, }, } m["he"] = { "Hebrew", 9288, "sem-can", "Hebr, Phnx, Brai, Samr", ancestors = "he-med", -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] -- Samr strip_diacritics, sort_key in [[Module:scripts/data]] -- Phnx translit in [[Module:scripts/data]] (NOTE: not present before, presumably an accidental omission) } m["hi"] = { "Hindi", 1568, "inc-hnd", "Deva, Kthi, Newa", translit = { Deva = "hi-translit" }, standard_chars = { Deva = "अआइईउऊएऐओऔकखगघङचछजझञटठडढणतथदधनपफबभमयरलवशषसहत्रज्ञक्षक़ख़ग़ज़झ़ड़ढ़फ़काखागाघाङाचाछाजाझाञाटाठाडाढाणाताथादाधानापाफाबाभामायारालावाशाषासाहात्राज्ञाक्षाक़ाख़ाग़ाज़ाझ़ाड़ाढ़ाफ़ाकिखिगिघिङिचिछिजिझिञिटिठिडिढिणितिथिदिधिनिपिफिबिभिमियिरिलिविशिषिसिहित्रिज्ञिक्षिक़िख़िग़िज़िझ़िड़िढ़िफ़िकीखीगीघीङीचीछीजीझीञीटीठीडीढीणीतीथीदीधीनीपीफीबीभीमीयीरीलीवीशीषीसीहीत्रीज्ञीक्षीक़ीख़ीग़ीज़ीझ़ीड़ीढ़ीफ़ीकुखुगुघुङुचुछुजुझुञुटुठुडुढुणुतुथुदुधुनुपुफुबुभुमुयुरुलुवुशुषुसुहुत्रुज्ञुक्षुक़ुख़ुग़ुज़ुझ़ुड़ुढ़ुफ़ुकूखूगूघूङूचूछूजूझूञूटूठूडूढूणूतूथूदूधूनूपूफूबूभूमूयूरूलूवूशूषूसूहूत्रूज्ञूक्षूक़ूख़ूग़ूज़ूझ़ूड़ूढ़ूफ़ूकेखेगेघेङेचेछेजेझेञेटेठेडेढेणेतेथेदेधेनेपेफेबेभेमेयेरेलेवेशेषेसेहेत्रेज्ञेक्षेक़ेख़ेग़ेज़ेझ़ेड़ेढ़ेफ़ेकैखैगैघैङैचैछैजैझैञैटैठैडैढैणैतैथैदैधैनैपैफैबैभैमैयैरैलैवैशैषैसैहैत्रैज्ञैक्षैक़ैख़ैग़ैज़ैझ़ैड़ैढ़ैफ़ैकोखोगोघोङोचोछोजोझोञोटोठोडोढोणोतोथोदोधोनोपोफोबोभोमोयोरोलोवोशोषोसोहोत्रोज्ञोक्षोक़ोख़ोग़ोज़ोझ़ोड़ोढ़ोफ़ोकौखौगौघौङौचौछौजौझौञौटौठौडौढौणौतौथौदौधौनौपौफौबौभौमौयौरौलौवौशौषौसौहौत्रौज्ञौक्षौक़ौख़ौग़ौज़ौझ़ौड़ौढ़ौफ़ौक्ख्ग्घ्ङ्च्छ्ज्झ्ञ्ट्ठ्ड्ढ्ण्त्थ्द्ध्न्प्फ्ब्भ्म्य्र्ल्व्श्ष्स्ह्त्र्ज्ञ्क्ष्क़्ख़्ग़्ज़्झ़्ड़्ढ़्फ़्।॥०१२३४५६७८९॰", c.punc }, } m["ho"] = { "Hiri Motu", 33617, "crp", "Latn", ancestors = "meu", } m["ht"] = { "Haitian Creole", 33491, "crp", "Latn", ancestors = "ht-sdm", sort_key = { from = { "oun", -- 3 chars "an", "ch", "è", "en", "ng", "ò", "on", "ou", "ui" -- 2 chars }, to = { "o" .. p[4], "a" .. p[1], "c" .. p[1], "e" .. p[1], "e" .. p[2], "n" .. p[1], "o" .. p[1], "o" .. p[2], "o" .. p[3], "u" .. p[1] } }, } m["hu"] = { "Hungarian", 9067, "urj-ugr", "Latn, Hung", ancestors = "ohu", sort_key = { Latn = { from = { "dzs", -- 3 chars "á", "cs", "dz", "é", "gy", "í", "ly", "ny", "ó", "ö", "ő", "sz", "ty", "ú", "ü", "ű", "zs", -- 2 chars }, to = { "d" .. p[2], "a" .. p[1], "c" .. p[1], "d" .. p[1], "e" .. p[1], "g" .. p[1], "i" .. p[1], "l" .. p[1], "n" .. p[1], "o" .. p[1], "o" .. p[2], "o" .. p[3], "s" .. p[1], "t" .. p[1], "u" .. p[1], "u" .. p[2], "u" .. p[3], "z" .. p[1], } }, }, standard_chars = { Latn = "AaÁáBbCcDdEeÉéFfGgHhIiÍíJjKkLlMmNnOoÓóÖöŐőPpQqRrSsTtUuÚúÜüŰűVvWwXxYyZz", c.punc }, } m["hy"] = { "Armenian", 8785, "hyx", "Armn, Brai", ancestors = "axm", -- Armn translit in [[Module:scripts/data]] override_translit = true, strip_diacritics = { Armn = { remove_diacritics = "՛՜՞՟", from = {"եւ", "<sup>յ</sup>", "<sup>ի</sup>", "<sup>է</sup>", "յ̵", "ՙ", "՚"}, to = {"և", "յ", "ի", "է", "ֈ", "ʻ", "’"} }, }, sort_key = { Armn = { from = { "ու", "եւ", -- 2 chars "և" -- 1 char }, to = { "ւ", "եվ", "եվ" } }, }, } m["hz"] = { "Herero", 33315, "bnt-swb", "Latn", } m["ia"] = { "Interlingua", 35934, "art", "Latn", } m["id"] = { "Indonesian", 9240, "poz-mly", "Latn", ancestors = "ms", standard_chars = "AaBbCcDdEeFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvWwXxYyZz" .. c.punc, } m["ie"] = { "Interlingue", 35850, "art", "Latn", type = "appendix-constructed", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ}, } m["ig"] = { "Igbo", 33578, "alv-igb", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.macron}, sort_key = { from = {"gb", "gh", "gw", "ị", "kp", "kw", "ṅ", "nw", "ny", "ọ", "sh", "ụ"}, to = {"g" .. p[1], "g" .. p[2], "g" .. p[3], "i" .. p[1], "k" .. p[1], "k" .. p[2], "n" .. p[1], "n" .. p[2], "n" .. p[3], "o" .. p[1], "s" .. p[1], "u" .. p[1]} }, } m["ii"] = { "Nuosu", 34235, "tbq-nlo", "Yiii", translit = "ii-translit", } m["ik"] = { "Inupiaq", 27183, "esx-inu", "Latn", sort_key = { from = { "ch", "ġ", "dj", "ḷ", "ł̣", "ñ", "ng", "r̂", "sr", "zr", -- 2 chars "ł", "ŋ", "ʼ" -- 1 char }, to = { "c" .. p[1], "g" .. p[1], "h" .. p[1], "l" .. p[1], "l" .. p[3], "n" .. p[1], "n" .. p[2], "r" .. p[1], "s" .. p[1], "z" .. p[1], "l" .. p[2], "n" .. p[2], "z" .. p[2] } }, } m["io"] = { "Ido", 35224, "art", "Latn", } m["is"] = { "Icelandic", 294, "gmq-ins", "Latn", sort_key = { from = {"á", "ð", "é", "í", "ó", "ú", "ý", "þ", "æ", "ö"}, to = {"a" .. p[1], "d" .. p[1], "e" .. p[1], "i" .. p[1], "o" .. p[1], "u" .. p[1], "y" .. p[1], "z" .. p[1], "z" .. p[2], "z" .. p[3]} }, standard_chars = "AaÁáBbDdÐðEeÉéFfGgHhIiÍíJjKkLlMmNnOoÓóPpRrSsTtUuÚúVvXxYyÝýÞþÆæÖö" .. c.punc, } m["it"] = { "Italian", 652, "roa-itr", "Latn", ancestors = "roa-oit", sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer .. c.ringabove}, standard_chars = "AaÀàBbCcDdEeÈèÉéFfGgHhIiÌìLlMmNnOoÒòPpQqRrSsTtUuÙùVvZz" .. c.punc, } m["iu"] = { "Inuktitut", 29921, "esx-inu", "Cans, Latn", translit = { Cans = "cr-translit" }, override_translit = true, } m["ja"] = { "Japanese", 5287, "jpx", "Jpan, Latn, Brai", ancestors = "ja-ear", translit = s["jpx-translit"], link_tr = true, display_text = s["jpx-displaytext"], strip_diacritics = s["jpx-stripdiacritics"], sort_key = s["jpx-sortkey"], } m["jv"] = { "Javanese", 33549, "poz", "Latn, Java, Arab", ancestors = "kaw", translit = { Java = "jv-translit" }, link_tr = true, strip_diacritics = { Latn = {remove_diacritics = c.circ} -- Modern jv don't use ê }, sort_key = { Latn = { from = {"å", "dh", "é", "è", "ng", "ny", "th"}, to = {"a" .. p[1], "d" .. p[1], "e" .. p[1], "e" .. p[2], "n" .. p[1], "n" .. p[2], "t" .. p[1]} }, }, } m["ka"] = { "Georgian", 8108, "ccs-gzn", "Geor, Geok, Hebr", -- Hebr is used to write Judeo-Georgian ancestors = "ka-mid", -- Geor, Geok translit in [[Module:scripts/data]] override_translit = true, strip_diacritics = { Geor = s["ka-stripdiacritics"], Geok = s["ka-stripdiacritics"], }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["kg"] = { "Kongo", 33702, "bnt-kng", "Latn", } m["ki"] = { "Kikuyu", 33587, "bnt-kka", "Latn", } m["kj"] = { "Kwanyama", 1405077, "bnt-ova", "Latn", } m["kk"] = { "Kazakh", 9252, "trk-kno", "Cyrl, Latn, Arab", translit = "kk-translit", -- override_translit = true, sort_key = { Cyrl = { from = {"ә", "ғ", "ё", "қ", "ң", "ө", "ұ", "ү", "һ", "і"}, to = {"а" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1], "у" .. p[2], "х" .. p[1], "ы" .. p[1]} }, }, standard_chars = { Cyrl = "АаӘәБбВвГгҒғДдЕеЁёЖжЗзИиЙйКкҚқЛлМмНнҢңОоӨөПпРрСсТтУуҰұҮүФфХхҺһЦцЧчШшЩщЪъЫыІіЬьЭэЮюЯя", c.punc }, } m["kl"] = { "Greenlandic", 25355, "esx-inu", "Latn", sort_key = { from = {"æ", "ø", "å"}, to = {"z" .. p[1], "z" .. p[2], "z" .. p[3]} } } m["km"] = { "Khmer", 9205, "mkh-kmr", "Khmr", ancestors = "xhm", translit = "km-translit", --This might yield unwanted result unless its entry has {{km-IPA}}. } m["kn"] = { "Kannada", 33673, "dra-kan", "Knda, Tutg", ancestors = "dra-mkn", -- Knda translit in [[Module:scripts/data]] } m["ko"] = { "Korean", 9176, "qfa-kor", "Kore, Brai", ancestors = "ko-ear", translit = { Kore = "ko-translit", }, -- Kore strip_diacritics in [[Module:scripts/data]] } m["kr"] = { "Kanuri", 36094, "ssa-sah", "Latn, Arab", -- the sortkey and strip_diacritics are only for standard Kanuri; when dialectal entries get added, someone will have to work out how the dialects should be represented orthographically strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.breve} }, sort_key = { Latn = { from = {"ǝ", "ny", "ɍ", "sh"}, to = {"e" .. p[1], "n" .. p[1], "r" .. p[1], "s" .. p[1]} }, }, } m["ks"] = { "Kashmiri", 33552, "inc-kas", "Aran, Deva, Shrd, Latn", translit = { Aran = "ks-Aran-translit", Deva = "ks-Deva-translit", -- Shrd translit in [[Module:scripts/data]] }, } -- "kv" is treated as "koi", "kpv", see [[WT:LT]] m["kw"] = { "Cornish", 25289, "cel-brs", "Latn", ancestors = "cnx", sort_key = { from = {"ch"}, to = {"c" .. p[1]} }, } m["ky"] = { "Kyrgyz", 9255, "trk-kkp", "Cyrl, Latn, Arab", translit = { Cyrl = "ky-translit" }, override_translit = true, sort_key = { Cyrl = { from = {"ё", "ң", "ө", "ү"}, to = {"е" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1]} }, }, } m["la"] = { "Latin", 397, "itc-laf", "Latn, Ital", ancestors = "itc-ola", -- Ital translit in [[Module:scripts/data]] (NOTE: formerly not present, probably an accidental omission) display_text = { Latn = s["itc-Latn-displaytext"] }, strip_diacritics = { Latn = s["itc-Latn-stripdiacritics"] }, sort_key = { Latn = s["itc-Latn-sortkey"] }, standard_chars = { Latn = "AaBbCcDdEeFfGgHhIiLlMmNnOoPpQqRrSsTtUuVvXx", c.punc }, } m["lb"] = { "Luxembourgish", 9051, "gmw-hgm", "Latn, Brai", ancestors = "gmw-cfr", sort_key = { Latn = { from = {"ä", "ë", "é"}, to = {"z" .. p[1], "z" .. p[2], "z" .. p[3]} }, }, } m["lg"] = { "Luganda", 33368, "bnt-nyg", "Latn", strip_diacritics = {remove_diacritics = c.acute .. c.circ}, sort_key = { from = {"ŋ"}, to = {"n" .. p[1]} }, } m["li"] = { "Limburgish", 102172, "gmw-frk", "Latn", ancestors = "dum", } m["ln"] = { "Lingala", 36217, "bnt-bmo", "Latn", sort_key = { remove_diacritics = c.acute .. c.circ .. c.caron, from = {"ɛ", "gb", "mb", "mp", "nd", "ng", "nk", "ns", "nt", "ny", "nz", "ɔ"}, to = {"e" .. p[1], "g" .. p[1], "m" .. p[1], "m" .. p[2], "n" .. p[1], "n" .. p[2], "n" .. p[3], "n" .. p[4], "n" .. p[5], "n" .. p[6], "n" .. p[7], "o" .. p[1]} }, } m["lo"] = { "Lao", 9211, "tai-swe", "Laoo", -- also Tai Noi/Lao Buhan script translit = "lo-translit", sort_key = "Laoo-sortkey", standard_chars = "0-9ກຂຄງຈຊຍດຕຖທນບປຜຝພຟມຢຣລວສຫອຮຯ-ໝ" .. c.punc, } m["lt"] = { "Lithuanian", 9083, "bat-eas", "Latn", ancestors = "olt", display_text = "lt-common", strip_diacritics = "lt-common", sort_key = "lt-common", standard_chars = "AaĄąBbCcČčDdEeĘęĖėFfGgHhIiĮįYyJjKkLlMmNnOoPpRrSsŠšTtUuŲųŪūVvZzŽž" .. c.punc, } m["lu"] = { "Luba-Katanga", 36157, "bnt-lub", "Latn", } m["lv"] = { "Latvian", 9078, "bat-eas", "Latn", strip_diacritics = { -- This attempts to convert vowels with tone marks to vowels either with or without macrons. Specifically, there should be no macrons if the vowel is part of a diphthong (including resonant diphthongs such pìrksts -> pirksts not #pīrksts). What we do is first convert the vowel + tone mark to a vowel + tilde in a decomposed fashion, then remove the tilde in diphthongs, then convert the remaining vowel + tilde sequences to macroned vowels, then delete any other tilde. We leave already-macroned vowels alone: Both e.g. ar and ār occur before consonants. FIXME: This still might not be sufficient. from = {"([Ee])" .. c.cedilla, "[" .. c.grave .. c.circ .. c.tilde .."]", "([aAeEiIoOuU])" .. c.tilde .."?([lrnmuiLRNMUI])" .. c.tilde .. "?([^aAeEiIoOuU])", "([aAeEiIoOuU])" .. c.tilde .."?([lrnmuiLRNMUI])" .. c.tilde .."?$", "([iI])" .. c.tilde .. "?([eE])" .. c.tilde .. "?", "([aAeEiIuU])" .. c.tilde, c.tilde}, to = {"%1", c.tilde, "%1%2%3", "%1%2", "%1%2", "%1" .. c.macron} }, sort_key = { from = {"ā", "č", "ē", "ģ", "ī", "ķ", "ļ", "ņ", "š", "ū", "ž"}, to = {"a" .. p[1], "c" .. p[1], "e" .. p[1], "g" .. p[1], "i" .. p[1], "k" .. p[1], "l" .. p[1], "n" .. p[1], "s" .. p[1], "u" .. p[1], "z" .. p[1]} }, standard_chars = "AaĀāBbCcČčDdEeĒēFfGgĢģHhIiĪīJjKkĶķLlĻļMmNnŅņOoPpRrSsŠšTtUuŪūVvZzŽž" .. c.punc, } m["mg"] = { "Malagasy", 7930, "poz-bre", "Latn, Arab", } m["mh"] = { "Marshallese", 36280, "poz-mic", "Latn", sort_key = { from = {"ā", "ļ", "m̧", "ņ", "n̄", "o̧", "ō", "ū"}, to = {"a" .. p[1], "l" .. p[1], "m" .. p[1], "n" .. p[1], "n" .. p[2], "o" .. p[1], "o" .. p[2], "u" .. p[1]} }, } m["mi"] = { "Māori", 36451, "poz-pep", "Latn", sort_key = { remove_diacritics = c.macron, from = {"ng", "wh"}, to = {"n" .. p[1], "w" .. p[1]} }, } m["mk"] = { "Macedonian", 9296, "zls", "Cyrl, Polyt", ancestors = "cu", translit = { Cyrl = "mk-translit", -- FIXME: formerly no translit specified for Polyt; unclear if the default [[Module:grc-translit]] is -- acceptable, so we disable it for now Polyt = false, }, strip_diacritics = { Cyrl = { remove_diacritics = c.acute, remove_exceptions = {"Ѓ", "ѓ", "Ќ", "ќ"} }, }, sort_key = { Cyrl = { remove_diacritics = c.grave, remove_exceptions = {"ѓ", "ќ"}, from = {"ѓ", "ѕ", "ј", "љ", "њ", "ќ", "џ"}, to = {"д" .. p[1], "з" .. p[1], "и" .. p[1], "л" .. p[1], "н" .. p[1], "т" .. p[1], "ч" .. p[1]} }, }, -- Polyt display_text, strip_diacritics, sort_key in [[Module:scripts/data]] standard_chars = { Cyrl = "АаБбВвГгДдЃѓЕеЖжЗзЅѕИиЈјКкЛлЉљМмНнЊњОоПпРрСсТтЌќУуФфХхЦцЧчЏџШш", c.punc }, } m["ml"] = { "Malayalam", 36236, "dra-mal", "Mlym", override_translit = true, -- Mlym translit in [[Module:scripts/data]] } m["mn"] = { "Mongolian", 9246, "xgn-cen", "Cyrl, Mong, Latn, Brai", ancestors = "cmg", translit = { Cyrl = "mn-translit", -- Mong translit in [[Module:scripts/data]] }, override_translit = true, -- Mong display_text and strip_diacritics in [[Module:scripts/data]] strip_diacritics = { Cyrl = {remove_diacritics = c.grave .. c.acute}, }, sort_key = { Cyrl = { remove_diacritics = c.grave, from = {"ё", "ө", "ү"}, to = {"е" .. p[1], "о" .. p[1], "у" .. p[1]} }, }, standard_chars = { Cyrl = "АаБбВвГгДдЕеЁёЖжЗзИиЙйЛлМмНнОоӨөРрСсТтУуҮүХхЦцЧчШшЫыЬьЭэЮюЯя—", Brai = c.braille, c.punc }, } -- "mo" is treated as "ro", see [[WT:LT]] m["mr"] = { "Marathi", 1571, "inc-sou", "Deva, Modi", ancestors = "omr", translit = { Deva = "mr-translit", Modi = "mr-Modi-translit", }, strip_diacritics = { Deva = { from = {"च़", "ज़", "झ़"}, to = {"च", "ज", "झ"} }, }, } m["ms"] = { "Malay", 9237, "poz-mly", "Latn, Arab", ancestors = "ms-cla", standard_chars = { Latn = "AaBbCcDdEeFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvWwXxYyZz", c.punc }, } m["mt"] = { "Maltese", 9166, "sem-arb", "Latn", display_text = { from = {"'"}, to = {"’"} }, strip_diacritics = { from = {"’"}, to = {"'"}, }, ancestors = "sqr", sort_key = { from = { "ċ", "ġ", "ż", -- Convert into PUA so that decomposed form does not get caught by the next step. "([cgz])", -- Ensure "c" comes after "ċ", "g" comes after "ġ" and "z" comes after "ż". "g" .. p[1] .. "ħ", -- "għ" after initial conversion of "g". p[3], p[4], "ħ", "ie", p[5] -- Convert "ċ", "ġ", "ħ", "ie", "ż" into final output. }, to = { p[3], p[4], p[5], "%1" .. p[1], "g" .. p[2], "c", "g", "h" .. p[1], "i" .. p[1], "z" } }, } m["my"] = { "Burmese", 9228, "tbq-brm", "Mymr", ancestors = "obr", translit = "my-translit", override_translit = true, sort_key = { from = {"ျ", "ြ", "ွ", "ှ", "ဿ"}, to = {"္ယ", "္ရ", "္ဝ", "္ဟ", "သ္သ"} }, } m["na"] = { "Nauruan", 13307, "poz-mic", "Latn", } m["nb"] = { "Norwegian Bokmål", 25167, "gmq", "Latn", wikimedia_codes = "no", ancestors = "gmq-mno, da", -- da as an (but not the) ancestor of nb was agreed on - do not change without discussion sort_key = s["no-sortkey"], standard_chars = s["no-standardchars"], } m["nd"] = { "Northern Ndebele", 35613, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } m["ne"] = { "Nepali", 33823, "inc-pae", "Deva, Newa", translit = { Deva = "ne-translit" }, } m["ng"] = { "Ndonga", 33900, "bnt-ova", "Latn", } m["nl"] = { "Dutch", 7411, "gmw-frk", "Latn, Brai", ancestors = "dum", sort_key = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.diaer .. c.ringabove .. c.cedilla .. "'"}, }, standard_chars = { Latn = "AaBbCcDdEeFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvWwXxYyZzÄäËëÏïÖöÜü", Brai = c.braille, c.punc }, } m["nn"] = { "Norwegian Nynorsk", 25164, "gmq-wes", "Latn", ancestors = "gmq-mno", strip_diacritics = { remove_diacritics = c.grave .. c.acute, }, sort_key = s["no-sortkey"], standard_chars = s["no-standardchars"], } m["no"] = { "Norwegian", 9043, "gmq-wes", "Latn", ancestors = "gmq-mno", sort_key = s["no-sortkey"], standard_chars = s["no-standardchars"], } m["nr"] = { "Southern Ndebele", 36785, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } m["nv"] = { "Navajo", 13310, "apa", "Latn, Brai", sort_key = { remove_diacritics = c.acute .. c.ogonek, from = { "chʼ", "tłʼ", "tsʼ", -- 3 chars "ch", "dl", "dz", "gh", "hw", "kʼ", "kw", "sh", "tł", "ts", "zh", -- 2 chars "ł", "ʼ" -- 1 char }, to = { "c" .. p[2], "t" .. p[2], "t" .. p[4], "c" .. p[1], "d" .. p[1], "d" .. p[2], "g" .. p[1], "h" .. p[1], "k" .. p[1], "k" .. p[2], "s" .. p[1], "t" .. p[1], "t" .. p[3], "z" .. p[1], "l" .. p[1], "z" .. p[2] } }, } m["ny"] = { "Chichewa", 33273, "bnt-nys", "Latn", strip_diacritics = {remove_diacritics = c.acute .. c.circ}, sort_key = { from = {"ng'"}, to = {"ng"} }, } m["oc"] = { "Occitan", 14185, "roa-ocr", "Latn, Hebr", ancestors = "pro", sort_key = { Latn = { remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla, from = {"([lns])·h"}, to = {"%1h"} }, }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["oj"] = { "Ojibwe", 33875, "alg", "Cans, Latn", sort_key = { Latn = { from = {"aa", "ʼ", "ii", "oo", "sh", "zh"}, to = {"a" .. p[1], "h" .. p[1], "i" .. p[1], "o" .. p[1], "s" .. p[1], "z" .. p[1]} }, }, } m["om"] = { "Oromo", 33864, "cus-eas", "Latn, Ethi", } m["or"] = { "Odia", 33810, "inc-eas", "Orya", ancestors = "inc-mor", translit = "or-translit", } m["os"] = { "Ossetian", 33968, "xsc-sar", "Cyrl, Geor, Latn", ancestors = "oos", translit = { Cyrl = "os-translit", -- Geor translit in [[Module:scripts/data]] }, override_translit = true, display_text = { Cyrl = { from = {"æ", "Æ"}, to = {"ӕ", "Ӕ"} }, Latn = { from = {"ӕ", "Ӕ"}, to = {"æ", "Æ"} }, }, strip_diacritics = { Cyrl = { remove_diacritics = c.grave .. c.acute, from = {"æ", "Æ"}, to = {"ӕ", "Ӕ"} }, Latn = { from = {"ӕ", "Ӕ"}, to = {"æ", "Æ"} }, }, sort_key = { Cyrl = { from = {"ӕ", "гъ", "дж", "дз", "ё", "къ", "пъ", "тъ", "хъ", "цъ", "чъ"}, to = {"а" .. p[1], "г" .. p[1], "д" .. p[1], "д" .. p[2], "е" .. p[1], "к" .. p[1], "п" .. p[1], "т" .. p[1], "х" .. p[1], "ц" .. p[1], "ч" .. p[1]} }, }, } m["pa"] = { "Punjabi", 58635, "inc-pan", "Guru, Aran", translit = { Guru = "Guru-translit", Aran = "pa-Aran-translit", }, strip_diacritics = { Aran = { remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna, from = {"ݨ", "ࣇ"}, to = {"ن", "ل"} }, }, } m["pi"] = { "Pali", 36727, "inc-mid", "Latn, Brah, Deva, Beng, Sinh, Mymr, Thai, Lana, Laoo, Khmr, Cakm", --and also Khom ancestors = "sa", translit = { -- Brah translit in [[Module:scripts/data]] Deva = "sa-translit", Beng = "pi-translit", Sinh = "si-translit", Mymr = "pi-translit", Thai = "pi-translit", Lana = "pi-translit", Laoo = "pi-translit", Khmr = "pi-translit", Cakm = "Cakm-translit", }, strip_diacritics = { Thai = { from = {"ึ", u(0xF700), u(0xF70F)}, -- FIXME: Not clear what's going on with the PUA characters here. to = {"ิํ", "ฐ", "ญ"} }, Mymr = { remove_diacritics = c.VS01, }, }, sort_key = { -- FIXME: This needs to be converted into the current standardized format. from = {"ā", "ī", "ū", "ḍ", "ḷ", "m[" .. c.dotabove .. c.dotbelow .. "]", "ṅ", "ñ", "ṇ", "ṭ", "ॐ", "([เโ])([ก-ฮ])", "([ເໂ])([ກ-ຮ])", "ᩔ", "ᩕ", "ᩖ", "ᩘ", "([ᨭ-ᨱ])ᩛ", "([ᨷ-ᨾ])ᩛ", "ᩤ", u(0xFE00), u(0x200D)}, to = {"a~", "i~", "u~", "d~", "l~", "m~", "n~", "n~~", "n~~~", "t~", "ओँ", "%2%1", "%2%1", "ᩈ᩠ᩈ", "᩠ᩁ", "᩠ᩃ", "ᨦ᩠", "%1᩠ᨮ", "%1᩠ᨻ", "ᩣ"} }, } m["pl"] = { "Polish", 809, "zlw-lch", "Latn", ancestors = "zlw-mpl", sort_key = { from = {"ą", "ć", "ę", "ł", "ń", "ó", "ś", "ź", "ż"}, to = {"a" .. p[1], "c" .. p[1], "e" .. p[1], "l" .. p[1], "n" .. p[1], "o" .. p[1], "s" .. p[1], "z" .. p[1], "z" .. p[2]} }, standard_chars = "AaĄąBbCcĆćDdEeĘęFfGgHhIiJjKkLlŁłMmNnŃńOoÓóPpRrSsŚśTtUuWwYyZzŹźŻż" .. c.punc, } m["ps"] = { "Pashto", 58680, "ira-pat", "Arab", strip_diacritics = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.zwarakay .. c.superalef}, } m["pt"] = { "Portuguese", 5146, "roa-gap", "Latn, Brai", sort_key = { Latn = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.diaer .. c.cedilla, from = {"ª", "æ", "º", "œ"}, to = {"a", "ae", "o", "oe"} }, }, standard_chars = { Latn = "AaÁáÂâÃãBbCcÇçDdEeÉéÊêFfGgHhIiÍíJjLlMmNnOoÓóÔôÕõPpQqRrSsTtUuÚúVvXxZz", Brai = c.braille, c.punc }, } m["qu"] = { "Quechua", 5218, "qwe", "Latn", } m["rm"] = { "Romansh", 13199, "roa-rhe", ancestors = "rm-old", "Latn", sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer .. c.small_e}, } m["ro"] = { "Romanian", 7913, "roa-eas", "Latn, Cyrl, Cyrs", translit = { Cyrl = "ro-translit" }, display_text = { Latn = { from = {"'", "ţ", "Ţ"}, to = {"’", "ț", "Ț"} }, }, strip_diacritics = { Latn = { from = {"’", "ţ", "Ţ"}, to = {"'", "ț", "Ț"} }, }, sort_key = { Latn = { remove_diacritics = c.grave .. c.acute, from = {"ă", "â", "î", "ș", "[țţ]"}, to = {"a" .. p[1], "a" .. p[2], "i" .. p[1], "s" .. p[1], "t" .. p[1]} }, Cyrl = { from = {"ӂ"}, to = {"ж" .. p[1]} }, }, -- Cyrs strip_diacritics, sort_key in [[Module:scripts/data]]; presumably not present standard_chars = { Latn = "AaĂăÂâBbCcDdEeFfGgHhIiÎîJjLlMmNnOoPpRrSsȘșTtȚțUuVvXxZz", Cyrl = "АаБбВвГгДдЕеЖжӁӂЗзИиЙйКкЛлМмНнОоПпРрСсТтУуФфХхЦцЧчШшЫыЬьЭэЮюЯя", c.punc }, } m["ru"] = { "Russian", 7737, "zle", "Cyrl, Brai", ancestors = "zle-mru", translit = { Cyrl = "ru-translit" }, display_text = { Cyrl = { from = {"'"}, to = {"’"} }, }, strip_diacritics = { Cyrl = { remove_diacritics = c.grave .. c.acute .. c.diaer, remove_exceptions = {"Ё", "ё", "Ѣ̈", "ѣ̈", "Я̈", "я̈"}, from = {"’"}, to = {"'"}, }, }, sort_key = { Cyrl = { remove_diacritics = c.grave .. c.acute .. c.diaer, from = { "і", "ѣ", "ѳ", "ѵ" }, to = { "и" .. p[1], "ь" .. p[1], "я" .. p[2], "я" .. p[3] } }, }, standard_chars = { Cyrl = "АаБбВвГгДдЕеЁёЖжЗзИиЙйКкЛлМмНнОоПпРрСсТтУуФфХхЦцЧчШшЩщЪъЫыЬьЭэЮюЯя—", Brai = c.braille, (c.punc:gsub("'", "")) -- Exclude apostrophe. }, } m["rw"] = { "Rwanda-Rundi", 3217514, "bnt-glb", "Latn", strip_diacritics = {remove_diacritics = c.acute .. c.circ .. c.macron .. c.caron}, } m["sa"] = { "Sanskrit", 11059, "inc", "as-Beng, Bali, Beng, Bhks, Brah, Mymr, xwo-Mong, Deva, Gujr, Guru, Gran, Hani, Java, Kthi, Knda, Kawi, Khar, Khmr, Laoo, Mlym, mnc-Mong, Marc, Modi, Mong, Nand, Newa, Orya, Phag, Ranj, Saur, Shrd, Sidd, Sinh, Soyo, Lana, Takr, Taml, Tang, Telu, Thai, Tibt, Tutg, Tirh, Zanb", --and also Khom; script codes sorted by canonical name rather than code for [[MOD:sa-convert]] translit = { Beng = "sa-Beng-translit", ["as-Beng"] = "sa-Beng-translit", -- Brah translit in [[Module:scripts/data]] Deva = "sa-translit", Gujr = "sa-Gujr-translit", Guru = "sa-Guru-translit", Java = "sa-Java-translit", Kthi = "sa-Kthi-translit", Khmr = "pi-translit", Knda = "sa-Knda-translit", Lana = "pi-translit", Laoo = "pi-translit", Mlym = "sa-Mlym-translit", Modi = "sa-Modi-translit", -- Mong, mnc-Mong, xwo-Mong translit in [[Module:scripts/data]] -- NOTE: Formerly used xal-translit for transliterating xwo-Mong but that only handles Cyrillic; it has -- code to transliterate xwo-Mong but it's broken so I've replaced it with the default xwo-translit. Mymr = "pi-translit", Orya = "sa-Orya-translit", -- Shrd translit in [[Module:scripts/data]] -- Sidd translit in [[Module:scripts/data]] Sinh = "si-translit", Taml = "sa-Taml-translit", Telu = "sa-Telu-translit", Thai = "pi-translit", -- Tibt translit in [[Module:scripts/data]] }, -- Mong display_text and strip_diacritics in [[Module:scripts/data]] -- Tibt display_text, strip_diacritics, sort_key in [[Module:scripts/data]] strip_diacritics = { Deva = s["sa-Deva-stripdiacritics"], Mymr = { remove_diacritics = c.VS01, }, Thai = { from = {"ึ", u(0xF700), u(0xF70F)}, -- FIXME: Not clear what's going on with the PUA characters here. to = {"ิํ", "ฐ", "ญ"} }, }, sort_key = { Deva = s["sa-Deva-stripdiacritics"], -- until we have a proper Sanskrit sorting algorithm. Lana = { -- Tai Tham from = {"ᩔ", "ᩕ", "ᩖ", "ᩘ", "([ᨭ-ᨱ])ᩛ", "([ᨷ-ᨾ])ᩛ", "ᩤ"}, to = {"ᩈ᩠ᩈ", "᩠ᩁ", "᩠ᩃ", "ᨦ᩠", "%1᩠ᨮ", "%1᩠ᨻ", "ᩣ"}, }, Laoo = "Laoo-sortkey", Latn = { from = {"ā", "ī", "ū", "ḍ", "ḷ", "ḹ", "m[" .. c.dotabove .. c.dotbelow .. "]", "ṅ", "ñ", "ṇ", "ṛ", "ṝ", "ś", "ṣ", "ṭ"}, to = {"a~", "i~", "u~", "d~", "l~", "l~~", "m~", "n~", "n~~", "n~~~", "r~", "r~~", "s~", "s~~", "t~"}, }, Mymr = { remove_diacritics = c.VS01, }, Thai = "Thai-sortkey", -- FIXME: The previous sort key which mixed all scripts removed ZWJ; I don't know which script(s) this was -- intended for and there are no other languages which remove it in the sort key AFAIK. If it needs to be -- removed, specify the script(s) it needs to be removed under or add handling for the "all" script that applies -- regardless of script. --all = { -- remove_diacritics = c.ZWJ, --}, }, } m["sc"] = { "Sardinian", 33976, "roa-sou", "Latn", ancestors = "sc-old", } m["sd"] = { "Sindhi", 33997, "inc-snd", "Arab, Deva, Sind, Khoj", translit = { Sind = "Sind-translit", Arab = "sd-Arab-translit" }, strip_diacritics = { Arab = { remove_diacritics = c.kashida .. c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.superalef, from = {"ٱ"}, to = {"ا"} }, }, } m["se"] = { "Northern Sami", 33947, "smi", "Latn", display_text = { from = {"'"}, to = {"ˈ"} }, strip_diacritics = {remove_diacritics = c.macron .. c.dotbelow .. "'ˈ"}, sort_key = { from = {"á", "č", "đ", "ŋ", "š", "ŧ", "ž"}, to = {"a" .. p[1], "c" .. p[1], "d" .. p[1], "n" .. p[1], "s" .. p[1], "t" .. p[1], "z" .. p[1]} }, standard_chars = "AaÁáBbCcČčDdĐđEeFfGgHhIiJjKkLlMmNnŊŋOoPpRrSsŠšTtŦŧUuVvZzŽž" .. c.punc, } m["sg"] = { "Sango", 33954, "crp", "Latn", ancestors = "ngb", } m["sh"] = { "Serbo-Croatian", 9301, "zls", "Latn, Cyrl, Glag, Arab", ietf_subtag = "hbs", -- ISO 639-3 code, since "sh" is deprecated from ISO 639-1 wikimedia_codes = "sh, bs, hr, sr", strip_diacritics = { Latn = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {"Ć", "ć", "Ś", "ś", "Ź", "ź"} }, Cyrl = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {"З́", "з́", "С́", "с́"} }, }, sort_key = { Latn = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {"ć", "ś", "ź"}, from = {"č", "ć", "dž", "đ", "lj", "nj", "š", "ś", "ž", "ź"}, to = {"c" .. p[1], "c" .. p[2], "d" .. p[1], "d" .. p[2], "l" .. p[1], "n" .. p[1], "s" .. p[1], "s" .. p[2], "z" .. p[1], "z" .. p[2]} }, Cyrl = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {"з́", "с́"}, from = {"ђ", "з́", "ј", "љ", "њ", "с́", "ћ", "џ"}, to = {"д" .. p[1], "з" .. p[1], "и" .. p[1], "л" .. p[1], "н" .. p[1], "с" .. p[1], "т" .. p[1], "ч" .. p[1]} }, }, standard_chars = { Latn = "AaBbCcČčĆćDdĐđEeFfGgHhIiJjKkLlMmNnOoPpRrSsŠšTtUuVvZzŽž", Cyrl = "АаБбВвГгДдЂђЕеЖжЗзИиЈјКкЛлЉљМмНнЊњОоПпРрСсТтЋћУуФфХхЦцЧчЏџШш", c.punc }, } m["si"] = { "Sinhalese", 13267, "inc-ins", "Sinh", translit = "si-translit", override_translit = true, } m["sk"] = { "Slovak", 9058, "zlw", "Latn", ancestors = "zlw-osk", sort_key = {remove_diacritics = c.acute .. c.circ .. c.diaer .. c.caron}, standard_chars = "AaÁáÄäBbCcČčDdĎďEeÉéFfGgHhIiÍíJjKkLlĹ弾MmNnŇňOoÓóÔôPpRrŔŕSsŠšTtŤťUuÚúVvYyÝýZzŽž" .. c.punc, } m["sl"] = { "Slovene", 9063, "zls", "Latn", strip_diacritics = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.dgrave .. c.invbreve .. c.dotbelow, remove_exceptions = {"Ć", "ć", "Ǵ", "ǵ", "Ś", "ś", "Ź", "ź"}, from = {"Ə", "ə", "Ł", "ł"}, to = {"E", "e", "L", "l"}, }, sort_key = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.dotabove .. c.ringabove .. c.dgrave .. c.invbreve .. c.dotbelow .. c.ringbelow .. c.ogonek, remove_exceptions = {"ć", "ǵ", "ś", "ź"}, from = {"ä", "č", "ć", "đ", "ə", "ë", "ǧ", "ǵ", "ï", "ł", "ö", "š", "ś", "ü", "ž", "ź"}, to = {"a" .. p[1], "c" .. p[1], "c" .. p[2], "d" .. p[1], "e", "e" .. p[1], "g" .. p[1], "g" .. p[2], "i" .. p[1], "l", "o" .. p[1], "s" .. p[1], "s" .. p[2], "u" .. p[1], "z" .. p[1], "z" .. p[2]}, }, standard_chars = "AaBbCcČčDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsŠšTtUuVvZzŽž" .. c.punc, } m["sm"] = { "Samoan", 34011, "poz-pnp", "Latn", } m["sn"] = { "Shona", 34004, "bnt-sho", "Latn", strip_diacritics = {remove_diacritics = c.acute}, } m["so"] = { "Somali", 13275, "cus-som", "Latn, Arab, Osma", strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ} }, } m["sq"] = { "Albanian", 8748, "sqj", "Latn, Grek, Arab, Elba, Todr, Vith", translit = { Elba = "Elba-translit", Vith = "Vith-translit", }, -- Grek display_text, strip_diacritics, sort_key in [[Module:scripts/data]] strip_diacritics = { Latn = { remove_diacritics = c.acute .. c.circ .. c.macron, from = {'^[ie] (%w)', '^të (%w)'}, to = {'%1', '%1'}, }, }, sort_key = { Latn = { remove_diacritics = c.acute .. c.circ .. c.macron .. c.tilde .. c.breve .. c.caron, from = {'^[ie] (%w)', '^të (%w)', 'ç', 'dh', 'ë', 'gj', 'll', 'nj', 'rr', 'sh', 'th', 'xh', 'zh'}, to = {'%1', '%1', 'c'..p[1], 'd'..p[1], 'e'..p[1], 'g'..p[1], 'l'..p[1], 'n'..p[1], 'r'..p[1], 's'..p[1], 't'..p[1], 'x'..p[1], 'z'..p[1]}, } -- TODO: Grek if the default sort key is unsuitable }, standard_chars = { Latn = "AaBbCcÇçDdEeËëFfGgHhIiJjKkLlMmNnOoPpQqRrSsTtUuVvXxYyZz", c.punc }, } m["ss"] = { "Swazi", 34014, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } m["st"] = { "Sotho", 34340, "bnt-sts", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } m["su"] = { "Sundanese", 34002, "poz-msa", "Latn, Sund, Arab", ancestors = "osn", translit = { Sund = "Sund-translit" }, } m["sv"] = { "Swedish", 9027, "gmq-eas", "Latn", ancestors = "gmq-osw-lat", sort_key = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron .. c.dacute .. c.caron .. c.cedilla .. "':", remove_exceptions = {"å"}, from = {"ø", "æ", "œ", "ß", "ꜩ", "å", "aͤ", "oͤ"}, to = {"ö", "ae", "oe", "ss", "tz", "z" .. p[1], "ä", "ö"} }, standard_chars = "AaBbCcDdEeFfGgHhIiJjKkLlMmNnOoPpRrSsTtUuVvXxYyÅåÄäÖö" .. c.punc, } m["sw"] = { "Swahili", 7838, "bnt-swh", "Latn, Arab", sort_key = { Latn = { from = {"ng'"}, to = {"ng" .. p[1]} }, }, } m["ta"] = { "Tamil", 5885, "dra-tam", "Taml", ancestors = "ta-mid", translit = "ta-translit", override_translit = true, } m["te"] = { "Telugu", 8097, "dra-tel", "Telu", translit = "te-translit", override_translit = true, } m["tg"] = { "Tajik", 9260, "ira-swi", "Cyrl, Arab, Latn", ancestors = "fa-cls", translit = { Cyrl = "tg-translit" }, override_translit = true, strip_diacritics = { Cyrl = s["tg-stripdiacritics"], Latn = s["tg-stripdiacritics"], }, sort_key = { Cyrl = { from = {"ғ", "ё", "ӣ", "қ", "ӯ", "ҳ", "ҷ"}, to = {"г" .. p[1], "е" .. p[1], "и" .. p[1], "к" .. p[1], "у" .. p[1], "х" .. p[1], "ч" .. p[1]} }, }, } m["th"] = { "Thai", 9217, "tai-swe", "Thai, Khomt, Brai", translit = { Thai = "th-translit" }, sort_key = { Thai = "Thai-sortkey" }, } m["ti"] = { "Tigrinya", 34124, "sem-eth", "Ethi", translit = "Ethi-translit", } m["tk"] = { "Turkmen", 9267, "trk-ogz", "Latn, Cyrl, Arab", strip_diacritics = { Latn = s["tk-stripdiacritics"], Cyrl = s["tk-stripdiacritics"], }, sort_key = { Latn = { from = {"ç", "ä", "ž", "ň", "ö", "ş", "ü", "ý"}, to = {"c" .. p[1], "e" .. p[1], "j" .. p[1], "n" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1], "y" .. p[1]} }, Cyrl = { from = {"ё", "җ", "ң", "ө", "ү", "ә"}, to = {"е" .. p[1], "ж" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1], "э" .. p[1]} }, }, ancestors = "trk-eog", } m["tl"] = { "Tagalog", 34057, "phi", "Latn, Tglg", translit = { Tglg = "tl-translit" }, override_translit = true, strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ} }, standard_chars = { Latn = "AaBbKkDdEeGgHhIiLlMmNnOoPpRrSsTtUuWwYy", c.punc }, sort_key = { Latn = "tl-sortkey", }, } m["tn"] = { "Tswana", 34137, "bnt-sts", "Latn", } m["to"] = { "Tongan", 34094, "poz-ton", "Latn", strip_diacritics = {remove_diacritics = c.acute}, sort_key = {remove_diacritics = c.macron}, } m["tr"] = { "Turkish", 256, "trk-ogz", "Latn", ancestors = "ota", dotted_dotless_i = true, sort_key = { from = { -- Ignore circumflex, but account for capital Î wrongly becoming ı + circ due to dotted dotless I logic. "ı" .. c.circ, c.circ, "i", -- Ensure "i" comes after "ı". "ç", "ğ", "ı", "ö", "ş", "ü" }, to = { "i", "", "i" .. p[1], "c" .. p[1], "g" .. p[1], "i", "o" .. p[1], "s" .. p[1], "u" .. p[1] } }, standard_chars = "AaÂâBbCcÇçDdEeFfGgĞğHhIıİiÎîJjKkLlMmNnOoÖöPpRrSsŞşTtUuÛûÜüVvYyZz" .. c.punc, } m["ts"] = { "Tsonga", 34327, "bnt-tsr", "Latn", } m["tt"] = { "Tatar", 25285, "trk-kbu", "Cyrl, Latn, Arab", translit = { Cyrl = "tt-translit", Arab = "tt-translit" }, --override_translit = true, -- enable override until Module code can detect Russian loans such as [[аэропорт]] dotted_dotless_i = true, sort_key = { Cyrl = { from = {"ә", "ў", "ғ", "ё", "җ", "қ", "ң", "ө", "ү", "һ"}, to = {"а" .. p[1], "в" .. p[1], "г" .. p[1], "е" .. p[1], "ж" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1], "х" .. p[1]} }, Latn = { from = { "i", -- Ensure "i" comes after "ı". "ä", "ə", "ç", "ğ", "ı", "ñ", "ŋ", "ö", "ɵ", "ş", "ü" }, to = { "i" .. p[1], "a" .. p[1], "a" .. p[2], "c" .. p[1], "g" .. p[1], "i", "n" .. p[1], "n" .. p[2], "o" .. p[1], "o" .. p[2], "s" .. p[1], "u" .. p[1] } }, }, } -- "tw" is treated as "ak", see [[WT:LT]] m["ty"] = { "Tahitian", 34128, "poz-pep", "Latn", } m["ug"] = { "Uyghur", 13263, "trk-kar", "Arab, Latn, Cyrl", ancestors = "chg", translit = { Arab = "ug-translit", Cyrl = "ug-translit", }, override_translit = true, } m["uk"] = { "Ukrainian", 8798, "zle", "Cyrl", ancestors = "zle-muk", translit = "uk-translit", strip_diacritics = {remove_diacritics = c.grave .. c.acute}, sort_key = { remove_diacritics = c.grave .. c.acute, from = { "ї", -- 2 chars "ґ", "є", "і" -- 1 char }, to = { "и" .. p[2], "г" .. p[1], "е" .. p[1], "и" .. p[1] } }, standard_chars = "АаБбВвГгДдЕеЄєЖжЗзИиІіЇїЙйКкЛлМмНнОоПпРрСсТтУуФфХхЦцЧчШшЩщЬьЮюЯя" .. c.punc:gsub("'", ""), -- Exclude apostrophe. } m["ur"] = { "Urdu", 1617, "inc-hnd", "Aran, Hebr", translit = { Aran = "ur-translit" }, strip_diacritics = { Aran = { -- character "ۂ" code U+06C2 to "ه" and "هٔ" (U+0647 + U+0654) to "ه"; hamzatu l-waṣli to a regular alif from = {"هٔ", "ۂ", "ٱ"}, to = {"ہ", "ہ", "ا"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] standard_chars = { Aran = "ایببپتثجچحخدذرزژسشصضطظعغفقکگلࣇڷمنݨوؤہھئٹڈڑآے", c.punc, }, } m["uz"] = { "Uzbek", 9264, "trk-kar", "Latn, Cyrl, Arab", ancestors = "chg", translit = { Cyrl = "uz-translit" }, sort_key = { Latn = { from = {"oʻ", "gʻ", "sh", "ch", "ng"}, to = {"z" .. p[1], "z" .. p[2], "z" .. p[3], "z" .. p[4], "z" .. p[5]} }, Cyrl = { from = {"ё", "ў", "қ", "ғ", "ҳ"}, to = {"е" .. p[1], "я" .. p[1], "я" .. p[2], "я" .. p[3], "я" .. p[4]} }, }, strip_diacritics = { Arab = "ar-stripdiacritics", }, } m["ve"] = { "Venda", 32704, "bnt-bso", "Latn", } m["vi"] = { "Vietnamese", 9199, "mkh-vie", "Latn, Hani", ancestors = "mkh-mvi", sort_key = { Latn = "vi-sortkey", Hani = "Hani-sortkey", }, } m["vo"] = { "Volapük", 36986, "art", "Latn", } m["wa"] = { "Walloon", 34219, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["wo"] = { "Wolof", 34257, "alv-fwo", "Latn, Arab, Gara", } m["xh"] = { "Xhosa", 13218, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } m["yi"] = { "Yiddish", 8641, "gmw-hgm", "Hebr, Latn", ancestors = "gmh", translit = { Hebr = "yi-translit", }, -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["yo"] = { "Yoruba", 34311, "alv-yor", "Latn, Arab", strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.macron} }, sort_key = { Latn = { from = {"ẹ", "ɛ", "gb", "ị", "kp", "ọ", "ɔ", "ṣ", "sh", "ụ"}, to = {"e" .. p[1], "e" .. p[1], "g" .. p[1], "i" .. p[1], "k" .. p[1], "o" .. p[1], "o" .. p[1], "s" .. p[1], "s" .. p[1], "u" .. p[1]} }, }, } m["za"] = { "Zhuang", 13216, "tai", "Latn, Hani", sort_key = { Latn = "za-sortkey", Hani = "Hani-sortkey", }, } m["zh"] = { "Chinese", 7850, "zhx", "Hants, Latn, Bopo, Nshu, Brai", ancestors = "ltc", generate_alternants = "zh-generatealternants", translit = { Hani = "zh-translit", Bopo = "zh-translit", }, sort_key = { Hani = "Hani-sortkey" }, } m["zu"] = { "Zulu", 10179, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.macron .. c.caron}, } return require("Module:languages").finalizeData(m, "language") l27lbmo6rtou10i7smdrsffh2aw86gg Module:romance utilities 828 8051 54855 2026-09-26T23:23:01Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local languages_module = "Module:languages" local parse_utilities_module = "Module:parse utilities" local string_pattern_escape_module = "Module:string/patternEscape" local string_replacement_escape_module = "Module:string/replacementEscape" local string_utilities_module = "Module:string utilities" local concat = table.concat local insert = table.insert local ipairs = ipairs local pairs = pairs local require = require local sort = table.sort l...' 54855 Scribunto text/plain local export = {} local languages_module = "Module:languages" local parse_utilities_module = "Module:parse utilities" local string_pattern_escape_module = "Module:string/patternEscape" local string_replacement_escape_module = "Module:string/replacementEscape" local string_utilities_module = "Module:string utilities" local concat = table.concat local insert = table.insert local ipairs = ipairs local pairs = pairs local require = require local sort = table.sort local type = type local function escape_wikicode(...) escape_wikicode = require(parse_utilities_module).escape_wikicode return escape_wikicode(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function pattern_escape(...) pattern_escape = require(string_pattern_escape_module) return pattern_escape(...) end local function replacement_escape(...) replacement_escape = require(string_replacement_escape_module) return replacement_escape(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function ugsub(...) ugsub = require(string_utilities_module).gsub return ugsub(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end export.allowed_special_indicators = { ["first"] = true, ["first-second"] = true, ["first-last"] = true, ["second"] = true, ["last"] = true, ["each"] = true, ["+"] = true, -- requests the default behavior with preposition handling } --[==[ Check for special indicators (values such as {"+first"} or {"+first-last"} that are used in a `pl`, `f`, etc. argument and indicate how to inflect a multiword term). If `form` is such an indicator, the return value is `form` minus the initial `+` sign; otherwise, if form begins with a `+` sign, an error is thrown; otherwise the return value is nil. ]==] function export.get_special_indicator(form) if form:find("^%+") then form = form:gsub("^%+", "") if not export.allowed_special_indicators[form] then local indicators = {} for indic, _ in pairs(export.allowed_special_indicators) do insert(indicators, "+" .. indic) end sort(indicators) error("Special inflection indicator beginning with '+' can only be " .. mw.text.listToText(indicators) .. ": +" .. form) end return form end return nil end local function add_endings(bases, endings) local retval = {} if type(bases) ~= "table" then bases = {bases} end if type(endings) ~= "table" then endings = {endings} end for _, base in ipairs(bases) do for _, ending in ipairs(endings) do insert(retval, base .. ending) end end return retval end --[==[ Inflect a possibly multiword or hyphenated term `form` using the function `inflect`, which is a function of one argument that is called on a single word to inflect and should return either the inflected word or a list of inflected words. `special` indicates how to inflect the multiword term and should be e.g. {"first"| to inflect only the first word, {"first-last"} to inflect the first and last words, {"each"} to inflect each word, etc. See `allowed_special_indicators` above for the possibilities. If `special` is `+`, or is omitted and the term is multiword (i.e. containing a space character), the function checks for multiword or hyphenated terms containing the prepositions in `prepositions`, e.g. Italian [[senso di marcia]] or [[medaglia d'oro]] or Portuguese [[tartaruga-do-mar]]. If such a term is found, only the first word is inflected. Otherwise, the default is {"first-last"}. `prepositions` is a list of regular expressions matching prepositions. The regular expressions will automatically have the separator character (space or hyphen) added to the left side but not the right side, so they should contain a space character (which will automatically be converted to the appropriate separator) on the right side unless the preposition is joined on the right side with an apostrophe. Examples of preposition regular expressions for Italian are {"di "}, {"sull'"} and {"d?all[oae] "} (which matches {"dallo "}, {"dalle "}, {"alla "}, etc.). The return value is always either a list of inflected multiword or hyphenated terms, or nil if `special` is omitted and `form` is not multiword. (If `special` is specified and `form` is not multiword or hyphenated, an error results.) ]==] function export.handle_multiword(form, special, inflect, prepositions, sep) sep = sep or form:find(" ") and " " or "%-" local raw_sep = sep == " " and " " or "-" -- Used to add regex version of separator in the replacement portion of ugsub() or :gsub() local sep_replacement = sep == " " and " " or "%%-" -- Given a Lua pattern, replace space with the appropriate separator. local function hack_re(re) if sep == " " then return re end return (re:gsub(" ", sep_replacement)) end if special == "first" then local first, rest = form:match(hack_re("^(.-)( .*)$")) if not first then error("Special indicator 'first' can only be used with a multiword term: " .. form) end return add_endings(inflect(first), rest) elseif special == "second" then local first, second, rest = form:match(hack_re("^([^ ]+ )([^ ]+)( .*)$")) if not first then error("Special indicator 'second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings({first}, inflect(second)), rest) elseif special == "first-second" then local first, space, second, rest = form:match(hack_re("^([^ ]+)( )([^ ]+)( .*)$")) if not first then error("Special indicator 'first-second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings(add_endings(inflect(first), space), inflect(second)), rest) elseif special == "each" then local terms = split(form, sep) if #terms < 2 then error("Special indicator 'each' can only be used with a multiword term: " .. form) end for i, term in ipairs(terms) do terms[i] = inflect(term) if i > 1 then terms[i] = add_endings(raw_sep, terms[i]) end end local result = "" for _, term in ipairs(terms) do result = add_endings(result, term) end return result elseif special == "first-last" then local first, middle, last = form:match(hack_re("^(.-)( .* )(.-)$")) if not first then first, middle, last = form:match(hack_re("^(.-)( )(.*)$")) end if not first then error("Special indicator 'first-last' can only be used with a multiword term: " .. form) end return add_endings(add_endings(inflect(first), middle), inflect(last)) elseif special == "last" then local rest, last = form:match(hack_re("^(.* )(.-)$")) if not rest then error("Special indicator 'last' can only be used with a multiword term: " .. form) end return add_endings(rest, inflect(last)) elseif special and special ~= "+" then error("Unrecognized special=" .. special) end -- Only do default behavior if special indicator '+' explicitly given or separator is space; otherwise we will -- break existing behavior with hyphenated words. if (special == "+" or sep == " ") and form:find(sep) then -- check for prepositions in the middle of the word; do it this way so we can handle -- more than one word before the preposition (and usually inflect each word) for _, prep in ipairs(prepositions) do local first, space_prep_rest = umatch(form, hack_re("^(.-)( " .. prep .. ".*)$")) if first then return add_endings(inflect(first), space_prep_rest) end end -- multiword or hyphenated expressions default to first-last; we need to pass in the separator to avoid -- problems with multiword terms containing hyphens in the individual words return export.handle_multiword(form, "first-last", inflect, prepositions, sep) end return nil end -- Auto-add links to a word that should not have spaces but may have hyphens and/or apostrophes. We split off final -- punctuation, then split on hyphens if `splithyph` is given, and also split on apostrophes. We only split on hyphens -- and apostrophes if they are in the middle of the word, not at the beginning of end (hyphens at the beginning or end -- indicate suffixes or prefixes, respectively, and apostrophes at the beginning or end are also possible, as in -- Italian [['ndrangheta]] or [[po']]). The apostrophe is included in the link to its left (so we auto-split French -- [[l'eau]] as [[l']][[eau]]). See `add_links_to_multiword_term()` for the explanation of `no_split_apostrophe_words` -- and `include_hyphen_prefixes`. local function add_single_word_links(space_word, splithyph, no_split_apostrophe_words, include_hyphen_prefixes) local space_word_no_punct, punct = space_word:match("^(.*)([,;:?!])$") space_word_no_punct = space_word_no_punct or space_word punct = punct or "" local words -- don't split prefixes and suffixes if not splithyph or space_word_no_punct:find("^%-") or space_word_no_punct:find("%-$") then words = {space_word_no_punct} else words = split(space_word_no_punct, "%-") end local linked_words = {} for j, word in ipairs(words) do if j < #words and include_hyphen_prefixes and include_hyphen_prefixes[word] then word = "[[" .. word .. "-]]" else -- Don't split on apostrophes if the word is in `no_split_apostrophe_words` or begins or ends with an apostrophe -- (e.g. [['ndrangheta]] or [[po']]). Handle multiple apostrophes correctly, e.g. [[l'altr'ieri]]. if (not no_split_apostrophe_words or not no_split_apostrophe_words[word]) and word:find("'") and not word:find("^'") and not word:find("'$") then local apostrophe_parts = split(word, "'") for i, apostrophe_part in ipairs(apostrophe_parts) do if i == #apostrophe_parts then apostrophe_parts[i] = "[[" .. apostrophe_part .. "]]" else apostrophe_parts[i] = "[[" .. apostrophe_part .. "']]" end end word = concat(apostrophe_parts) else word = "[[" .. word .. "]]" end if j < #words then word = word .. "-" end end insert(linked_words, word) end return concat(linked_words) .. punct end --[==[ Auto-add links to a multiword term. Links are not added to single-word terms. We split on spaces, and also on hyphens if `splithyph` is given or the word has no spaces. In addition, we split on apostrophes, including the apostrophe in the link to its left (so we auto-split {"de l'eau"} {"[[de]] [[l']][[eau]]"}). We don't always split on hyphens because of cases like {"boire du petit-lait"} where {"petit-lait"} should be linked as a whole, but provide the option to do it for cases like {"croyez-le ou non"}. If there's no space, however, then it makes sense to split on hyphens by default (e.g. for {"avant-avant-hier"}). Cases where only some of the hyphens should be split can always be handled by explicitly specifying the head (e.g. {"Nord-Pas-de-Calais"} given as `head=[[Nord]]-[[Pas-de-Calais]]`). `no_split_apostrophe_words` and `include_hyphen_prefixes` allow for special-case handling of particular words and are as described in the comment above `add_single_word_links()`. `no_split_apostrophe_words`, if given, is a set of words that contain apostrophes but which should not be split on the apostrophes, such as French [[c'est]] and [[quelqu'un]]. `include_hyphen_prefixes`, if given, is a set of prefixes (not including the final hyphen) where we should include the final hyphen in the prefix. Hence, e.g. if {"anti"} is in the set, a Portuguese word like [[anti-herói]] "anti-hero" will be split [[anti-]][[herói]] (whereas a word like [[código-fonte]] "source code" will be split as [[código]]-[[fonte]]). ]==] function export.add_links_to_multiword_term(term, splithyph, no_split_apostrophe_words, include_hyphen_prefixes) if not term:find(" ", nil, true) then splithyph = true end local words = split(term, " ") local linked_words = {} for _, word in ipairs(words) do insert(linked_words, add_single_word_links(word, splithyph, no_split_apostrophe_words, include_hyphen_prefixes)) end local retval = concat(linked_words, " ") -- If we ended up with a single link consisting of the entire term, -- remove the link. return retval:match("^%[%[([^%[%]]*)%]%]$") or retval end --[==[ Given a `linked_term` that is the output of add_links_to_multiword_term(), apply modifications as given in `modifier_spec` to change the link destination of subterms (normally single-word non-lemma forms; sometimes collections of adjacent words). This is usually used to link non-lemma forms to their corresponding lemma, but can also be used to replace a span of adjacent separately-linked words to a single multiword lemma. The format of `modifier_spec` is one or more semicolon-separated subterm specs, where each such spec is of the form `SUBTERM:DEST`, where `SUBTERM` is one or more words in the `linked_term` but without brackets in them, and `DEST` is the corresponding link destination to link the subterm to. Any occurrence of `~` in `DEST` is replaced with `SUBTERM`. Alternatively, a single modifier spec can be of the form `BEGIN[FROM:TO]`, which is equivalent to writing `BEGINFROM:BEGINTO` (see example below). For example, given the source phrase [[il bue che dice cornuto all'asino]] "the pot calling the kettle black" (literally "the ox that calls the donkey horned/cuckolded"), the result of calling `add_links_to_multiword_term()` is [[il]] [[bue]] [[che]] [[dice]] [[cornuto]] [[all']][[asino]]. With a modifier_spec of `dice:dire`, the result is [[il]] [[bue]] [[che]] [[dire|dice]] [[cornuto]] [[all']][[asino]]. Here, based on the modifier spec, the non-lemma form [[dice]] is replaced with the two-part link [[dire|dice]]. Another example: given the source phrase [[chi semina vento raccoglie tempesta]] "sow the wind, reap the whirlwind" (literally "(he) who sows wind gathers [the] tempest"). The result of calling `add_links_to_multiword_term()` is [[chi]] [[semina]] [[vento]] [[raccoglie]] [[tempesta]], and with a modifier_spec of `semina:~re; raccoglie:~re`, the result is [[chi]] [[seminare|semina]] [[vento]] [[raccogliere|raccoglie]] [[tempesta]]. Here we use the `~` notation to stand for the non-lemma form in the destination link. A more complex example is [[se non hai altri moccoli puoi andare a letto al buio]], which becomes [[se]] [[non]] [[hai]] [[altri]] [[moccoli]] [[puoi]] [[andare]] [[a]] [[letto]] [[al]] [[buio]] after calling `add_links_to_multiword_term()`. With the following modifier_spec: `hai:avere; altr[i:o]; moccol[i:o]; puoi: potere; andare a letto:~; al buio:~`, the result of applying the spec is [[se]] [[non]] [[avere|hai]] [[altro|altri]] [[moccolo|moccoli]] [[potere|puoi]] [[andare a letto]] [[al buio]]. Here, we rely on the alternative notation mentioned above for e.g. `altr[i:o]`, which is equivalent to `altri:altro`, and link multiword subterms using e.g. `andare a letto:~`. (The code knows how to handle multiword subexpressions properly, and if the link text and destination are the same, only a single-part link is formed.) ]==] function export.apply_link_modifiers(linked_term, modifier_spec) local split_modspecs = split(modifier_spec, "%s*;%s*") for j, modspec in ipairs(split_modspecs) do local subterm, dest, otherlang local begin_from, begin_to, rest, end_from, end_to = modspec:match("^%[(.-):(.*)%]([^:]*)%[(.-):(.*)%]$") if begin_from then subterm = begin_from .. rest .. end_from dest = begin_to .. rest .. end_to end if not subterm then rest, end_from, end_to = modspec:match("^([^:]*)%[(.-):(.*)%]$") if rest then subterm = rest .. end_from dest = rest .. end_to end end if not subterm then begin_from, begin_to, rest = modspec:match("^%[(.-):(.*)%]([^:]*)$") if begin_from then subterm = begin_from .. rest dest = begin_to .. rest end end if not subterm then subterm, dest = modspec:match("^(.-)%s*:%s*(.*)$") if subterm and subterm ~= "^" and subterm ~= "$" then local langdest -- Parse off an initial language code (e.g. 'en:Higgs', 'la:minūtia' or 'grc:σκατός'). Also handle -- Wikipedia prefixes ('w:Abatemarco' or 'w:it:Colle Val d'Elsa'). otherlang, langdest = dest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if otherlang == "w" then local foreign_wikipedia, foreign_term = langdest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if foreign_wikipedia then otherlang = otherlang .. ":" .. foreign_wikipedia langdest = foreign_term end dest = ("%s:%s"):format(otherlang, langdest) otherlang = nil elseif otherlang then otherlang = get_lang(otherlang, true, "allow etym") dest = langdest end end end if not subterm then error(("Single modifier spec %s should be of the form SUBTERM:DEST where SUBTERM is one or more words " .. "a multiword term and DEST is the destination to link the subterm to (possibly prefixed by a " .. "language code); or of the form BEGIN[FROM:TO], which is equivalent to BEGINFROM:BEGINTO; or " .. "similarly [FROM:TO]END, which is equivalent to FROMEND:TOEND"):format(modspec)) end if subterm == "^" then linked_term = dest:gsub("_", " ") .. linked_term elseif subterm == "$" then linked_term = linked_term .. dest:gsub("_", " ") else if subterm:find("%[") then error(("Subterm '%s' in modifier spec '%s' cannot have brackets in it"):format( escape_wikicode(subterm), escape_wikicode(modspec))) end local escaped_subterm = pattern_escape(subterm) local subterm_re = "%[%[" .. escaped_subterm:gsub("(%%?[ '%-])", "%%]*%1%%[*") .. "%]%]" local expanded_dest if dest:find("~") then expanded_dest = dest:gsub("~", replacement_escape(subterm)) else expanded_dest = dest end if otherlang then expanded_dest = expanded_dest .. "#" .. otherlang:getCanonicalName() end local subterm_replacement if expanded_dest:find("%[") then -- Use the destination directly if it has brackets in it (e.g. to put brackets around parts of a word). subterm_replacement = expanded_dest elseif expanded_dest == subterm then subterm_replacement = "[[" .. subterm .. "]]" else subterm_replacement = "[[" .. expanded_dest .. "|" .. subterm .. "]]" end local replaced_linked_term = ugsub(linked_term, subterm_re, replacement_escape(subterm_replacement)) if replaced_linked_term == linked_term then error(("Subterm '%s' could not be located in %slinked expression %s, or replacement same as subterm" ):format(subterm, j > 1 and "intermediate " or "", escape_wikicode(linked_term))) else linked_term = replaced_linked_term end end end return linked_term end return export k9f9aae2ymvx65t4ayfp2c66bwmjy2p Module:en-utilities 828 8052 54856 2026-09-26T23:23:29Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local add_suffix -- Defined below. local find = string.find local is_regular_plural -- Defined below. local match = string.match local remove_possessive -- Defined below. local reverse = string.reverse local sub = string.sub local toNFD = mw.ustring.toNFD local ugsub = mw.ustring.gsub local ulower = mw.ustring.lower local umatch = mw.ustring.match local usub = mw.ustring.sub local uupper = mw.ustring.upper local vowels = "aæᴀᴁɐɑɒ@eᴇ...' 54856 Scribunto text/plain local export = {} local add_suffix -- Defined below. local find = string.find local is_regular_plural -- Defined below. local match = string.match local remove_possessive -- Defined below. local reverse = string.reverse local sub = string.sub local toNFD = mw.ustring.toNFD local ugsub = mw.ustring.gsub local ulower = mw.ustring.lower local umatch = mw.ustring.match local usub = mw.ustring.sub local uupper = mw.ustring.upper local vowels = "aæᴀᴁɐɑɒ@eᴇǝⱻəɛɘɜɞɤiıɪɨᵻoøœᴏɶɔᴐɵuᴜʉᵾɯꟺʊʋʌyʏ" local hyphens = "%-‐‑‒–—" --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local diacritics local function get_diacritics() diacritics, get_diacritics = mw.loadData("Module:headword/data").page.comb_chars.diacritics_all .. "+", nil return diacritics end -- Normalize a string, so that case and diacritics are ignored. By default, "gu" -- and "qu" are normalized to "g" and "q", because they behave like consonants -- under certain conditions (e.g. final "y" does not usually have the plural -- "ies" after a vowel, but it's regular for "quy" to become "quies". The flag -- `not_gu` prevents this happening to "gu", and is needed because terms ending -- "-guy" are almost always compounds of "guy" (→ "guys"). local function normalize(str, followed_by, not_gu) if not followed_by then followed_by = "" end str = ugsub(toNFD(str) .. followed_by, "([" .. (not_gu and "" or "Gg") .. "Qq])u([".. vowels .. "])", "%1%2") return ulower(ugsub(sub(str, 1, #str - #followed_by), diacritics or get_diacritics(), "")) end local function epenthetic_e_default(stem) return sub(stem, -1) ~= "e" end local function epenthetic_e_for_s(stem, term) -- If the stem is different, it must be from "y" → "i". if stem ~= term then return true end local final if match(stem, "^[^\128-\255]*$") then final = sub(stem, -1) else stem = ugsub(toNFD(stem), diacritics or get_diacritics(), "") final = usub(stem, -1) end -- Epenthetic "e" is added after a sibilant or sibilant-affricate. The vast -- majority of these are spelled "s", "x", "z", "ch" and "sh", but "dg" -- (→ "dge") and "ß" (→ "ss") can be found in obsolete spellings, "shh" in -- onomatopoeia, and "zh", "dj", "jj" (and more) in loanwords. return ( final == "g" and sub(stem, -2, -2) == "d" or final == "h" and match(stem, "[csz]h+$") or final == "j" and umatch(stem, "[^" .. vowels .. "]j$") or final == "s" or final == "u" and umatch(stem, "%f[%w']u$") or final == "x" or final == "z" or final == "ß" ) end function export.remove_possessive(stem) return match(stem, "^(.*)'s$") or match(stem, "^(.*s)'$") or stem end remove_possessive = export.remove_possessive local suffixes = {} suffixes["'s"] = { truncated = function(stem) return sub(stem, -1) == "s" and "'" or "'s" end, } suffixes["s.plural"] = { final_y_is_i = true, epenthetic_e = epenthetic_e_for_s, modifies_possessive = true, } suffixes["s.verb"] = { final_y_is_i = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_for_s } suffixes["ing"] = { final_consonant_is_doubled = true, remove_silent_e = true, } suffixes["d"] = { final_y_is_i = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_default, } suffixes["dst"] = suffixes["d"] suffixes["st.verb"] = suffixes["d"] suffixes["th"] = suffixes["d"] suffixes["n"] = { final_y_is_i = true, final_y_is_i_after_vowel = true, final_guy_is_gui = true, final_consonant_is_doubled = true, -- No epenthetic "e" after an "e", or an "i", "r" or "w" preceded by a vowel. epenthetic_e = function(stem) return not ( sub(stem, -1) == "e" or umatch(normalize(stem), "[" .. vowels .. "][irw]$") ) end, } suffixes["r"] = { final_y_is_i = true, final_ey_is_i = true, final_guy_is_gui = true, final_consonant_is_doubled = true, epenthetic_e = epenthetic_e_default } suffixes["st.superlative"] = suffixes["r"] -- Returns the stem used for suffixes that sometimes convert final "y" into "i", -- such as "-es" ("-ies"), e.g. "penny" → "penni" ("pennies"). If -- `final_ey_is_i` is true, final "ey" may also be converted, e.g. "plaguey" → -- "plagui"; this is needed for "-er" ("-ier") and "-est" ("-iest"). If `not_gu` -- is true, then normalize() will be called with the `not_gu` flag (see there -- for more info); this is true in most cases. local function convert_final_y_to_i(str, not_gu, final_ey_is_i, final_y_is_i_after_vowel) local final3 = usub(str, -3) -- Special case: treat "eey" as "ee" + "y" (e.g. "treey" → "treeiest"). -- "oey" and "uey" are usually vowel + "ey", but examples of "oe" + "y" and -- "ue" = "y" do also exist: compare "go" → "goey" → "goier" with "doe" → -- "doey" → "doeier"; "flu" → "fluey" → "fluiest" and "flue" → "fluey" → -- "flueiest" form a theoretically possible minimal pair. if final3 == "eey" then return sub(str, 1, -2) .. "i" end local final2 = usub(str, -2) -- If `final_ey_is_i` is true, treat final "-ey" can also be reduced. if final_ey_is_i and final2 == "ey" then -- Remove "ey" to get the base stem. local base_stem = sub(str, 1, -3) -- Special case: allow final "-ey" ("potato-ey" → "potato-iest"). if umatch(final3, "[" .. hyphens .. "]ey") then return base_stem .. "i" end -- Final "ey" becomes "i" iff the term is polysyllabic (e.g. not -- "grey"). "ey" is common if the base stem ends in a vowel ("echo → -- "echoey"), so the presence of a vowel anywhere in the base stem is -- sufficient to deem it polysyllabic. ("echoey" → "echo" → "echoiest", -- "beigey" → "beig" → "beigiest", but "grey" → "gr" → "greyest"). The -- first "y" in "-yey" can be treated as a vowel as long as it's -- preceded by something ("clayey" → "clay" → "clayiest", "cryey" → -- "cry" → "cryiest", but "*yey" → "*y" → "*yeyest"), so it needs to be -- treated as a special case. local normalized = normalize(base_stem, "ey") if sub(normalized, -1) == "y" then if umatch(normalized, "[%w@][yY]$") then return base_stem .. "i" end elseif umatch(normalized, "[" .. vowels .. "%d]%w*$") then return base_stem .. "i" end -- Special cases: -- Final "quy" ("soliloquy" → "soliloquies"). -- Final "guy" iff `not_gu` is false ("roguy" → "roguiest"). -- Final "y" after a vowel iff `final_y_is_i_after_vowel` is true ("slay" → -- "slain"). -- Final "-y" ("bro-y" → "bro-iest"), accounting for hyphen variation. elseif umatch(final2, "[" .. hyphens .. "]y") then -- Replace final "y" with "i". return sub(str, 1, -2) .. "i" -- Otherwise, final "y" becomes "i" iff it's not preceded by a vowel -- ("shy" → "shiest", "horsy" → "horsies", but "day" → "days", "coy" → -- "coyest"). else -- Remove "y" to get the base stem. local base_stem = sub(str, 1, -2) if umatch(normalize(base_stem, "y", not_gu), "[^%s%p" .. (final_y_is_i_after_vowel and "" or vowels) .. "]$") then return base_stem .. "i" end end return str end local function double_final_consonant(str, final) local initial = umatch(normalize(sub(str, 1, -2), final), "^.*%f[^%z%s" .. hyphens .. "…]([%l%p]*)[" .. vowels .. "]$") return initial and ( initial == "" or initial == "y" or match(initial, "^.[\128-\191]*$") and umatch(initial, "[^" .. vowels .. "]") or umatch(initial, "^[^" .. vowels .. "]*%f[^%l]$") ) and (str .. final) or str end local function remove_silent_e(str) local final2 = sub(str, -2) if final2 == "ie" then -- Replace "ie" with "y", unless it follows another "y" (e.g. -- "spulyie" → "spulyieing"). return ugsub(str, "([^yY%s%p])ie$", "%1y") end local base_stem = sub(str, 1, -2) -- Silent "e" occurs after "u" or a consonant (cluster) preceded by a vowel. return ( final2 == "ue" or umatch(normalize(base_stem, "e"), "[" .. vowels .. "][^" .. vowels .. "]+$") ) and base_stem or str end function export.add_suffix(term, suffix, pos) local data, possessive = suffixes[suffix] -- If modifies_possessive is set, check for and remove any possessive -- suffix, which will be re-added again at the end. if data.modifies_possessive then local new = remove_possessive(term) if new ~= term then term, possessive = new, true end end suffix = match(suffix, "^([^.]*)") local final, stem = sub(term, -1) -- Proper nouns don't have a final "y" changed to "i" (e.g. "the Gettys", -- "the public Ivys"). if data.final_y_is_i and final == "y" and pos ~= "proper noun" then stem = convert_final_y_to_i(term, not data.final_guy_is_gui, data.final_ey_is_i, data.final_y_is_i_after_vowel) elseif data.remove_silent_e and final == "e" then stem = remove_silent_e(term) else stem = term end local epenthetic_e = data.epenthetic_e if epenthetic_e and epenthetic_e(stem, term) then suffix = "e" .. suffix end if ( data.final_consonant_is_doubled and match(final, "^[bcdfgjklmnpqrstvz]$") and -- Only double regular consonants. umatch(suffix, "^[" .. vowels .. "]") ) then stem = double_final_consonant(term, final) end local truncated = data.truncated if truncated then suffix = truncated(stem) end local output = stem .. suffix -- Re-add the possessive suffix, if applicable. if possessive then output = add_suffix(output, "'s", pos) end return output end add_suffix = export.add_suffix --[==[ Pluralize a word in a smart fashion, according to normal English rules. # If the word ends in a consonant or "qu" + "-y", replace "-y" with "-ies". # If the word ends in "s", "x", "z", "ch", "sh" or "zh", add "-es". # Otherwise, add "-s". This handles links correctly: # If a piped link, change the second part appropriately. # If a non-piped link and rule #1 above applies, convert to a piped link with the second part containing the plural. # If a non-piped link and rules #2 or #3 above apply, add the plural outside the link. ]==] function export.pluralize(str) -- Treat as a link if a "[[" is present and the string ends with "]]". if not (find(str, "[[", 1, true) and sub(str, -2) == "]]") then return add_suffix(str, "s.plural") end -- Find the last "[[" (in case there is more than one) by reversing -- the string. local str_rev = reverse(str) local open = find(str_rev, "[[", 3, true) -- If the last "[[" is followed by a "]]" which isn't at the end, -- then the final "]]" is just plaintext (e.g. "[[foo]]bar]]"). local bad_close = find(str_rev, "]]", 3, true) -- Note: the bad "]]" will have a lower index than the last "[[" in -- the reversed string. if bad_close and bad_close < open then return add_suffix(str, "s.plural") end open = #str - open + 2 -- Get the target and display text by searching from just after "[[". local target, display = match(str, "([^|]*)|?(.*)%]%]$", open) display = add_suffix(display ~= "" and display or target, "s.plural") -- If the link target is a substring of the display text, then -- use a trail (e.g. "[[foo]]" → "[[foo]]s", since "foo" is a substring -- of "foos"). local index, trail = find(display, target, 1, true) if index == 1 then return sub(str, 1, open - 1) .. target .. "]]" .. sub(display, trail + 1) end -- Otherwise, return a piped link. return sub(str, 1, open - 1) .. target .. "|" .. display .. "]]" end --[==[ Returns true if `plural` is an expected, regular plural of `term`. The optional parameter `pos` can be used to specify the part of speech, which is necessary because proper nouns do not change a {"-y"} suffix to {"-ies"} (e.g. {"Abby"} → {"Abbys"}). By default, `pos` is set to {"noun"}. In addition to {"proper noun"}, it can also take the special value {"noun+"}, which means that the function will first attempt the check with the {"noun"} setting, and will then attempt it with the {"proper noun"} setting iff the term begins with a capital letter. ]==] function export.is_regular_plural(plural, term, pos) local init_plural, init_term, try_as_proper_noun = plural, term if pos == "noun+" then pos, try_as_proper_noun = "noun", true end -- Ignore any final punctuation that occurs in both forms, which is common -- in abbreviations (e.g. "abbr." → "abbrs."). local final_punc = umatch(term, "%p*$") local final_punc_len = #final_punc if sub(plural, -final_punc_len) == final_punc then term = sub(term, 1, -final_punc_len - 1) plural = sub(plural, 1, -final_punc_len - 1) end if plural == add_suffix(term, "s.plural", pos) then return true end local final = sub(term, -1) if ( -- Doubled final consonants in "s" and "z". final == "s" and plural == term .. "ses" or -- e.g. "busses" final == "z" and plural == term .. "zes" or -- e.g. "quizzes" -- convert_final_y_to_i() without the `not_gu` flag set, to catch -- "-guy" → "-guies", but not "day" → "daies". final == "y" and plural == convert_final_y_to_i(term) .. "es" or -- Capitalized terms like "$DEITY" → "$DEITIES (should we treat this as regular?) final == "Y" and ulower(plural) == convert_final_y_to_i(ulower(term)) .. "es" ) then return true elseif try_as_proper_noun then local init = umatch(init_term, "^[^%w%s]*(%w)") return init and uupper(init) == init and ulower(init) ~= init and is_regular_plural(init_plural, init_term, "proper noun") or false end return false end is_regular_plural = export.is_regular_plural do local function do_singularize(str) local sing = match(str, "^(.-)ies$") if sing then return sing .. "y" end -- Handle cases like "[[parish]]es" return match(str, "^(.-[cs]h%]*)es$") or -- not -zhes -- Handle cases like "[[box]]es" match(str, "^(.-x%]*)es$") or -- not -ses or -zes -- Handle regular plurals match(str, "^(.-)s$") or -- Otherwise, return input str end local function collapse_link(link, linktext) if link == linktext then return "[[" .. link .. "]]" end return "[[" .. link .. "|" .. linktext .. "]]" end --[==[ Singularize a word in a smart fashion, according to normal English rules. Works analogously to {pluralize()}. '''NOTE''': This doesn't always work as well as {pluralize()}. Beware. It will mishandle cases like "passes" -> "passe", "eyries" -> "eyry". # If word ends in -ies, replace -ies with -y. # If the word ends in -xes, -shes, -ches, remove -es. [Does not affect -ses, cf. "houses", "impasses".] # Otherwise, remove -s. This handles links correctly: # If a piped link, change the second part appropriately. Collapse the link to a simple link if both parts end up the same. # If a non-piped link, singularize the link. # A link like "[[parish]]es" will be handled correctly because the code that checks for -shes etc. allows ] characters between the 'sh' etc. and final -es. ]==] function export.singularize(str) if type(str) == "table" then -- allow calling from a template str = str.args[1] end -- Check for a link. This pattern matches both piped and unpiped links. -- If the link is not piped, the second capture (linktext) will be empty. local beginning, link, linktext = match(str, "^(.*)%[%[([^|%]]+)%|?(.-)%]%]$") if not link then return do_singularize(str) elseif linktext ~= "" then return beginning .. collapse_link(link, do_singularize(linktext)) end return beginning .. "[[" .. do_singularize(link) .. "]]" end end --[==[ Return the appropriate indefinite article to prefix to `str`. Correctly handles links and capitalized text. Does not correctly handle words like [[union]], [[uniform]] and [[university]] that take "a" despite beginning with a 'u'. The returned article will have its first letter capitalized if `ucfirst` is specified, otherwise lowercase. ]==] function export.get_indefinite_article(str, ucfirst) str = str or "" -- If there's a link at the beginning, examine the first letter of the -- link text. This pattern matches both piped and unpiped links. -- If the link is not piped, the second capture (linktext) will be empty. local link, linktext = match(str, "^%[%[([^|%]]+)%|?(.-)%]%]") if match(link and (linktext ~= "" and linktext or link) or str, "^()[AEIOUaeiou]") then return ucfirst and "An" or "an" end return ucfirst and "A" or "a" end get_indefinite_article = export.get_indefinite_article --[==[ Prefix `text` with the appropriate indefinite article to prefix to `text`. Correctly handles links and capitalized text. Does not correctly handle words like [[union]], [[uniform]] and [[university]] that take "a" despite beginning with a 'u'. The returned article will have its first letter capitalized if `ucfirst` is specified, otherwise lowercase. ]==] function export.add_indefinite_article(text, ucfirst) return get_indefinite_article(text, ucfirst) .. " " .. text end export.vowels = vowels export.vowel = "[" .. vowels .. "]" return export qmuwy34gf3az49fu17xn4hzrc6y4prc Module:headword utilities 828 8053 54857 2026-09-26T23:23:54Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local require_when_needed = require("Module:utilities/require when needed") local affix_module = "Module:affix" local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local fun_is_callable_module = "Module:fun/isCallable" local headword_module = "Module:headword" local headword_data_module = "Module:headword/data" local languages_module = "Module:language...' 54857 Scribunto text/plain local export = {} local require_when_needed = require("Module:utilities/require when needed") local affix_module = "Module:affix" local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local fun_is_callable_module = "Module:fun/isCallable" local headword_module = "Module:headword" local headword_data_module = "Module:headword/data" local languages_module = "Module:languages" local links_module = "Module:links" local parameters_module = "Module:parameters" local parse_interface_module = "Module:parse interface" local parse_utilities_module = "Module:parse utilities" local string_pattern_escape_module = "Module:string/patternEscape" local string_replacement_escape_module = "Module:string/replacementEscape" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local yesno_module = "Module:yesno" local dump = mw.dumpObject local unpack = unpack or table.unpack -- Lua 5.2 compatibility local insert = table.insert local concat = table.concat local remove = table.remove local sort = table.sort local deep_equals = require_when_needed(table_module, "deepEquals") local extend = require_when_needed(table_module, "extend") local insert_if_not = require_when_needed(table_module, "insertIfNot") local list_to_set = require_when_needed(table_module, "listToSet") local serial_comma_join = require_when_needed(table_module, "serialCommaJoin") local shallow_copy = require_when_needed(table_module, "shallowCopy") local split = require_when_needed(string_utilities_module, "split") local ugsub = require_when_needed(string_utilities_module, "gsub") local umatch = require_when_needed(string_utilities_module, "match") local pattern_escape = require_when_needed(string_pattern_escape_module) local replacement_escape = require_when_needed(string_replacement_escape_module) local escape_wikicode = require_when_needed(parse_utilities_module, "escape_wikicode") local parse_inline_modifiers = require_when_needed(parse_utilities_module, "parse_inline_modifiers") local term_contains_top_level_html = require_when_needed(parse_utilities_module, "term_contains_top_level_html") local get_lang_by_code = require_when_needed(languages_module, "getByCode") local is_callable = require_when_needed(fun_is_callable_module) local format_decorations = require_when_needed(decorations_module, "format_decorations") local function split_on_comma(val) if val:find(",") then return require(parse_interface_module).split_on_comma(val) else return {val} end end local function ine(val) if val == "" then return nil else return val end end --[=[ Add decorations to a term. `termobj` is the object describing the term, which should optionally contain: * left qualifiers in `q`, an array of strings; * right qualifiers in `qq`, an array of strings; * left labels in `l`, an array of strings; * right labels in `ll`, an array of strings; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; `text` is the text of the term itself, and `lang` is the language object. ]=] local function add_decorations(text, termobj, lang) local function field_non_empty(field) local list = termobj[field] if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `termobj.%s`=%s, should be \"table\""):format( field, mw.dumpObject(list))) end return list[1] end if field_non_empty("q") or field_non_empty("qq") or field_non_empty("l") or field_non_empty("ll") or field_non_empty("refs") then text = format_decorations { lang = lang, text = text, q = termobj.q, qq = termobj.qq, l = termobj.l, ll = termobj.ll, refs = termobj.refs, } end return text end local param_mods = { id = {}, -- disabled when `is_head = true` q = {type = "qualifier"}, qq = {type = "qualifier"}, l = {type = "labels"}, ll = {type = "labels"}, -- [[Module:headword]] expects part references in `.refs`. ref = {item_dest = "refs", type = "references", store = "insert-flattened"}, } local optional_param_mods = { g = {item_dest = "genders", type = "genders"}, alt = {}, lang = {type = "language"}, sc = {type = "script"}, t = {item_dest = "gloss"}, gloss = {}, pos = {}, lit = {}, tr = {}, ts = {}, face = {}, nolinkinfl = {type = "boolean"}, } local optional_headword_param_mods = { sc = {type = "script"}, tr = {}, ts = {}, } --[==[ Parse a single inflection or headword form or list of such forms. In either case, inline modifiers may be attached. `data` is an object with the following fields: * `val`: The raw value to parse. Required. * `paramname`: The name of the parameter from which the value was taken; used in error messages. Required. * `is_head`: We are parsing a headword parameter (a value which goes into the `heads` field of `data`). This changes the allowed modifiers, disabling the `id` modifier and only allowing a subset of optional modifiers. * `frob`: An optional function of one value to apply to the form after inline modifiers have been removed (i.e. to apply to the `.term` field of the returned object). * `include_mods`: List of extra inline modifiers to include, besides the default ones (see below). Each list item is either a string specifying a recognized extra inline modifier (see `optional_param_mods` in the code), or a two-item list of modifier name and modifier spec, where the spec should follow the syntax for modifier specs in `parse_inline_modifiers` in [[Module:parse utilities]]. * `exclude_mods`: List of default inline modifiers to not include. * `splitchar`: If specified, the value in `val` can be a list of forms to parse, separated by the value of `splitchar` (which is a Lua pattern, as in `parse_inline_modifiers` in [[Module:parse utilities]]). Most commonly, `splitchar` is a single comma and the values are comma-separated (in this case, splitting will not happen if a space follows the comma). * `parse_lang_prefix`: If specified, allow a language prefix to precede a form, and if found, store into the `.lang` field of the returned object. * `preserve_splitchar`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_inline_modifiers` in [[Module:parse utilities]]. Returns an object suitable for storing as one element of one of the lists in `headdata.inflections`, where `headdata` is the structure passed to [[Module:headword]]. If `splitchar` is specified, howeve, the return value is a list of such objects. The following default inline modifiers are currently recognized: * `q`: Left qualifier. * `qq`: Right qualifier. * `l`: Comma-separated list of left labels. No space should follow the comma. * `ll`: Comma-separated list of right labels. No space should follow the comma. * `ref`: Reference or references. See {{tl|IPA}} for the syntax. * `id`: Sense ID, in case there are multiple senses. See {{tl|l}}. The following are the recognized additional inline modifiers: * `g`: Comma-separated list of genders. * `alt`: Display text. * `lang`: Language code of language of the form, if different from the language of the headword. * `sc`: Script code of script of the form. Almost never needed. * `t`: Gloss for the form. * `gloss`: Gloss for the form (alias for `t`). * `pos`: Part of speech of the form. * `lit`: Literal meaning of the form. * `tr`: Manual transliteration of the form. * `ts`: Transcription of the form, for languages where the transliteration differs markedly from the pronunciation. * `face`: Face to display the form in, e.g. {"hypothetical"} for a hypothetical form (unlinkable and displayed in italics). * `nolinkinfl`: Make the form unlinkable. ]==] function export.parse_term_with_modifiers(data) local paramname, val, frob = data.paramname, data.val, data.frob local function generate_obj(term, parse_err) if frob then term = frob(term, parse_err) end if data.parse_lang_prefix and term:find(":") then return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix { term = term, paramname = paramname, parse_lang_prefix = true, parse_err = parse_err, } else return {term = term} end end -- Check for inline modifier, e.g. מרים<tr:Miryem>. But exclude top-level HTML entry with <span ...>, -- <sup> or similar in it. if (val:find("<", nil, true) or data.splitchar) and not term_contains_top_level_html(val) and -- don't parse inline modifiers if is_head and the value begins with a ~ (link modifier syntax) (not data.is_head or not val:find("^~")) then local param_mods = param_mods if data.is_head then param_mods = shallow_copy(param_mods) param_mods.id = nil end if data.include_mods or data.exclude_mods then if not data.is_head then -- already copied when data.is_head param_mods = shallow_copy(param_mods) end if data.include_mods then local optional_mods = data.is_head and optional_headword_param_mods or optional_param_mods for _, mod in ipairs(data.include_mods) do if type(mod) == "table" then if #mod ~= 2 then error(("Internal error: Modifier spec %s in `include_mods` should be of length 2"):format( dump(mod))) end local modkey, modvalue = unpack(mod) param_mods[modkey] = modvalue elseif not optional_mods[mod] then error(("Internal error: Unrecognized modifier spec %s in `include_mods`"):format( dump(mod))) else param_mods[mod] = optional_mods[mod] end end end if data.exclude_mods then for _, mod in ipairs(data.exclude_mods) do if not param_mods[mod] then error(("Internal error: Modifier spec %s in `exclude_mods` not found among existing modifiers" ):format(dump(mod))) else param_mods[mod] = nil end end end end return parse_inline_modifiers(val, { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, splitchar = data.splitchar, preserve_splitchar = data.preserve_splitchar, delimiter_key = data.delimiter_key, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, pre_normalize_modifiers = data.pre_normalize_modifiers, }) else local retval = generate_obj(val) if data.splitchar then retval = {retval} end return retval end end --[==[ Parse a list of inflection forms that may have inline modifiers attached. `data` is an object with the following fields: * `forms`: The list of raw values to parse. Required. * `paramname`: The name of the first parameter from which the value was taken; used in error messages. If this is a two-element list, the first element is the first parameter and the second element is the prefix of the remaining parameters. Parameter names that are numbers are handled correctly, as are those with \1 in it marking where the parameter index goes. Required. * `qualifiers`: If specified, a possibly gappy list of left qualifiers to add to the parsed terms (for compatibility purposes). * `splitchar`: As in `parse_term_with_modifiers()`. The resulting per-term lists will be flattened. * `frob`, `include_mods`, `exclude_mods`, `is_head`, `preserve_splitchar`, `parse_lang_prefix`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_with_modifiers()`. Returns a list of objects, suitable for storing as one of the lists in `headdata.inflections` (once a label is added), where `headdata` is the structure passed to [[Module:headword]]. ]==] function export.parse_term_list_with_modifiers(data) local paramname, forms = data.paramname, data.forms local qualifiers = data.qualifiers local first, restpref if type(paramname) == "table" then first = paramname[1] restpref = paramname[2] else first = paramname restpref = paramname end local terms = {} data = shallow_copy(data) for i, val in ipairs(forms) do data.paramname = i == 1 and first or type(restpref) == "number" and restpref + i - 1 or restpref:find("\1", nil, true) and restpref:gsub("\1", tostring(i)) or restpref .. i data.val = val local parsed = export.parse_term_with_modifiers(data) if qualifiers and qualifiers[i] then if data.splitchar then for _, term in ipairs(parsed) do term.q = {qualifiers[i]} end else parsed.q = {qualifiers[i]} end end if data.splitchar then extend(terms, parsed) else terms[i] = parsed end end return terms end --[==[ Construct a link to [[Appendix:Glossary]] for `entry`. If `text` is specified, it is the display text; otherwise, `entry` is used. ]==] function export.glossary_link(entry, text) text = text or entry return "[[Appendix:Glossary#" .. entry .. "|" .. text .. "]]" end function export.replace_glossary_links_in_label(label) if label:find("<<", nil, true) then label = label:gsub("<<(.-)|(.-)>>", export.glossary_link):gsub("<<(.-)>>", export.glossary_link) end return label end --[==[ Insert a fixed inflection (a label not associated with any inflection values) into an `inflections` field. The `inflections` field will be initialized if needed. `data` is an object with the following fields: * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs to be set for nested inflections, which are specified for an inflection object rather than the headword data structure as a whole. * `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required. * `originating_term`: The term object from which this label is derived. If specified, decorations will be taken from this object. ]==] function export.insert_fixed_inflection(data) local headdata, origterm, label = data.headdata, data.originating_term, data.label local inflobj = data.inflobj or headdata inflobj.inflections = inflobj.inflections or {} if not origterm then insert(inflobj.inflections, { label = export.replace_glossary_links_in_label(label) }) else if origterm.id then error(("It doesn't make sense to pass in an ID '%s' for label '%s' in conjunction with a term value '%s'" ):format(origterm.id, label, origterm.term)) end origterm = shallow_copy(origterm) -- Preserve decorations origterm.term = nil origterm.label = export.replace_glossary_links_in_label(label) insert(inflobj.inflections, origterm) end end --[==[ Insert previously-parsed terms into an `inflections` field. The `inflections` field will be initialized if needed. `data` is an object with the following fields: * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs to be set for nested inflections, which are specified for an inflection object rather than the headword data structure as a whole. * `terms`: The list of parsed terms. If {nil} or omitted, nothing happens unless `request` is set. * `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required. * `no_label`: If the term is {"-"} and there are no other terms, insert a fixed label with this value. Defaults to {"no "} plus the label. * `usually_no_label`: If the term is {"-"} and there are other terms, insert a fixed label with this value. Defaults to {"usually no "} plus the label. * `cats`: List of categories to insert when terms are given that are not {"-"}. Each category is a string naming a full category to insert (including the appropriate language name prefixed). * `no_cats`: List of categories to insert when a term is given as {"-"}. * `usually_no_cats`: List of categories to insert when a term is given as {"-"} and additional terms are specified as well (representing, e.g. for the inflection {"plural"}, a term which usually has no plural but does under some circumstances). If omitted, both the categories in `cats` and `no_cats` are inserted. * `accel`: If specified, a full accelerator object to add to the inflections. * `request`: If specified and no terms are given, insert a label with a request for inflections to be given. * `enable_auto_translit`: If specified and terms are given, display automatic transliteration of the terms. The return value indicates whether the inflection exists and how many terms are in it. It is an object with the following fields: * `exists`: {"yes"} if one or more terms were specified; {"no"} if the value was given as {"-"}; {"usually no"} if the first value was given as {"-"} but additional terms were supplied; otherwise {nil}, indicating that the status is unspecified. * `numterms`: Number of terms in the inflection. Will be 0 unless `exists` has the value {"yes"} or {"usually no"}. * `request`: True if no terms were specified but a term request was inserted into the inflection (because `data.request` was specified). Otherwise {nil}. ]==] function export.insert_inflection(data) local headdata, terms, label = data.headdata, data.terms, data.label local inflobj = data.inflobj or headdata local retval = {} local accel = data.accel if data.accel_form then if accel then error("Internal error: can't specify both data.accel and data.accel_form") end if headdata.heads then local lemmas = {} local lemma_translits = {} for i, headobj in ipairs(headdata.heads) do lemmas[i] = headobj.term if lemmas[i] == "+" then error("Internal error: If you use data.accel_form, you should have resolved all occurrences of + in heads appropriately") end lemma_translits[i] = headobj.tr end accel = { lemma = lemmas, lemma_translit = lemma_translits, form = data.accel_form, } else accel = { form = data.accel_form, } end end local function insert_cats(cats) for _, cat in ipairs(cats) do insert(headdata.categories, cat) end end if terms and terms[1] then terms = shallow_copy(terms) if terms[1].term == "-" then if terms[2] then export.insert_fixed_inflection { headdata = headdata, inflobj = inflobj, originating_term = terms[1], label = data.usually_no_label or "usually no " .. label, } remove(terms, 1) retval.numterms = #terms retval.exists = "usually no" if data.usually_no_cats then insert_cats(data.usually_no_cats) else if data.no_cats then insert_cats(data.no_cats) end if data.cats then insert_cats(data.cats) end end else export.insert_fixed_inflection { headdata = headdata, inflobj = inflobj, originating_term = terms[1], label = data.no_label or "no " .. label, } retval.numterms = 0 retval.exists = "no" if data.no_cats then insert_cats(data.no_cats) end return retval end else retval.numterms = #terms retval.exists = "yes" if data.cats then insert_cats(data.cats) end end if data.check_missing then error("Internal error: check_missing support removed; use checkredlinks=true in [[Module:headword]]") end terms.label = export.replace_glossary_links_in_label(label) if accel then terms.accel = accel end terms.enable_auto_translit = data.enable_auto_translit inflobj.inflections = inflobj.inflections or {} insert(inflobj.inflections, terms) elseif data.request then inflobj.inflections = inflobj.inflections or {} insert(inflobj.inflections, { label = export.replace_glossary_links_in_label(label), request = true, }) retval.numterms = 0 -- retval.exists = nil retval.request = true else retval.numterms = 0 -- retval.exists = nil end return retval end --[==[ Parse raw arguments from `forms` for inline modifiers, and insert the resulting terms (which should not require significant additional processing) into `headdata.inflections`. `data` is an object with the following fields: * `forms`: The list of raw values to parse. If {nil} or omitted, nothing happens. * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `paramname`: As in `parse_term_list_with_modifiers()`. Required. * `label`: As in `insert_inflection()`. Required. * `qualifiers`, `frob`, `include_mods`, `exclude_mods`, `is_head`, `splitchar`, `preserve_splitchar`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_list_with_modifiers()`. * `accel`: As in `insert_inflection()`. Return value is as in `insert_inflection()`. ]==] function export.parse_and_insert_inflection(data) local forms = data.forms if forms and forms[1] then data = shallow_copy(data) data.forms = forms data.terms = export.parse_term_list_with_modifiers(data) return export.insert_inflection(data) end return { numterms = 0 } end --[==[ Canonicalize a single term or term-like object or a list of either into a list of term-like objects. `abterms` is the term or list to canonicalize, and `field` is the name of the field holding the term (defaulting to {"term"}). This does the minimal work necessary, meaning that the return value may partly or completely share memory with the value passed in. As a special case, if `abterms` is {nil}, {nil} is returned. If `origin_val` is specified, add a field `origin` containing the value of `origin_val` to each resulting term-like object (in this case, the object will be copied a necessary to avoid side-effecting the passed-in objects). ]==] function export.canonicalize_termobj_list(abterms, field, origin_val) if abterms == nil then return nil end field = field or "term" if type(abterms) == "string" then return {{[field] = abterms, origin = origin_val}} elseif not abterms[1] then if origin_val ~= nil then abterms = shallow_copy(abterms) abterms.origin = origin_val end return {abterms} else -- Check if already in full list term and return directly if so (unless `origin_val` is given, in which case we -- need to shallow-copy both the list and each term in it). local must_convert = false for _, term in ipairs(abterms) do if type(term) == "string" then must_convert = true break end end if not must_convert then if origin_val ~= nil then abterms = shallow_copy(abterms) for i, abterm in ipairs(abterms) do abterms[i] = shallow_copy(abterm) abterms[i].origin = origin_val end end return abterms end end local retval = {} for _, term in ipairs(abterms) do if type(term) == "string" then insert(retval, {[field] = term, origin = origin_val}) else if origin_val ~= nil then term = shallow_copy(term) term.origin = origin_val end insert(retval, term) end end return retval end --[==[ Combine two sets of decorations. If either is {nil}, just return the other, and if both are {nil}, return {nil}. ]==] function export.combine_decorations(decs1, decs2) if not decs1 and not decs2 then return nil end if not decs1 then return decs2 end if not decs2 then return decs1 end local combined = shallow_copy(decs1) for _, dec in ipairs(decs2) do insert_if_not(combined, dec) end return combined end function export.combine_qualifiers_or_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use combine_decorations instead") end --[==[ Combine the decorations (qualifiers, labels, references) and ID's of two term objects. `destobj` is the "destination term object" into which the combined properties are written, and `srcobj` is the "source object" into which the properties are merged. `destobj` is side-effected (but the lists inside of `destobj` are not); if this is undesirable, make sure to shallow-copy `destobj` first. If both objects have values for a given decoration, the values of `destobj` come first. If both objects have a value for `id`, the values must match or an error is thrown; otherwise, the resulting value of `id` comes from whichever one is defined. '''NOTE:''' This may not be the correct behavior when deduplicating a list of term objects. See `insert_termobj_combining_duplicates` for a different approach. ]==] function export.combine_termobj_decorations(destobj, srcobj) destobj.q = export.combine_decorations(destobj.q, srcobj.q) destobj.qq = export.combine_decorations(destobj.qq, srcobj.qq) destobj.l = export.combine_decorations(destobj.l, srcobj.l) destobj.ll = export.combine_decorations(destobj.ll, srcobj.ll) destobj.refs = export.combine_decorations(destobj.refs, srcobj.refs) if destobj.id and srcobj.id and destobj.id ~= srcobj.id then -- FIXME: We probably want to pass in an error function error(("Can't specify two different ID's %s and %s when combining objects"):format(srcobj.id, destobj.id)) end destobj.id = destobj.id or srcobj.id return destobj end function export.combine_termobj_qualifiers_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use combine_termobj_decorations instead") end function export.termobj_has_decorations(obj) return obj.q and obj.q[1] or obj.qq and obj.qq[1] or obj.l and obj.l[1] or obj.ll and obj.ll[1] or obj.refs and obj.refs[1] end function export.termobj_has_qualifiers_or_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use termobj_has_decorations instead") end local function one_decoration_equal(prop1, prop2) local prop1_is_nil = not prop1 or not prop1[1] local prop2_is_nil = not prop2 or not prop2[1] if prop1_is_nil and prop2_is_nil then return true end if prop1_is_nil or prop2_is_nil then return false end return deep_equals(prop1, prop2) end function export.termobj_decorations_equal(obj1, obj2) return one_decoration_equal(obj1.q, obj2.q) and one_decoration_equal(obj1.qq, obj2.qq) and one_decoration_equal(obj1.l, obj2.l) and one_decoration_equal(obj1.ll, obj2.ll) and one_decoration_equal(obj1.refs, obj2.refs) and obj1.id == obj2.id end function export.termobj_ancillary_properties_equal(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use termobj_decorations_equal instead") end function export.convert_termobj_to_formobj(termobj) local formobj = { form = termobj.term, translit = termobj.tr, } local footnotes local function mods_to_footnote(mod_prefix, mod_vals) if mod_vals and mod_vals[1] then footnotes = footnotes or {} for _, val in ipairs(mod_vals) do insert(footnotes, "[" .. mod_prefix .. ":" .. val .. "]") end end end mods_to_footnote("q", termobj.q) mods_to_footnote("qq", termobj.qq) mods_to_footnote("l", termobj.l) mods_to_footnote("ll", termobj.ll) mods_to_footnote("ref", termobj.refs) mods_to_footnote("id", termobj.id and {termobj.id} or nil) formobj.footnotes = footnotes return formobj end local recognized_multi_mods = { q = "q", qq = "qq", l = "l", ll = "ll", ref = "refs", } local recognized_single_mods = { id = "id", } function export.add_footnote_to_termobj(termobj, footnote) local stripped_footnote = footnote:match("^%[(.*)%]$") if not stripped_footnote then error("Internal error: Footnote should be surrounded by brackets at this stage: " .. footnote) end local prefix, rest = stripped_footnote:match("^([a-z]+):(.+)$") local field, is_multi if prefix then if recognized_multi_mods[prefix] then field = recognized_multi_mods[prefix] is_multi = true elseif recognized_single_mods[prefix] then field = recognized_single_mods[prefix] is_multi = false end end if not field then rest = stripped_footnote field = "l" is_multi = true end if is_multi then if not termobj[field] then termobj[field] = {} end insert(termobj[field], rest) else if termobj[field] and termobj[field] ~= rest then error(("Can't set two values for '%s': '%s' and '%s'"):format(field, termobj[field], rest)) end termobj[field] = rest end end function export.convert_formobj_to_termobj(formobj) local termobj = { term = formobj.form, tr = formobj.translit, } if formobj.footnotes then for _, footnote in ipairs(formobj.footnotes) do export.add_footnote_to_termobj(termobj, footnote) end end return termobj end local function extract_termobj_field_modifiers(fieldval) return fieldval:match("^([*+]?)(.*)$") end function export.remove_termobj_field_modifiers(termobj) local function remove_field_modifiers(field) if termobj[field] and termobj[field][1] then local any_field_modifiers = false for _, val in ipairs(termobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods ~= "" then any_field_modifiers = true break end end local new_field = {} if any_field_modifiers then for _, val in ipairs(termobj[field]) do local _, field_without_mods = extract_termobj_field_modifiers(val) insert_if_not(new_field, field_without_mods) end termobj[field] = new_field end end end remove_field_modifiers("q") remove_field_modifiers("qq") remove_field_modifiers("l") remove_field_modifiers("ll") remove_field_modifiers("refs") end function export.insert_termobj_combining_duplicates(destobjs, termobj) for _, destobj in ipairs(destobjs) do if destobj.term == termobj.term and destobj.tr == termobj.tr then -- Form already present; maybe combine footnotes. local function combine_field_values(field) if termobj[field] and termobj[field][1] then -- Check to see if there are existing values with *; if so, remove them. if destobj[field] and destobj[field][1] then local any_values_with_asterisk = false for _, val in ipairs(destobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods:find("%*") then any_values_with_asterisk = true break end end if any_values_with_asterisk then local filtered_values = {} for _, val in ipairs(destobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if not field_mods:find("%*") then insert(filtered_values, val) end end if filtered_values[1] then destobj[field] = filtered_values else destobj[field] = nil end end end local any_footnotes_with_plus = false for _, val in ipairs(termobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods:find("%+") then any_footnotes_with_plus = true break end end if any_footnotes_with_plus then if not destobj[field] then destobj[field] = {} else destobj[field] = shallow_copy(destobj[field]) end for _, val in ipairs(termobj[field]) do local already_seen = false local field_mods, field_without_mods = extract_termobj_field_modifiers(val) if field_mods:find("%+") then for _, existing_val in ipairs(destobj[field]) do local _, existing_field_without_mods = extract_termobj_field_modifiers(existing_val) if existing_field_without_mods == field_without_mods then already_seen = true break end end if not already_seen then insert(destobj[field], val) end end end end end end combine_field_values("q") combine_field_values("qq") combine_field_values("l") combine_field_values("ll") combine_field_values("refs") if destobj.id and termobj.id and destobj.id ~= termobj.id then -- FIXME: We probably want to pass in an error function error(("Can't specify two different ID's %s and %s when combining objects"):format(termobj.id, destobj.id)) end destobj.id = destobj.id or termobj.id return end end insert(destobjs, termobj) end export.allowed_special_indicators = { ["first"] = true, ["first-second"] = true, ["first-last"] = true, ["second"] = true, ["last"] = true, ["each"] = true, ["+"] = true, -- requests the default behavior with preposition handling } --[==[ Check for special indicators (values such as {"+first"} or {"+first-last"} that are used in a `pl`, `f`, etc. argument and indicate how to inflect a multiword term). If `form` is such an indicator, the return value is `form` minus the initial `+` sign; otherwise, if form begins with a `+` sign, an error is thrown; otherwise the return value is nil. ]==] function export.get_special_indicator(form, noerror) if form:find("^%+") then form = form:gsub("^%+", "") if not export.allowed_special_indicators[form] then if noerror then return nil end local indicators = {} for indic, _ in pairs(export.allowed_special_indicators) do insert(indicators, "+" .. indic) end sort(indicators) error("Special inflection indicator beginning with '+' can only be " .. mw.text.listToText(indicators) .. ": +" .. form) end return form end return nil end local function add_endings(bases, endings) local retval = {} if type(bases) ~= "table" then bases = {bases} end if type(endings) ~= "table" then endings = {endings} end for _, base in ipairs(bases) do for _, ending in ipairs(endings) do insert(retval, base .. ending) end end return retval end --[==[ Inflect a possibly multiword or hyphenated term `form` using the function `inflect`, which is a function of one argument that is called on a single word to inflect and should return either the inflected word or a list of inflected words. `special` indicates how to inflect the multiword term and should be e.g. {"first"} to inflect only the first word, {"first-last"} to inflect the first and last words, {"each"} to inflect each word, etc. See `allowed_special_indicators` above for the possibilities. If `special` is `+`, or is omitted and the term is multiword (i.e. containing a space character), and `prepositions` is supplied, the function checks for multiword or hyphenated terms containing the prepositions in `prepositions`, e.g. Italian [[senso di marcia]] or [[medaglia d'oro]] or Portuguese [[tartaruga-do-mar]]. If such a term is found, only the first word is inflected. Otherwise, the default is {"first-last"}. `prepositions` is a list of Lua patterns matching prepositions. The patterns will automatically have the separator character (space or hyphen) added to the left side but not the right side, so they should contain a space character (which will automatically be converted to the appropriate separator) on the right side unless the preposition is joined on the right side with an apostrophe. Examples of preposition patterns for Italian are {"di "}, {"sull'"} and {"d?all[oae] "} (which matches {"dallo "}, {"dalle "}, {"alla "}, etc.). The return value is always either a list of inflected multiword or hyphenated terms, or nil if `special` is omitted and `form` is not multiword. (If `special` is specified and `form` is not multiword or hyphenated, an error results.) ]==] function export.handle_multiword(form, special, inflect, prepositions, sep) sep = sep or form:find(" ") and " " or "%-" local raw_sep = sep == " " and " " or "-" -- Used to add regex version of separator in the replacement portion of ugsub() or :gsub() local sep_replacement = sep == " " and " " or "%%-" -- Given a Lua pattern, replace space with the appropriate separator. local function hack_re(re) if sep == " " then return re end return (re:gsub(" ", sep_replacement)) end if special == "first" then local first, rest = form:match(hack_re("^(.-)( .*)$")) if not first then error("Special indicator 'first' can only be used with a multiword term: " .. form) end return add_endings(inflect(first), rest) elseif special == "second" then local first, second, rest = form:match(hack_re("^([^ ]+ )([^ ]+)( .*)$")) if not first then error("Special indicator 'second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings({first}, inflect(second)), rest) elseif special == "first-second" then local first, space, second, rest = form:match(hack_re("^([^ ]+)( )([^ ]+)( .*)$")) if not first then error("Special indicator 'first-second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings(add_endings(inflect(first), space), inflect(second)), rest) elseif special == "each" then local terms = split(form, sep) if #terms < 2 then error("Special indicator 'each' can only be used with a multiword term: " .. form) end for i, term in ipairs(terms) do terms[i] = inflect(term) if i > 1 then terms[i] = add_endings(raw_sep, terms[i]) end end local result = "" for _, term in ipairs(terms) do result = add_endings(result, term) end return result elseif special == "first-last" then local first, middle, last = form:match(hack_re("^(.-)( .* )(.-)$")) if not first then first, middle, last = form:match(hack_re("^(.-)( )(.*)$")) end if not first then error("Special indicator 'first-last' can only be used with a multiword term: " .. form) end return add_endings(add_endings(inflect(first), middle), inflect(last)) elseif special == "last" then local rest, last = form:match(hack_re("^(.* )(.-)$")) if not rest then error("Special indicator 'last' can only be used with a multiword term: " .. form) end return add_endings(rest, inflect(last)) elseif special and special ~= "+" then error("Unrecognized special=" .. special) end -- Only do default behavior if special indicator '+' explicitly given or separator is space; otherwise we will -- break existing behavior with hyphenated words. if (special == "+" or sep == " ") and form:find(sep) then if prepositions then -- check for prepositions in the middle of the word; do it this way so we can handle -- more than one word before the preposition (and usually inflect each word) for _, prep in ipairs(prepositions) do local first, space_prep_rest = umatch(form, hack_re("^(.-)( " .. prep .. ".*)$")) if first then return add_endings(inflect(first), space_prep_rest) end end end -- multiword or hyphenated expressions default to first-last; we need to pass in the separator to avoid -- problems with multiword terms containing hyphens in the individual words return export.handle_multiword(form, "first-last", inflect, prepositions, sep) end return nil end local function link_hyphen_split_component(word, data) if data.link_hyphen_split_component then return data.link_hyphen_split_component(word) else return "[[" .. word .. "]]" end end -- Default function to split a word on apostrophes. Don't split apostrophes at the beginning or end of a word (e.g. -- [['ndrangheta]] or [[po']]). Handle multiple apostrophes correctly, e.g. [[l'altr'ieri]] -> [[l']][altr']][[ieri]]. function export.default_split_apostrophe(word, data) local apostrophe_parts = split(word, "'", true, true) local linked_apostrophe_parts = {} local apostrophes_at_beginning = "" local i = 1 -- Apostrophes at beginning get attached to the first word after (which will always exist but may -- be blank if the word consists only of apostrophes). while i < #apostrophe_parts do -- <, not <=, in case the word consists only of apostrophes local apostrophe_part = apostrophe_parts[i] i = i + 1 if apostrophe_part == "" then apostrophes_at_beginning = apostrophes_at_beginning .. "'" else break end end apostrophe_parts[i] = apostrophes_at_beginning .. apostrophe_parts[i] -- Now, do the remaining parts. A blank part indicates more than one apostrophe in a row; we join -- all of them to the preceding word. while i <= #apostrophe_parts do local apostrophe_part = apostrophe_parts[i] if apostrophe_part == "" then linked_apostrophe_parts[#linked_apostrophe_parts] = linked_apostrophe_parts[#linked_apostrophe_parts] .. "'" elseif i == #apostrophe_parts then insert(linked_apostrophe_parts, apostrophe_part) else insert(linked_apostrophe_parts, apostrophe_part .. "'") end i = i + 1 end for j, tolink in ipairs(linked_apostrophe_parts) do linked_apostrophe_parts[j] = link_hyphen_split_component(tolink, data) end return concat(linked_apostrophe_parts) end --[=[ Auto-add links to a word that should not have spaces but may have hyphens and/or apostrophes. We split off final punctuation, then split on hyphens if `data.split_hyphen` is given, and also split on apostrophes if `data.split_apostrophe` is given. We only split on hyphens if they are in the middle of the word, not at the beginning or end (hyphens at the beginning or end indicate suffixes or prefixes, respectively). `include_hyphen_prefixes`, if given, is a set of prefixes (not including the final hyphen) where we should include the final hyphen in the prefix. Hence, e.g. if "anti" is in the set, a Portuguese word like [[anti-herói]] "anti-hero" will be split [[anti-]][[herói]] (whereas a word like [[código-fonte]] "source code" will be split as [[código]]-[[fonte]]). If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French [[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word to split, and the return value should be the split and linked word. ]=] local function add_single_word_links(space_word, data, term_has_spaces) local space_word_no_punct, punct local punct_pattern = data.punctuation if punct_pattern and is_callable(punct_pattern) then space_word_no_punct, punct = punct_pattern(space_word) else if punct_pattern == nil then punct_pattern = "[,;:?!]" end space_word_no_punct, punct = umatch(space_word, "^(.*)(" .. punct_pattern .. ")$") end space_word_no_punct = space_word_no_punct or space_word punct = punct or "" local words if space_word_no_punct:sub(1, 1) == "-" or space_word_no_punct:sub(-1) == "-" then -- don't split prefixes and suffixes words = {space_word_no_punct} else local splitter if term_has_spaces then splitter = data.split_hyphen_when_space else splitter = data.split_hyphen_when_no_space end if is_callable(splitter) then words = splitter(space_word_no_punct) if type(words) == "string" then return words .. punct end end end if not words then local split_hyphen if term_has_spaces then split_hyphen = data.split_hyphen_when_space else split_hyphen = data.split_hyphen_when_no_space if split_hyphen == nil then -- default to true; use `false` to avoid this split_hyphen = true end end if split_hyphen then words = split(space_word_no_punct, "-", true, true) else words = {space_word_no_punct} end end local linked_words = {} for j, word in ipairs(words) do if j < #words and data.include_hyphen_prefixes and data.include_hyphen_prefixes[word] then word = "[[" .. word .. "-]]" elseif j > 1 and data.include_hyphen_suffixes and data.include_hyphen_suffixes[word] then word = "[[-" .. word .. "]]" else -- Don't split on apostrophes if the word is in `no_split_apostrophe_words`. if (not data.no_split_apostrophe_words or not data.no_split_apostrophe_words[word]) and data.split_apostrophe and word:find("'", nil, true) then if data.split_apostrophe == true then word = export.default_split_apostrophe(word, data) else -- custom apostrophe splitter/linker word = data.split_apostrophe(word) end elseif word ~= "" then -- avoid -[[]]- (e.g. f--k) word = link_hyphen_split_component(word, data) end if j < #words then word = word .. "-" end end insert(linked_words, word) end return concat(linked_words) .. punct end --[=[ Auto-add links to a multiword term. `data` contains fields customizing how to do this. By default we proceed as follows: (1) If the term already has embedded links in it, they are left unchanged. (2) Otherwise, if there are spaces present, we split on spaces and link each word separately. (3) If a given space-separated component ends in punctuation (defaulting to [,;:?!]), it is separated off, the remainder of the algorithm run, and the punctuation pasted back on. (4) If there are hyphens in a given space-separated component, we may link each hyphenated term separately depending on the settings in `data`. Normally the hyphens are not included in the linked terms, but this can be overridden for specific prefixes and/or suffixes. By default, if there are spaces in the multiword term, we do not link hyphenated components (because of cases like "boire du petit-lait" where "petit-lait" should be linked as a whole), but do so otherwise (e.g. for "avant-avant-hier"); this can overridden for cases like "croyez-le ou non". Cases where only some of the hyphens should be split can always be handled by explicitly specifying the head (e.g. "Nord-Pas-de-Calais" given as head=[[Nord]]-[[Pas-de-Calais]]). (5) If there are apostrophes in a given component, we may link each apostrophe-separated term separately depending on the settings in `data`, including the apostrophe in the link to its left (so we split "de l'eau" as "[[de]] [[l']][[eau]]"). The settings in `data` are as follows: `split_hyphen_when_no_space`: Whether to split on hyphens when the term has no spaces. Defaults to true if set to `nil`. This can be a function of one argument, to implement a custom splitting algorithm for hyphen-separated terms. If this returns [FIXME: FINISH ME ...] If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French [[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word to split, and the return value should be the split and linked word. We don't always split on hyphens because of cases like "boire du petit-lait" where "petit-lait" should be linked as a whole, but provide the option to do it for cases like "croyez-le ou non". If there's no space, however, then it makes sense to split on hyphens by `no_split_apostrophe_words` and `include_hyphen_prefixes` allow for special-case handling of particular words and are as described in the comment above add_single_word_links(). ]=] function export.add_links_to_multiword_term(term, data) if term:match("[%[%]]") then return term end local words = split(term, " ", true, true) local term_has_spaces = #words > 1 local linked_words = {} for _, word in ipairs(words) do insert(linked_words, add_single_word_links(word, data, term_has_spaces)) end local retval = concat(linked_words, " ") -- If we ended up with a single link consisting of the entire term, -- remove the link. return retval:match("^%[%[([^%[%]]*)%]%]$") or retval end local function canonicalize_begin_end_spec(spec) local from, to = spec:match("^(.-):(.*)$") if not from then from = spec to = "" end return from, to end --[==[ Given a `linked_term` that is the output of add_links_to_multiword_term(), apply modifications as given in `modifier_spec` to change the link destination of subterms (normally single-word non-lemma forms; sometimes collections of adjacent words). This is usually used to link non-lemma forms to their corresponding lemma, but can also be used to replace a span of adjacent separately-linked words to a single multiword lemma. The format of `modifier_spec` is one or more semicolon-separated subterm specs, where each such spec is of the form SUBTERM:DEST, where SUBTERM is one or more words in the `linked_term` but without brackets in them, and DEST is the corresponding link destination to link the subterm to. Any occurrence of ~ in DEST is replaced with SUBTERM. Alternatively, a single modifier spec can be of the form BEGIN[FROM:TO], which is equivalent to writing BEGINFROM:BEGINTO (see example below). For example, given the source phrase [[il bue che dice cornuto all'asino]] "the pot calling the kettle black" (literally "the ox that calls the donkey horned/cuckolded"), the result of calling add_links_to_multiword_term() is [[il]] [[bue]] [[che]] [[dice]] [[cornuto]] [[all']][[asino]]. With a modifier_spec of 'dice:dire', the result is [[il]] [[bue]] [[che]] [[dire|dice]] [[cornuto]] [[all']][[asino]]. Here, based on the modifier spec, the non-lemma form [[dice]] is replaced with the two-part link [[dire|dice]]. Another example: given the source phrase [[chi semina vento raccoglie tempesta]] "sow the wind, reap the whirlwind" (literally (he) who sows wind gathers [the] tempest"). The result of calling add_links_to_multiword_term() is [[chi]] [[semina]] [[vento]] [[raccoglie]] [[tempesta]], and with a modifier_spec of 'semina:~re; raccoglie:~re', the result is [[chi]] [[seminare|semina]] [[vento]] [[raccogliere|raccoglie]] [[tempesta]]. Here we use the ~ notation to stand for the non-lemma form in the destination link. A more complex example is [[se non hai altri moccoli puoi andare a letto al buio]], which becomes [[se]] [[non]] [[hai]] [[altri]] [[moccoli]] [[puoi]] [[andare]] [[a]] [[letto]] [[al]] [[buio]] after calling add_links_to_multiword_term(). With the following modifier_spec: 'hai:avere; altr[i:o]; moccol[i:o]; puoi: potere; andare a letto:~; al buio:~', the result of applying the spec is [[se]] [[non]] [[avere|hai]] [[altro|altri]] [[moccolo|moccoli]] [[potere|puoi]] [[andare a letto]] [[al buio]]. Here, we rely on the alternative notation mentioned above for e.g. 'altr[i:o]', which is equivalent to 'altri:altro', and link multiword subterms using e.g. 'andare a letto:~'. (The code knows how to handle multiword subexpressions properly, and if the link text and destination are the same, only a single-part link is formed.) ]==] function export.apply_link_modifiers(linked_term, modifier_spec, lang) local split_modspecs = split(modifier_spec, "%s*;%s*") for j, modspec in ipairs(split_modspecs) do local id if modspec:find("<") then local rest rest, id = modspec:match("^(.*)<id:(.-)>$") if rest then modspec = rest end end local subterm, dest, otherlang local begin_spec, rest, end_spec = modspec:match("^%[(.-)%]([^:]*)%[(.-)%]$") if begin_spec then local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec) local end_from, end_to = canonicalize_begin_end_spec(end_spec) subterm = begin_from .. rest .. end_from dest = begin_to .. rest .. end_to end if not subterm then rest, end_spec = modspec:match("^([^:]*)%[(.-)%]$") if rest then local end_from, end_to = canonicalize_begin_end_spec(end_spec) subterm = rest .. end_from dest = rest .. end_to end end if not subterm then begin_spec, rest = modspec:match("^%[(.-)%]([^:]*)$") if begin_spec then local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec) subterm = begin_from .. rest dest = begin_to .. rest end end if not subterm then subterm, dest = modspec:match("^(.-)%s*:%s*(.*)$") if subterm and subterm ~= "^" and subterm ~= "$" then local langdest -- Parse off an initial language code (e.g. 'en:Higgs', 'la:minūtia' or 'grc:σκατός'). Also handle -- Wikipedia prefixes ('w:Abatemarco' or 'w:it:Colle Val d'Elsa'). otherlang, langdest = dest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if otherlang == "w" then local foreign_wikipedia, foreign_term = langdest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if foreign_wikipedia then otherlang = otherlang .. ":" .. foreign_wikipedia langdest = foreign_term end dest = ("%s:%s"):format(otherlang, langdest) otherlang = nil elseif otherlang then otherlang = get_lang_by_code(otherlang, true, "allow etym") dest = langdest end end end if not subterm then if modspec == "?" or modspec == "!" then subterm = "$" dest = modspec elseif modspec == "..." or modspec == "...?" then subterm = "$" dest = " " .. modspec elseif modspec:find("^[A-Z]$") then -- X, Y, etc. by themselves are unlinked, to help with snowclones subterm = modspec dest = "_" else subterm = modspec dest = "~" end end if subterm == "^" then linked_term = dest:gsub("_", " ") .. linked_term elseif subterm == "$" then linked_term = linked_term .. dest:gsub("_", " ") else if subterm:find("[", nil, true) then error(("Subterm '%s' in modifier spec '%s' cannot have brackets in it"):format( escape_wikicode(subterm), escape_wikicode(modspec))) end local escaped_subterm = pattern_escape(subterm) local subterm_re = "%[%[" .. escaped_subterm:gsub("(%%?[ ',%-])", "%%]*%1%%[*") .. "%]%]" local expanded_dest if dest:find("~", nil, true) then expanded_dest = dest:gsub("~", replacement_escape(subterm)) else expanded_dest = dest end if otherlang then expanded_dest = expanded_dest .. "#" .. otherlang:getCanonicalName() end local subterm_replacement if expanded_dest == "_" then subterm_replacement = subterm if id then error("Can't supply <id:...> with an unlinked subterm") end if otherlang then error("Can't supply prefixed language with an unlinked subterm") end elseif id or otherlang then if id and expanded_dest:find("[", nil, true) then error("Can't supply <id:...> with destination with embedded brackets") end subterm_replacement = require(links_module).language_link { lang = otherlang or lang, term = expanded_dest, alt = subterm, id = id, } elseif expanded_dest:find("[", nil, true) then -- Use the destination directly if it has brackets in it (e.g. to put brackets around parts of a word). subterm_replacement = expanded_dest elseif expanded_dest == subterm then subterm_replacement = "[[" .. subterm .. "]]" else subterm_replacement = "[[" .. expanded_dest .. "|" .. subterm .. "]]" end local escaped_subterm_replacement = replacement_escape(subterm_replacement) local replaced_linked_term = ugsub(linked_term, subterm_re, escaped_subterm_replacement) if replaced_linked_term == linked_term then mw.log(("Attempted to replace %s with %s in %s"):format(subterm_re, escaped_subterm_replacement, linked_term)) error(("Subterm '%s' could not be located in %slinked expression %s, or replacement same as subterm"):format( subterm, j > 1 and "intermediate " or "", escape_wikicode(linked_term))) else linked_term = replaced_linked_term end end end return linked_term end local inflection_to_cats = { plural = { filter_plpos = function(plpos) -- plurals also occur with determiners, adjectives etc. and we don't want to generate categories like -- 'countable determiners', 'countable adjectives', etc. Note that the passed-in `plpos` has `proper nouns` -- converted to `nouns`. return plpos == "nouns" end, cats = {"countable {plpos}"}, no_cats = {"uncountable {plpos}"}, }, comparative = { cats = {"comparable {plpos}"}, no_cats = {"uncomparable {plpos}"}, }, ["female equivalent"] = { cats = {"{plpos} with other-gender equivalents"}, }, ["male equivalent"] = { cats = {"{plpos} with other-gender equivalents"}, }, } --[=[ Validate the items in `items` against the list or set of valid items in `valid_items`. If `field` is given, fetch the item to check from that-named field of each object in `items`; otherwise use the items in `items` directly. If an error occurs, `item_type` specifies the type of item to mention in the error message, which will also list the allowed items (either taken directly from `valid_items` if a list, or from the sorted keys if a set). ]=] local function validate_items(data) local items, field, valid_items, item_type = data.items, data.field, data.valid_items, data.item_type local valid_set if valid_items[1] then valid_set = list_to_set(valid_items) else valid_set = valid_items end for _, item in ipairs(items) do if field then item = item[field] end if not valid_set[item] then local valid_list if valid_items[1] then valid_list = valid_items else valid_list = {} for valid_item, _ in pairs(valid_items) do insert(valid_list, valid_item) end table.sort(valid_list) end error(("Invalid %s: %s; expected one of %s"):format(item_type, item, mw.text.listToText(valid_list))) end end end local Headdata = {} function Headdata:get_canonicalized_plpos() return (self.pos_category:gsub("proper noun", "noun")) end --[==[ Canonicalize a category. The category string will have the full language name (i.e. the name of the L2 language under which an entry is inserted, which may a parent language if the language in question is an etymology-only language) prepended to it, and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech (with some canonicalization; specifically, `proper nouns` is converted to `nouns` when replacing `{plpos}`). To specify a full category and not have the language name prepended to it, precede it with {"Category:"}, which will be removed. ]==] function Headdata:canonicalize_category(category) if category:find("{plpos}") then local plpos = self:get_canonicalized_plpos() category = category:gsub("{plpos}", plpos) end if category:find("^Category:") then return (category:gsub("^Category:", "")) else return self.langfullname .. " " .. category end end --[==[ Canonicalize a list of categories according to the process described in `Headdata:canonicalize_category`. This simply loops over each category in `categories` and calls `Headdata:canonicalize_category` on each one. ]==] function Headdata:canonicalize_categories(categories) if not categories then return categories end local canon_cats = {} for _, cat in ipairs(categories) do insert(canon_cats, self:canonicalize_category(cat)) end return canon_cats end --[==[ Insert a category into the `categories` list in the headword `data` structure. `category` is normally a string naming the category, which will have the full language name prepended to it and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech (with some canonicalization; specifically, `proper nouns` is converted to `nouns` when replacing `{plpos}`). To specify a full category and not have the language name prepended to it, precede it with {"Category:"}. ]==] function Headdata:insert_category(category) insert(self.categories, self:canonicalize_category(category)) end --[==[ Validate the genders in `genders` (a list of gender spec objects, as produced by {type = "genders"} in [[Module:parameters]] and accepted by [[Module:gender and number]]), checking that all specified genders are in the list given in `valid_genders`. Optional `props` controls how the validation happens. In particular, unless `props.no_augment` is given, then for any gender beginning with `m`, if a corresponding gender beginning with `f` occurs, analogous genders beginning with `mf`, `mfbysense` and `mfequiv` are also allowed. For example, if `m-d` (masculine dual) and `f-d` (feminine dual) both occur, genders `mf-d`, `mfbysense-d` and `mfequiv-d` are also allowed. If a disallowed gender is given, an error occurs, giving the disallowed gender along with the list of all allowed genders. ]==] function Headdata:validate_genders(genders, valid_genders, props) if not genders then return end props = props or {} local gender_type, no_augment = props.gender_type, props.no_augment gender_type = gender_type or "headword" local valid_gender_set = list_to_set(valid_genders) local augmented_gender_set if no_augment then augmented_gender_set = valid_gender_set else augmented_gender_set = {} for g, _ in pairs(valid_gender_set) do augmented_gender_set[g] = true if g:find("^m") and not g:find("^mf") and valid_gender_set[g:gsub("^m", "f")] then augmented_gender_set[g:gsub("^m", "mf")] = true augmented_gender_set[g:gsub("^m", "mfbysense")] = true augmented_gender_set[g:gsub("^m", "mfequiv")] = true end end end validate_items { items = genders, field = "spec", valid_items = augmented_gender_set, item_type = ("%s gender"):format(gender_type), } end --[==[ Parse an inflection specified in `field`, the name of a parameter holding an inflection. If the parameter is numeric, the field should be given as a number (as with the `params` structure passed to [[Module:parameters]]), not a string containing the representation of a number. The field can specify multiple comma-separated terms, and each term can have associated inline modifiers that will be parsed (unless there is top-level HTML in the parameter, i.e. HTML not contained inside an inline modifier, e.g. as may be generated by using {{tl|l}} or similar template inside a parameter). This is a wrapper around the top-level `parse_term_with_modifiers()` function. `props` is an optional structure containing additional properties, including all additional properties documented for the top-level `parse_term_with_modifiers()` function. If the parameter in `field` is unspecified, the return value of this function will be an empty list, not {nil}, so it is always safe to iterate over the return value. By default, the allowed modifiers are the same as for `parse_term_with_modifiers()`, except that (normally) the `<tr:...>` modifier will be allowed if `include_tr` was specified in the original call to `process_headword()`; likewise for the `<ts:...>` modifier if `include_ts` was specified and the `<sc:...>` modifier if `include_sc` was specified. If If you pass in your own `include_mods` list of additional allowed modifiers, it will (normally) automatically be augmented with {"tr"}, {"ts"} and/or {"sc"} if `include_tr`, `include_ts` and/or `include_sc` was specified when calling `process_headword()`. To disable automatic augmentation of these modifiers (whether or not you specify an `include_mods` property), specify {no_augment_include_mods = true} in `props`. ]==] function Headdata:parse_inflection(field, props) local val = self.process_props.args[field] if not val then return {} end props = props and shallow_copy(props) or {} local include_mods = props.include_mods local data = self.process_props.data if not props.no_augment_include_mods and (data.include_tr or data.include_ts or data.include_sc) then include_mods = include_mods and shallow_copy(include_mods) or {} if data.include_tr then insert_if_not(include_mods, "tr") end if data.include_ts then insert_if_not(include_mods, "ts") end if data.include_sc then insert_if_not(include_mods, "sc") end end props.val = val props.paramname = field props.splitchar = props.splitchar or "," props.include_mods = include_mods return export.parse_term_with_modifiers(props) or {} end --[==[ Insert previously-parsed terms into the `inflections` of the headword `data` structure. This is a wrapper around the top-level `insert_inflection()` function. `terms` is the list of parsed terms. (If {nil}, nothing happens unless `request` is set in `props`.) `label` is the the label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) `props` is an optional structure containing additional properties, including all additional properties documented for the top-level `insert_inflection()` function. Unless `no_auto_cats` is given in `props`, certain labels automatically trigger the insertion of additional categories in specific circumstances. This is controlled by the `inflection_to_cats` structure in [[Module:headword utilities]]. For example, if the part of speech is {"nouns"} or {"proper nouns"} and the label (after removing any links and `<<...>>` glossary specs) is {"plural"}, an additional category <code><var>lang</var> countable nouns</code> will be added if a plural value is given (i.e. the value is not {"-"}). If the value is {"-"} (which indicates that there is no plural and triggers the insertion of the fixed inflection label {"no plural"}), <code><var>lang</var> uncountable nouns</code> will be inserted instead, and if both {"-"} and a value are given (which triggers the insertion of the {"usually no plural"} fixed inflection label), both categories are added. Similar categories are inserted when a comparative is given (with a label {"comparative"}), and if the label is {"female equivalent"} or {"male equivalent"} and the value is not {"-"}, a category such as <code><var>lang</var> nouns with other-gender equivalents</code> is inserted. ]==] function Headdata:insert_inflection(terms, label, props) props = props and shallow_copy(props) or {} if not props.no_auto_cats then local bare_label = label if bare_label:find("[[", nil, true) then bare_label = require(links_module).remove_links(bare_label) end if bare_label:find("<<", nil, true) then bare_label = bare_label:gsub("<<.-|(.-)>>", "%1"):gsub("<<(.-)>>", "%1") end local cats = inflection_to_cats[bare_label] if cats then if not cats.filter_plpos or cats.filter_plpos(self:get_canonicalized_plpos()) then if props.cats == nil then props.cats = self:canonicalize_categories(cats.cats) end if props.usually_no_cats == nil then props.usually_no_cats = self:canonicalize_categories(cats.usually_no_cats) end if props.no_cats == nil then props.no_cats = self:canonicalize_categories(cats.no_cats) end end end end props.headdata = self props.terms = terms props.label = label return export.insert_inflection(props) end --[==[ Insert a "fixed" inflection (a label without associated values) into the `inflections` table of the headword `data` structure, labeled according to `label` (which can have glossary links in it specified using `<<...>>`, exactly as for `:insert_inflection()`). An example label (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"}, specifying that the noun belongs to the hidden-''g'' declension. This is a direct wrapper around the top-level function `insert_fixed_inflection()`; see that function for more details on optional `props`. ]==] function Headdata:insert_fixed_inflection(label, props) props = props and shallow_copy(props) or {} props.headdata = self props.label = label export.insert_fixed_inflection(props) end --[==[ Parse the inflection(s) specified in `field` and insert them into the `inflections` table of the headword `data` structure, labeled according to `label`. This is equivalent to calling {terms = data:parse_inflection(field, props)} followed by {return data:insert_inflection(terms, label, props)} and behaves the same as the combination of those two functions. See their documentation for more details. ]==] function Headdata:parse_and_insert_inflection(field, label, props) local terms = self:parse_inflection(field, props) return self:insert_inflection(terms, label, props) end --[==[ Generate an inflection that may be specified explicitly or defaulted (which involves looping over the specified or defaulted heads and determining the script of each one, since the formation of the default depends on the script). `data` is the data object passed into the POS handler. `terms` is the list of terms to process. Those where the term itself is not `+` will be returned unchanged, while those where the term is `+` will be handled by generating the appropriate inflections from the headwords using `make_inflection` (which is passed three arguments, `head`, `tr` and `sccode`, i.e. the script code of `head`) and should return two values, term and translit, either of which can be nil. A nil head will be ignored, and otherwise the decorations specified on the `+` term will be combined with the decorations specified on the head. The return value is a list of inflections where no requests for the default inflection remain. ]==] function Headdata:resolve_special(terms, handle_special, props) props = props or {} local infls = {} local is_special = props.is_special or function(_data, infl) return infl.term == "+" end for _, termobj in ipairs(terms) do if not is_special(self, termobj) then insert(infls, termobj) else for _, headobj in ipairs(self.heads) do local head = headobj.term or self.pagename local head_no_links if props.with_links then head = head:find("%[") and head or require(headword_module).add_multiword_links(head, not headobj.term) head_no_links = require(links_module).remove_links(head) else head = require(links_module).remove_links(head) head_no_links = head end local newterms = handle_special { head = head, tr = headobj.tr, infl = termobj, sc = self.lang:findBestScript(head_no_links), } if newterms then newterms = export.canonicalize_termobj_list(newterms, "term", "resolve_special") for _, newterm in ipairs(newterms) do if not props.no_combine_handle_special_retval_with_origin then export.combine_termobj_decorations(newterm, termobj) end if not props.no_combine_handle_special_retval_with_head then export.combine_termobj_decorations(newterm, headobj) end insert(infls, newterm) end end end end end return infls end --[==[ Add the current page to a tracking page named `Wiktionary:Tracking/``lang``-headword/``page```, where ``lang`` is the language code of the current language. For example, if the current language is `mak` and `page` is {"redundant-lon"}, the current page will get added to the tracking page `Wiktionary:Tracking/mak-headword/redundant-lon`. All pages added to that tracking page can be seen by going to [[Special:WhatLinksHere/Wiktionary:Tracking/mak-headword/redundant-lon]]. This is typically used to track issues occurring in user-specified parameters that do not rise to the level of errors (e.g. redundant parameters, deprecated usages or other dispreferred values). ]==] function Headdata:track(page) return require(debug_track_module)(self.langcode .. "-headword/" .. page) end local boolean_param = {type = "boolean"} --[==[ Process an arbitrary headword in an arbitrary language, handling generic and language-specific arguments and calling `full_headword()` in [[Module:headword]]. This is intended for use in implementing headword modules (e.g. [[Module:uz-headword]] for Uzbek, [[Module:mn-headword]] for Uzbek, [[Module:gsw-headword]] for Alemannic German, etc.) and provides a general implementation of such modules. On input, `data` is an object with the following fields: * `lang`: The language object of the language being handled. '''Required.''' Use the special value {false} to indicate that the language is specified by the user in {{para|1}}. * `frame`: The frame object passed into the `show()` function of your module, which implements headword-handling for all parts of speech in the module, including a generic POS-handling template (e.g. {{tl|uz-head}} or {{tl|mn-head}}), which allows arbitrary parts of speech to be handled. '''Required.''' * `pos_functions`: A table listing, for each part of speech requiring special handling, the extra parameters (if any) that the part of speech accepts, along with how to handle them. See the examples below. '''Required.''' * `validate_lang`: If {lang = true} is specified, this is a function of one argument (a language object, based on the language specified in {{para|1}}) that should throw an error if the language object is disallowed. If omitted, all languages are allowed. * `numbered_head`: If true, explicit headwords are specified in a numbered param instead of in {{para|head}}. The param used is usually {{para|1}}, but is {{para|2}} for generic POS templates such as {{tl|mn-head}} or for POS-specific templates when {lang = true} is specified (e.g. {{tl|arb-noun}}), and is {{para|3}} for generic POS templates when {lang = true} is spcified (e.g. {{tl|arb-head}}). * `include_tr`: If true, allow explicit transliterations to be specified. The transliteration(s) for the headword(s) themselves is/are specified in {{para|tr}} or through the {{cd|<tr:...>}} inline modifier on headwords, and transliterations of inflections are specified through the {{cd|<tr:...>}} inline modifier. This should generally be given when a headword for the language may be in a script other than Latin. * `include_ts`: If true, allow explicit transcriptions to be specified. The transcription(s) for the headword(s) themselves is/are specified in {{para|ts}} or through the {{cd|<ts:...>}} inline modifier on headwords, and transcriptions of inflections are specified through the {{cd|<ts:...>}} inline modifier. This should generally only be given for certain languages where the spelling is radically different from the pronunciation (e.g. in cuneiform languages such as Hittite and Akkadian, and potentially in Tibetan), and represents a pronunciation-based rendering (usually not direct IPA). * `include_sc`: If true, allow an explicit script code to be specified. The overall script code for the headword(s) themselves can be specified using {{para|sc}}, and per-headword or per-inflection script codes are specified using the {{cd|<sc:...>}} inline modifier. This should generally be given when a language supports multiple scripts. * `infls`: An inflections structure specifying extra generic parameters that apply to all parts of speech and how to handle them. The format is the same as for the `infls` structure in `pos_functions`. * `augment_params`: A callback function to add extra generic parameters, or modify existing generic parameters in the `params` structure; but in general, extra generic parameters should be added through the `infls` structure instead. This callback should not be used to add part-of-speech-specific parameters; those are handled through the appropriate setting in `pos_functions`. It is passed two single arguments, the `headdata` object and the `params` table to be augmented. See below for extra fields stored in the `process_props` structure of the `headdata` object. This function is called after initializing the `params` table and processing the overall `infls` structure, but just before adding part-of-speech-specific parameters (from `pos_functions`) to `params`. Thus, it can override any generic parameters but may itself be overridden by a part-of-speech-specific parameter. * `augment_headdata`: A callback function to modify the `headdata` object passed to `full_headword()` in [[Module:headword]]. This should not be used to for part-of-speech-specific parameter handling; this is handled through the appropriate setting in `pos_functions`. It is passed a single argument, the `headdata` object, as for `augment_params`; but the `process_props` structure and other fields will be more filled out, as this callback is called later. This can be used, for example, to override the value of a generic setting in `headdata` (e.g. [[Module:uz-headword]] uses this to mark non-Latin terms as variant forms by setting `headdata.var`, being careful not to override a value already set by the user). This function is called after initializing the `headdata` table with all information taken from generic parameters and processing generic parameters specified in the overall `infls` structure, and just before processing the appropriate part-of-speech-specific `infls` structure in `pos_functions` (which in turn is followed by any handler function in `pos_functions`). Thus, it can override any value set during generic parameter processing but may itself be overridden by a part-of-speech-specific parameter or handler. * `force_cat`: If true, add the headword to the appropriate categories even on non-mainspace pages. This can be used for testing category handling in sample template calls on userspace test pages or template documentation pages. It should not be set in production code. * `enable_auto_translit`: If true, turn on automatic transliteration of inflections at a global level (i.e. applying to all inflections). This has no effect on headwords, which are automatically transliterated by default if in a non-Latin script and automatic transliteration is available for the language. You can also set this value for particular inflections in the `insert_inflection()` function. The `headdata` headword data structure has an extra field in it called `process_props` that is specific to the `process_headword()` function, containing various extra properies. As the operation of `process_headword()` proceeds, this object gets filled out with more fields. For example, once parameter parsing happens, the resulting values are available in the `args` field of `process_props`. The following fields are found in `process_props` (note that `poscat`, the canonicalized plural part of speech of the headword being processed, is *not* present here; it's directly on `headdata`): * `namespace`: The name of the current namespace; an empty string for the mainspace. This references the namespace of the actual page and isn't affected by the {{para|pagename}} parameter. * `indexing_poscat`: The canonicalized part of speech of the headword used to index into `pos_functions`. This is the same as `poscat` for specific part-of-speech templates such as {{tl|uz-noun}}, but has the value {"head"} for generic part-of-speech templates such as {{tl|uz-head}}. (Note that `poscat` is directly available on `headdata`.) * `generic_pos_template`: True if a generic POS templates like {{tl|uz-head}} or {{tl|mn-head}} was used. (This is signaled by omitting the invocation parameter {{para|1}} to `process_headword`.) * `lang_in_1`: True if the language code is to be fetched from {{para|1}}. * `pos_param`: The parameter holding the part of speech, if a generic POS tempalte like {{tl|mn-head}} is being processed (i.e. `generic_pos_template` is set). In such a case, it will have the value of {1} or {2}, depending on whether the language code is being fetched from {{para|1}} (see `lang_in_1`). Otherwise it will be {nil}. * `head_param`: The parameter holding the explicit headword. If `numbered_head` was specified (as for Mongolian headword templates), this has the value {1}, {2} or {3} depending on whether the language code is being fetched from {{para|1}} (see `lang_in_1`) and whether a generic POS template like {{tl|mn-head}} is being processed (see `generic_pos_template`). Otherwise, it has the value {"head"}. Also see the `lang` and `numbered_head` properties in the `data` structure sent to `process_headword()`. * `is_suffix`: True if the current term is a suffix. This is set when processing the `suffix`, `nosuffix` and `clitic` parameters; it is always {false} beforehand (i.e. during `augment_params` and processing of the general `infls` structure). * `insert_specs`: This is a table mapping parameter names to the return value of `Headdata:insert_inflection()`, filled out as parameter values are processed. This lets a given parameter processing function in `infls` gain access to the result of calling `insert_inflection()` on previous parameters (which indicates the number of items inserted as well as whether `-` was specified). The `pos_functions` table contains an entry for each part of speech needing special handling, where the key is the canonical plural part of speech (e.g. {"adverbs"} or {"proper nouns"}). The value associated with each key is a table normally containing a field `infls`, listing the extra part-of-speech-specific inflection and other parameters along with how to handle them. The specs in `infls` are used in three ways: # to augment the `params` object passed to the `process()` function in [[Module:parameters]], specifying how to parse the appropriate inflectional parameters; # to specify how to process any inflectional parameters given and insert them into the `headdata` object passed to `full_headword()` in [[Module:headword]]; # to generate appropriate documentation for the parameters and other changes made by the headword template (e.g. inserting categories). Alternatively, you can separately control the augmentation of the `params` object and the procesing of the resulting arguments. This is done by specifying two fields in place of `infls`, named `params` and `func`. `params` is a table containing extra parameters to add to the overall `params` object passed to the `process()` function in [[Module:parameters]]. `func` is a function of two arguments, normally called `data` (the headword data structure `headdata`) and `args` (the processed arguments table). However, this alternative method is not normally recommended because it leads to duplication between the `params and `func` fields and the documentation, which must be manually specified. A simple example, as used to handle pronouns for Turoyo, is { local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"} pos_functions["pronouns"] = { infls = { {2, type = "genders", validate = valid_genders}, {"f", label = "feminine"}, {"pl", label = "plural"}, }, } } The equivalent using `params` and `func` is { local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"} pos_functions["pronouns"] = { params = { [2] = {type = "genders"}, f = true, pl = true, }, func = function(data, args) data:validate_genders(args[2], valid_genders) data.genders = args[2] data:parse_and_insert_inflection("f", "feminine") data:parse_and_insert_inflection("pl", "plural") end } } Note how the version with separate `params and `func` is longer and splits information on the parameters between the two fields. The `params` structure sets extra user-specifiable parameters {{para|2}} for genders (since the headword is in {{para|1}}) as well a {{para|f}} and {{para|pl}}, and the `func` handler processes those parameters. Note how this is done by calling methods on the headword `data` structure. Each such parameter can have multiple comma-separated values, and each value can have inline modifiers attached to it to specify further properties of the value. The `infls` version ends up making the same method calls, but does it for you instead of you having to do it yourself. These methods are implemented through a metatable set on the headword `data` structure, which is removed before calling `full_headword()` in [[Module:headword]]. The methods access extra information related to headword processing (such as the `args` table) that is stored in the `process_props` field of the headword `data` strucuture. This field is also removed prior to calling `full_headword()`. The methods available on the headword `data` structure are as follows. Each one also has its own documentation. * {parse_inflection(field, props)}: Parse value(s) specified in `field` (a user-specified parameter in the `args` table) and return a list of term objects. Optional `props` specifies additional properties controlling the parsing. * {insert_inflection(terms, label, props)}: Insert the terms in `terms` (a list of term objects as returned by `parse_inflection()`) into the `inflections` list in the headword `data` structure, giving the inflection the label as specified in `label`. Optional `props` specifies additional properties controlling the parsing. * {parse_and_insert_inflection(field, label, props)}: A combination of `parse_inflection()` and `insert_inflection()`, if no further processing of the parsed values needs to be done before insertion. * {insert_fixed_inflection(label, props)}: Insert a "fixed" inflection (a label without associated values) into the `inflections` table. An example (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"}, specifying that the noun belongs to the hidden-''g'' declension. * {resolve_special(terms, handle_special, props)}: Resolve "special" indicators as specified by the user in an inflection parameter. A typical example is {"+"}, requesting a default value. `terms` is the list of parsed term objects and `handle_special` is a handler function to process special indicators and convert them to their actual values. * {validate_genders(genders, valid_genders, props)}: Validate that the user-specified genders in `genders` all belong to the list given in `valid_genders`, throwing an error if not. * {insert_category(category)}: Insert a category into the `categories` list in the headword `data` structure. `category` is normally a string naming the category, which will have the language prepended to it and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech. ]==] function export.process_headword(data) local lang, frame, pos_functions, validate_lang, numbered_head, include_tr, include_ts, include_sc, force_cat, enable_auto_translit, infls, augment_params, augment_headdata = data.lang, data.frame, data.pos_functions, data.validate_lang, data.numbered_head, data.include_tr, data.include_ts, data.include_sc, data.force_cat, data.enable_auto_translit, data.infls, data.augment_params, data.augment_headdata local iparams = { [1] = true, def = true, } local iargs = require(parameters_module).process(frame.args, iparams) local parargs = frame:getParent().args local langcode if not lang then error("Internal error: `data.lang` must be specified; either a language object or `true` for a user-specified language") end local lang_in_1 if lang == true then lang_in_1 = true langcode = ine(parargs[1]) if langcode then langcode = mw.text.trim(langcode) lang = require(languages_module).getByCode(langcode, 1, true) if validate_lang then validate_lang(lang) end else error("Language code (see [[WT:Language codes]]) must be specified in 1=") end else langcode = lang:getCode() if validate_lang then error("Internal error: `data.validate_lang` must not be specified if a language code is given in `data.lang`") end end local poscat = iargs[1] local generic_pos_template = not poscat local pos_param if generic_pos_template then pos_param = lang_in_1 and 2 or 1 poscat = ine(parargs[pos_param]) or mw.title.getCurrentTitle().fullText == ("Template:%s-head"):format(langcode) and "interjection" or error(("Part of speech must be specified in %s="):format(pos_param)) poscat = require(headword_module).canonicalize_pos(poscat) end local head_param = numbered_head and (generic_pos_template and lang_in_1 and 3 or (generic_pos_template or lang_in_1) and 2 or 1) or "head" local indexing_poscat = generic_pos_template and "head" or poscat local namespace = mw.loadData(headword_data_module).page.namespace -- Partly initialize headdata now for use in generic infls callbacks. Will be further initialized later after -- processing parameters. local headdata = { lang = lang, langcode = langcode, langfullcode = lang:getFullCode(), langname = lang:getCanonicalName(), langfullname = lang:getFullName(), process_props = { namespace = namespace, data = data, indexing_poscat = indexing_poscat, generic_pos_template = generic_pos_template, lang_in_1 = lang_in_1, pos_param = pos_param, head_param = head_param, is_suffix = false, insert_specs = {}, }, pos_category = poscat, orig_poscat = poscat, -- preserve user-specified poscat in case pos_category is changed to 'suffixes' categories = {}, inflections = {enable_auto_translit = enable_auto_translit}, force_cat_output = force_cat, no_redundant_head_cat = true, } setmetatable(headdata, {__index = Headdata}) local params = { [head_param] = {template_default = iargs.def}, head2 = {replaced_by = false, instead = ("use comma-separated |%s="):format(head_param)}, id = true, sort = true, cat = true, nolink = boolean_param, nolinkhead = {type = "boolean", alias_of = "nolink"}, suffix = boolean_param, nosuffix = boolean_param, clitic = true, addlpos = true, var = {type = "boolean", allow = {"both"}}, json = boolean_param, pagename = true, -- for testing } if include_sc then params.sc = {type = "script"} end if include_tr then params.tr = true params.tr2 = {replaced_by = false, instead = "use comma-separated |tr= or <tr:...> inline modifier on head"} end if include_ts then params.ts = true params.ts2 = {replaced_by = false, instead = "use comma-separated |ts= or <ts:...> inline modifier on head"} end if lang_in_1 then params[1] = {required = true} -- required but ignored as already processed above end if generic_pos_template then params[pos_param] = {required = true} -- required but ignored as already processed above end local function resolve_prop(prop, ...) if type(prop) == "function" then prop = prop(headdata, ...) end return prop end local function augment_params_from_infls(infls) infls = resolve_prop(infls) for _, infl in ipairs(infls) do local function interr(txt) error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl))) end local param = infl[1] if param then param = resolve_prop(param) if type(param) ~= "string" and type(param) ~= "number" then interr(("Parameter name %s must be a string or number"):format(dump(param))) end -- We handle defaults as well as validation ourselves. local typ = resolve_prop(infl.type) or "string" if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then -- FIXME: Handle more types. interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format( dump(typ))) end params[param] = {type = typ, required = resolve_prop(infl.required), template_default = resolve_prop(infl.template_default)} if typ ~= "boolean" and type(param) == "string" then params[param .. "2"] = {replaced_by = false, instead = ("use comma-separated |%s="):format(param)} end end end end if infls then augment_params_from_infls(infls) end if augment_params then augment_params(headdata, params) end if pos_functions[indexing_poscat] then local pos_infls = pos_functions[indexing_poscat].infls if pos_infls then augment_params_from_infls(pos_infls) end local pos_params = pos_functions[indexing_poscat].params if pos_params then for key, val in pairs(pos_params) do params[key] = val end end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData(headword_data_module).pagename local sc = args.sc or lang:findBestScript(pagename) headdata.pagename = pagename headdata.process_props.args = args headdata.sc = sc headdata.id = args.id headdata.sort = args.sort -- No redundant script cat unless the user explicitly gave sc= headdata.no_script_code_cat = not args.sc headdata.var = args.var local extra_term_mods = {} if include_tr then insert(extra_term_mods, "tr") end if include_ts then insert(extra_term_mods, "ts") end if include_sc then insert(extra_term_mods, "sc") end if not extra_term_mods[1] then extra_term_mods = nil end local trs = args.tr and split_on_comma(args.tr) or {} local num_trs = #trs local tss = args.ts and split_on_comma(args.ts) or {} local num_tss = #tss local heads = args[head_param] and export.parse_term_with_modifiers { val = args[head_param], paramname = head_param, splitchar = ",", is_head = true, include_mods = extra_term_mods, } or {} local num_heads = #heads if num_heads > 0 and num_trs > 0 and num_heads ~= num_trs then error(("%s head%s specified explicitly but %s translit%s; they must match; use '+' to stand for the default head (the pagename) or default automatic translit and '-' to stand for no translit"):format( num_heads, num_heads > 1 and "s" or "", num_trs, num_trs > 1 and "s" or "")) end if num_heads > 0 and num_tss > 0 and num_heads ~= num_tss then error(("%s head%s specified explicitly but %s transcription%s; they must match; use '+' to stand for the default head (the pagename) and '-' to stand for no transcription"):format( num_heads, num_heads > 1 and "s" or "", num_tss, num_tss > 1 and "s" or "")) end if num_trs > 0 and num_tss > 0 and num_trs ~= num_tss then error(("%s translit%s specified explicitly but %s transcription%s; they must match; use '+' to stand for default automatic translit and '-' to stand for no translit or transcription"):format( num_trs, num_trs > 1 and "s" or "", num_tss, num_tss > 1 and "s" or "")) end -- Be careful here not to overwrite user_specified_heads if it's empty so we can later check user_specified_heads -- to see if the user provided any heads. local max_tr_ts = math.max(num_trs, num_tss) if num_heads == 0 and max_tr_ts > 0 then heads = {} for i = 1, max_tr_ts do heads[i] = {term = "+"} end end if not heads[1] then heads = {{term = "+"}} end for i, headobj in ipairs(heads) do if headobj.tr and trs[i] then if headobj.tr ~= trs[i] then error(("Saw two different translits '%s' and '%s' for head #%s"):format( headobj.tr, trs[i], i)) end else headobj.tr = headobj.tr or trs[i] end if headobj.tr == "+" then headobj.tr = nil end if headobj.ts and tss[i] then if headobj.ts ~= tss[i] then error(("Saw two different transcriptions '%s' and '%s' for head #%s"):format( headobj.ts, tss[i], i)) end else headobj.ts = headobj.ts or tss[i] end if headobj.ts == "-" then headobj.ts = nil end if headobj.term == "+" then headobj.term = args.nolink and pagename or nil if headobj.term and namespace == "Reconstruction" then headobj.term = "*" .. headobj.term end end end headdata.heads = heads local function pagename_is_suffix() if sc:getCode() == "Latn" then -- shortcut Latin terms to avoid unnecessarily loading [[Module:affix]] return pagename:find("^%-") and not pagename:find("%-$") else local affix_type, _, _, _ = require(affix_module).parse_term_for_affixes(pagename, lang, sc) return affix_type == "suffix" end end local clitic_label if args.clitic then clitic_label = require(yesno_module)(args.clitic, args.clitic) end if clitic_label == true then clitic_label = "clitic" end if clitic_label then headdata:insert_category("clitics") headdata:insert_fixed_inflection(clitic_label) elseif args.suffix or ( not args.nosuffix and pagename_is_suffix() and poscat ~= "suffixes" and poscat ~= "suffix forms" ) then headdata.process_props.is_suffix = true local function handle_suffix_pos(pos, is_first) local form_type = pos:match("^(.*) forms$") local actual_poscat if form_type then headdata:insert_category(("%s suffix forms"):format(form_type)) headdata:insert_fixed_inflection(form_type .. " suffix form") else local singular_pos = require(en_utilities_module).singularize(pos) headdata:insert_category(("%s-forming suffixes"):format(singular_pos)) headdata:insert_fixed_inflection(singular_pos .. "-forming suffix") end local postype = require(headword_module).pos_lemma_or_nonlemma(pos) if not postype then error(("Unrecognized canonicalized part of speech '%s' in addlpos=, cannot determine whether lemma or non-lemma form"):format( pos )) end actual_poscat = postype == "lemma" and "suffixes" or "suffix forms" if is_first then headdata.pos_category = actual_poscat elseif headdata.pos_category ~= actual_poscat then error(("Cannot mix suffixes and suffix forms using addlpos=; '%s' is a %s while overall POS '%s' is a %s; use separate POS headers for the two"): format(pos, actual_poscat, poscat, headdata.pos_category)) end end handle_suffix_pos(poscat, true) if args.addlpos then for _, addlpos in ipairs(split(args.addlpos, "%s*,%s*")) do addlpos = require(headword_module).canonicalize_pos(addlpos) handle_suffix_pos(addlpos, false) end end end if args.cat then for _, cat in ipairs(split_on_comma(args.cat)) do headdata:insert_category(cat) end end local function augment_headdata_from_infls(infls) infls = resolve_prop(infls) for _, infl in ipairs(infls) do local function interr(txt) error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl))) end local function process_labelobjs(labelobjs, originating_term, handle_labelobj) if labelobjs == nil then return end if type(labelobjs) ~= "string" and type(labelobjs) ~= "table" then interr(("Wrong type '%s' for label object(s) %s, expected string or table"):format( type(labelobjs), dump(labelobjs) )) end if type(labelobjs) == "string" or type(labelobjs) == "table" and not labelobjs[1] then labelobjs = {labelobjs} end for _, labelobj in ipairs(labelobjs) do local label, termobj if type(labelobj) == "string" then label = labelobj termobj = originating_term elseif type(labelobj) ~= "table" then interr(("Wrong type '%s' for label object %s, expected string or table"):format( type(labelobj), dump(labelobj) )) label = labelobj.term if type(label) ~= "string" then interr(("Wrong type '%s' for label %s from label object %s, expected string"):format( type(label), dump(label), dump(labelobj) )) end termobj = labelobj end handle_labelobj(label, termobj) end end local param = infl[1] if param then local vals -- Fetch the param and make sure it's a string or number. param = resolve_prop(param) if type(param) ~= "string" and type(param) ~= "number" then interr(("Parameter name %s must be a string or number"):format(dump(param))) end -- Fetch the type and validate. local typ = resolve_prop(infl.type) if typ == nil then typ = "string" end if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then -- FIXME: Handle more types. interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format( dump(typ))) end -- Fetch the value(s). if typ == "genders" or typ == "boolean" then vals = args[param] elseif typ == "string" then local parse_inflection_props = resolve_prop(infl.parse_inflection_props) local include_mods = resolve_prop(infl.include_mods) local no_augment_include_mods = resolve_prop(infl.no_augment_include_mods) if include_mods ~= nil or no_augment_include_mods ~= nil then if parse_inflection_props == nil then parse_inflection_props = {} else parse_inflection_props = shallow_copy(parse_inflection_props) end if include_mods ~= nil then parse_inflection_props.include_mods = include_mods end if no_augment_include_mods ~= nil then parse_inflection_props.no_augment_include_mods = no_augment_include_mods end end vals = headdata:parse_inflection(param, parse_inflection_props) -- Convert an empty list to nil for consistent checking below. if not vals[1] then vals = nil end else interr(("Unrecognized type '%s"):format(typ)) end -- If value(s) nil, fetch the default. if vals == nil and infl.default ~= nil then local default = resolve_prop(infl.default) if typ == "genders" then vals = export.canonicalize_termobj_list(default, "spec", "default") elseif typ == "boolean" then vals = default elseif typ == "string" then vals = export.canonicalize_termobj_list(default, "term", "default") else interr(("Unrecognized type '%s"):format(typ)) end end -- Resolve "special" values (special signals a string values, such as requesting the default with "+"). if vals ~= nil and infl.resolve_special then if typ ~= "string" then interr(("Cannot specify resolve_special= for type %s"):format(dump(typ))) end local resolve_special_props = resolve_prop(infl.resolve_special_props, vals) if infl.is_special ~= nil then if resolve_special_props == nil then resolve_special_props = {} else resolve_special_props = shallow_copy(resolve_special_props) end resolve_special_props.is_special = infl.is_special end vals = headdata:resolve_special(vals, infl.resolve_special, resolve_special_props) end -- Validate the value(s). if vals ~= nil and infl.validate ~= nil then if typ == "boolean" then interr('Cannot specify validate= when type is "boolean"') elseif type(infl.validate) == "function" then infl.validate(headdata, vals) elseif typ == "genders" then headdata:validate_genders(vals, infl.validate) elseif typ == "string" then validate_items { items = vals, field = "term", valid_items = infl.validate, item_type = ("values in |%s="):format(param), } else interr(("Unrecognized type '%s"):format(typ)) end end -- Run the process_after_parse handler, if it exists. if vals ~= nil and infl.process_after_parse ~= nil then local intentionally_nil vals, intentionally_nil = infl.process_after_parse(headdata, vals) if vals == nil and not intentionally_nil then interr("If you return nil from process_after_parse, you must return a second non-nil return " .. "value to indicate that the nil return value was intentional") end end -- "Implement" the values, if non-falsy (i.e. we don't want to fire on boolean false or empty list). -- If a fixed label is specified, insert it. Then, depending on the type, attach the values to a label -- as an inflection, set the `genders` field, or do nothing if boolean (throwing an error if there was -- no fixed label). if vals == true or type(vals) == "table" and vals[1] then if infl.fixed_label and infl.all_fixed_label then interr("Cannot specify both fixed_label= and all_fixed_label=; specify one or the other") end local function check_fixed_label_references_val(label) if type(label) == "table" and label[1] then for _, lab in ipairs(label) do if check_fixed_label_references_val(lab) then return true end end return false end if type(label) == "table" then if not label.term then interr(("Fixed label structure %s does not have a value for `.term`"):format(dump(label))) end label = label.term end if type(label) ~= "string" then interr(("Wrong type for fixed label %s, should be string"):format(type(label))) end return not not label:find("{val}") end local fixed_label = infl.fixed_label local all_fixed_label = infl.all_fixed_label -- If the value being processed is boolean, there's only one value so treat a fixed_label as an -- all_fixed_label and output only once; likewise if the caller specified a fixed_label without -- {val} in it. if fixed_label and (typ == "boolean" or type(fixed_label) ~= "function" and not check_fixed_label_references_val(fixed_label)) then all_fixed_label = fixed_label fixed_label = nil end local inserted_fixed_label if fixed_label then if typ == "boolean" then interr("Boolean fixed_label values should have been converted to all_fixed_label") end for _, valobj in ipairs(vals) do local labelobjs = resolve_prop(fixed_label, valobj) process_labelobjs(labelobjs, valobj, function(label, termobj) if label:find("{val}") then if typ ~= "string" then interr(('Cannot specify {val} in fixed_label %s when type is "%s"'):format(dump(label), typ)) end label = label:gsub("{val}", replacement_escape(valobj.term)) end headdata:insert_fixed_inflection(label, { originating_term = termobj }) inserted_fixed_label = true end) end elseif all_fixed_label then local labelobjs = resolve_prop(all_fixed_label, vals) process_labelobjs(labelobjs, nil, function(label, termobj) if label:find("{vals}") then if typ ~= "string" then interr(('Cannot specify {vals} in all_fixed_label %s when type is "%s"'):format(dump(label), typ)) end local formatted_labels = {} for _, valobj in ipairs(vals) do insert(formatted_labels, add_decorations(valobj.term, valobj, lang)) end label = label:gsub("{vals}", replacement_escape(serial_comma_join(formatted_labels))) end headdata:insert_fixed_inflection(label, termobj) inserted_fixed_label = true end) end local inserted_vals if infl.label ~= nil then if typ ~= "string" then interr(("label=%s can only be specified for type 'string', not '%s'"):format( dump(infl.label), typ )) end local label = resolve_prop(infl.label, vals) if label ~= nil then local insert_inflection_props = resolve_prop(infl.insert_inflection_props, vals) local no_auto_cats = resolve_prop(infl.no_auto_cats, vals) if no_auto_cats ~= nil then if insert_inflection_props == nil then insert_inflection_props = {} else insert_inflection_props = shallow_copy(insert_inflection_props) end insert_inflection_props.no_auto_cats = infl.no_auto_cats end local insert_spec = headdata:insert_inflection(vals, label, insert_inflection_props) headdata.process_props.insert_specs[param] = insert_spec inserted_vals = true end end if typ == "genders" then headdata.genders = vals end local inserted_cat if infl.cat then local allcats = {} for _, valobj in ipairs(vals) do local cats = resolve_prop(infl.cat, valobj) if type(cats) == "string" then cats = {cats} end if cats ~= nil then for _, cat in ipairs(cats) do if cat:find("{val}") then cat = cat:gsub("{val}", replacement_escape(valobj.term)) end insert_if_not(allcats, cat) end end end for _, cat in ipairs(allcats) do headdata:insert_category(cat) inserted_cat = true end end if typ == "boolean" then if not inserted_fixed_label and not inserted_cat then interr(("User set boolean setting for %s= but no fixed label added and no category " .. "inserted; if you took action in process_after_parse(), make sure to return " .. "`nil, true`"):format(param)) end elseif typ == "string" then if not inserted_vals and not inserted_fixed_label then interr(("User set value(s) %s for %s= but no inflection inserted and no fixed label " .. "added; if you took action in process_after_parse(), make sure to return " .. "`nil, true`"):format(dump(vals), param)) end end end else -- no param specified if infl.label or infl.all_fixed_label then interr("Cannot have label= or all_fixed_label= without specifying a param") end if infl.fixed_label then local labelobjs = resolve_prop(infl.fixed_label) process_labelobjs(labelobjs, nil, function(label, termobj) if label:find("{val}") then interr("Cannot specify {val} in a fixed_label= value without specifying a param") end headdata:insert_fixed_inflection(label, { originating_term = termobj }) end) end if infl.cat then local cats = resolve_prop(infl.cat) if type(cats) == "string" then cats = {cats} end if cats ~= nil then for _, cat in ipairs(cats) do if cat:find("{val}") then interr("Cannot specify {val} in a cat= value without specifying a param") end headdata:insert_category(cat) end end end end end end if infls then augment_headdata_from_infls(infls) end if augment_headdata then augment_headdata(headdata, args) end if pos_functions[indexing_poscat] then local pos_infls = pos_functions[indexing_poscat].infls if pos_infls then augment_headdata_from_infls(pos_infls) end local func = pos_functions[indexing_poscat].func if func then func(headdata, args) end end setmetatable(headdata, nil) if args.json then return require("Module:JSON").toJSON(headdata) end headdata.process_props = nil return require(headword_module).full_headword(headdata) end return export bn9v6026ivyevkq8uxmr7qwxoaz1tnb Module:affix 828 8054 54860 2026-09-26T23:31:13Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local debug_force_cat = false -- if set to true, always display categories even on userspace pages local m_links = require("Module:links") local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local etymology_module = "Module:etymology" local scripts_module = "Module:scripts" local utilities_module = "Modul...' 54860 Scribunto text/plain local export = {} local debug_force_cat = false -- if set to true, always display categories even on userspace pages local m_links = require("Module:links") local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local etymology_module = "Module:etymology" local scripts_module = "Module:scripts" local utilities_module = "Module:utilities" -- Export this so the category code in [[Module:category tree/etymology]] can access it. export.affix_lang_data_module_prefix = "Module:affix/lang-data/" local ulen = m_str_utils.len local rfind = m_str_utils.find local rmatch = m_str_utils.match local pluralize = require(en_utilities_module).pluralize local singularize = require(en_utilities_module).singularize local u = m_str_utils.char local ucfirst = m_str_utils.ucfirst local unpack = unpack or table.unpack -- Lua 5.2 compatibility function export.affix_variants(canonical, variants) local mappings = {} for _, variant in ipairs(variants) do mappings[variant] = canonical end return mappings end function export.id_mapping(default, ids) local mapping = { default = default } if ids then for id, target in pairs(ids) do mapping[id] = target end end return mapping end function export.id_mapping_with_affix_variants(base, id_variants) local mappings = {} for id, variants in pairs(id_variants) do for _, variant in ipairs(variants) do mappings[variant] = export.id_mapping(base, {[id] = base}) end end return mappings end function export.merge_tables(...) local result = {} for i = 1, select('#', ...) do local t = select(i, ...) if t then for k, v in pairs(t) do result[k] = v end end end return result end -- Export this so the category code in [[Module:category tree/etymology]] can access it. export.langs_with_lang_specific_data = { ["az"] = true, ["fi"] = true, ["fr"] = true, ["izh"] = true, ["la"] = true, ["sah"] = true, ["tr"] = true, ["trk-pro"] = true, } local default_pos = "term" local function pluralize_pos(pos) return pluralize(singularize(pos)) end --[==[ intro: ===About different types of hyphens ("template", "display" and "lookup"):=== * The "template hyphen" is the per-script hyphen character that is used in template calls to indicate that a term is an affix. This is always a single Unicode char, but there may be multiple possible hyphens for a given script. Normally this is just the regular hyphen character "-", but for some non-Latin-script languages (currently only right-to-left languages), it is different. * The "display hyphen" is the string (which might be an empty string) that is added onto a term as displayed and linked, to indicate that a term is an affix. Currently this is always either the same as the template hyphen or an empty string, but the code below is written generally enough to handle arbitrary display hyphens. Specifically: *# For East Asian languages, the display hyphen is always blank. *# For Arabic-script languages, either tatweel (ـ) or ZWNJ (zero-width non-joiner) are allowed as template hyphens, where ZWNJ is supported primarily for Farsi, because some suffixes have non-joining behavior. The display hyphen corresponding to tatweel is also tatweel, but the display hyphen corresponding to ZWNJ is blank (tatweel is also the default display hyphen, for calls to {{tl|prefix}}/{{tl|suffix}}/etc. that don't include an explicit hyphen). * The "lookup hyphen" is the hyphen that is used when looking up language-specific affix mappings. (These mappings are discussed in more detail below when discussing link affixes.) It depends only on the script of the affix in question. Most scripts (including East Asian scripts) use a regular hyphen "-" as the lookup hyphen, but Hebrew and Arabic have their own lookup hyphens (respectively maqqef and tatweel). Note that for Arabic in particular, there are three possible template hyphens that are recognized (tatweel, ZWNJ and regular hyphen), but mappings must use tatweel. ===About different types of affixes ("template", "display", "link", "lookup" and "category"):=== * A "template affix" is an affix in its source form as it appears in a template call. Generally, a template affix has an attached template hyphen (see above) to indicate that it is an affix and indicate what type of affix it is (prefix, suffix, interfix or circumfix), but some of the older-style templates such as {{tl|suffix}}, {{tl|prefix}}, {{tl|confix}}, etc. have "positional" affixes where the presence of the affix in a certain position (e.g. the second or third parameter) indicates that it is a certain type of affix, whether or not it has an attached template hyphen. * A "display affix" is the corresponding affix as it is actually displayed to the user. The display affix may differ from the template affix for various reasons: *# The display affix may be specified explicitly using the {{para|alt<var>N</var>}} parameter, the `<alt:...>` inline modifier or a piped link of the form e.g. `<nowiki>[[-kas|-käs]]</nowiki>` (here indicating that the affix should display as `-käs` but be linked as `-kas`). Here, the template affix is arguably the entire piped link, while the display affix is `-käs`. *# Even in the absence of {{para|alt<var>N</var>}} parameters, `<alt:...>` inline modifiers and piped links, certain languages have differences between the "template hyphen" specified in the template (which always needs to be specified somehow or other in templates like {{tl|affix}}, to indicate that the term is an affix and what type of affix it is) and the display hyphen (see above), with corresponding differences between template and display affixes. * A (regular) "link affix" is the affix that is linked to when the affix is shown to the user. The link affix is usually the same as the display affix, but will differ in one of three circumstances: *# The display and link affixes are explicitly made different using {{para|alt<var>N</var>}} parameters, `<alt:...>` inline modifiers or piped links, as described above under "display affix". *# For certain languages, certain affixes are mapped to canonical form using language-specific mappings. For example, in Finnish, the adjective-forming suffix {{m|fi|-kas}} appears as {{m|fi|-käs}} after front vowels, but logically both forms are the same suffix and should be linked and categorized the same. Similarly, in Latin, the negative and intensive prefixes spelled {{m|la|in-}} (etymologically two distinct prefixes) appear variously as {{m|la|il-}}, {{m|la|im-}} or {{m|la|ir-}} before certain consonants. Mappings are supplied in [[Module:affix/lang-data/LANGCODE]] to convert Finnish {{m|fi|-käs}} to {{m|fi|-kas}} for linking and categorization purposes. Note that the affixes in the mappings use "lookup hyphens" to indicate the different types of affixes, which is usually the same as the template hyphen but differs for Arabic scripts, because there are multiple possible template hyphens recognized but only one lookup hyphen (tatweel). The form of the affix as used to look up in the mapping tables is called the "lookup affix"; see below. * A "stripped link affix" is a link affix that has been passed through the language's `stripDiacritics()` function, which may strip certain diacritics: e.g. macrons in Latin and Old English (indicating length); acute and grave accents in Russian and various other Slavic languages (indicating stress); vowel diacritics in most Arabic-script languages; and also tatweel in some Arabic-script languages (currently, for example, Persian, Arabic and Urdu strip tatweel, but Ottoman Turkish does not). Stripped link affixes are currently what are used in category names. * A "lookup affix" is the form of the affix as it is looked up in the language-specific lookup mappings described above under link affixes. There are actually two lookup stages: *# First, the affix is looked up in a modified display form (specifically, the same as the display affix but using lookup hyphens). Note that this lookup does not occur if an explicit display form is given using {{para|alt<var>N</var>}} or an `<alt:...>` inline modifier, or if the template affix contains a piped or embedded link. *# If no entry is found, the affix is then looked up in a modified link form (specifically, the modified display form passed through the language's `stripDiacritics()` function, which strips out certain diacritics, but with the lookup hyphen re-added if it was stripped out, as in the case of tatweel in many Arabic-script languages). The reason for this double lookup procedure is to allow for mappings that are sensitive to the extra diacritics, but also allow for mappings that are not sensitive in this fashion (e.g. Russian {{m|ru|-ливый}} occurs both stressed and unstressed, but is the same prefix either way). * A "category affix" is the affix as it appears in categories such as [[:Category:Finnish terms suffixed with -kas| Category:Finnish terms suffixed with ''-kas'']]. The category affix is currently always the same as the stripped link affix. This means that for Arabic-script languages, it may or may not have a tatweel, even if the correponding display affix and regular link affix have a tatweel. As mentioned above, stripDiacritics() strips tatweel for Arabic, Persian and Urdu, but not for Ottoman Turkish. Hence affix categories for Arabic, Persian and Urdu will be missing the tatweel, but affix categories for Ottoman Turkish will have it. An additional complication is that if the template affix contains a ZWNJ, the display (and hence the link and category affixes) will have no hyphen attached in any case. ]==] ----------------------------------------------------------------------------------------- -- Template and display hyphens -- ----------------------------------------------------------------------------------------- --[=[ Per-script template hyphens. The template hyphen is what appears in the {{affix}}/{{prefix}}/{{suffix}}/etc. template (in the wikicode). See above. They key below is a script code, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong' and 'xwo-Mong' will match 'Mong'. The value below is a string consisting of one or more hyphen characters. If there is more than one character, the default hyphen must come last and a non-default function must be specified for the script in display_hyphens[] so the correct display hyphen will be specified when no template hyphen is given (in {{suffix}}/{{prefix}}/etc.). Script detection is normally done when linking, but we need to do it earlier. However, under most circumstances we don't need to do script detection. Specifically, we only need to do script detection for a given language if (a) the language has multiple scripts; and (b) at least one of those scripts is listed below or in display_hyphens. ]=] local ZWNJ = u(0x200C) -- zero-width non-joiner local template_hyphens = { -- This covers all Arabic scripts. See above. ["Arab"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen ["Aran"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen ["Hebr"] = "־", -- Hebrew-specific hyphen termed "maqqef" ["Mong"] = "᠊", -- FIXME! What about the following right-to-left scripts? -- Adlm (Adlam) -- Armi (Imperial Aramaic) -- Avst (Avestan) -- Cprt (Cypriot) -- Khar (Kharoshthi) -- Mand (Mandaic/Mandaean) -- Mani (Manichaean) -- Mend (Mende/Mende Kikakui) -- Narb (Old North Arabian) -- Nbat (Nabataean/Nabatean) -- Nkoo (N'Ko) -- Orkh (Orkhon runes) -- Phli (Inscriptional Pahlavi) -- Phlp (Psalter Pahlavi) -- Phlv (Book Pahlavi) -- Phnx (Phoenician) -- Prti (Inscriptional Parthian) -- Rohg (Hanifi Rohingya) -- Samr (Samaritan) -- Sarb (Old South Arabian) -- Sogd (Sogdian) -- Sogo (Old Sogdian) -- Syrc (Syriac) -- Thaa (Thaana) } -- Hyphens used when looking up an affix in a lang-specific affix mapping. Defaults to regular hyphen (-). The keys -- are script codes, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong' and 'xwo-Mong' -- will match 'Mong'. The value should be a single character. local lookup_hyphens = { ["Hebr"] = "־", -- This covers all Arabic scripts. See above. ["Arab"] = "ـ", ["Aran"] = "ـ", } -- Default display-hyphen function. local function default_display_hyphen(script, hyph) if not hyph then return template_hyphens[script] or "-" end return hyph end local function arab_get_display_hyphen(_script, hyph) if not hyph then return "ـ" -- tatweel elseif hyph == ZWNJ then return "" else return hyph end end local function no_display_hyphen(_script, _hyph) return "" end -- Per-script function to return the correct display hyphen given the script and template hyphen. The function should -- also handle the case where the passed-in template hyphen is nil, corresponding to the situation in -- {{prefix}}/{{suffix}}/etc. where no template hyphen is specified. The key is the script code after removing a hyphen -- and anything preceding, so 'mnc-Mong', 'xwo-Mong' etc. will match 'Mong'. local display_hyphens = { -- This covers all Arabic scripts. See above. ["Arab"] = arab_get_display_hyphen, ["Aran"] = arab_get_display_hyphen, ["Bopo"] = no_display_hyphen, ["Hani"] = no_display_hyphen, ["Hans"] = no_display_hyphen, ["Hant"] = no_display_hyphen, -- The following is a mixture of several scripts. Hopefully the specs here are correct! ["Jpan"] = no_display_hyphen, ["Jurc"] = no_display_hyphen, ["Kitl"] = no_display_hyphen, ["Kits"] = no_display_hyphen, ["Laoo"] = no_display_hyphen, ["Nshu"] = no_display_hyphen, ["Shui"] = no_display_hyphen, ["Tang"] = no_display_hyphen, ["Thaa"] = no_display_hyphen, ["Thai"] = no_display_hyphen, ["Tibt"] = no_display_hyphen, } ----------------------------------------------------------------------------------------- -- Basic Utility functions -- ----------------------------------------------------------------------------------------- local function glossary_link(entry, text) text = text or entry return "[[Appendix:Glossary#" .. entry .. "|" .. text .. "]]" end local function track(page) if type(page) == "table" then for i, pg in ipairs(page) do page[i] = "affix/" .. pg end else page = "affix/" .. page end require("Module:debug/track")(page) end local function ine(val) return val ~= "" and val or nil end ----------------------------------------------------------------------------------------- -- Compound types -- ----------------------------------------------------------------------------------------- local function make_compound_type(typ, alttext) return { text = glossary_link(typ, alttext) .. " compound", cat = typ .. " compounds", } end -- Make a compound type entry with a simple rather than glossary link. -- These should be replaced with a glossary link when the entry in the glossary -- is created. local function make_non_glossary_compound_type(typ, alttext) local link = alttext and "[[" .. typ .. "|" .. alttext .. "]]" or "[[" .. typ .. "]]" return { text = link .. " compound", cat = typ .. " compounds", } end local function make_raw_compound_type(typ, alttext) return { text = glossary_link(typ, alttext), cat = pluralize(typ), } end local function make_borrowing_type(typ, alttext) return { text = glossary_link(typ, alttext), borrowing_type = pluralize(typ), } end export.etymology_types = { ["adapted borrowing"] = make_borrowing_type("adapted borrowing"), ["adap"] = "adapted borrowing", ["abor"] = "adapted borrowing", ["alliterative"] = make_non_glossary_compound_type("alliterative"), ["allit"] = "alliterative", ["antonymous"] = make_non_glossary_compound_type("antonymous"), ["ant"] = "antonymous", ["bahuvrihi"] = make_compound_type("bahuvrihi", "bahuvrīhi"), ["bahu"] = "bahuvrihi", ["bv"] = "bahuvrihi", ["coordinative"] = make_compound_type("coordinative"), ["coord"] = "coordinative", ["descriptive"] = make_compound_type("descriptive"), ["desc"] = "descriptive", ["determinative"] = make_compound_type("determinative"), ["det"] = "determinative", ["dvandva"] = make_compound_type("dvandva"), ["dva"] = "dvandva", ["dvigu"] = make_compound_type("dvigu"), ["dvi"] = "dvigu", ["endocentric"] = make_compound_type("endocentric"), ["endo"] = "endocentric", ["exocentric"] = make_compound_type("exocentric"), ["exo"] = "exocentric", ["izafet I"] = make_compound_type("izafet I"), ["iz1"] = "izafet I", ["izafet II"] = make_compound_type("izafet II"), ["iz2"] = "izafet II", ["izafet III"] = make_compound_type("izafet III"), ["iz3"] = "izafet III", ["karmadharaya"] = make_compound_type("karmadharaya", "karmadhāraya"), ["karma"] = "karmadharaya", ["kd"] = "karmadharaya", ["kenning"] = make_raw_compound_type("kenning"), ["ken"] = "kenning", ["rhyming"] = make_non_glossary_compound_type("rhyming"), ["rhy"] = "rhyming", ["synonymous"] = make_non_glossary_compound_type("synonymous"), ["syn"] = "synonymous", ["tatpurusa"] = make_compound_type("tatpurusa", "tatpuruṣa"), ["tat"] = "tatpurusa", ["tp"] = "tatpurusa", } local function process_etymology_type(typ, nocap, notext, has_parts) local text_sections = {} local categories = {} local borrowing_type if typ then local typdata = export.etymology_types[typ] if type(typdata) == "string" then typdata = export.etymology_types[typdata] end if not typdata then error("Internal error: Unrecognized type '" .. typ .. "'") end local text = typdata.text if not nocap then text = ucfirst(text) end local cat = typdata.cat borrowing_type = typdata.borrowing_type local oftext = typdata.oftext or " of" if not notext then table.insert(text_sections, text) if has_parts then table.insert(text_sections, oftext) table.insert(text_sections, " ") end end if cat then table.insert(categories, cat) end end return text_sections, categories, borrowing_type end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- -- Iterate an array up to the greatest integer index found. local function ipairs_with_gaps(t) local indices = m_table.numKeys(t) local max_index = #indices > 0 and math.max(unpack(indices)) or 0 local i = 0 return function() if i < max_index then i = i + 1 return i, t[i] end end end export.ipairs_with_gaps = ipairs_with_gaps --[==[ Join formatted parts (in `parts_formatted`) together with any overall {{para|lit}} spec (in `lit`) plus categories, which are formatted by prepending the language name as found in `lang`. The value of an entry in `categories` can be either a string (which is formatted using `sort_key`) or a table of the form `{ {cat=<var>category</var>, sort_key=<var>sort_key</var>, sort_base=<var>sort_base</var>}`, specifying the sort key and sort base to use when formatting the category. If `nocat` is given, no categories are added; otherwise, `force_cat` causes categories to be added even on userspace pages. ]==] function export.join_formatted_parts(data) local cattext local lang = data.data.lang local force_cat = data.data.force_cat or debug_force_cat if data.data.nocat then cattext = "" else for i, cat in ipairs(data.categories) do if type(cat) == "table" then data.categories[i] = require(utilities_module).format_categories(lang:getFullName() .. " " .. cat.cat, lang, cat.sort_key, cat.sort_base, force_cat) else data.categories[i] = require(utilities_module).format_categories(lang:getFullName() .. " " .. cat, lang, data.data.sort_key, nil, force_cat) end end cattext = table.concat(data.categories) end local result = table.concat(data.parts_formatted, not data.separator_already_added and " +&lrm; " or nil) .. (data.data.lit and ", literally " .. m_links.mark(data.data.lit, "gloss") or "") local q = data.data.q local qq = data.data.qq local l = data.data.l local ll = data.data.ll if q and q[1] or qq and qq[1] or l and l[1] or ll and ll[1] then result = require(decorations_module).format_decorations { lang = lang, text = result, q = q, qq = qq, l = l, ll = ll, } end return result .. cattext end -- Remove links and call lang:stripDiacritics(term). local function strip_diacritics_no_links(lang, term) return lang:stripDiacritics(m_links.remove_links(term)) end --[=[ Convert a raw part as passed into an entry point into a part ready for linking. `lang` and `sc` are the overall language and script objects. This uses the overall language and script objects as defaults for the part and parses off any fragment from the term. We need to do the latter so that fragments don't end up in categories and so that we correctly do affix mapping even in the presence of fragments. ]=] local function canonicalize_part(part, lang, sc) if not part then return end -- Save the original (user-specified, part-specific) value of `lang`. If such a value is specified, we don't insert -- a '*fixed with' category, and we format the part using format_derived() in [[Module:etymology]] rather than -- full_link() in [[Module:links]]. part.part_lang = part.lang part.lang = part.lang or lang part.sc = part.sc or sc local term = part.term if not term then return elseif not part.fragment then part.term, part.fragment = m_links.get_fragment(term) else part.term = m_links.get_fragment(term) end end --[==[ Construct a single linked part based on the information in `part`, for use by `show_affix()` and other entry points. This should be called after `canonicalize_part()` is called on the part. This is a thin wrapper around `full_link()` in [[Module:links]] unless `part.part_lang` is specified (indicating that a part-specific language was given), in which case `format_derived()` in [[Module:etymology]] is called to display a term in a language other than the language of the overall term (specified in `data.lang`). `data` contains the entire object passed into the entry point and is used to access information for constructing the categories added by `format_derived()`. ]==] function export.link_term(part, data, include_separator) local result if part.part_lang then result = require(etymology_module).format_derived { lang = data.lang, terms = {part}, sources = {part.lang}, sort_key = data.sort_key, nocat = data.nocat, template_name = "affix", decorations_on_outside = true, borrowing_type = data.borrowing_type, force_cat = data.force_cat or debug_force_cat, } else result = m_links.full_link(part, "term") end if include_separator and part.separator then return part.separator .. result else return result end end local function canonicalize_script_code(scode) -- Convert 'mnc-Mong', 'xwo-Mong' etc. to 'Mong'. return (scode:gsub("^.*%-", "")) end ----------------------------------------------------------------------------------------- -- Affix-handling functions -- ----------------------------------------------------------------------------------------- -- Figure out the appropriate script for the given affix and language (unless the script is explicitly passed in), and -- return the values of template_hyphens[], display_hyphens[] and lookup_hyphens[] for that script, substituting -- default values as appropriate. Four values are returned: -- DETECTED_SCRIPT, TEMPLATE_HYPHEN, DISPLAY_HYPHEN, LOOKUP_HYPHEN local function detect_script_and_hyphens(text, lang, sc) local scode -- 1. If the script is explicitly passed in, use it. if sc then scode = sc:getCode() else local possible_script_codes = lang:getScriptCodes() -- YUCK! `possible_script_codes` comes from loadData() so #possible_scripts doesn't work (always returns 0). local num_possible_script_codes = m_table.length(possible_script_codes) if num_possible_script_codes == 0 then -- This shouldn't happen; if the language has no script codes, -- the list {"None"} should be returned. error("Something is majorly wrong! Language " .. lang:getCanonicalName() .. " has no script codes.") end if num_possible_script_codes == 1 then -- 2. If the language has only one possible script, use it. scode = possible_script_codes[1] else -- 3. Check if any of the possible scripts for the language have non-default values for template_hyphens[] -- or display_hyphens[]. If so, we need to do script detection on the text. If not, just use "Latn", -- which may not be technically correct but produces the right results because Latn has all default -- values for template_hyphens[] and display_hyphens[]. local may_have_nondefault_hyphen = false for _, script_code in ipairs(possible_script_codes) do script_code = canonicalize_script_code(script_code) if template_hyphens[script_code] or display_hyphens[script_code] then may_have_nondefault_hyphen = true break end end if not may_have_nondefault_hyphen then scode = "Latn" else scode = lang:findBestScript(text):getCode() end end end scode = canonicalize_script_code(scode) local template_hyphen = template_hyphens[scode] or "-" local lookup_hyphen = lookup_hyphens[scode] or "-" local display_hyphen = display_hyphens[scode] or default_display_hyphen return scode, template_hyphen, display_hyphen, lookup_hyphen end --[=[ Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string, specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen, or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix). ]=] local function reconstruct_term_per_hyphens(term, affix_type, scode, thyph_re, new_hyphen) local function get_hyphen(hyph) if type(new_hyphen) == "string" then return new_hyphen end return new_hyphen(scode, hyph) end if affix_type == "non-affix" then return term elseif affix_type == "circumfix" then local before, before_hyphen, after_hyphen, after = rmatch(term, "^(.*)" .. thyph_re .. " " .. thyph_re .. "(.*)$") if not before or ulen(term) <= 3 then -- Unlike with other types of affixes, don't try to add hyphens in the middle of the term to convert it to -- a circumfix. Also, if the term is just hyphen + space + hyphen, return it. return term end return before .. get_hyphen(before_hyphen) .. " " .. get_hyphen(after_hyphen) .. after elseif affix_type == "infix" or affix_type == "interfix" then local before_hyphen, middle, after_hyphen = rmatch(term, "^" .. thyph_re .. "(.*)" .. thyph_re .. "$") if before_hyphen and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return get_hyphen(before_hyphen) .. (middle or term) .. get_hyphen(after_hyphen) elseif affix_type == "prefix" then local middle, after_hyphen = rmatch(term, "^(.*)" .. thyph_re .. "$") if middle and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return (middle or term) .. get_hyphen(after_hyphen) elseif affix_type == "suffix" then local before_hyphen, middle = rmatch(term, "^" .. thyph_re .. "(.*)$") if before_hyphen and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return get_hyphen(before_hyphen) .. (middle or term) else error(("Internal error: Unrecognized affix type '%s'"):format(affix_type)) end end --[=[ Look up a mapping from a given affix variant to the canonical form used in categories and links. The lookup tables are language-specific according to `lang`, and may be ID-specific according to `affix_id`. The affixes as they appear in the lookup tables (both the variant and the canonical form) are in "lookup affix" format (approximately speaking, they use a regular hyphen for most scripts, but a tatweel for Arabic-script entries and a maqqef for Hebrew-script entries), but the passed-in `affix` param is in "template affix" format (which differs from the lookup affix for Arabic-script entries, because more types of hyphens are allowed in template affixes; see the comments at the top of the file). The remaining parameters to this function are used to convert from template affixes to lookup affixes; see the reconstruct_term_per_hyphens() function above. If the affix contains brackets, no lookup is done. Otherwise, a two-stage process is used, first looking up the affix directly and then stripping diacritics and looking it up again. The reason for this is documented above in the comments at the top of the file (specifically, the comments describing lookup affixes). The value of a mapping can either be a string (do the mapping regardless of affix ID) or a table indexed by affix ID (where the special value `false` indicates no affix ID). The values of entries in this table can also be strings, or tables with keys `affix` and `id` (again, use `false` to indicate no ID). This allows an affix mapping to map from one ID to another (for example, this is used in English to map the [[an-]] prefix with no ID to the [[a-]] prefix with the ID 'not'). The Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string, specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen, or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix). ]=] local function lookup_affix_mapping(affix, affix_type, lang, scode, thyph_re, lookup_hyph, affix_id) local function do_lookup(afx) -- Ensure that the affix uses lookup hyphens regardless of whether it used a different type of hyphens before -- or no hyphens. local lookup_affix = reconstruct_term_per_hyphens(afx, affix_type, scode, thyph_re, lookup_hyph) local function do_lookup_for_langcode(langcode) if export.langs_with_lang_specific_data[langcode] then local langdata = mw.loadData(export.affix_lang_data_module_prefix .. langcode) if langdata.affix_mappings then local mapping = langdata.affix_mappings[lookup_affix] if mapping then if type(mapping) == "table" then mapping = mapping[affix_id] or mapping.default or mapping[affix_id or false] if mapping then return mapping end else return mapping end end end end end -- If `lang` is an etymology-only language, look for a mapping both for it and its full parent. local langcode = lang:getCode() local mapping = do_lookup_for_langcode(langcode) if mapping then return mapping end local full_langcode = lang:getFullCode() if full_langcode ~= langcode then mapping = do_lookup_for_langcode(full_langcode) if mapping then return mapping end end return nil end if affix:find("%[%[") then return nil end return do_lookup(affix) or do_lookup(lang:stripDiacritics(affix)) or nil end --[==[ For a given template term in a given language (see the definition of "template affix" near the top of the file), possibly in an explicitly specified script `sc` (but usually nil), return the term's affix type ({"prefix"}, {"interfix"}, {"suffix"}, {"circumfix"} or {"non-affix"}) along with the corresponding link and display affixes (see definitions near the top of the file); also the corresponding lookup affix (if `return_lookup_affix` is specified). The term passed in should already have any fragment (after the # sign) parsed off of it. Four values are returned: `affix_type`, `link_term`, `display_term` and `lookup_term`. The affix type can be passed in instead of autodetected; in this case, the template term need not have any attached hyphens, and the appropriate hyphens will be added in the appropriate places. If `do_affix_mapping` is specified, look up the affix in the lang-specific affix mappings, as described in the comment at the top of the file; otherwise, the link and display terms will always be the same. (They will be the same in any case if the template term has a bracketed link in it or is not an affix.) If `return_lookup_affix` is given, the fourth return value contains the term with appropriate lookup hyphens in the appropriate places; otherwise, it is the same as the display term. (This functionality is used in [[Module:category tree/affixes and compounds]] to convert link affixes into lookup affixes so that they can be looked up in the affix mapping tables.) Exported because used by [[Module:headword utilities]] to determine the affix type of a given pagename. ]==] function export.parse_term_for_affixes(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) if not term then return "non-affix", nil, nil, nil end if term == "^" then -- Indicates a null term to emulate the behavior of {{suffix|foo||bar}}. term = "" return "non-affix", term, term, term end if term:find("^%^") then -- HACK! ^ at the beginning of Korean languages has a special meaning, triggering capitalization of the -- transliteration. Don't interpret it as "force non-affix" for those languages. local langcode = lang:getCode() if langcode ~= "ko" and langcode ~= "okm" and langcode ~= "jje" then -- Formerly we allowed ^ to force non-affix type; this is now handled using an inline modifier -- <naf>, <root>, etc. Throw an error for the moment when the old way is encountered. error("Use of ^ to force non-affix status is no longer supported; use an inline modifier <naf> or <root> " .. "after the component") end end -- Remove an asterisk if the morpheme is reconstructed and add it back at the end. local reconstructed = "" if term:find("^%*") then reconstructed = "*" term = term:gsub("^%*", "") end local scode, thyph, dhyph, lhyph = detect_script_and_hyphens(term, lang, sc) thyph = "([" .. thyph .. "])" if not affix_type then if rfind(term, thyph .. " " .. thyph) then affix_type = "circumfix" else local has_beginning_hyphen = rfind(term, "^" .. thyph) local has_ending_hyphen = rfind(term, thyph .. "$") if has_beginning_hyphen and has_ending_hyphen then affix_type = "interfix" elseif has_ending_hyphen then affix_type = "prefix" elseif has_beginning_hyphen then affix_type = "suffix" else affix_type = "non-affix" end end end local link_term, display_term, lookup_term if affix_type == "non-affix" then link_term = term display_term = term lookup_term = term else display_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, dhyph) if do_affix_mapping then link_term = lookup_affix_mapping(term, affix_type, lang, scode, thyph, lhyph, affix_id) -- The return value of lookup_affix_mapping() may be an affix mapping with lookup hyphens if a mapping -- was found, otherwise nil if a mapping was not found. We need to convert to display hyphens in -- either case, but in the latter case we can reuse the display term, which has already been converted. if link_term then link_term = reconstruct_term_per_hyphens(link_term, affix_type, scode, thyph, dhyph) else link_term = display_term end else link_term = display_term end if return_lookup_affix then lookup_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, lhyph) else lookup_term = display_term end end link_term = reconstructed .. link_term display_term = reconstructed .. display_term lookup_term = reconstructed .. lookup_term return affix_type, link_term, display_term, lookup_term end --[==[ Add a hyphen to a term in the appropriate place, based on the specified affix type, stripping off any existing hyphens in that place. For example, if `affix_type` == {"prefix"}, we'll add a hyphen onto the end if it's not already there (or is of the wrong type). Three values are returned: the link term, display term and lookup term. This function is a thin wrapper around `parse_term_for_affixes`; see the comments above that function for more information. Note that this function is exposed externally because it is called by [[Module:category tree/affixes and compounds]]; see the comment in `parse_term_for_affixes` for more information. ]==] function export.make_affix(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) if not (affix_type == "prefix" or affix_type == "suffix" or affix_type == "circumfix" or affix_type == "infix" or affix_type == "interfix" or affix_type == "non-affix") then error("Internal error: Invalid affix type " .. (affix_type or "(nil)")) end local _, link_term, display_term, lookup_term = export.parse_term_for_affixes(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) return link_term, display_term, lookup_term end ----------------------------------------------------------------------------------------- -- Main entry points -- ----------------------------------------------------------------------------------------- --[==[ Core categorization logic for affixes. This is shared between show_affix(), show_compound_like() and get_affix_categories_only(). Returns the categories array and other metadata needed for formatting. ]==] local function generate_affix_categories(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) local text_sections, categories, borrowing_type = process_etymology_type(data.type, data.surface_analysis or data.nocap, data.notext, #data.parts > 0) data.borrowing_type = borrowing_type -- Process each part local whole_words = 0 local is_affix_or_compound = false -- Canonicalize and generate links for all the parts first; then do categorization in a separate step, because when -- processing the first part for categorization, we may access the second part and need it already canonicalized. for i, part in ipairs_with_gaps(data.parts) do part = part or {} data.parts[i] = part canonicalize_part(part, data.lang, data.sc) -- Determine affix type and get link and display terms (see text at top of file). Store them in the part -- (in fields that won't clash with fields used by full_link() in [[Module:links]] or link_term()), so they -- can be used in the loop below when categorizing. part.affix_type, part.affix_link_term, part.affix_display_term = export.parse_term_for_affixes(part.term, part.lang, part.sc, part.type, not part.alt, nil, part.id) -- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with inline -- modifiers. The intention in either case is not to link the term. part.term = ine(part.affix_link_term) -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. part.alt = part.alt or (part.affix_display_term ~= part.affix_link_term and part.affix_display_term) or nil end if not data.noaffixcat then -- Now do categorization. for i, part in ipairs_with_gaps(data.parts) do local affix_type = part.affix_type if affix_type ~= "non-affix" then is_affix_or_compound = true -- Make a sort key. For the first part, use the second part as the sort key; the intention is that if the -- term has a prefix, sorting by the prefix won't be very useful so we sort by what follows, which is -- presumably the root. local part_sort_base = nil local part_sort = part.sort or data.sort_key if i == 1 and data.parts[2] and data.parts[2].term then local part2 = data.parts[2] -- If the second-part link term is empty, the user requested an unlinked term; avoid a wikitext error -- by using the alt value if available. part_sort_base = ine(part2.affix_link_term) or ine(part2.alt) if part_sort_base then part_sort_base = strip_diacritics_no_links(part2.lang, part_sort_base) end end if part.pos and rfind(part.pos, "patronym") then table.insert(categories, {cat = "patronymics", sort_key = part_sort, sort_base = part_sort_base}) end if data.pos ~= "terms" and part.pos and rfind(part.pos, "diminutive") then table.insert(categories, {cat = "diminutive " .. data.pos, sort_key = part_sort, sort_base = part_sort_base}) end -- Don't add a '*fixed with' category if the link term is empty or is in a different language. if ine(part.affix_link_term) and not part.part_lang then table.insert(categories, {cat = data.pos .. " " .. affix_type .. "ed with " .. strip_diacritics_no_links(part.lang, part.affix_link_term) .. (part.id and " (" .. part.id .. ")" or ""), sort_key = part_sort, sort_base = part_sort_base}) end else whole_words = whole_words + 1 if whole_words == 2 then is_affix_or_compound = true table.insert(categories, "compound " .. data.pos) end end end -- Make sure there was either an affix or a compound (two or more non-affix terms). if not is_affix_or_compound and not data.allow_no_affixes_or_compounds then error("The parameters did not include any affixes, and the term is not a compound. Please provide at least one affix.") end end return text_sections, categories, borrowing_type end --[==[ Implementation of {{tl|affix}} and {{tl|surface analysis}}. `data` contains all the information describing the affixes to be displayed, and contains the following: * `.lang` ('''required'''): Overall language object. Different from term-specific language objects (see `.parts` below). * `.sc`: Overall script object (usually omitted). Different from term-specific script objects. * `.parts` ('''required'''): List of objects describing the affixes to show. The general format of each object is as would be passed to `full_link()`, except that the `.lang` field should be missing unless the term is of a language different from the overall `.lang` value (in such a case, the language name is shown along with the term and an additional "derived from" category is added). '''WARNING''': The data in `.parts` will be destructively modified. * `.pos`: Overall part of speech (used in categories, defaults to {"terms"}). Different from term-specific part of speech. * `.sort_key`: Overall sort key. Normally omitted except e.g. in Japanese. * `.type`: Type of compound, if the parts in `.parts` describe a compound. Strictly optional, and if supplied, the compound type is displayed before the parts (normally capitalized, unless `.nocap` is given). * `.nocap`: Don't capitalize the first letter of text displayed before the parts (relevant only if `.type` or `.surface_analysis` is given). * `.notext`: Don't display any text before the parts (relevant only if `.type` or `.surface_analysis` is given). * `.nocat`: Disable all categorization. * `.noaffixcat`: Disable affix (and compound) categorization. Relevant for e.g. blends, which may otherwise be incorrectly categorized as compound terms. * `.lit`: Overall literal definition. Different from term-specific literal definitions. * `.force_cat`: Always display categories, even on userspace pages. * `.surface_analysis`: Implement {{surface analysis}}; adds `By surface analysis, ` before the parts. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_affix(data) local text_sections, categories, _ = generate_affix_categories(data) -- Process each part for display local parts_formatted = {} for i, part in ipairs_with_gaps(data.parts) do -- Make a link for the part table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if data.surface_analysis then local text = "by " .. glossary_link("surface analysis") .. ", " if not data.nocap then text = ucfirst(text) end table.insert(text_sections, 1, text) end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Get only the categories that would be generated by show_affix(), without any text output or formatting. This is used by Module:etymon to get affix categorization. Returns an array of category objects, where each entry is either a string (simple category name) or a table with keys `cat`, `sort_key`, and `sort_base` for more complex categorization. `data` should have the same structure as passed to show_affix(): * `.lang` (required): Overall language object * `.parts` (required): Array of affix part objects with `.term`, `.lang`, `.id`, etc. * `.pos`: Part of speech (defaults to "terms") * `.sort_key`: Overall sort key for categories '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.get_affix_categories_only(data) local _, categories, _ = generate_affix_categories(data) return categories end function export.show_surface_analysis(data) data.surface_analysis = true data.allow_no_affixes_or_compounds = true return export.show_affix(data) end --[==[ Implementation of {{tl|compound}}. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_compound(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) local text_sections, categories, borrowing_type = process_etymology_type(data.type, data.nocap, data.notext, #data.parts > 0) data.borrowing_type = borrowing_type local parts_formatted = {} table.insert(categories, "compound " .. data.pos) -- Make links out of all the parts local whole_words = 0 for i, part in ipairs(data.parts) do canonicalize_part(part, data.lang, data.sc) -- Determine affix type and get link and display terms (see text at top of file). local affix_type, link_term, display_term = export.parse_term_for_affixes(part.term, part.lang, part.sc, part.type, not part.alt, nil, part.id) -- If the term is an interfix or the type was explicitly given, recognize it as such (which means e.g. that we -- will display the term without hyphens for East Asian languages). Otherwise, ignore the fact that it looks -- like an affix and display as specified in the template (but pay attention to the detected affix type for -- certain tracking purposes). if affix_type == "interfix" or (part.type and part.type ~= "non-affix") then -- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with -- inline modifiers. The intention in either case is not to link the term. Don't add a '*fixed with' -- category in this case, or if the term is in a different language. -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. if link_term and link_term ~= "" and not part.part_lang then table.insert(categories, {cat = data.pos .. " " .. affix_type .. "ed with " .. strip_diacritics_no_links(part.lang, link_term), sort_key = part.sort or data.sort_key}) end part.term = link_term ~= "" and link_term or nil part.alt = part.alt or (display_term ~= link_term and display_term) or nil else if affix_type ~= "non-affix" then local langcode = data.lang:getCode() -- If `data.lang` is an etymology-only language, track both using its code and its full parent's code. track { affix_type, affix_type .. "/lang/" .. langcode } local full_langcode = data.lang:getFullCode() if langcode ~= full_langcode then track(affix_type .. "/lang/" .. full_langcode) end else whole_words = whole_words + 1 end end table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if whole_words == 1 then track("one whole word") elseif whole_words == 0 then track("looks like confix") end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Implementation of {{tl|blend}}, {{tl|univerbation}} and similar "compound-like" templates. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_compound_like(data) data.allow_no_affixes_or_compounds = true local text_sections, categories, _ = generate_affix_categories(data) if data.cat then table.insert(categories, data.cat) end -- Process each part for display local parts_formatted = {} for i, part in ipairs_with_gaps(data.parts) do -- Make a link for the part table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if #data.parts > 0 and data.oftext then table.insert(text_sections, 1, " " .. data.oftext .. " ") end if data.text then table.insert(text_sections, 1, data.text) end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Make `part` (a structure holding information on an affix part) into an affix of type `affix_type`, and apply any relevant affix mappings. For example, if the desired affix type is "suffix", this will (in general) add a hyphen onto the beginning of the term, alt, tr and ts components of the part if not already present. The hyphen that's added is the "display hyphen" (see above) and may be script-specific. (In the case of East Asian scripts, the display hyphen is an empty string whereas the template hyphen is the regular hyphen, meaning that any regular hyphen at the beginning of the part will be effectively removed.) `lang` and `sc` hold overall language and script objects. Note that this also applies any language-specific affix mappings, so that e.g. if the language is Finnish and the user specified [[-käs]] in the affix and didn't specify an `.alt` value, `part.term` will contain [[-kas]] and `part.alt` will contain [[-käs]]. This function is used by the "legacy" templates ({{tl|prefix}}, {{tl|suffix}}, {{tl|confix}}, etc.) where the nature of the affix is specified by the template itself rather than auto-determined from the affix, as is the case with {{tl|affix}}. '''WARNING''': This destructively modifies `part`. ]==] local function make_part_into_affix(part, lang, sc, affix_type) canonicalize_part(part, lang, sc) local link_term, display_term = export.make_affix(part.term, part.lang, part.sc, affix_type, not part.alt, nil, part.id) part.term = link_term -- When we don't specify `do_affix_mapping` to make_affix(), link and display terms (first and second retvals of -- make_affix()) are the same. -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. part.alt = part.alt and export.make_affix(part.alt, part.lang, part.sc, affix_type) or (display_term ~= link_term and display_term) or nil local Latn = require(scripts_module).getByCode("Latn") part.tr = export.make_affix(part.tr, part.lang, Latn, affix_type) part.ts = export.make_affix(part.ts, part.lang, Latn, affix_type) end local function track_wrong_affix_type(template, part, expected_affix_type) if part and not part.type then local affix_type = export.parse_term_for_affixes(part.term, part.lang, part.sc) if affix_type ~= expected_affix_type then local part_name = expected_affix_type or "base" local langcode = part.lang:getCode() local full_langcode = part.lang:getFullCode() require("Module:debug/track") { template, template .. "/" .. part_name, template .. "/" .. part_name .. "/" .. (affix_type or "none"), template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. langcode } -- If `part.lang` is an etymology-only language, track both using its code and its full parent's code. if full_langcode ~= langcode then require("Module:debug/track")( template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. full_langcode ) end end end end local function insert_affix_category(categories, pos, affix_type, part, sort_key, sort_base) -- Don't add a '*fixed with' category if the link term is empty or is in a different language. if part.term and not part.part_lang then local cat = pos .. " " .. affix_type .. "ed with " .. strip_diacritics_no_links(part.lang, part.term) .. (part.id and " (" .. part.id .. ")" or "") if sort_key or sort_base then table.insert(categories, {cat = cat, sort_key = sort_key, sort_base = sort_base}) else table.insert(categories, cat) end end end --[==[ Implementation of {{tl|circumfix}}. '''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`. ]==] function export.show_circumfix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.prefix, data.lang, data.sc, "prefix") make_part_into_affix(data.suffix, data.lang, data.sc, "suffix") track_wrong_affix_type("circumfix", data.prefix, "prefix") track_wrong_affix_type("circumfix", data.base, nil) track_wrong_affix_type("circumfix", data.suffix, "suffix") -- Create circumfix term. local circumfix = nil if data.prefix.term and data.suffix.term then circumfix = data.prefix.term .. " " .. data.suffix.term data.prefix.alt = data.prefix.alt or data.prefix.term data.suffix.alt = data.suffix.alt or data.suffix.term data.prefix.term = circumfix data.suffix.term = circumfix end -- Make links out of all the parts. local parts_formatted = {} local categories = {} local sort_base if data.base.term then sort_base = strip_diacritics_no_links(data.base.lang, data.base.term) end table.insert(parts_formatted, export.link_term(data.prefix, data)) table.insert(parts_formatted, export.link_term(data.base, data)) table.insert(parts_formatted, export.link_term(data.suffix, data)) -- Insert the categories, but don't add a '*fixed with' category if the link term is in a different language. if not data.prefix.part_lang then table.insert(categories, {cat=data.pos .. " circumfixed with " .. strip_diacritics_no_links(data.prefix.lang, circumfix), sort_key=data.sort_key, sort_base=sort_base}) end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|confix}}. '''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`. ]==] function export.show_confix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.prefix, data.lang, data.sc, "prefix") make_part_into_affix(data.suffix, data.lang, data.sc, "suffix") track_wrong_affix_type("confix", data.prefix, "prefix") track_wrong_affix_type("confix", data.base, nil) track_wrong_affix_type("confix", data.suffix, "suffix") -- Make links out of all the parts. local parts_formatted = {} local prefix_sort_base if data.base and data.base.term then prefix_sort_base = strip_diacritics_no_links(data.base.lang, data.base.term) elseif data.suffix.term then prefix_sort_base = strip_diacritics_no_links(data.suffix.lang, data.suffix.term) end -- Insert the categories and parts. local categories = {} table.insert(parts_formatted, export.link_term(data.prefix, data)) insert_affix_category(categories, data.pos, "prefix", data.prefix, data.sort_key, prefix_sort_base) if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) end table.insert(parts_formatted, export.link_term(data.suffix, data)) -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "suffix", data.suffix) return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|infix}}. '''WARNING''': This destructively modifies both `data` and `.base` and `.infix`. ]==] function export.show_infix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.infix, data.lang, data.sc, "infix") track_wrong_affix_type("infix", data.base, nil) track_wrong_affix_type("infix", data.infix, "infix") -- Make links out of all the parts. local parts_formatted = {} local categories = {} table.insert(parts_formatted, export.link_term(data.base, data)) table.insert(parts_formatted, export.link_term(data.infix, data)) -- Insert the categories. -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "infix", data.infix) return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|prefix}}. '''WARNING''': This destructively modifies both `data` and the structures within `.prefixes`, as well as `.base`. ]==] function export.show_prefix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. for i, prefix in ipairs(data.prefixes) do make_part_into_affix(prefix, data.lang, data.sc, "prefix") end for i, prefix in ipairs(data.prefixes) do track_wrong_affix_type("prefix", prefix, "prefix") end track_wrong_affix_type("prefix", data.base, nil) -- Make links out of all the parts. local parts_formatted = {} local first_sort_base = nil local categories = {} if data.prefixes[2] then first_sort_base = ine(data.prefixes[2].term) or ine(data.prefixes[2].alt) if first_sort_base then first_sort_base = strip_diacritics_no_links(data.prefixes[2].lang, first_sort_base) end elseif data.base then first_sort_base = ine(data.base.term) or ine(data.base.alt) if first_sort_base then first_sort_base = strip_diacritics_no_links(data.base.lang, first_sort_base) end end for i, prefix in ipairs(data.prefixes) do table.insert(parts_formatted, export.link_term(prefix, data)) insert_affix_category(categories, data.pos, "prefix", prefix, data.sort_key, i == 1 and first_sort_base or nil) end if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) else table.insert(parts_formatted, "") end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|suffix}}. '''WARNING''': This destructively modifies both `data` and the structures within `.suffixes`, as well as `.base`. ]==] function export.show_suffix(data) local categories = {} data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. for i, suffix in ipairs(data.suffixes) do make_part_into_affix(suffix, data.lang, data.sc, "suffix") end track_wrong_affix_type("suffix", data.base, nil) for i, suffix in ipairs(data.suffixes) do track_wrong_affix_type("suffix", suffix, "suffix") end -- Make links out of all the parts. local parts_formatted = {} if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) else table.insert(parts_formatted, "") end for i, suffix in ipairs(data.suffixes) do table.insert(parts_formatted, export.link_term(suffix, data)) end -- Insert the categories. for i, suffix in ipairs(data.suffixes) do -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "suffix", suffix) if suffix.pos and rfind(suffix.pos, "patronym") then table.insert(categories, "patronymics") end end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end return export tffe29xfh32ju1jfs6x92q3tam80d7z Bysen:en-noun 10 8055 54865 2026-09-26T23:37:41Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:en-headword|show|nouns}}<noinclude>{{documentation}}</noinclude>' 54865 wikitext text/x-wiki {{#invoke:en-headword|show|nouns}}<noinclude>{{documentation}}</noinclude> 23mge28vi7mm8w5znn9hiq2wsewc8op Module:en-headword 828 8056 54866 2026-09-26T23:38:06Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local pos_functions = {} --[==[ Author from 2020 on: mostly Benwing2, with significant contributions from Theknightwho. Based on a prior version by Rua (by now mostly rewritten), with contributions from Erutuon and others (see history for full attribution). ]==] local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local require = require local require_when_needed = require("Module:require when needed") lo...' 54866 Scribunto text/plain local export = {} local pos_functions = {} --[==[ Author from 2020 on: mostly Benwing2, with significant contributions from Theknightwho. Based on a prior version by Rua (by now mostly rewritten), with contributions from Erutuon and others (see history for full attribution). ]==] local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local require = require local require_when_needed = require("Module:require when needed") local en_utilities_module = "Module:en-utilities" local headword_utilities_module = "Module:headword utilities" local headword_module = "Module:headword" local inflection_utilities_module = "Module:inflection utilities" local parse_utilities_module = "Module:parse utilities" local JSON_module = "Module:JSON" local labels_module = "Module:labels" local links_module = "Module:links" local parameters_module = "Module:parameters" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local yesno_module = "Module:yesno" local iut = require_when_needed(inflection_utilities_module) local put = require_when_needed(parse_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local add_links_to_multiword_term = require_when_needed(headword_utilities_module, "add_links_to_multiword_term") local add_suffix = require_when_needed(en_utilities_module, "add_suffix") local apply_link_modifiers = require_when_needed(headword_utilities_module, "apply_link_modifiers") local concat = table.concat local deepEquals = require_when_needed(table_module, "deepEquals") local dump = mw.dumpObject local format_categories = require_when_needed(utilities_module, "format_categories") local full_headword = require_when_needed(headword_module, "full_headword") local get_label_info = require_when_needed(labels_module, "get_label_info") local get_link_page = require_when_needed(links_module, "get_link_page") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local insert = table.insert local insertIfNot = require_when_needed(table_module, "insertIfNot") local ipairs = ipairs local is_regular_plural = require_when_needed(en_utilities_module, "is_regular_plural") local list_to_set = require_when_needed(table_module, "listToSet") local pairs = pairs local process_params = require_when_needed(parameters_module, "process") local remove = table.remove local remove_links = require_when_needed(links_module, "remove_links") local replacement_escape = require_when_needed(string_utilities_module, "replacement_escape") local shallowCopy = require_when_needed(table_module, "shallowCopy") local singularize = require_when_needed(en_utilities_module, "singularize") local split = require_when_needed(string_utilities_module, "split") local toJSON = require_when_needed(JSON_module, "toJSON") local toNFD = mw.ustring.toNFD local type = type local ulen = require_when_needed(string_utilities_module, "len") local ulower = require_when_needed(string_utilities_module, "lower") local umatch = require_when_needed(string_utilities_module, "match") local u = require_when_needed(string_utilities_module, "char") local ugsub = require_when_needed(string_utilities_module, "gsub") local lang = require("Module:languages").getByCode("en") local langname = lang:getCanonicalName() local list_param = {list = true, disallow_holes = true} local list_allow_holes = {list = true, allow_holes = true} local boolean_param = {type = "boolean"} local function ine(val) if val == "" then return nil else return val end end local function track(page) require("Module:debug/track")("en-headword/" .. page) return true end ------------------------------------------- UTILITY FUNCTIONS ------------------------------------------ -- Parse and return an inflection not requiring additional processing. The raw arguments come from `args[field]`, which -- is parsed for inline modifiers. local function parse_inflection(args, field, is_head) local argfield = field if type(argfield) == "table" then argfield = argfield[1] end return m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[argfield], splitchar = ",", is_head = is_head, } end -- Insert the parsed inflections in `terms` (as parsed by `parse_inflection`) into `data.inflections`, with label -- `label` and optional accelerator spec `accel`. local function insert_inflection(data, terms, label, accel, no_label) for _, termobj in ipairs(terms) do m_headword_utilities.remove_termobj_field_modifiers(termobj) end m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, no_label = no_label, accel = accel and {form = accel} or nil, } end -- Insert a fixed label `label` into the inflections for `data`. If `originating_term` is supplied, copy the decorations -- from it into the fixed label. local function insert_fixed_inflection(data, label, originating_term) m_headword_utilities.insert_fixed_inflection { headdata = data, originating_term = originating_term, label = label, } end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- These functions are used directly in the <> format as well as in the utility functions #2 below. local function compute_double_last_cons_stem(term) local last_cons = term:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("Verb stem '" .. term .. "' must end in a consonant to use ++") end return term .. last_cons end local function compute_plusplus_s_form(term, default_s_form) if term:find("[szx]$") then -- regas -> regasses, derez -> derezzes return compute_double_last_cons_stem(term) .. "es" else return default_s_form end end -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local iparams = { [1] = true, } local iargs = require("Module:parameters").process(frame.args, iparams) local parargs = frame:getParent().args local poscat = iargs[1] local pos_in_1 = not poscat if pos_in_1 then poscat = ine(parargs[1]) or mw.title.getCurrentTitle().fullText == "Template:en-head" and "interjection" or error("Part of speech must be specified in 1=") poscat = require(headword_module).canonicalize_pos(poscat) end local indexing_poscat = pos_in_1 and "head" or poscat local params = { ["head"] = list_param, ["id"] = true, ["json"] = boolean_param, ["sort"] = true, ["splithyph"] = boolean_param, ["nosplithyph"] = boolean_param, ["hyphspace"] = boolean_param, ["nolink"] = boolean_param, ["nolinkhead"] = {type = "boolean_param", alias_of = "nolink"}, ["suffix"] = boolean_param, ["nosuffix"] = boolean_param, ["nomultiwordcat"] = boolean_param, ["abbr"] = list_param, ["the"] = true, ["def"] = {alias_of = "the"}, ["pagename"] = true, -- for testing } if pos_in_1 then params[1] = {required = true} -- required but ignored as already processed above end local pos_data = pos_functions[indexing_poscat] local pos_func if pos_data then local pos_params = pos_data.params if pos_params then for key, val in pairs(pos_params) do params[key] = val end end pos_func = pos_data.func end local args = process_params(parargs, params) -- Account for unsupported titles, e.g. 'C|N>K' instead of 'Unsupported titles/C through N to K'. local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = parse_inflection(args, "head", "is_head") local heads = user_specified_heads local autohead if args.nolink or not pagename:find("[ '%-]") then autohead = pagename else local en_no_split_apostrophe_words = list_to_set { "one's", "someone's", "he's", "she's", "it's", } local en_include_hyphen_prefixes = list_to_set { -- We don't include things that are also words even though they are often (perhaps mostly) prefixes, e.g. -- "be", "counter", "cross", "extra", "half", "mid", "over", "pan", "under". "acro", "acousto", "Afro", "agro", "anarcho", "angio", "Anglo", "ante", "anti", "arch", "auto", "bi", "bio", "cis", "co", "cryo", "crypto", "de", "demi", "eco", "electro", "Euro", "ex", "Greco", "hemi", "hydro", "hyper", "hypo", "infra", "Indo", "inter", "intra", "Judeo", "macro", "meta", "micro", "mini", "multi", "neo", "neuro", "non", "para", "peri", "post", "pre", "pro", "proto", "pseudo", "re", "semi", "sub", "super", "trans", "un", "vice", } local function is_english(term) local title = mw.title.new(term) if title and title.exists then local content = title:getContent() if content and content:find("==English==\n") then return true end end return false end local function en_split_hyphen_when_space(word) if not word:find("-", nil, true) then return nil end if args.hyphspace then return "[[" .. word:gsub("%-+", " ") .. "|" .. word .. "]]" end if args.nosplithyph then return "[[" .. word .. "]]" end if not args.splithyph then local space_word = word:gsub("%-+", " ") if is_english(space_word) then return "[[" .. space_word .. "|" .. word .. "]]" end if is_english(word) then return "[[" .. word .. "]]" end end return nil end local function en_split_apostrophe(word) local base = word:match("^(.*)'s$") if base then return "[[" .. base .. "]][[-'s|'s]]" end -- Only treat final apostrophe as possessive if preceded by something that looks like a plural ending in /z/. -- In particular we don't want to do it for words like [[truckin']]. base = word:match("^(.*[sxz])'$") if base then if base:find("s$") then local sg = singularize(base) if is_english(sg) then return "[[" .. sg .. "|" .. base .. "]][[-'|']]" end end return "[[" .. base .. "]][[-'|']]" end return "[[" .. word .. "]]" end autohead = add_links_to_multiword_term(pagename, { split_hyphen_when_space = en_split_hyphen_when_space, split_apostrophe = en_split_apostrophe, no_split_apostrophe_words = en_no_split_apostrophe_words, include_hyphen_prefixes = en_include_hyphen_prefixes, }) end if not heads[1] then heads = {{term = autohead}} else for _, headobj in ipairs(heads) do local head = headobj.term if head:find("^~") then head = apply_link_modifiers(autohead, head:sub(2), lang) headobj.term = head elseif head:find("^[!?]$") then -- If explicit head= just consists of ! or ?, add it to the end of the default head. headobj.term = autohead .. head end if head == autohead then track("redundant-head") end end end -- handle the=/def= if args.the == "~" then local newheads = {} for _, headobj in ipairs(heads) do local barehead = shallowCopy(headobj) insert(newheads, barehead) headobj.term = "the " .. headobj.term insert(newheads, headobj) end heads = newheads elseif args.the then local the = require(yesno_module)(args.the) if the then for _, headobj in ipairs(heads) do headobj.term = "the " .. headobj.term end end end local data = { lang = lang, pos_category = poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, -- We use our own splitting algorithm so the redundant head cat will be inaccurate. no_redundant_head_cat = true, inflections = {}, nomultiwordcat = args.nomultiwordcat, sort_key = args.sort, pagename = pagename, id = args.id, force_cat_output = force_cat, } local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local is_suffix = false if args.suffix or not args.nosuffix and pagename:find("^%-") and not pagename:find("^%-%-") and poscat ~= "suffix forms" then is_suffix = true data.pos_category = "suffixes" local singular_poscat = singularize(poscat) inscat(singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end if pos_func then pos_func(args, data, is_suffix) end local extra_categories = {} if pagename:find("[Qq]") then -- Check for q not followed by u. We want to exclude things like [[13q deletion syndrome]] and [[BFOQ]] that -- don't have a lowercase letter on either side, as well as things like [[& seq.]] and [[acq.]] that are -- abbreviations for words containing a following u. -- -- Approximate range of combining diacritics; we want to remove them so the checks below for -- a lowercase letter next to the q aren't tripped up by diacritics on the letter. local u300 = u(0x0300) local u36F = u(0x036F) local pagename_no_diacritics = ugsub(toNFD(pagename), "[" .. u300 .. "-" .. u36F .. "]", "") if pagename_no_diacritics:find("[Qq][a-tv-z]") or pagename_no_diacritics:find("[a-z]q[^u.]") or pagename_no_diacritics:find("[a-z]q$") then inscat("words containing Q not followed by U") end end -- toNFD performs decomposition, so letters that decompose to an ASCII -- vowel and a diacritic, such as é, are counted as vowels and do not do not -- need to be included in the pattern. if not umatch(ulower(toNFD(pagename)), "[aeiouyæœøəªºαεηιουω]") then inscat("words spelled without vowels") end if pagename:find("yre$") then inscat('words ending in "-yre"') end if not pagename:find(" ") and ulen(pagename) >= 25 then insert(extra_categories, "Long " .. langname .. " words") end if pagename:find("^[^aeiou ]*a[^aeiou ]*e[^aeiou ]*i[^aeiou ]*o[^aeiou ]*u[^aeiou ]*$") then inscat("words that use all vowels in alphabetical order") end parse_and_insert_inflection(data, args, "abbr", "abbreviation") if args.json then return toJSON(data) end return full_headword(data) .. (extra_categories[1] and format_categories(extra_categories, lang, args.sort) or "") end local function make_default_comparative(word) if word == "good" or word == "well" then return {"better"} elseif word == "bad" or word == "badly" then return {"worse"} elseif word == "far" then return {"further", "farther"} else return {add_suffix(word, "r")} end end local function make_default_superlative(word) if word == "good" or word == "well" then return {"best"} elseif word == "bad" or word == "badly" then return {"worst"} elseif word == "far" then return {"furthest", "farthest"} else return {add_suffix(word, "st.superlative")} end end -- This function does the common work between adjectives and adverbs. local function process_comparative_args(data, args, plpos) local pagename = data.pagename local comps = parse_inflection(args, 1) local sups = parse_inflection(args, "sup") local outcomps, outsups if args.componly then if comps[1] then error("Can't specify comparatives of comparative-only " .. plpos) end insert(data.inflections, {label = glossary_link("comparative") .. " form only"}) insert(data.categories, langname .. " comparative-only " .. plpos) -- Set to empty list so we don't get any comparatives output, but process superlatives if specified. outcomps = {} if not sups[1] then -- Set to empty list so we don't get any superlatives output unless explicitly given. outsups = {} end elseif args.suponly then if comps[1] or sups[1] then error("Can't specify comparatives or superlatives of or superlative-only " .. plpos) end insert(data.inflections, {label = glossary_link("superlative") .. " form only"}) insert(data.categories, langname .. " superlative-only " .. plpos) return end -- If the first parameter is ?, then don't show anything, just return. if comps[1] and comps[1].term == "?" then if comps[2] then error("Can't specify additional comparatives along with '?'") end if sups[1] then error("Can't specify superlatives along with '?' for the comparative") end return end if comps[1] and comps[1].term == "-" then local hyphencomp = remove(comps, 1) -- Remove the "-" but retain for decorations. -- Not (generally) comparable; may occasionally have a comparative if comps[1] then insert_fixed_inflection(data, "not generally <<comparable>>", hyphencomp) elseif not sups[1] then insert_fixed_inflection(data, "not <<comparable>>", hyphencomp) insert(data.categories, langname .. " uncomparable " .. plpos) return else -- No comparative, but a superlative. insert_inflection() will correctly generate 'no comparative' if we -- pass in "-" as the value. outcomps = {hyphencomp} end elseif not comps[1] then comps = {{term = "more"}} end if not outcomps then -- not if we set `outcomps` to "-" above or processed a comparative-only term outcomps = {} -- Go over each parameter given and create a comparative and superlative form. for _, compobj in ipairs(comps) do local comp = compobj.term if comp == "-" then error("Comparative of '-' only allowed as first comparative") end if comp == "+" then comp = "+more" elseif comp == "more" and pagename ~= "many" and pagename ~= "much" then comp = "+more" elseif comp == "further" and pagename ~= "far" then comp = "+further" elseif comp == "better" and pagename ~= "good" and pagename ~= "well" then comp = "+better" elseif comp:find("~") then comp = comp:gsub("~", replacement_escape(pagename)) end compobj.origterm = comp if comp == "+more" then comp = "more [[" .. pagename .. "]]" elseif comp == "+further" then comp = {"further [[" .. pagename .. "]]", "farther [[" .. pagename .. "]]"} elseif comp == "+better" then comp = "better [[" .. pagename .. "]]" elseif comp == "er" then -- Add -er. comp = add_suffix(pagename, "r") elseif comp == "ier" then if pagename:sub(-1) ~= "y" then error("Can't specify 'ier' comparative unless the term ends with 'y': " .. pagename) end comp = pagename:gsub("e?y$", "ier") elseif comp:find("^%+") then local special = m_headword_utilities.get_special_indicator(comp, "noerror") if special then comp = m_headword_utilities.handle_multiword(pagename, special, make_default_comparative) end end if type(comp) == "table" and not comp[2] then comp = comp[1] end if type(comp) == "table" then for i = 1, #comp - 1 do local outobj = shallowCopy(compobj) outobj.term = comp[i] insert(outcomps, outobj) end compobj.term = comp[#comp] insert(outcomps, compobj) else compobj.term = comp insert(outcomps, compobj) end end end if sups[1] and sups[1].term == "-" then if sups[2] then error("Can't specify '-' as superlative followed by further values") end -- No superlative. insert_inflection() will correctly generate 'no superlative' if we pass in "-" as the value. outsups = sups else if not sups[1] then sups = {{term = "+"}} end end -- `outsups` will be set if we set `outsups` to "-" above or processed a comparative-only term without superlatives. if not outsups then outsups = {} local function process_sup(sup, special, supobj, compobj) if special then sup = m_headword_utilities.handle_multiword(pagename, special, make_default_superlative) elseif sup == "-" or sup == "+" then error(("Internal error: Superlative value of '%s' should have been handled earlier"):format(sup)) elseif sup == "+most" then sup = "most [[" .. pagename .. "]]" elseif sup == "+furthest" then sup = {"furthest [[" .. pagename .. "]]", "farthest [[" .. pagename .. "]]"} elseif sup == "+best" then sup = "best [[" .. pagename .. "]]" elseif sup == "est" then -- Add -est. sup = add_suffix(pagename, "st.superlative") elseif sup == "iest" then if pagename:sub(-1) ~= "y" then error("Can't specify 'iest' superlative unless the term ends with 'y': " .. pagename) end sup = pagename:gsub("e?y$", "iest") end if type(sup) == "table" and not sup[2] then sup = sup[1] end if compobj then supobj = shallowCopy(supobj) supobj = m_headword_utilities.combine_termobj_decorations(supobj, compobj) end if type(sup) == "table" then for i = 1, #sup - 1 do local outobj = shallowCopy(supobj) outobj.term = sup[i] insert(outsups, outobj) end supobj.term = sup[#sup] insert(outsups, supobj) else supobj.term = sup insert(outsups, supobj) end end for _, supobj in ipairs(sups) do local sup = supobj.term if sup == "-" then error("Superlative of '-' only allowed as first superlative") end if sup == "+" then if not comps[1] then error("Superlative of '+' can't be specified when there are no comparatives") end for _, compobj in ipairs(comps) do local comp = compobj.origterm local special if comp == "+more" then sup = "+most" elseif comp == "+further" then sup = "+furthest" elseif comp == "+better" then sup = "+best" elseif comp == "er" then sup = "est" elseif comp == "ier" then sup = "iest" else if comp:find("^%+") then special = m_headword_utilities.get_special_indicator(comp, "noerror") end if not special then -- If the full comparative was given, then derive the superlative by replacing -er with -- -est. if comp:sub(-2) == "er" then sup = comp:sub(1, -3) .. "est" else error(("The superlative cannot be derived automatically from comparative '%s' because it doesn't end in -er"):format(comp)) end end end process_sup(sup, special, supobj, compobj) end else local special = m_headword_utilities.get_special_indicator(sup, "noerror") -- Do some work here rather than in process_sup() so we don't end up double-processing a term with a '~' -- in it or a term that happens to be 'most' or similar after substitution of ~ in the comparative. if not special then if sup == "most" and pagename ~= "many" and pagename ~= "much" then sup = "+most" elseif sup == "furthest" and pagename ~= "far" then sup = "+furthest" elseif sup == "best" and pagename ~= "good" and pagename ~= "well" then sup = "+best" elseif sup:find("~") then sup = sup:gsub("~", replacement_escape(pagename)) end end process_sup(sup, special, supobj) end end end insert_inflection(data, outcomps, "<<comparative>>", "comparative") insert_inflection(data, outsups, "<<superlative>>", "superlative") end pos_functions["adjectives"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adjectives") end, } pos_functions["adverbs"] = { params = { [1] = list_param, ["comp_qual"] = {list = "comp\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the comparative value", }, ["sup"] = list_param, ["sup_qual"] = {list = "sup\1_qual", allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the superlative value", }, ["componly"] = boolean_param, ["suponly"] = boolean_param, }, func = function(args, data) -- Process the comparatives and superlatives. process_comparative_args(data, args, "adverbs") end, } local function escape(str) return (str:gsub("\\([:#])", "\\\\%1") :gsub("[:#]", "\\%0")) end local function canonicalize_plural(pl, pagename, pos) if pl == "+" then return escape(add_suffix(pagename, "s.plural", pos)) elseif pl == "++" then return escape(compute_plusplus_s_form(pagename, add_suffix(pagename, "s.plural", pos))) elseif pl == "*" then return escape(pagename) elseif pl == "ies" then if pagename:sub(-1) == "y" then return escape(pagename:gsub("e?y$", pl)) end error("Can't specify 'ies' plural unless the term ends with 'y'.") elseif pl == "s" or pl == "es" or pl == "'s" then return escape(pagename .. pl) end end local function do_nouns(args, data, pos) local pagename = data.pagename pos = pos or "noun" local plurals = parse_inflection(args, 1) local function insert_plurale_tantum_inflections(is_plural_only, originating_label) if args.sg[1] then insert_fixed_inflection(data, "normally plural", originating_label) parse_and_insert_inflection(data, args, "sg", "singular") elseif is_plural_only then insert_fixed_inflection(data, "plural only", originating_label) end if args.attr[1] then parse_and_insert_inflection(data, args, "attr", "attributive") end end local function first_pl_term() return plurals[1] and plurals[1].term or nil end if first_pl_term() == "p" then -- plurale tantum if plurals[2] then error("With plurale tantum noun, can't specify more than one plural") end data.genders = {"p"} -- this should auto-insert the correct 'pluralia tantum' category insert_plurale_tantum_inflections("plural only", plurals[1]) return end local function inscat(cat) insert(data.categories, langname .. " " .. cat) end local need_default_plural = pos == "noun" if first_pl_term() == "sp" then -- construed as singular or plural sp = remove(plurals, 1) -- Remove the "sp" but retain it for its decorations. inscat("nouns construed as singular or plural") data.genders = {"s", "p"} -- this should auto-insert the correct 'pluralia tantum' category insert_plurale_tantum_inflections(nil, sp) need_default_plural = false elseif first_pl_term() == "-" then -- Uncountable noun; may occasionally have a plural local hyphpl = remove(plurals, 1) -- Remove the "-" but retain for decorations. inscat("uncountable nouns") -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert_fixed_inflection(data, "usually <<uncountable>>", hyphpl) else insert_fixed_inflection(data, "<<uncountable>>", hyphpl) end need_default_plural = false elseif first_pl_term() == "#" then -- Usually countable (e.g., "grilled cheese") local hashpl = remove(plurals, 1) -- Remove the "#" but retain for decorations. insert_fixed_inflection(data, "usually <<countable>>", hashpl) inscat("uncountable nouns") inscat("countable nouns") -- If no plural was given, add a default one now if not plurals[1] then plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end elseif first_pl_term() == "~" then -- Mixed countable/uncountable noun, always has a plural local tildepl = remove(plurals, 1) -- Remove the "~" but retain for decorations. insert_fixed_inflection(data, "<<countable>> and <<uncountable>>", tildepl) inscat("uncountable nouns") inscat("countable nouns") -- If no plural was given, add a default one now if not plurals[1] then plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end end -- Plural is unknown if first_pl_term() == "?" then local questionpl = remove(plurals, 1) -- Remove the "?" but retain for decorations. -- Not desired; see [[Wiktionary:Tea_room/2021/August#"Plural unknown or uncertain"]] -- insert_fixed_inflection(data, "plural unknown or uncertain", questionpl) inscat("nouns with unknown or uncertain plurals") if plurals[1] then error("Can't specify explicit plurals along with '?' for unknown/uncertain plural") end return end -- Plural is not attested if first_pl_term() == "!" then local exclampl = remove(plurals, 1) -- Remove the "!" but retain for decorations. insert_fixed_inflection(data, "plural not attested", exclampl) inscat("nouns with unattested plurals") if plurals[1] then error("Can't specify explicit plurals along with '!' for unattested plural") end return end -- If no plural was given, maybe add a default one, otherwise (when "-" was given or proper noun) return. if not plurals[1] then if not need_default_plural then inscat("uncountable nouns") return end plurals[1] = {term = escape(add_suffix(pagename, "s.plural", pos))} end -- There are plural forms to show, so show them. inscat("countable nouns") local irregular, indeclinable for i, pl in ipairs(plurals) do local canon_pl = canonicalize_plural(pl.term, pagename, pos) if canon_pl then pl.term = canon_pl end local pl_term = get_link_page(pl.term, lang) if not (pagename:find(" ") or is_regular_plural(pl_term, pagename)) then irregular = true if pl_term == pagename then indeclinable = true end end end if irregular then inscat("nouns with irregular plurals") end if indeclinable then inscat("indeclinable nouns") end insert_inflection(data, plurals, "plural", "p") end -- Return the parameters to be used for nouns and proper nouns. Currently the same. local noun_params = { [1] = list_param, ["pl\1qual"] = {list = true, allow_holes = true, replaced_by = false, instead = "use <l:...> or <q:...> inline modifier on the plural", }, -- The following four only used for pluralia tantum (1=p) ["sg"] = list_param, ["attr"] = list_param, } pos_functions["nouns"] = { params = noun_params, func = do_nouns, } pos_functions["proper nouns"] = { params = noun_params, func = function(args, data) return do_nouns(args, data, "proper noun") end, } local function base_default_verb_forms(verb) return escape(add_suffix(verb, "s.verb")), escape(add_suffix(verb, "ing")), escape(add_suffix(verb, "d")) end local function default_verb_forms(verb) local full_s_form, full_ing_form, full_ed_form = base_default_verb_forms(verb) if verb:find(" ") then local first, rest = verb:match("^(.-)( .*)$") local first_s_form, first_ing_form, first_ed_form = base_default_verb_forms(first) return full_s_form, full_ing_form, full_ed_form, first_s_form .. rest, first_ing_form .. rest, first_ed_form .. rest, first, rest else return full_s_form, full_ing_form, full_ed_form, nil, nil, nil, nil, nil end end local function compute_double_last_cons_stem_of_split_verb(verb, ending) local first, rest = verb:match("^(.-)( .*)$") if not first then error("Verb '" .. verb .. "' must have a space in it to use **") end local last_cons = first:match("([bcdfghjklmnpqrstvwxyzBCDFGHJKLMNPQRSTVWXYZ])$") if not last_cons then error("First word '" .. first .. "' must end in a consonant to use **") end return first .. last_cons .. ending .. rest end local function check_non_nil_star_form(form, pagename) if form == nil then error("Verb '" .. pagename .. "' must have a space in it to use *, **, *l, *! or *'") end return form end local function sub_tilde(form, pagename) if not form then return nil end if form:find("~") then form = form:gsub("~", replacement_escape(pagename)) end return form end local deprecated_qual_replaced_by_inline_modifier = { list = true, allow_holes = true, replaced_by = false, instead = "use an inline modifier <q:...> or <l:...> on the value" } pos_functions["verbs"] = { params = { [1] = {list = "pres_3sg", disallow_holes = true}, ["pres_3sg\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [2] = {list = "pres_ptc", disallow_holes = true}, ["pres_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [3] = {list = "past", disallow_holes = true}, ["past\1_qual"] = deprecated_qual_replaced_by_inline_modifier, [4] = {list = "past_ptc", allow_holes = true}, ["past_ptc\1_qual"] = deprecated_qual_replaced_by_inline_modifier, ["noautolinkverb"] = boolean_param, ["angle_bracket"] = boolean_param, }, func = function(args, data) -- Get parameters local par1s local par2s = parse_inflection(args, {2, "pres_ptc"}) local par3s = parse_inflection(args, {3, "past"}) local par4s = parse_inflection(args, {4, "past_ptc"}) local pres_3sgs, pres_ptcs, pasts, past_ptcs local pagename = data.pagename ------------------------------------------- UTILITY FUNCTIONS #2 ------------------------------------------ -- These functions are used in both in the separate-parameter format and in the override params such as past_ptc2=. local full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed local lemma local function set_lemma_and_default_forms(the_lemma) lemma = the_lemma full_default_s, full_default_ing, full_default_ed, split_default_s, split_default_ing, split_default_ed, lemma_first, lemma_rest = default_verb_forms(the_lemma) end local function canonicalize_s_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_s elseif form == "*" then return check_non_nil_star_form(split_default_s, lemma) elseif form == "++" then return compute_plusplus_s_form(lemma, full_default_s) elseif form == "**" then if lemma:find("^[^ ]*[szx] ") then return compute_double_last_cons_stem_of_split_verb(lemma, "es") else return check_non_nil_star_form(split_default_s, lemma) end elseif form == "+!" then return lemma .. "s" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "s" .. lemma_rest elseif form == "+'" then return lemma .. "'s" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'s" .. lemma_rest elseif form == "+l" then if lemma:find("[szx]$") then return {{term = full_default_s, l = {"US"}}, {term = compute_plusplus_s_form(lemma, full_default_s), l = {"UK"}}} else return compute_plusplus_s_form(lemma, full_default_s) end elseif form == "*l" then if lemma:find("^[^ ]*[szx] ") then return {{term = check_non_nil_star_form(split_default_s, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "es"), l = {"UK"}}} else return check_non_nil_star_form(split_default_s, lemma) end else return sub_tilde(form, lemma) end end local function canonicalize_ing_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ing elseif form == "*" then return check_non_nil_star_form(split_default_ing, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ing" elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ing") elseif form == "+!" then return lemma .. "ing" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ing" .. lemma_rest elseif form == "+'" then return lemma .. "'ing" elseif form == "*'" then return check_non_nil_star_form(lemma_first) .. "'ing" .. lemma_rest elseif form == "+l" then return {{term = full_default_ing, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ing", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ing, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ing"), l = {"UK"}}} else return sub_tilde(form, lemma) end end local function canonicalize_ed_form(form) if form == "+" then error("Internal error: Should not see '+' here") elseif form == "^" then return full_default_ed elseif form == "*" then return check_non_nil_star_form(split_default_ed, lemma) elseif form == "++" then return compute_double_last_cons_stem(lemma) .. "ed" elseif form == "+!" then return lemma .. "ed" elseif form == "*!" then return check_non_nil_star_form(lemma_first) .. "ed" .. lemma_rest elseif form == "+'" then return {{term = lemma .. "'d"}, {term = lemma .. "'ed"}} elseif form == "*'" then return {{term = check_non_nil_star_form(lemma_first) .. "'d" .. lemma_rest}, {term = check_non_nil_star_form(lemma_first) .. "'ed" .. lemma_rest}} elseif form == "**" then return compute_double_last_cons_stem_of_split_verb(lemma, "ed") elseif form == "+l" then return {{term = full_default_ed, l = {"US"}}, {term = compute_double_last_cons_stem(lemma) .. "ed", l = {"UK"}}} elseif form == "*l" then return {{term = check_non_nil_star_form(split_default_ed, lemma), l = {"US"}}, {term = compute_double_last_cons_stem_of_split_verb(lemma, "ed"), l = {"UK"}}} else return sub_tilde(form, lemma) end end -- FIXME: options should be "+", "*", "++", "**", "+n", "*n", "++n" and "**n", but not "n" local function canonicalize_en_form(form) if form == "n" then track("n4") return add_suffix(lemma, "n") end return canonicalize_ed_form(form) end --------------------------------- MAIN PARSING/CONJUGATING CODE -------------------------------- local is_angle_bracket = args.angle_bracket if is_angle_bracket then if par2s[1] or par3s[1] or par4s[1] then error("Can't specify explicit values for 2=, 3= or 4= along with the angle-bracket format") end elseif is_angle_bracket == nil and not par2s[1] and not par3s[1] and not par4s[1] and not args[1][2] and args[1][1] and args[1][1]:find("<") then if put.term_contains_top_level_html(args[1][1]) then -- Often, term_contains_top_level_html() returns true on the angle-bracket format, which would -- make the pcall() below succeed but leave the angle brackets as-is. Check for this and only do the -- pcall() if term_contains_top_level_html() returns false. is_angle_bracket = true else -- If it's ambiguous whether it's an angle-bracket format or separate params with an inline modifier, -- try to parse as the latter. If an error occurs, treat as the former. local ok ok, par1s = pcall(parse_inflection, args, {1, "pres_3sg"}) if not ok then par1s = nil is_angle_bracket = true end end end if is_angle_bracket then -------------------------- ANGLE-BRACKET FORMAT -------------------------- -- (0) Expand multiword term with angle brackets just on the first word. local arg11 = args[1][1] if arg11:find("^<.*>$") and pagename:find(" ") then local first, rest = pagename:match("^(.-)( .*)$") arg11 = first .. arg11 .. rest end -- (1) Parse the indicator specs inside of angle brackets. local function parse_indicator_spec(angle_bracket_spec) local inside = angle_bracket_spec:match("^<(.*)>$") assert(inside) local segments = put.parse_balanced_segment_run(inside, "[", "]") local comma_separated_groups = put.split_alternating_runs(segments, ",") if #comma_separated_groups > 4 then error("Too many comma-separated parts in indicator spec, expected at most 4: " .. angle_bracket_spec) end local function fetch_footnotes(separated_group) local footnotes for j = 2, #separated_group - 1, 2 do if separated_group[j + 1] ~= "" then error("Extraneous text after bracketed footnotes: '" .. concat(separated_group) .. "'") end if not footnotes then footnotes = {} end insert(footnotes, separated_group[j]) end return footnotes end local function fetch_specs(comma_separated_group) if not comma_separated_group then return {{term = "+"}} end local specs = {} local colon_separated_groups = put.split_alternating_runs(comma_separated_group, ":") for _, colon_separated_group in ipairs(colon_separated_groups) do local form = colon_separated_group[1] if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error("*, **, *l, *! and *' not allowed inside of indicator specs: " .. angle_bracket_spec) end if form == "" then form = "+" end local termobj = { term = form } local footnotes = fetch_footnotes(colon_separated_group) if footnotes then for _, footnote in ipairs(footnotes) do m_headword_utilities.add_footnote_to_termobj(termobj, footnote) end end insert(specs, termobj) end return specs end local s_specs = fetch_specs(comma_separated_groups[1]) local ing_specs = fetch_specs(comma_separated_groups[2]) local ed_specs = fetch_specs(comma_separated_groups[3]) local en_specs = fetch_specs(comma_separated_groups[4]) return { forms = {}, s_specs = s_specs, ing_specs = ing_specs, ed_specs = ed_specs, en_specs = en_specs, } end local parse_props = { parse_indicator_spec = parse_indicator_spec, } local alternant_multiword_spec = iut.parse_inflected_text(arg11, parse_props) -- (2) Check for user-specified brackets; remove any links from the lemma, but remember the original -- form so we can use it below in the 'lemma_linked' form. -- Check to see if there are brackets in the pre-text or post-text. If so, use the linked lemma (with the -- verb autolinked unless noautolinkverb is given). Otherwise, use the default headword algorithm. local function check_bracket(val) if val:find("%[%[") then alternant_multiword_spec.saw_bracket = true end end for _, alternant_or_word_spec in ipairs(alternant_multiword_spec.alternant_or_word_specs) do check_bracket(alternant_or_word_spec.before_text) if alternant_or_word_spec.alternants then for _, multiword_spec in ipairs(alternant_or_word_spec.alternants) do for _, word_spec in ipairs(multiword_spec.word_specs) do check_bracket(word_spec.before_text) end check_bracket(multiword_spec.post_text) end end end check_bracket(alternant_multiword_spec.post_text) iut.map_word_specs(alternant_multiword_spec, function(base) if base.lemma == "" then base.lemma = pagename end base.orig_lemma = base.lemma base.lemma = remove_links(base.lemma) if args.noautolinkverb or base.orig_lemma:find("%[%[") then base.linked_lemma = base.orig_lemma else base.linked_lemma = "[[" .. base.orig_lemma .. "]]" end end) -- (3) Conjugate the verbs according to the indicator specs parsed above. local all_verb_slots = { lemma = "infinitive", lemma_linked = "infinitive", s_form = "3|s|pres", ing_form = "pres|ptcp", ed_form = "past", en_form = "past|ptcp", } local function conjugate_verb(base) local function process_specs(slot, specs, canon_func, default_values, default_already_formobj) local function insert_termobj_into_slot(termobj) local formobj = m_headword_utilities.convert_termobj_to_formobj(termobj) -- If the form is -, don't insert any forms, which will result in there being no overall forms -- (in fact it will be nil). We check for that down below and substitute a single "-" as the -- form, which in turn gets turned into special labels like "no present participle". if formobj.form == "-" then if formobj.footnotes then error("Unable to preserve footnotes specified on missing form '-': FIXME: " .. dump(formobj.footnotes)) end else iut.insert_form(base.forms, slot, formobj) end end local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg insert_termobj_into_slot(arg) else for _, canon in ipairs(canon_arg) do m_headword_utilities.combine_termobj_decorations(canon, arg) insert_termobj_into_slot(canon) end end end for _, arg in ipairs(specs) do if arg.term == "+" then if default_values then -- will be nil if past tense specified as - and no past ptc given for _, val in ipairs(default_values) do val = shallowCopy(val) if default_already_formobj then local argformobj = m_headword_utilities.convert_termobj_to_formobj(arg) val.footnotes = iut.combine_footnotes(val.footnotes, argformobj.footnotes) iut.insert_form(base.forms, slot, val) else m_headword_utilities.combine_termobj_decorations(val, arg) canonicalize_and_insert(val) end end end else canonicalize_and_insert(arg) end end end set_lemma_and_default_forms(base.lemma) local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then error(("Internal error: '%s' should have already thrown an error"):format(form)) end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_specs("s_form", base.s_specs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_specs("ing_form", base.ing_specs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_specs("ed_form", base.ed_specs, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_specs("en_form", base.en_specs, function(arg) return canonicalize_en_form(arg.term) end, base.forms.ed_form, "default already formobj") iut.insert_form(base.forms, "lemma", {form = base.lemma}) -- Add linked version of lemma for use in head=. We write this in a general fashion in case -- there are multiple lemma forms (which isn't possible currently at this level, although it's -- possible overall using the ((...,...)) notation). iut.insert_forms(base.forms, "lemma_linked", iut.map_forms(base.forms.lemma, function(form) if form == base.lemma and base.linked_lemma:find("%[%[") then return base.linked_lemma else return form end end)) end local inflect_props = { slot_table = all_verb_slots, inflect_word_spec = conjugate_verb, } iut.inflect_multiword_or_alternant_multiword_spec(alternant_multiword_spec, inflect_props) -- (4) Fetch the forms and put the conjugated lemmas in data.heads if not explicitly given. local function fetch_termobjs(slot) local forms = alternant_multiword_spec.forms[slot] -- See above. This should only occur if the user explicitly used - for a spec. if not forms or not forms[1] then return {{term = "-"}} end local termobjs = {} for _, formobj in ipairs(forms) do insert(termobjs, m_headword_utilities.convert_formobj_to_termobj(formobj)) end return termobjs end pres_3sgs = fetch_termobjs("s_form") pres_ptcs = fetch_termobjs("ing_form") pasts = fetch_termobjs("ed_form") past_ptcs = fetch_termobjs("en_form") -- Use the "linked" form of the lemma as the head if no head= explicitly given and the user specified -- brackets in one of the lemmas. Otherwise we use the default headword-linking algorithm. if not data.user_specified_heads[1] and alternant_multiword_spec.saw_bracket then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.lemma_linked) do insert(data.heads, m_headword_utilities.convert_formobj_to_termobj(lemma_obj)) end end else -------------------------- SEPARATE-PARAM FORMAT -------------------------- set_lemma_and_default_forms(pagename) par1s = par1s or parse_inflection(args, {1, "pres_3sg"}) pres_3sgs = {} pres_ptcs = {} pasts = {} past_ptcs = {} if not par1s[1] then par1s = {{term = "+"}} end if not par2s[1] then par2s = {{term = "+"}} end if not par3s[1] then par3s = {{term = "+"}} end if not par4s[1] then par4s = {{term = "+"}} end local function process_argument(args, dest, canon_func, default_values, default_already_canonicalized) local function canonicalize_and_insert(arg) local canon_arg = canon_func(arg) if type(canon_arg) == "string" then arg.term = canon_arg m_headword_utilities.insert_termobj_combining_duplicates(dest, arg) else for _, canon in ipairs(canon_arg) do m_headword_utilities.combine_termobj_decorations(canon, arg) m_headword_utilities.insert_termobj_combining_duplicates(dest, canon) end end end for _, arg in ipairs(args) do if arg.term == "+" then for _, val in ipairs(default_values) do val = shallowCopy(val) m_headword_utilities.combine_termobj_decorations(val, arg) if default_already_canonicalized then m_headword_utilities.insert_termobj_combining_duplicates(dest, val) else canonicalize_and_insert(val) end end else canonicalize_and_insert(arg) end end end local all_part_default_specs = {} local function process_and_canonicalize_s_form(arg) local form = arg.term if form == "+" then error("Internal error: '+' should have been converted to '^' by now") end if form == "^" or form == "++" or form == "+l" or form == "+!" or form == "+'" or form == "*" or form == "**" or form == "*l" or form == "*!" or form == "*'" then insert(all_part_default_specs, shallowCopy(arg)) end return canonicalize_s_form(form) end process_argument(par1s, pres_3sgs, process_and_canonicalize_s_form, {{term = "^"}}) if not all_part_default_specs[1] then all_part_default_specs[1] = {term = "^"} end process_argument(par2s, pres_ptcs, function(arg) return canonicalize_ing_form(arg.term) end, all_part_default_specs) process_argument(par3s, pasts, function(arg) return canonicalize_ed_form(arg.term) end, all_part_default_specs) process_argument(par4s, past_ptcs, function(arg) return canonicalize_en_form(arg.term) end, pasts, "default already canonicalized") end ------------------------------------------- INSERT INFLECTIONS ------------------------------------------ insert_inflection(data, pres_3sgs, "third-person singular simple present", "s-verb-form") insert_inflection(data, pres_ptcs, "present participle", "ing-form") if deepEquals(pasts, past_ptcs) then insert_inflection(data, pasts, "simple past and past participle", "ed-form", "no simple past or past participle") else insert_inflection(data, pasts, "simple past", "spast") insert_inflection(data, past_ptcs, "past participle", "past|part") end if pagename:find(" ") then -- Check for placeholder "it" local words = split(pagename, " ") for _, word in ipairs(words) do if word == "it" or word == "its" or word == "it's" then insert(data.categories, langname .. ' terms with placeholder "it"') break end end -- Check for phrasal verbs local phrasal_adverbs = list_to_set{ -- NOTE: This should only contain common phrasal adverbs, not random words like [[low]], -- [[adrift]], etc. "aback", "about", "above", "across", "after", "against", "ahead", "along", "apart", "around", "as", "aside", "at", "away", "back", "before", "behind", "below", "between", "beyond", "by", "down", "for", "forth", "from", "in", "into", "of", "off", "on", "onto", "out", "over", "past", "round", "through", "to", "together", "towards", "under", "up", "upon", "with", "without", } local allowed_non_adverb_words = list_to_set{ "it", "one", "oneself", "someone", } local base = pagename local seen_adverbs = {} -- Only consider a verb to be phrasal if it consists of a single base verb followed exclusively by either -- adverbs from `phrasal_adverbs` or placeholder words from `allowed_non_adverb_words`, where at -- least one following word is from `phrasal_adverbs` (hence [[can it]] is not a phrasal verb). while true do local prev, word = base:match("^(.+) (.-)$") if not prev then break end if phrasal_adverbs[word] then insert(seen_adverbs, word) elseif allowed_non_adverb_words[word] then -- do nothing else break end base = prev end if not base:find(" ") and seen_adverbs[1] then insert(data.categories, langname .. " phrasal verbs") for i = #seen_adverbs, 1, -1 do insert(data.categories, langname .. ' phrasal verbs formed with "' .. seen_adverbs[i] .. '"') end end end end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["suffix forms"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, }, func = function(args, data, is_suffix) local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. m_table.serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export 16kpxx6u8rw55bh0m9lw3iuegl8ptae Bysen:ang-decl-noun-a-n 10 8057 54867 2026-09-26T23:39:19Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''a''-stem<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}{{#if:{{{vowel|}}}||e}}s<!-- -->|7={{{datsg|{{{1}}}{{#if:{{{vowel|}}}||e}}}}}<!-- -->|2={{#if:{{{short|}}}|{{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}|{{{nomsg|{{{2|{{{1}}}}}}}}}}}<!-- -->|4={{#if:{{{short|}}}|{{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}|{{{nomsg|{{{2|{{{1}}}}}}}}}}}<!...' 54867 wikitext text/x-wiki {{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''a''-stem<!-- -->|1={{{nomsg|{{{1}}}}}}<!-- -->|3={{{nomsg|{{{1}}}}}}<!-- -->|5={{{1}}}{{#if:{{{vowel|}}}||e}}s<!-- -->|7={{{datsg|{{{1}}}{{#if:{{{vowel|}}}||e}}}}}<!-- -->|2={{#if:{{{short|}}}|{{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}|{{{nomsg|{{{2|{{{1}}}}}}}}}}}<!-- -->|4={{#if:{{{short|}}}|{{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}|{{{nomsg|{{{2|{{{1}}}}}}}}}}}<!-- -->|6={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}|n}}a<!-- -->|8={{{2|{{{1}}}}}}{{#if:{{{vowel|}}}||u}}m{{#if:{{{vowel|}}}|,{{{2|{{{1}}}}}}um}}<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|neuter a-stem nouns}}<!-- --><noinclude>{{documentation}}</noinclude> 7lekjiluuj7dd7p3vflzj1f7a4envzy æx 0 8058 54868 2026-09-26T23:50:22Z Deadend0914 7211 Gesceop tramet þe hafaþ '=={{sprǣc|ang}}== ===Āwendednessa=== * æces * acas * eax ===Rihtstefn=== * IPA: /æks/ ===Wordstǣr=== Of Ealdoric Germanisce *akusi ===Wiflic Nama=== {{ang-noun|wif}} # trēow fiellende tōl, wiþ weċġ-blæd on ānum ende, handel on þām oðrum ====Declīnung==== {{ang-decl-noun-i-f|æx}} ====Wendunga==== Nīwenglisc: [[axe]]' 54868 wikitext text/x-wiki =={{sprǣc|ang}}== ===Āwendednessa=== * æces * acas * eax ===Rihtstefn=== * IPA: /æks/ ===Wordstǣr=== Of Ealdoric Germanisce *akusi ===Wiflic Nama=== {{ang-noun|wif}} # trēow fiellende tōl, wiþ weċġ-blæd on ānum ende, handel on þām oðrum ====Declīnung==== {{ang-decl-noun-i-f|æx}} ====Wendunga==== Nīwenglisc: [[axe]] 3xfffrij51sa585184mjo8a2cu0u01i Bysen:ang-decl-noun-i-f 10 8059 54869 2026-09-26T23:50:44Z Deadend0914 7211 Gesceop tramet þe hafaþ '{{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''i''-stem<!-- -->|1={{{1}}}{{#if:{{{short|}}}|e}}<!-- -->|3={{#if:{{{short|}}}|{{{1}}}e|{{{1}}},{{{1}}}e}}<!-- -->|5={{{1}}}e<!-- -->|7={{{1}}}e<!-- -->|2={{{1}}}e,{{{1}}}a<!-- -->|4={{{1}}}e,{{{1}}}a<!-- -->|6={{{1}}}a<!-- -->|8={{{1}}}um<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|i-stem nouns}}<!-- --><noinclude>{{document...' 54869 wikitext text/x-wiki {{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strong ''i''-stem<!-- -->|1={{{1}}}{{#if:{{{short|}}}|e}}<!-- -->|3={{#if:{{{short|}}}|{{{1}}}e|{{{1}}},{{{1}}}e}}<!-- -->|5={{{1}}}e<!-- -->|7={{{1}}}e<!-- -->|2={{{1}}}e,{{{1}}}a<!-- -->|4={{{1}}}e,{{{1}}}a<!-- -->|6={{{1}}}a<!-- -->|8={{{1}}}um<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|i-stem nouns}}<!-- --><noinclude>{{documentation}}</noinclude> s9jo2waelt0u6j4bw9b3wc7ud6kfjqu 54872 54869 2026-09-26T23:52:36Z Deadend0914 7211 54872 wikitext text/x-wiki {{#invoke:checkparams|error}}<!-- Validate template parameters -->{{ang-decl-noun<!-- -->|type=strang ''i''-stefn<!-- -->|1={{{1}}}{{#if:{{{short|}}}|e}}<!-- -->|3={{#if:{{{short|}}}|{{{1}}}e|{{{1}}},{{{1}}}e}}<!-- -->|5={{{1}}}e<!-- -->|7={{{1}}}e<!-- -->|2={{{1}}}e,{{{1}}}a<!-- -->|4={{{1}}}e,{{{1}}}a<!-- -->|6={{{1}}}a<!-- -->|8={{{1}}}um<!-- -->|num={{{num|}}}<!-- -->|title={{{title|}}}<!-- -->}}<!-- -->{{cln|ang|i-stem nouns}}<!-- --><noinclude>{{documentation}}</noinclude> m47jb8xoepc1myu9tdv9ic8i6x7vhtl