Wikiwordbōc angwiktionary https://ang.wiktionary.org/wiki/H%C4%93afodtramet MediaWiki 1.47.0-wmf.21 case-sensitive Media Syndrig Mōtung Brūcend Brūcendmōtung Wikiwordbōc Wikiwordbōcmōtung Ymele Ymelmōtung MediaWiki MediaWikimōtung Bysen Bysenmōtung Help Helpmōtung Flocc Floccmōtung Ætēaca Ætēacmōtung TimedText TimedText talk Module Module talk Event Event talk Wikiwordbōc:Requests for adminship 4 917 54885 48763 2026-09-27T22:30:33Z Deadend0914 7211 /* Ābiddunga Bewitanhāda */ 54885 wikitext text/x-wiki ==Bewitan== #[[User:Espreon|Espreon]] ==Ābiddunga Bewitanhāda== Current requests for adminship: ===[[User:Espreon|Espreon]]=== Given that no one really contributes to this Wiktionary anymore, the fact that it has become an unstandardized mess, and that no one really seems to care about it enough to clean it up right now, I've recently decided to devote a lot of my free time to cleaning it up. Though, of course, without powers such as the ability to delete pages (mostly in order to remove redirects left behind after moving pages), I can only do so much. Additionally, having someone with power here would help get the occasional spam eliminated more quickly. I therefore ask that I be granted adminship so that this Wiktionary will get all the care that it needs ''now''. I thank you for your time and consideration. — [[User:Espreon|Espreon]] ([[User talk:Espreon|talk]]) 22:16, 17 Mǣdmōnaþ 2013 (UTC) * '''Support''', and you can always ask me if you need help with something specific (as I'm a global sysop). [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 23:23, 17 Mǣdmōnaþ 2013 (UTC) * '''Support''', I know Espreon to do good work in Old English. And this Wiktionary could do with a cleanup. [[User:Gottistgut|Gottistgut]] ([[User talk:Gottistgut|talk]]) 22:06, 18 Mǣdmōnaþ 2013 (UTC) ===[[User:Espreon|Espreon]] (renewal)=== He has done a good job, and is still cleaning up and reorganizing the project. We recently made a lot of progress with localization and configuration changes, including the introduction of an "Appendix" (Æteaca:) namespace. He should continue his work here. [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 04:28, 24 Blōtmōnaþ 2013 (UTC) * '''Support''' as nominator. [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 04:28, 24 Blōtmōnaþ 2013 (UTC) *'''Support''' Espreon is the driving force behind the advancement of ang Wikiwordboc at the moment. [[User:Gottistgut|Gottistgut]] ([[User talk:Gottistgut|talk]]) 20:31, 24 Blōtmōnaþ 2013 (UTC) :[//meta.wikimedia.org/w/index.php?title=Steward_requests/Permissions&diff=6566614&oldid=6566171 Done.] [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 15:37, 1 Gēolmōnaþ 2013 (UTC) ===[[User:Espreon|Espreon]] (renewal 2)=== Same as last time. He should be able to remain an admin, since he is really the only one maintaining this wiki. [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 03:42, 2 Ēastermōnaþ 2014 (UTC) * '''Support''' as nom. [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 03:42, 2 Ēastermōnaþ 2014 (UTC) * '''Support'''. [[User:Gottistgut|Gottistgut]] ([[User talk:Gottistgut|talk]]) 22:18, 13 Ēastermōnaþ 2014 (UTC) ===[[User:Espreon|Espreon]] (renewal 3)=== I am nominating myself for renewed adminship; I am still around, and I am still the only one working on this. [[User:Espreon|Espreon]] ([[User talk:Espreon|talk]]) 18:55, 19 Winterfylleþ 2014 (UTC) * '''Support'''. [[User:PiRSquared17|PiRSquared17]] ([[User talk:PiRSquared17|talk]]) 21:13, 19 Winterfylleþ 2014 (UTC) * '''Support'''. Espreon is the main active editor here. [[User:Gottistgut|Gottistgut]] ([[User talk:Gottistgut|talk]]) 20:17, 23 Winterfylleþ 2014 (UTC) ===[[User:Blackkdark|Blackkdark]]=== I'd like a chance at Adminship. I think I'm going to create a structure we should follow for all entries. I'm going to start putting it in motion. --[[Brūcend:Blackkdark|Timoði Pætricus Snīðer]] ([[Brūcendmōtung:Blackkdark|mōtung]]) 06:20, 4 Solmōnaþ 2016 (UTC) * '''Comment''': I would much rather you follow the structure used on the Modern English Wiktionary (as I have done) and that you not leave behind so many empty sections, as you have done in the past (... not to mention what you did last week at [[fremman]]). [[Brūcend:Espreon|Espreon]] ([[Brūcendmōtung:Espreon|mōtung]]) 18:42, 13 Solmōnaþ 2016 (UTC) ===[[User:Deadend0914|Deadend0914]]=== I'd like to be the admin here. The format of a lot of pages is weird and I really think this wiki relies way too heavily on modern English (ie. how every page is supposed to have a spot where modern English/Nīwenglisc translations must be). I'd like to standardize from a dictionary made for English speakers to one of Engliscum sprecum, add more scripts (from the English wiki) to make adding pages easier, and get the word Wikiwordboc into the mouths and brains of many more people out there. ==Sēo ēac== [[m:requests for permissions]] fky3476erb1ou9z70di8xbuo8dj5vxc Module:parameter utilities 828 7860 54880 54378 2026-09-27T21:40:11Z Deadend0914 7211 54880 Scribunto text/plain local export = {} local dump = mw.dumpObject local parameters_module = "Module:parameters" local parse_utilities_module = "Module:parse utilities" local table_module = "Module:table" local function track(page, track_module) return require("Module:debug/track")((track_module or "parameter utilities") .. "/" .. page) end -- Throw an error prefixed with the words "Internal error" (and suffixed with a dumped version of `spec`, if provided). -- This is for logic errors in the code itself rather than template user errors. local function internal_error(msg, spec) if spec then msg = ("%s: %s"):format(msg, dump(spec)) end error(("Internal error: %s"):format(msg)) end -- Table listing the recognized special separator arguments and how they display. local special_separators = { [";"] = "; ", ["_"] = " ", ["~"] = " ~ ", } --[==[ intro: The purpose of this module is to facilitate implementation of a template that takes a list of items with associated properties, which can be specified either through separate parameters (e.g. {{para|t2}}, {{para|pos3}}) or inline modifiers (`<t:...>`, `<pos:...>`, etc.). Some examples of templates that work this way are {{tl|alter}}/{{tl|alt}}; {{tl|synonyms}}/{{tl|syn}}, {{tl|antonyms}}/{{tl|ant}}, and other "nyms" templates; {{tl|col}}, {{tl|col2}}, {{tl|col3}}, {{tl|col4}} and other columns templates; {{tl|descendant}}/{{tl|desc}}; {{tl|affix}}/{{tl|af}}, {{tl|prefix}}/{{tl|pre}} and related *fix templates; {{tl|affixusex}}/{{tl|afex}} and related templates; {{tl|IPA}}; {{tl|homophones}}; {{tl|rhymes}}; and several others. This module can be thought of as a combination of [[Module:parameters]] (which parses template parameters, and in particular handles the separate parameter versions of the properties) and `parse_inline_modifiers()` in [[Module:parse utilities]] (which parses inline modifiers). The main entry point is `process_list_arguments()`, which takes an object specifying various properties and returns a list of objects, one per item specified by the user, where the individual objects are much like the objects returned by `parse_inline_modifiers()`. However, there are other functions provided, in particular to initialize the `param_mods` structured that is passed to `process_list_arguments()`. The typical workflow for using this module looks as follows (a slightly simplified version of the code in [[Module:nyms]]): { local export = {} local parameter_utilities_module = "Module:parameter utilities" ... -- Entry point to be invoked from a template. function export.show(frame) local parent_args = frame:getParent().args -- Parameters that don't have corresponding inline modifiers. Note in particular that the parameter corresponding to -- the items themselves must be specified this way, and must specify either `allow_holes = true` (if the user can -- omit terms, typically by specifying the term using |altN= or <alt:...> so that they remain unlinked) or -- `disallow_holes = true` (if omitting terms is not allowed). (If neither `allow_holes` nor `disallow_holes` is -- specified, an error is thrown in process_list_arguments().) local params = { [1] = {required = true, type = "language", default = "und"}, [2] = {list = true, allow_holes = true, required = true, default = "term"}, } local m_param_utils = require(parameter_utilities_module) -- This constructs the `param_mods` structure by adding well-known groups of parameters (such as all the parameters -- associated with based on full_link() in [[Module:links]], with default properties that can be overridden. This is -- easier and less error-prone than manually specifying the `param_mods` structure (see below for how this would -- look). Here, we specify the group "link" (consisting of all the link parameters for use with full_link()), group -- "ref" (which adds the "ref" parameter for specifying references), group "l" (which adds the "l" and "ll" -- parameters for specifying labels) and group "q" (which adds the "q" and "qq" parameters for specifying regular -- qualifiers). By default, labels and qualifiers have `separate_no_index` set so that e.g. |q1= is distinct from -- |q=, the former specifying the left qualifier for the first item and the latter specifying the overall left -- qualifier. For compatibility, we override the `separate_no_index` setting for the group "q", which causes |q= and -- |q1= to be the same, and likewise for |qq= and |qq1=. Finally, also for compatibility, we add an "lb" parameter -- that is an alias of "ll" (in all respects; |lb= is the same as |ll=, |lb1= is the same as |ll1=, <lb:...> is the -- same as <ll:...>, etc.). local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l"}}, {group = "q", separate_no_index = false}, {param = "lb", alias_of = "ll"}, } -- This processes the raw arguments in `parent_args`, parses inline modifiers and creates corresponding objects -- containing the property values specified either through inline modifiers or separate parameters. local items, args = m_param_utils.process_list_arguments { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, parse_lang_prefix = true, track_module = "nyms", lang = 1, sc = "sc.default", } local lang = args[1] -- Now do the actual implementation of the template. Generally this should be split into a separate function, often -- in a separate module (if the implementation goes in [[Module:foo]], the template interface code goes in -- [[Module:foo/templates]]). ... } The `param_mods` structure controls the properties that can be specified by the user for a given item, and is conceptually very similar to the `param_mods` structure used by `parse_inline_modifiers()`. The key is the name of the parameter (e.g. {"t"}, {"pos"}) and the value is a table with optional elements as follows: * `item_dest`, `store`: Same as the corresponding fields in the `param_mods` structure passed to `parse_inline_modifiers()`. * `type`, `set`, `sublist`, `convert` and associated fields such as `family` and `method`: These control parsing and conversion of the raw values specified by the user and have the same meaning as in [[Module:parameters]] and also in `parse_inline_modifiers()` (which delegates the actual conversion to [[Module:parameters]]). These fields — and for that matter, all fields other than `item_dest`, `store` and `overall` — are forwarded to the `process()` function in [[Module:parameters]]. * `alias_of`: This parameter is an alias of some other parameter. This spec is recognized only by `process()` in [[Module:parameters]], and not by `parse_inline_modifiers()`; to set up an alias in `parse_inline_modifiers()`, you need to make sure (using `item_dest`) that both the alias and aliasee modifiers store their values in the same location, and you need to copy the remaining properties from the aliasee's spec to the aliasing modifier's spec. All of this happens automatically if you generate the `param_mods` structure using `construct_param_mods()`. * `require_index`: This means that the non-indexed parameter version of the property is not recognized. E.g. in the case of the {"sc"} property, use of the {{para|sc}} parameter would result in an error, while {{para|sc1}} is recognized and specifies the {"sc"} property for the first item. The default, if neither `require_index` nor `separate_no_index` is given, is for {{para|sc}} and {{para|sc1}} to mean the same thing (both would specify the {"sc"} property of the first item). Note that `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `separate_no_index`: This means that e.g. the {{para|sc}} parameter is distinct from the {{para|sc1}} parameter (and thus from the `<sc:...>` inline modifier on the first item). This is typically used to distinguish an overall version of a property from the corresponding item-specific property on the first item. (In this case, for example, {{para|sc}} overrides the script code for all items, while {{para|sc1}} overrides the script code only for the first item.) If not given, and if `require_index` is not given, {{para|sc}} and {{para|sc1}} would have the same meaning and refer to the item-specific property on the first item. When this is given, the overall value can be accessed using the `.default` field of the property value in `args`, e.g. in this case `args.sc.default`. Note that (as mentioned above) `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `list`, `allow_holes`, `disallow_holes`: These should '''not''' be given. `list` and `allow_holes` are automatically set for all parameter specs added to the `params` structure used by `process()` in [[Module:parameters]], and `disallow_holes` clashes with `allow_holes`. For the above workflow example, the call to `construct_param_mods()` generates the following `param_mods` structure: { local param_mods = { -- the parameters generated by group "link" alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", sublist = true, }, pos = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, -- the parameters generated by group "ref" ref = { item_dest = "refs", type = "references", }, -- the parameters generated by group "l" l = { type = "labels", separate_no_index = true, }, ll = { type = "labels", separate_no_index = true, }, -- the parameters generated by group "q"; note that `separate_no_index = true` would be set, but is overridden -- (specifying `separate_no_index = false` in the `param_mods` structure is equivalent to not specifying it at all) q = { type = "qualifier", separate_no_index = false, }, qq = { type = "qualifier", separate_no_index = false, }, -- the parameter generated by the individual "lb" parameter spec; note that only `alias_of` was explicitly given, -- while `item_dest` is automatically set so that inline modifier <lb:...> stores into the same place as <ll:...>, -- and the other specs are copied from the `ll` spec so `lb` works like `ll` in all regards lb = { alias_of = "ll", item_dest = "ll", type = "labels", separate_no_index = true, }, } } ]==] local qualifier_spec = { type = "qualifier", separate_no_index = true, } local label_spec = { type = "labels", separate_no_index = true, } local recognized_param_mod_groups = { link = { alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", sublist = true, }, pos = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, }, lang = { lang = { require_index = true, type = "language", }, }, q = { q = qualifier_spec, qq = qualifier_spec, }, a = { a = label_spec, aa = label_spec, }, l = { l = label_spec, ll = label_spec, }, ref = { ref = { item_dest = "refs", type = "references", }, }, } local function merge_param_mod_settings(orig, additions) local merged = require(table_module).shallowCopy(orig) for k, v in pairs(additions) do merged[k] = v if k == "require_index" then merged.separate_no_index = nil elseif k == "separate_no_index" then merged.require_index = nil end end merged.default = nil merged.group = nil merged.param = nil merged.exclude = nil merged.include = nil return merged end local function verify_type(spec, param, typ1, typ2) if not spec[param] then return end local val = spec[param] if type(val) ~= typ1 and (not typ2 or type(val) ~= typ2) then internal_error(("Parameter `%s` must be a %s%s but saw a %s"):format(param, typ1, typ2 and " or " .. typ2 or "", type(val)), spec) end end local function verify_well_constructed_spec(spec) local num_control = (spec.default and 1 or 0) + (spec.group and 1 or 0) + (spec.param and 1 or 0) if num_control == 0 then internal_error( "Spec passed to construct_param_mods() must have either the `default`, `group` or `param` keys set", spec) end if num_control > 1 then internal_error( "Exactly one of `default`, `group` or `param` must be set in construct_param_mods() spec", spec) end if spec.list or spec.allow_holes then -- FIXME: We need to support list = "foo" for list parameters that are stored in e.g. 2=, foo2=, foo3=, etc. internal_error("`list` and `allow_holes` may not be set; they are automatically set when constructing the " .. "corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.disallow_holes then internal_error("`disallow_holes` may not be set; it conflicts with `allow_holes`, which is automatically " .. "set when constructing the corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.include and spec.exclude then internal_error("Saw both `include` and `exclude` in the same spec", spec) end if (spec.include or spec.exclude) and not spec.group then internal_error( "`include` and `exclude` can only be specified along with `group`, not with `default` or `param`", spec) end verify_type(spec, "group", "string", "table") verify_type(spec, "param", "string", "table") verify_type(spec, "include", "table") verify_type(spec, "exclude", "table") end --[==[ Construct the `param_mods` structure used in parsing arguments and inline modifiers from a list of specifications. A sample invocation (a slightly simplified version of the actual invocation associated with {{tl|affix}} and related templates) looks like this: { local param_mods = require("Module:parameter utilities").construct_param_mods { -- We want to require an index for all params (or use separate_no_index, which also requires an index for the -- param corresponding to the first item). {default = true, require_index = true}, {group = {"link", "ref", "lang", "q", "l"}}, -- Override these two to have separate_no_index. {param = {"lit", "pos"}, separate_no_index = true}, } } Each specification either sets the default value for further parameter specs or adds one or more parameters. Parameters can be added directly using `param`, or groups of predefined parameters can be added using `group`. Specifications are one of three types: # Those that set the default properties for future-added parameters. These contain {default = true} as one of the properties of the spec. Specs are processed in order and you can change the defaults mid-way through. # Those that add the parameters associated with one or more pre-defined groups. These contain {group = "group"} or {group = {"group1", "group2", ...}}. The pre-defined parameter groups and their associated properties are listed below. The pre-defined properties of parameters in a group override properties associated with a {default = true} spec, and are in turn overridden by any properties given directly in the spec itself. Note as well that setting the `separate_no_index` property will automatically cause the `require_index` property to be unset and vice-versa, as the two are mutually exclusive. (This happens in the example above, where the {separate_no_index = true} setting associated with the params {"lit"} and {"pos"} cancels out the {require_index = true} default setting, as well as less obviously with the pre-defined {"sc"} property of the {"link"} group, the {"q"} and {"qq"} properties of the {"q"} group, and the {"l"} and {"ll"} properties of the {"l"} group, all of which have an associated pre-defined property {separate_no_index = true}, which overrides and cancels out the {require_index = true} default setting. Finally, when adding the parameters of a group, you can request the only a subset of the parameters be added using either the `include` or `exclude` properties, each of whose values is a list of parameters that specify (respectively) the parameters to include (all other parameters of the group are excluded) or to exclude (all other parameters of the group are included). This is used, for example, in [[Module:romance etymology]] and [[Module:it-etymology]], which specify {group = "link", exclude = {"tr", "ts", "sc"}} to exclude link parameters that aren't relevant to Latin-script languages such as the Romance languages, and conversely in [[Module:IPA/templates]], which specifies {group = "link", include = {"t", "gloss", "pos"}} to include only the specified parameters for use with {{tl|IPA}}. # Those that add individual parameters. These contain {param = "param"} or {param = {"param1", "param2", ...}}, the latter syntax used to control a set of parameters together. The resulting spec is formed by initializing the parameter's settings with any previously-specified default properties (using a spec containing {default = true}) if the parameter hasn't already been initialized, and then overriding the resulting settings with any settings given directly in the specification. In the above example, the {"lit"} and {"pos"} parameters were previously initialized through the {"link"} group (specified in the second of the three specifications) but ended up with {require_index = true} due to the {default = true} spec (the first of the three specifications). We override these two parameters to have {separate_no_index = true} (which, as mentioned above, cancels out {require_index = true}). This is done so that {{tl|affix}} and related templates have {{para|pos}} and {{para|lit}} parameters distinct from {{para|pos1}} and {{para|lit1}}, which are used to specify an overall part of speech (which applies to all parts of the affix, as opposed to applying to just one element of the expression) or a literal definition for the entire expression (instead of just for one element of the expression). The built-in parameter groups are as follows: {|class="wikitable" ! Group !! Group meaning !! Parameter !! Parameter meaning !! Default properties |- | rowspan=10| `link` | rowspan=10| link parameters; same as those available on {{tl|l}}, {{tl|m}} and other linking templates | `alt` || display text, overriding the term's display form || — |- | `t` || gloss (translation) of a non-English term || {item_dest = "gloss"} |- | `gloss` || gloss (translation); same as `t` || {alias_of = "t"} |- | `tr` || transliteration of a non-Latin-script term; only needed if the automatic transliteration is incorrect or unavailable (e.g. in Hebrew, which doesn't have automatic transliteration) || — |- | `ts` || transcription of a non-Latin-script term, if the transliteration is markedly different from the actual pronunciation; should not be used for IPA pronunciations || — |- | `g` || comma-separated list of genders; whitespace may surround the comma and will be ignored || {item_dest = "genders", sublist = true} |- | `pos` || part of speech for the term || — |- | `lit` || literal meaning (translation) of the term || — |- | `id` || a sense ID for the term, which links to anchors on the page set by the {{tl|senseid}} template || — |- | `sc` || the script code (see [[Wiktionary:Scripts]]) for the script that the term is written in; rarely necessary, as the script is autodetected (in most cases, correctly) || {separate_no_index = true, type = "script"} |- | rowspan=2| `q` | rowspan=2| left and right normal qualifiers (as displayed using {{tl|q}}) | `q` || left normal qualifier || {separate_no_index = true, type = "qualifier"} |- | `qq` || right normal qualifier || {separate_no_index = true, type = "qualifier"} |- | rowspan=2| `a` | rowspan=2| left and right accent qualifiers (as displayed using {{tl|a}}) | `a` || comma-separated list of left accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `aa` || comma-separated list of right accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | rowspan=2| `l` | rowspan=2| left and right labels (as displayed using {{tl|lb}}, but without categorizing) | `l` || comma-separated list of left labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ll` || comma-separated list of right labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ref` | reference(s) (in the format accepted by [[Module:references]]; see also the documentation for the {{para|ref}} parameter to {{tl|IPA}}) | `ref` || one or more references, in the format accepted by [[Module:references]] || {item_dest = "refs", type = "references"} |- | `lang` | language for an individual term (provided for compatibility; it is preferred to specify languages for individual terms using language prefixes instead) | `lang` || language code (see [[Wiktionary:Languages]]) for the term || {require_index = true, type = "language"} |} ]==] function export.construct_param_mods(specs) local param_mods = {} local default_specs = {} for _, spec in ipairs(specs) do verify_well_constructed_spec(spec) if spec.default then -- This will have an extra `default` field in it, but it will be erased by merge_param_mod_settings() default_specs = spec else if spec.group then local groups = spec.group if type(groups) ~= "table" then groups = {groups} end local include_set if spec.include then include_set = require(table_module).listToSet(spec.include) end local exclude_set if spec.exclude then exclude_set = require(table_module).listToSet(spec.exclude) end for _, group in ipairs(groups) do local group_specs = recognized_param_mod_groups[group] if not group_specs then internal_error(("Unrecognized built-in param mod group '%s'"):format(group), spec) end for group_param, group_param_settings in pairs(group_specs) do local include_param if include_set then include_param = include_set[group_param] elseif exclude_set then include_param = not exclude_set[group_param] else include_param = true end if include_param then local merged_settings = merge_param_mod_settings(merge_param_mod_settings( param_mods[group_param] or default_specs, group_param_settings), spec) param_mods[group_param] = merged_settings end end end end if spec.param then local params = spec.param if type(params) ~= "table" then params = {params} end for _, param in ipairs(params) do local settings = merge_param_mod_settings(param_mods[param] or default_specs, spec) -- If this parameter is an alias of another parameter, we need to copy the specs from the other -- parameter, since parse_inline_modifiers() doesn't know about `alias_of` and having the specs -- duplicated won't cause problems for [[Module:parameters]]. We also need to set `item_dest` to -- point to the `item_dest` of the aliasee (defaulting to the aliasee's value itself), so that -- both modifiers write to the same location. Note that this works correctly in the common case of -- <t:...> with `item_dest = "gloss"` and <gloss:...> with `alias_of = "t"`, because both will end -- up with `item_dest = "gloss"`. local aliasee = settings.alias_of if aliasee then local aliasee_settings = param_mods[aliasee] if not aliasee_settings then internal_error(("Undefined aliasee '%s'"):format(aliasee), spec) end for k, v in pairs(aliasee_settings) do if settings[k] == nil then settings[k] = v end end if settings.item_dest == nil then settings.item_dest = aliasee end end param_mods[param] = settings end end end end return param_mods end -- Return true if `k` is a "built-in" (specially recognized) key in a `param_mod` specification. All other keys -- are forwarded to the structure passed to [[Module:parameters]]. local function param_mod_spec_key_is_builtin(k) return k == "item_dest" or k == "overall" or k == "store" end --[==[ Convert the properties in `param_mods` into the appropriate structures for use by `process()` in [[Module:parameters]] and store them in `params`. If `overall_only` is given, only store the properties in `param_mods` that correspond to overall (non-item-specific) parameters. Currently this only happens when `separate_no_index` is specified. ]==] function export.augment_params_with_modifiers(params, param_mods, overall_only) if overall_only then for param_mod, param_mod_spec in pairs(param_mods) do if param_mod_spec.separate_no_index then local param_spec = {} for k, v in pairs(param_mod_spec) do if k ~= "separate_no_index" and not param_mod_spec_key_is_builtin(k) then param_spec[k] = v end end params[param_mod] = param_spec end end else local list_with_holes = { list = true, allow_holes = true } -- Add parameters for each term modifier. for param_mod, param_mod_spec in pairs(param_mods) do local has_extra_specs = false for k, _ in pairs(param_mod_spec) do if not param_mod_spec_key_is_builtin(k) then has_extra_specs = true break end end if not has_extra_specs then params[param_mod] = list_with_holes else local param_spec = mw.clone(list_with_holes) for k, v in pairs(param_mod_spec) do if not param_mod_spec_key_is_builtin(k) then param_spec[k] = v end end params[param_mod] = param_spec end end end end --[==[ Return true if `k`, a key in an item, refers to a property of the item (is not one of the specially stored values). Note that `lang` and `sc` are considered properties of the item, although `lang` is set when there's a language prefix and both `lang` and `sc` may be set from default values specified in the `data` structure passed into `process_list_arguments()`. If you don't want these treated as property keys, you need to check for them yourself. ]==] function export.item_key_is_property(k) return k ~= "term" and k ~= "termlang" and k ~= "termlangs" and k ~= "itemno" and k ~= "orig_index" and k ~= "separator" end -- Fetch the argument in `args` corresponding to `index_or_value`, which may be a string of the form "foo.default" -- (requesting the value of `args["foo"].default`); a string or number (requesting the value at that key); a function of -- one argument (`args`), which returns the argument value; or the value itself. local function fetch_argument(args, index_or_value) if type(index_or_value) == "string" then local index_without_default = index_or_value:match("^(.*)%.default$") if index_without_default then local arg_obj = fetch_argument(args, index_without_default) if type(arg_obj) ~= "table" then internal_error(("Requested that the '.default' key of argument `%s` be fetched, but argument value is undefined or not a table"): format(index_without_default), arg_obj) end return arg_obj.default end if index_or_value:find("^[0-9]+$") then index_or_value = tonumber(index_or_value) end return args[index_or_value] elseif type(index_or_value) == "number" then return args[index_or_value] elseif type(index_or_value) == "function" then return index_or_value(args) else return index_or_value end end --[==[ Parse inline modifiers and create corresponding item objects containing the property values specified either through inline modifiers or separate parameters. `data` is an object containing the following properties: * `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from {frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. * `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both) of `raw_args` and `processed_args` must be set. * `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified manually. * `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters, '''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically "augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the "augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by [[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be correctly handled.) * `termarg` ('''required'''): The argument containing the first item with attached inline modifiers to be parsed. Usually a numeric value such as {1} or {2}. * `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used internally to track pages containing template invocations with certain properties. Example properties tracked are missing items with corresponding properties as well as missing items without corresponding properties (which are skipped entirely). To find out the exact properties tracked and the name of the tracking pages, read the code. * `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed arguments. It should make modifications in-place. * `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language prefixes are stripped off. Defaults to {"term"}. * `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in [[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a comma-separated list of language prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the `termlangs` field, and the first specified object is stored in the `lang` field). * `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple comma-separated language code prefixes can be given. See `parse_lang_prefix` above. * `allow_bad_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an error and will tend to remain unfixed. * `lang`: The language object for the language of the items, or the name of the argument to fetch the object from. In general it is not necessary to specify this as `process_list_arguments()` only initializes items based on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used for certain purposes: *# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language prefix). *# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same language code may appear in several items, such as with {{tl|col}} and related templates). The value of `lang` can be any of the following: * If a string of the form "foo.default", it is assumed to be requesting the value of `args["foo"].default`. * Otherwise, if a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching (this also happens with a spec like "2.default"). * Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed to the function as its only argument. * Otherwise, it is used directly. * `sc`: The script object for the items, or the name of the argument to fetch the object from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned items if not otherwise set (e.g. by the {{para|sc<var>N</var>}} parameter or `<sc:...>` inline modifier). * `disallow_custom_separators`: If specified, disallow specifying a bare semicolon as an item value to indicate that the item's previous separator should be a semicolon. By default, the previous separator of each item is considered to be an empty string (for the first item) and otherwise a comma + space, unless either the preceding item is a bare semicolon (which causes the following item's previous separator to be a semicolon + space) or an item has an embedded comma in it (which causes ''all'' items other than the first to have their previous separator be a semicolon + space). The previous separator of each item is set on the item's `separator` property. Bare semicolons do not count when indexing items using separate parameters. For example, the following is correct: ** {{tl|template|lang|item 1|q1=qualifier 1|;|item 2|q2=qualifier 2}} If `disallow_custom_separators` is specified, however, the `separator` property is not set and bare semicolons do not get any special treatment. * `dont_skip_items`: Normally, items that are completely unspecified (have no term and no properties) are skipped and not inserted into the returned list of items. (Such items cannot occur if `disallow_holes = true` is set on the term specification in the `params` structure passed to `process()` in [[Module:parameters]]. It is generally recommended to do so unless a specific meaning is associated the term value being missing.) If `dont_skip_items` is set, however, items are never skipped, and completely unspecified items will be returned along with others. (They will not have the term or any properties set, but will have the normal non-property fields set; see below.) * `stop_when`: If specified, a function to determine when to prematurely stop processing items. It is passed a single argument, an object containing the following fields: ** `term`: The raw term, prior to parsing off language prefixes and inline modifiers (since the processing of `stop_when` happens before parsing the term). ** `any_param_at_index`: True if any separate property parameters exist for this item. ** `orig_index`: Same as `orig_index` below. ** `itemno`: Same as `itemno` below. ** `stored_itemno`: The index where this item will be stored into the returned items table. This may differ from `itemno` due to skipped items (it will never be different if `dont_skip_items` is set). The function should return true to stop processing items and return the ones processed so far (not including the item currently being processed). This is used, for example, in [[Module:alternative forms]], where an unspecified item signal the end of items and the start of labels. Two values are returned, the list of items and the processed `args` structure. In each returned item, there will be one field set for each specified property (either through inline modifiers or separate parameters). In addition, the following fields may be set: * `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given. * `orig_index`: The original index into the item in the items table returned by `process()` in [[Module:parameters]]. This may differ from `itemno` if there are raw semiclons and `disallow_custom_separators` is not given. * `itemno`: The logical index of the item. The index of separate parameters corresponds to this index. This may be different from `orig_index` in the presence of raw semicolons; see above. * `separator`: The separator to display before the term. Always set unless `disallow_custom_separators` is given, in which case it is not set. * `termlang`: If there is a language prefix, the corresponding language object is stored here (only if `parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set). * `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are set, the list of corresponding language objects is stored here. * `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified; (c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value. * `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b) `sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value. ]==] function export.process_list_arguments(data) local args if not data.termarg then internal_error("Required value `data.termarg` not specified") end if not data.param_mods then internal_error("Required value `data.param_mods` not specified") end if data.raw_args then -- FIXME, remove support for `data.args` in favor of `data.processed_args` if data.processed_args or data.args then internal_error("Only one of `data.raw_args` and `data.processed_args` can be specified") end if not data.params then internal_error("When `data.raw_args` is specified, so must `data.params`, so that the raw arguments can be parsed") end local termarg_spec = data.params[data.termarg] if not termarg_spec then internal_error("There must be a spec in `data.params` corresponding to `data.termarg`") end if not termarg_spec.list then internal_error("Term spec in `data.params` must have `list` set", termarg_spec) end if not termarg_spec.allow_holes and not termarg_spec.disallow_holes then internal_error("Term spec in `data.params` must have either `allow_holes` or `disallow_holes` set", termarg_spec) end export.augment_params_with_modifiers(data.params, data.param_mods) args = require(parameters_module).process(data.raw_args, data.params) else args = data.processed_args or data.args if not args then internal_error("Either `data.raw_args` or `data.processed_args` must be specified") end if data.params then internal_error("When `data.processed_args` is specified, `data.params` should not be specified") end end if data.process_args_before_parsing then data.process_args_before_parsing(args) end -- Find the maximum index among any of the list parameters. local term_args = args[data.termarg] -- As a special case, the term args might not have a `maxindex` field because they might have -- been declared with `disallow_holes = true`, so fall back to the actual length of the list. local maxmaxindex = term_args.maxindex or #term_args for k, v in pairs(args) do if type(v) == "table" and v.maxindex and v.maxindex > maxmaxindex then maxmaxindex = v.maxindex end end local items = {} local ind = 0 local lang = fetch_argument(args, data.lang) local sc = fetch_argument(args, data.sc) local lang_cache = {} if lang then lang_cache[lang:getCode()] = lang end local use_semicolon = false local term_dest = data.term_dest or "term" local itemno = 0 for i = 1, maxmaxindex do local term = term_args[i] if data.disallow_custom_separators or not special_separators[term] then itemno = itemno + 1 -- Compute whether any of the separate indexed params exist for this index. local any_param_at_index = term ~= nil if not any_param_at_index then for k, v in pairs(args) do -- Look for named list parameters. We check: -- (1) key is a string (excludes the term param, which is a number); -- (2) value is a table, i.e. a list; -- (3) v.maxindex is set (i.e. allow_holes was used); -- (4) the value has an entry at index `itemno` (the current logical index). if type(k) == "string" and type(v) == "table" and v.maxindex and v[itemno] then any_param_at_index = true break end end end if data.stop_when and data.stop_when { term = term, any_param_at_index = any_param_at_index, orig_index = i, itemno = itemno, stored_itemno = #items + 1, } then break end -- If any of the params used for formatting this term is present, create a term and add it to the list. if not data.dont_skip_items and not any_param_at_index then track("skipped-term", data.track_module) else if not term then track("missing-term", data.track_module) end local termobj = { itemno = itemno, orig_index = i, } if not data.disallow_custom_separators then termobj.separator = i == 1 and "" or special_separators[term_args[i - 1]] end -- Parse all the term-specific parameters and store in `termobj`. for param_mod, param_mod_spec in pairs(data.param_mods) do local dest = param_mod_spec.item_dest or param_mod local arg = args[param_mod] and args[param_mod][itemno] if arg then termobj[dest] = arg end end local function generate_obj(term, parse_err) if data.parse_lang_prefix and term:find(":") then local actual_term, termlangs = require(parse_utilities_module).parse_term_with_lang { term = term, parse_err = parse_err, paramname = paramname, allow_bad = data.allow_bad_lang_prefix, allow_multiple = data.allow_multiple_lang_prefixes, lang_cache = lang_cache, } termobj[term_dest] = actual_term ~= "" and actual_term or nil if termlangs then -- If we couldn't parse a language code, don't overwrite an existing setting in `lang` -- that may have originated from a separate |langN= param. if data.allow_multiple_lang_prefixes then termobj.termlangs = termlangs termobj.lang = termlangs and termlangs[1] or nil else termobj.termlang = termlangs termobj.lang = termlangs end end else termobj[term_dest] = term ~= "" and term or nil end return termobj end -- Check for inline modifier, e.g. מרים<tr:Miryem>. But exclude top-level HTML entry with <span ...>, -- <br/> or similar in it, often caused by wrapping an argument in {{m|...}} or similar. if term and term:find("<") and not require(parse_utilities_module).term_contains_top_level_html(term) then require(parse_utilities_module).parse_inline_modifiers(term, { -- Add 1 because first term index starts at 2. paramname = data.termarg + i - 1, param_mods = data.param_mods, generate_obj = generate_obj, }) elseif term then generate_obj(term) end -- Set these after parsing inline modifiers, not in generate_obj(), otherwise we'll get an error in -- parse_inline_modifiers() if we try to use <lang:...> or <sc:...> as inline modifiers. termobj.lang = termobj.lang or lang termobj.sc = termobj.sc or sc if not data.disallow_custom_separators then -- If the displayed term (from .term/etc. or .alt) has an embedded comma, use a semicolon to join -- the terms. local term_text = termobj[term_dest] or termobj.alt if not use_semicolon and term_text then if term_text:find(",", 1, true) then use_semicolon = true end end end table.insert(items, termobj) end end end if not data.disallow_custom_separators then -- Set the default separator of all those items for which a separator wasn't explicitly given to comma -- (or semicolon if any items have embedded commas). for i, item in ipairs(items) do if not item.separator then item.separator = use_semicolon and "; " or ", " end end end return items, args end return export rg6jb31a0yz1r6lsdqyajb6epalu93q 54882 54880 2026-09-27T21:46:58Z Deadend0914 7211 54882 Scribunto text/plain local export = {} local debug_track_module = "Module:debug/track" local functions_module = "Module:fun" local parameters_module = "Module:parameters" local parse_interface_module = "Module:parse interface" local parse_utilities_module = "Module:parse utilities" local table_module = "Module:table" local dump = mw.dumpObject local error = error local insert = table.insert local ipairs = ipairs local next = next local pairs = pairs local require = require local tonumber = tonumber local type = type --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls. ]==] local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function is_callable(...) is_callable = require(functions_module).is_callable return is_callable(...) end local function list_to_set(...) list_to_set = require(table_module).listToSet return list_to_set(...) end local function parse_inline_modifiers(...) parse_inline_modifiers = require(parse_interface_module).parse_inline_modifiers return parse_inline_modifiers(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function table_len(...) table_len = require(table_module).length return table_len(...) end ----------------- end loaders ---------------- local function track(page, track_module) return debug_track((track_module or "parameter utilities") .. "/" .. page) end -- Throw an error prefixed with the words "Internal error" (and suffixed with a dumped version of `spec`, if provided). -- This is for logic errors in the code itself rather than template user errors. local function internal_error(msg, spec) if spec then msg = ("%s: %s"):format(msg, dump(spec)) end error(("Internal error: %s"):format(msg)) end -- Table listing the default recognized special separator arguments and how they display. export.default_special_separators = { [";"] = "; ", ["_"] = " ", ["~"] = " ~ ", ["→"] = " → ", } -- Table listing how subitem delimiters display. Unlike for `default_special_separators`, the presence of an item in -- this table does not mean that the delimiter is recognized; only those specified by `data.splitchar` are recognized. export.default_subitem_separator_map = { [";"] = "; ", [","] = ", ", ["/"] = "/", ["_"] = " ", ["~"] = " ~ ", ["→"] = " → ", } --[==[ intro: The purpose of this module is to facilitate implementation of templates that can have arguments specified either through inline modifiers or separate parameters. There are two types of templates supported: those that take a list of items with associated properties, which can be specified either through indexed separate parameters (e.g. {{para|t2}}, {{para|pos3}}) or inline modifiers (`<t:...>`, `<pos:...>`, etc.); and those that take a single term, whose properties can be specified through non-indexed separate parameters (e.g. {{para|t}} or {{para|pos}}) or inline modifiers. Both types of templates can optionally have subitems in the term parameter(s), where the subitems are typically (but not necessarily) separated with commas and each subitem can have its own inline modifiers. Some examples of templates that take a list of items are {{tl|alter}}/{{tl|alt}}; {{tl|synonyms}}/{{tl|syn}}, {{tl|antonyms}}/{{tl|ant}}, and other "nyms" templates; {{tl|col}}, {{tl|col2}}, {{tl|col3}}, {{tl|col4}} and other column templates; {{tl|descendant}}/{{tl|desc}}; {{tl|affix}}/{{tl|af}}, {{tl|prefix}}/{{tl|pre}} and related *fix templates; {{tl|affixusex}}/{{tl|afex}} and related templates; {{tl|IPA}}; {{tl|homophones}}; {{tl|rhymes}}; and several others. Examples of templates that take a single item are form-of templates ({{tl|inflection of}}/{{tl|infl of}}, {{tl|form of}}, and specific templates such as {{tl|alt form}}/{{tl|alternative form of}}, {{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others); for etymology templates ({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates like {{tl|ellipsis}}, {{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and other templates that take an argument structure similar to {{tl|l}} or {{tl|m}}. This module can be thought of as a combination of [[Module:parameters]] (which parses template parameters, and in particular handles the separate parameter versions of the properties) and `parse_inline_modifiers()` in [[Module:parse utilities]] (which parses inline modifiers). The two main entry points are `parse_list_with_inline_modifiers_and_separate_params()` (for templates that take a list of items) and `parse_term_with_inline_modifiers_and_separate_params()` (for templates that take a single item). However, there are other functions provided, e.g. to initialize the `param_mods` structure that is passed to the two entry points. The typical workflow for using `parse_list_with_inline_modifiers_and_separate_params()` looks as follows (a slightly simplified version of the code in [[Module:nyms]]): { local export = {} local parameter_utilities_module = "Module:parameter utilities" ... -- Entry point to be invoked from a template. function export.show(frame) local parent_args = frame:getParent().args -- Parameters that don't have corresponding inline modifiers. Note in particular that the parameter corresponding to -- the items themselves must be specified this way, and must specify either `allow_holes = true` (if the user can -- omit terms, typically by specifying the term using |altN= or <alt:...> so that they remain unlinked) or -- `disallow_holes = true` (if omitting terms is not allowed). (If neither `allow_holes` nor `disallow_holes` is -- specified, an error is thrown in parse_list_with_inline_modifiers_and_separate_params().) local params = { [1] = {required = true, type = "language", default = "und"}, [2] = {list = true, allow_holes = true, required = true, default = "term"}, } local m_param_utils = require(parameter_utilities_module) -- This constructs the `param_mods` structure by adding well-known groups of parameters (such as all the parameters -- associated with based on full_link() in [[Module:links]], with default properties that can be overridden. This is -- easier and less error-prone than manually specifying the `param_mods` structure (see below for how this would -- look). Here, we specify the group "link" (consisting of all the link parameters for use with full_link()), group -- "ref" (which adds the "ref" parameter for specifying references), group "l" (which adds the "l" and "ll" -- parameters for specifying labels) and group "q" (which adds the "q" and "qq" parameters for specifying regular -- qualifiers). By default, labels and qualifiers have `separate_no_index` set so that e.g. |q1= is distinct from -- |q=, the former specifying the left qualifier for the first item and the latter specifying the overall left -- qualifier. For compatibility, we override the `separate_no_index` setting for the group "q", which causes |q= and -- |q1= to be the same, and likewise for |qq= and |qq1=. Finally, also for compatibility, we add an "lb" parameter -- that is an alias of "ll" (in all respects; |lb= is the same as |ll=, |lb1= is the same as |ll1=, <lb:...> is the -- same as <ll:...>, etc.). local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l"}}, {group = "q", separate_no_index = false}, {param = "lb", alias_of = "ll"}, } -- This processes the raw arguments in `parent_args`, parses inline modifiers and creates corresponding objects -- containing the property values specified either through inline modifiers or separate parameters. local items, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, parse_lang_prefix = true, track_module = "nyms", lang = 1, sc = "sc.default", } local lang = args[1] -- Now do the actual implementation of the template. Generally this should be split into a separate function, often -- in a separate module (if the implementation goes in [[Module:foo]], the template interface code goes in -- [[Module:foo/templates]]). ... } The `param_mods` structure controls the properties that can be specified by the user for a given item, and is conceptually very similar to the `param_mods` structure used by `parse_inline_modifiers()`. The key is the name of the parameter (e.g. {"t"}, {"pos"}) and the value is a table with optional elements as follows: * `item_dest`, `store`: Same as the corresponding fields in the `param_mods` structure passed to `parse_inline_modifiers()`. * `type`, `set`, `sublist`, `convert` and associated fields such as `family` and `method`: These control parsing and conversion of the raw values specified by the user and have the same meaning as in [[Module:parameters]] and also in `parse_inline_modifiers()` (which delegates the actual conversion to [[Module:parameters]]). These fields — and for that matter, all fields other than `item_dest`, `store` and `overall` — are forwarded to the `process()` function in [[Module:parameters]]. * `alias_of`: This parameter is an alias of some other parameter. This spec is recognized only by `process()` in [[Module:parameters]], and not by `parse_inline_modifiers()`; to set up an alias in `parse_inline_modifiers()`, you need to make sure (using `item_dest`) that both the alias and aliasee modifiers store their values in the same location, and you need to copy the remaining properties from the aliasee's spec to the aliasing modifier's spec. All of this happens automatically if you generate the `param_mods` structure using `construct_param_mods()`. * `require_index`: This means that the non-indexed parameter version of the property is not recognized. E.g. in the case of the {"sc"} property, use of the {{para|sc}} parameter would result in an error, while {{para|sc1}} is recognized and specifies the {"sc"} property for the first item. The default, if neither `require_index` nor `separate_no_index` is given, is for {{para|sc}} and {{para|sc1}} to mean the same thing (both would specify the {"sc"} property of the first item). Note that `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `separate_no_index`: This means that e.g. the {{para|sc}} parameter is distinct from the {{para|sc1}} parameter (and thus from the `<sc:...>` inline modifier on the first item). This is typically used to distinguish an overall version of a property from the corresponding item-specific property on the first item. (In this case, for example, {{para|sc}} overrides the script code for all items, while {{para|sc1}} overrides the script code only for the first item.) If not given, and if `require_index` is not given, {{para|sc}} and {{para|sc1}} would have the same meaning and refer to the item-specific property on the first item. When this is given, the overall value can be accessed using the `.default` field of the property value in `args`, e.g. in this case `args.sc.default`. Note that (as mentioned above) `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `list`, `allow_holes`, `disallow_holes`: These should '''not''' be given. `list` and `allow_holes` are automatically set for all parameter specs added to the `params` structure used by `process()` in [[Module:parameters]], and `disallow_holes` clashes with `allow_holes`. For the above workflow example, the call to `construct_param_mods()` generates the following `param_mods` structure: { local param_mods = { -- the parameters generated by group "link" alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", type = "genders", }, pos = {}, ng = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, -- the parameters generated by group "ref" ref = { item_dest = "refs", type = "references", }, -- the parameters generated by group "l" l = { type = "labels", separate_no_index = true, }, ll = { type = "labels", separate_no_index = true, }, -- the parameters generated by group "q"; note that `separate_no_index = true` would be set, but is overridden -- (specifying `separate_no_index = false` in the `param_mods` structure is equivalent to not specifying it at all) q = { type = "qualifier", separate_no_index = false, }, qq = { type = "qualifier", separate_no_index = false, }, infl = { type = "form of tags", separate_no_index = true, }, -- the parameter generated by the individual "lb" parameter spec; note that only `alias_of` was explicitly given, -- while `item_dest` is automatically set so that inline modifier <lb:...> stores into the same place as <ll:...>, -- and the other specs are copied from the `ll` spec so `lb` works like `ll` in all regards lb = { alias_of = "ll", item_dest = "ll", type = "labels", separate_no_index = true, }, } } ]==] local qualifier_spec = { type = "qualifier", separate_no_index = true, } local label_spec = { type = "labels", separate_no_index = true, } local form_of_spec = { type = "form of tags", separate_no_index = true, } local recognized_param_mod_groups = { link = { alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", type = "genders", }, pos = {}, ng = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, }, lang = { lang = { require_index = true, type = "language", }, }, q = { q = qualifier_spec, qq = qualifier_spec, }, a = { a = label_spec, aa = label_spec, }, l = { l = label_spec, ll = label_spec, }, infl = { infl = form_of_spec, }, ref = { ref = { item_dest = "refs", type = "references", }, }, } local function merge_param_mod_settings(orig, additions) local merged = shallow_copy(orig) for k, v in pairs(additions) do merged[k] = v if k == "require_index" then merged.separate_no_index = nil elseif k == "separate_no_index" then merged.require_index = nil end end merged.default = nil merged.group = nil merged.param = nil merged.exclude = nil merged.include = nil return merged end local function verify_type(spec, param, typ1, typ2) if not spec[param] then return end local val = spec[param] if type(val) ~= typ1 and (not typ2 or type(val) ~= typ2) then internal_error(("Parameter `%s` must be a %s%s but saw a %s"):format(param, typ1, typ2 and " or " .. typ2 or "", type(val)), spec) end end local function verify_well_constructed_spec(spec) local num_control = (spec.default and 1 or 0) + (spec.group and 1 or 0) + (spec.param and 1 or 0) if num_control == 0 then internal_error( "Spec passed to construct_param_mods() must have either the `default`, `group` or `param` keys set", spec) end if num_control > 1 then internal_error( "Exactly one of `default`, `group` or `param` must be set in construct_param_mods() spec", spec) end if spec.list or spec.allow_holes then -- FIXME: We need to support list = "foo" for list parameters that are stored in e.g. 2=, foo2=, foo3=, etc. internal_error("`list` and `allow_holes` may not be set; they are automatically set when constructing the " .. "corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.disallow_holes then internal_error("`disallow_holes` may not be set; it conflicts with `allow_holes`, which is automatically " .. "set when constructing the corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.include and spec.exclude then internal_error("Saw both `include` and `exclude` in the same spec", spec) end if (spec.include or spec.exclude) and not spec.group then internal_error( "`include` and `exclude` can only be specified along with `group`, not with `default` or `param`", spec) end verify_type(spec, "group", "string", "table") verify_type(spec, "param", "string", "table") verify_type(spec, "include", "table") verify_type(spec, "exclude", "table") end --[==[ Construct the `param_mods` structure used in parsing arguments and inline modifiers from a list of specifications. A sample invocation (a slightly simplified version of the actual invocation associated with {{tl|affix}} and related templates) looks like this: { local param_mods = require("Module:parameter utilities").construct_param_mods { -- We want to require an index for all params (or use separate_no_index, which also requires an index for the -- param corresponding to the first item). {default = true, require_index = true}, {group = {"link", "ref", "lang", "q", "l"}}, -- Override these two to have separate_no_index. {param = {"lit", "pos"}, separate_no_index = true}, } } Each specification either sets the default value for further parameter specs or adds one or more parameters. Parameters can be added directly using `param`, or groups of predefined parameters can be added using `group`. Specifications are one of three types: # Those that set the default properties for future-added parameters. These contain {default = true} as one of the properties of the spec. Specs are processed in order and you can change the defaults mid-way through. # Those that add the parameters associated with one or more pre-defined groups. These contain {group = "group"} or {group = {"group1", "group2", ...}}. The pre-defined parameter groups and their associated properties are listed below. The pre-defined properties of parameters in a group override properties associated with a {default = true} spec, and are in turn overridden by any properties given directly in the spec itself. Note as well that setting the `separate_no_index` property will automatically cause the `require_index` property to be unset and vice-versa, as the two are mutually exclusive. (This happens in the example above, where the {separate_no_index = true} setting associated with the params {"lit"} and {"pos"} cancels out the {require_index = true} default setting, as well as less obviously with the pre-defined {"sc"} property of the {"link"} group, the {"q"} and {"qq"} properties of the {"q"} group, and the {"l"} and {"ll"} properties of the {"l"} group, all of which have an associated pre-defined property {separate_no_index = true}, which overrides and cancels out the {require_index = true} default setting. Finally, when adding the parameters of a group, you can request the only a subset of the parameters be added using either the `include` or `exclude` properties, each of whose values is a list of parameters that specify (respectively) the parameters to include (all other parameters of the group are excluded) or to exclude (all other parameters of the group are included). This is used, for example, in [[Module:romance etymology]] and [[Module:it-etymology]], which specify {group = "link", exclude = {"tr", "ts", "sc"}} to exclude link parameters that aren't relevant to Latin-script languages such as the Romance languages, and conversely in [[Module:IPA/templates]], which specifies {group = "link", include = {"t", "gloss", "pos"}} to include only the specified parameters for use with {{tl|IPA}}. # Those that add individual parameters. These contain {param = "param"} or {param = {"param1", "param2", ...}}, the latter syntax used to control a set of parameters together. The resulting spec is formed by initializing the parameter's settings with any previously-specified default properties (using a spec containing {default = true}) if the parameter hasn't already been initialized, and then overriding the resulting settings with any settings given directly in the specification. In the above example, the {"lit"} and {"pos"} parameters were previously initialized through the {"link"} group (specified in the second of the three specifications) but ended up with {require_index = true} due to the {default = true} spec (the first of the three specifications). We override these two parameters to have {separate_no_index = true} (which, as mentioned above, cancels out {require_index = true}). This is done so that {{tl|affix}} and related templates have {{para|pos}} and {{para|lit}} parameters distinct from {{para|pos1}} and {{para|lit1}}, which are used to specify an overall part of speech (which applies to all parts of the affix, as opposed to applying to just one element of the expression) or a literal definition for the entire expression (instead of just for one element of the expression). The built-in parameter groups are as follows: {|class="wikitable" ! Group !! Group meaning !! Parameter !! Parameter meaning !! Default properties |- | rowspan=11| `link` | rowspan=11| link parameters; same as those available on {{tl|l}}, {{tl|m}} and other linking templates | `alt` || display text, overriding the term's display form || — |- | `t` || gloss (translation) of a non-English term || {item_dest = "gloss"} |- | `gloss` || gloss (translation); same as `t` || {alias_of = "t"} |- | `tr` || transliteration of a non-Latin-script term; only needed if the automatic transliteration is incorrect or unavailable (e.g. in Hebrew, which doesn't have automatic transliteration) || — |- | `ts` || transcription of a non-Latin-script term, if the transliteration is markedly different from the actual pronunciation; should not be used for IPA pronunciations || — |- | `g` || comma-separated list of genders; whitespace may surround the comma and will be ignored || {item_dest = "genders", type = "genders"} |- | `pos` || part of speech for the term || — |- | `ng` || arbitrary non-gloss descriptive text for the term || — |- | `lit` || literal meaning (translation) of the term || — |- | `id` || a sense ID for the term, which links to anchors on the page set by the {{tl|senseid}} template || — |- | `sc` || the script code (see [[Wiktionary:Scripts]]) for the script that the term is written in; rarely necessary, as the script is autodetected (in most cases, correctly) || {separate_no_index = true, type = "script"} |- | rowspan=2| `q` | rowspan=2| left and right normal qualifiers (as displayed using {{tl|q}}) | `q` || left normal qualifier || {separate_no_index = true, type = "qualifier"} |- | `qq` || right normal qualifier || {separate_no_index = true, type = "qualifier"} |- | rowspan=2| `a` | rowspan=2| left and right accent qualifiers (as displayed using {{tl|a}}) | `a` || comma-separated list of left accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `aa` || comma-separated list of right accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | rowspan=2| `l` | rowspan=2| left and right labels (as displayed using {{tl|lb}}, but without categorizing) | `l` || comma-separated list of left labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ll` || comma-separated list of right labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ref` | reference(s) (in the format accepted by [[Module:references]]; see also the documentation for the {{para|ref}} parameter to {{tl|IPA}}) | `ref` || one or more references, in the format accepted by [[Module:references]] || {item_dest = "refs", type = "references"} |- | `lang` | language for an individual term (provided for compatibility; it is preferred to specify languages for individual terms using language prefixes instead) | `lang` || language code (see [[Wiktionary:Languages]]) for the term || {require_index = true, type = "language"} |} ]==] function export.construct_param_mods(specs) local param_mods = {} local default_specs = {} for _, spec in ipairs(specs) do verify_well_constructed_spec(spec) if spec.default then -- This will have an extra `default` field in it, but it will be erased by merge_param_mod_settings() default_specs = spec else if spec.group then local groups = spec.group if type(groups) ~= "table" then groups = {groups} end local include_set if spec.include then include_set = list_to_set(spec.include) end local exclude_set if spec.exclude then exclude_set = list_to_set(spec.exclude) end for _, group in ipairs(groups) do local group_specs = recognized_param_mod_groups[group] if not group_specs then internal_error(("Unrecognized built-in param mod group '%s'"):format(group), spec) end for group_param, group_param_settings in pairs(group_specs) do local include_param if include_set then include_param = include_set[group_param] elseif exclude_set then include_param = not exclude_set[group_param] else include_param = true end if include_param then local merged_settings = merge_param_mod_settings(merge_param_mod_settings( param_mods[group_param] or default_specs, group_param_settings), spec) param_mods[group_param] = merged_settings end end end end if spec.param then local params = spec.param if type(params) ~= "table" then params = {params} end for _, param in ipairs(params) do local settings = merge_param_mod_settings(param_mods[param] or default_specs, spec) -- If this parameter is an alias of another parameter, we need to copy the specs from the other -- parameter, since parse_inline_modifiers() doesn't know about `alias_of` and having the specs -- duplicated won't cause problems for [[Module:parameters]]. We also need to set `item_dest` to -- point to the `item_dest` of the aliasee (defaulting to the aliasee's value itself), so that -- both modifiers write to the same location. Note that this works correctly in the common case of -- <t:...> with `item_dest = "gloss"` and <gloss:...> with `alias_of = "t"`, because both will end -- up with `item_dest = "gloss"`. local aliasee = settings.alias_of if aliasee then local aliasee_settings = param_mods[aliasee] if not aliasee_settings then internal_error(("Undefined aliasee '%s'"):format(aliasee), spec) end for k, v in pairs(aliasee_settings) do if settings[k] == nil then settings[k] = v end end if settings.item_dest == nil then settings.item_dest = aliasee end end param_mods[param] = settings end end end end return param_mods end -- Return true if `k` is a "built-in" (specially recognized) key in a `param_mod` specification. All other keys -- are forwarded to the structure passed to [[Module:parameters]]. local function param_mod_spec_key_is_builtin(k) return k == "item_dest" or k == "overall" or k == "store" end --[==[ Convert the properties in `param_mods` into the appropriate structures for use by `process()` in [[Module:parameters]] and store them in `params`. If `overall_only` is given, only store the properties in `param_mods` that correspond to overall (non-item-specific) parameters. Currently this only happens when `separate_no_index` is specified. ]==] function export.augment_params_with_modifiers(params, param_mods, overall_only) if overall_only then for param_mod, param_mod_spec in pairs(param_mods) do if overall_only == "always" or param_mod_spec.separate_no_index then local param_spec = {} for k, v in pairs(param_mod_spec) do if k ~= "separate_no_index" and k ~= "require_index" and not param_mod_spec_key_is_builtin(k) then param_spec[k] = v end end params[param_mod] = param_spec end end else local list_with_holes -- Add parameters for each term modifier. for param_mod, param_mod_spec in pairs(param_mods) do local param_spec for k, v in pairs(param_mod_spec) do if not param_mod_spec_key_is_builtin(k) then if param_spec == nil then param_spec = {list = true} end param_spec[k] = v end end if param_spec == nil then if list_with_holes == nil then list_with_holes = {list = true, allow_holes = true} end param_spec = list_with_holes elseif param_spec.alias_of == nil then param_spec.allow_holes = true end params[param_mod] = param_spec end end end --[==[ Return true if `k`, a key in an item, refers to a property of the item (is not one of the specially stored values). Note that `lang` and `sc` are considered properties of the item, although `lang` is set when there's a language prefix and both `lang` and `sc` may be set from default values specified in the `data` structure passed into `parse_list_with_inline_modifiers_and_separate_params()` and `parse_term_with_inline_modifiers_and_separate_params()`. If you don't want these treated as property keys, you need to check for them yourself. ]==] function export.item_key_is_property(k) return k ~= "term" and k ~= "termlang" and k ~= "termlangs" and k ~= "itemno" and k ~= "orig_index" and k ~= "separator" end -- Fetch the argument in `args` corresponding to `index_or_value`, which may be a string of the form "foo.default" -- (requesting the value of `args["foo"].default`); a string or number (requesting the value at that key); a function of -- one argument (`args`), which returns the argument value; or the value itself. Return the resulting value and the -- parameter in `args` that the value came from, or nil if unknown (i.e. a function or direct value was specified). local function fetch_argument(args, index_or_value) if not index_or_value then return index_or_value, nil end local index_or_value_type = type(index_or_value) if index_or_value_type == "string" then if index_or_value:sub(-8) == ".default" then local index_without_default = index_or_value:sub(1, -9) local arg_obj = fetch_argument(args, index_without_default) if type(arg_obj) ~= "table" then internal_error(("Requested that the '.default' key of argument `%s` be fetched, but argument value is undefined or not a table"): format(index_without_default), arg_obj) end return arg_obj.default, index_without_default end if index_or_value:match("^%d+$") then index_or_value = tonumber(index_or_value) end return args[index_or_value], index_or_value elseif index_or_value_type == "number" then return args[index_or_value], index_or_value elseif is_callable(index_or_value) then return index_or_value(args), nil end return index_or_value, nil end function export.generate_obj_maybe_parsing_lang_prefix(data) return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix(data) end -- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and -- parse_term_with_inline_modifiers_and_separate_params(), validating certain argument-related fields that are shared -- among the two functions. local function validate_argument_related_fields(data) if not data.termarg then internal_error("`data.termarg` must be given, indicating which argument contains the terms to be parsed", data) end if not data.param_mods then internal_error("`data.param_mods` must be given, indicating the allowed inline modifiers and separate " .. "parameters to copy", data) end local subitem_param_handling = data.subitem_param_handling or "only" if subitem_param_handling ~= "only" and subitem_param_handling ~= "first" and subitem_param_handling ~= "last" then internal_error("Unrecognized value for `data.subitem_param_handling`, should be 'first', 'last' or 'only'", subitem_param_handling) end if data.raw_args then if data.processed_args then internal_error("Only one of `data.raw_args` and `data.processed_args` can be specified", data) end if not data.params then internal_error("When `data.raw_args` is specified, so must `data.params`, so that the raw arguments " .. "can be parsed", data) end if data.params[data.termarg] == nil then internal_error("There must be a spec in `data.params` corresponding to `data.termarg`", data) end else if not data.processed_args then internal_error("Either `data.raw_args` or `data.processed_args` must be specified", data) end if data.params then internal_error("When `data.processed_args` is specified, `data.params` should not be specified", data) end end end local function argval_missing(val) return val == nil or type(val) == "table" and next(val) == nil end -- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and -- parse_term_with_inline_modifiers_and_separate_params(). After parsing inline modifiers, copy the separate parameters -- to the generated object (or to the appropriate subobject if there are multiple). `data` contains the following -- fields: -- -- `args`: The separate-parameter argument structure. -- `param_mods`: The structure describing the inline modifiers. -- `itemno`: The logical item number of the term being processed, or nil if there's only a single term. -- `termobj`: The object to store the inline modifiers into. If there are subitems, they are in the `terms` field; -- otherwise the properties are stored directly into `termobj`. -- `has_subitems`: True if there are subitems. -- `subitem_separator_map`: If `has_subitems` and this is specified, controls the assignment of the `separator` field -- in subitems. If not specified or a delimiter is not in the map, it is copied unchanged. -- `lang`: Language object to store into all items. -- `sc`: Script object to store into all items, or nil. -- `subitem_param_handling`: "only", "first" or "last", indicating what to do if there are multiple subitems. -- `allow_conflicting_inline_mods_and_separate_params`: If true, specifying a value for both an inline modifier and -- corresponding separate parameter is allowed, and the inline modifier takes precedence. Otherwise, an error -- occurs. -- `postprocess_termobj`: Optional function called on all items at the end, to do any postprocessing. Called with one -- argument, the object to postprocess. -- `no_show_decorations`: If true, don't automatically set {show_decorations = true} on the object or subobject if there -- are decorations (i.e. qualifiers, labels or references) specified for the object. local function copy_separate_params_to_termobj_and_postprocess(data) local args, param_mods, itemno, termobj = data.args, data.param_mods, data.itemno, data.termobj local function set_lang_and_sc(termobj) -- Set these after parsing inline modifiers, not in generate_obj(), otherwise we'll get an error in -- parse_inline_modifiers() if we try to use <lang:...> or <sc:...> as inline modifiers. termobj.lang = termobj.lang or data.lang termobj.sc = termobj.sc or data.sc end local function set_show_decorations(termobj) -- Need to set after parsing inline modifiers. if not data.no_show_decorations and (termobj.q or termobj.qq or termobj.a or termobj.aa or termobj.l or termobj.ll or termobj.refs) then termobj.show_decorations = true end end local function fetch_separate_param(args, paramkey, itemno) local argval = args[paramkey] -- Careful with argument values that may be `false`. if argval and itemno then argval = argval[itemno] end return argval end -- Copy separate parameters to a given object. local function copy_separate_params_to_termobj(fetch_destobj) for param_mod, param_mod_spec in pairs(param_mods) do local dest = param_mod_spec.item_dest or param_mod -- Don't do anything with the `sc` param, which will get overwritten below; we don't -- want it to cause an error if there are multiple subitems. if dest ~= "sc" then local argval = fetch_separate_param(args, param_mod, itemno) if not argval_missing(argval) then local destobj = fetch_destobj(param_mod, param_mod_spec, dest) -- Don't overwrite a value already set by an inline modifier. if argval_missing(destobj[dest]) then destobj[dest] = argval elseif not data.allow_conflicting_inline_mods_and_separate_params then error(("Can't specify a value for separate parameter %s%s= because there is " .. "already an inline modifier <%s:...> specifying a value for the term"):format( param_mod, itemno or "", param_mod)) end end end end end if data.has_subitems then -- If there are any separate indexed parameters, we need to copy them to the first, last or only -- subitem, depending on the value of `data.subitem_param_handling` (which defaults to 'only', -- meaning it's an error if there are multiple subitems). Do this before calling -- postprocess_termobj() because the latter sets .lang and .sc and we want the user to be able to -- set separate langN= and scN= parameters. -- If there was no term, `termobj.terms` will not exist; make it exist to make the callers' lives easier. if not termobj.terms then termobj.terms = {} end -- Compute whether any of the separate indexed params exist for this index. local any_param_at_index for param_mod in pairs(param_mods) do local argval = fetch_separate_param(args, param_mod, itemno) if not argval_missing(argval) then any_param_at_index = true break end end -- If there was no term, but there's a separate parameter, we need to create an empty subitem. if any_param_at_index and not termobj.terms[1] then termobj.terms[1] = {} end local function fetch_destobj(param_mod, param_mod_spec, dest) if param_mod_spec.overall then return termobj end if data.subitem_param_handling == "only" and termobj.terms[2] then error(("Can't specify a value for separate parameter %s%s= because there are " .. "multiple subitems (%s) in the term; use an inline modifier"):format( param_mod, itemno or "", #termobj.terms)) end local termind -- q/a/l need to go at the beginning and qq/aa/ll/refs at the end, regardless; otherwise, respect -- `data.subitem_param_handling`. if dest == "q" or dest == "a" or dest == "l" then termind = 1 elseif dest == "qq" or dest == "aa" or dest == "ll" or dest == "refs" then termind = #termobj.terms elseif data.subitem_param_handling == "only" or data.subitem_param_handling == "first" then termind = 1 else termind = #termobj.terms end return termobj.terms[termind] end copy_separate_params_to_termobj(fetch_destobj) for i, subitem in ipairs(termobj.terms) do set_lang_and_sc(subitem) set_show_decorations(subitem) if subitem.delimiter then subitem.separator = i == 1 and "" or data.subitem_separator_map and data.subitem_separator_map[subitem.delimiter] or subitem.delimiter end if data.postprocess_termobj then data.postprocess_termobj(subitem, data) end end else -- Copy all the parsed term-specific parameters into `termobj`. copy_separate_params_to_termobj(function(param_mod, dest) return termobj end) set_lang_and_sc(termobj) set_show_decorations(termobj) if data.postprocess_termobj then data.postprocess_termobj(termobj, data) end end end local function postprocess_termobj(item, data) if not (data.disallow_custom_separators or data.use_semicolon) then if data.has_subitems and item.separator and item.separator:find(",", nil, true) then data.use_semicolon = true else -- If the displayed term (from .term/etc. or .alt) has an embedded comma, use a semicolon to -- join the terms. local term_text = item[data.term_dest] or item.alt if term_text and term_text:find(",", nil, true) then data.use_semicolon = true end end end end --[==[ Parse a list of terms, each of which may have properties specified using inline modifiers or separate parameters. This function is intended for parsing the arguments of templates like {{tl|syn}}, {{tl|ant}} and related ''*nym'' templates; alternative-form templates {{tl|alt}}/{{tl|alter}}; affix templates like {{tl|af}}/{{tl|affix}}, {{tl|com}}/{{tl|compound}}, etc.; affix usex templates like {{tl|afex}}/{{tl|affixusex}}; name templates like {{tl|name translit}}; column templates like {{tl|col}}; pronunciation templates like {{tl|rhyme}}/{{tl|rhymes}} and {{tl|hmp}}/{{tl|homophones}}; etc. In these templates there are one or more terms specified using numeric parameters, and associated separate parameters specifying per-term properties such as {{para|t1}}, {{para|t2}}, {{para|t3}}, ... for the gloss of the first, second, third, ... term respectively. All such properties can also be specified through inline modifiers attached directly to each term (`<t:...>`, `<pos:...>`, etc.). Normally it is an error if both an inline modifier and separate parameter for the same value are given, but this can be overridden (in which case inline modifiers take precedence over separate parameters when both occur). For an example of a typical workflow involving this function, see the comment at the top of this file. Some notable properties of this function: # Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]], containing only the term list itself along with any other parameters that are '''not''' term properties (such as a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is ''augmented'' with list parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the latter case where raw argument processing is done by the caller, they must build the partial `params` structure; augment it themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the processed arguments. In both cases, the return value of this function contains three values: a list of objects, one per term, specifying the term and all properties; the processed arguments structure, so that the non-term-property arguments can be processed as appropriate; and an object containing miscellaneous global computed properties (currently only `use_semicolon`; see below). # Optionally, each term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality is used, for example, by {{tl|col}} and variants, which allow each row to have comma-separated or tilde-separated subitems. When this feature is invoked, the format of the per-term object changes; instead of directly being an object describing the term and its properties, it is an object with a `terms` field containing a list of per-subitem objects along with other top-level fields describing per-term properties. By default, if there are separate parameters specified along with multiple subitems, an error occurs, but this is controllable; currently, you can request that the parameters be assigned to the first or last subitem. # By default, special ''separator'' arguments may be present, mixed in among regular term arguments. Examples of such separator arguments are (by default; this can be overridden) a bare semicolon, specifying that the terms on either side should be separated by a semicolon instead of a comma (indicating a higher-level grouping); a bare tilde, replacing the comma separator with a tilde (indicating that the terms on either side are alternants); and a bare underscore, replacing the comma separator with a space. Separator arguments are ignored when numbering the separate parameters. You disable the separator argument handling entirely if it doesn't make sense to have this (e.g. in {{tl|af}}/{{tl|affix}}, where the separator is always a {{cd|+}} sign). `data` is an object containing several possible fields. 1. Fields that are required or recommended (usually related to argument processing): * `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from {frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw arguments. * `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both) of `raw_args` and `processed_args` must be set. * `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified manually. * `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters, '''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically "augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the "augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by [[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be correctly handled.) * `termarg` ('''required'''): The argument containing the first item with attached inline modifiers to be parsed. Usually a numeric value such as {1} or {2}. * `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used internally to track pages containing template invocations with certain properties. Example properties tracked are missing items with corresponding properties as well as missing items without corresponding properties (which are skipped entirely). To find out the exact properties tracked and the name of the tracking pages, read the code. * `lang` ('''recommended'''): The language object for the language of the items, or the name of the argument to fetch the object from. It is not strictly necessary to specify this, as this function only initializes items based on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used for certain purposes: *# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language prefix). *# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same language code may appear in several items, such as with {{tl|col}} and related templates). The value of `lang` can be any of the following: * If a string of the form "foo.default", it is assumed to be requesting the value of `args["foo"].default`. * Otherwise, if a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching (this also happens with a spec like "2.default"). * Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed to the function as its only argument. * Otherwise, it is used directly. * `sc` ('''recommended'''): The script object for the items, or the name of the argument to fetch the object from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not strictly necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned items if not otherwise set (e.g. by the {{para|sc<var>N</var>}} parameter or `<sc:...>` inline modifier). The most common value is {"sc.default"}. 2. Other argument-related fields: * `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed arguments. It should make modifications in-place. * `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language prefixes are stripped off. Defaults to {"term"}. * `pre_normalize_modifiers`: As in `parse_inline_modifiers()`. * `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence. 3. Fields related to language prefixes: * `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in [[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the `termlangs` field, and the first specified object is stored in the `lang` field). * `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above. * `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an error and will tend to remain unfixed. 4. Fields related to custom/special separators: * `disallow_custom_separators`: If specified, disallow specifying custom separators (semicolon, underscore, tilde; see the internal `default_special_separators` table, or the `special_separators` field) as an item value to override the default separator. By default, the previous separator of each item is considered to be an empty string (for the first item) and otherwise the value of the field `default_separator` (normally a comma + space), unless either the preceding item is one of the values listed in `special_separators`, such as a bare semicolon (which causes the following item's previous separator to be a semicolon + space) or an item has an embedded comma in it (which causes ''all'' items other than the first to have their previous separator be a semicolon + space). The previous separator of each item is set on the item's `separator` property. Bare semicolons and other separator arguments do not count when indexing items using separate parameters. For example, the following is correct: ** {{tl|template|lang|item 1|q1=qualifier 1|;|item 2|q2=qualifier 2}} If `disallow_custom_separators` is specified, however, the `separator` property is not set and separator arguments are not recognized. * `default_separator`: Override the default separator (normally {", "}). * `special_separators`: Table giving the special/custom separators that can be given, and how they should display. If not specified, the default in `default_special_separators` is used. This is a table mapping separator values (such as {"~"}) to the corresponding display string (such as {" ~ "}). 5. Fields related to multiple subitems in a given term: * `splitchar`: A Lua pattern. If specified, each user-specified argument can consist of multiple delimiter-separated subitems, each of which may be followed by inline modifiers. In this case, each element in the returned list of items is no longer an object describing an item, but instead an object with a `terms` field, whose value is a list describing the subitems (whose format is the same as the normal format of an item in the top-level list when `splitchar` is not specified). Each subitem object will have a `delimiter` field holding the actual delimiter occurring before the subitem, which is useful in the case where `splitchar` matches multiple possible characters. In this case, it is possible to specify that a given modifier can only occur after the last subitem and effectively modifies the whole collection of subitems by setting {overall = true} on the modifier. In this case, the modifier's value will be stored in the top-level object (the object with the `terms` field specifying the subitems). Note that splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). In addition, the algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers or inside of square brackets, which do not trigger splitting (whether or not contained within protected sequences). * `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in [[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified (see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that comma+whitespace sequences won't be split. * `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems. The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored into all subitems. * `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator` field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar` are recognized. Delimiters not in this map display as-is. 6. Other fields: * `dont_skip_items`: Normally, items that are completely unspecified (have no term and no properties) are skipped and not inserted into the returned list of items. (Such items cannot occur if {disallow_holes = true} is set on the term specification in the `params` structure passed to `process()` in [[Module:parameters]]. It is generally recommended to do so unless a specific meaning is associated the term value being missing.) If `dont_skip_items` is set, however, items are never skipped, and completely unspecified items will be returned along with others. (They will not have the term or any properties set, but will have the normal non-property fields set; see below.) * `stop_when`: If specified, a function to determine when to prematurely stop processing items. It is passed a single argument, an object containing the following fields: ** `term`: The raw term, prior to parsing off language prefixes and inline modifiers (since the processing of `stop_when` happens before parsing the term). ** `any_param_at_index`: True if any separate property parameters exist for this item. ** `orig_index`: Same as `orig_index` below. ** `itemno`: Same as `itemno` below. ** `stored_itemno`: The index where this item will be stored into the returned items table. This may differ from `itemno` due to skipped items (it will never be different if `dont_skip_items` is set). The function should return true to stop processing items and return the ones processed so far (not including the item currently being processed). This is used, for example, in [[Module:alternative forms]], where an unspecified item signal the end of items and the start of labels. * `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set, causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem. If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations. Three values are returned: the list of items; the processed `args` structure; and an object of miscellaneous computed global values (currently only `use_semicolon`, indicating that commas were found in individual arguments and so the default separator should be a semicolon). In each returned item, there will be one field set for each specified property (either through inline modifiers or separate parameters). If subitems are not allowed, each item directly has fields set on it for the specified properties. If subitems ''are'' allowed, each item contains a `terms` field, which is a list of subitem objects, each of which has fields set on it for the specified properties of that subitem. In addition, the following fields may be set on each item or subitem: * `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given. * `orig_index`: The original index into the item in the items table returned by `process()` in [[Module:parameters]]. This may differ from `itemno` if there are raw semiclons and `disallow_custom_separators` is not given. * `itemno`: The logical index of the item. The index of separate parameters corresponds to this index. This may be different from `orig_index` in the presence of raw semicolons; see above. * `termlang`: If there is a language prefix, the corresponding language object is stored here (only if `parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set). * `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are set, the list of corresponding language objects is stored here. * `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified; (c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value. * `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b) `sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value. * `delimiter`: If subitems are allowed, this is set on subitems and specifies the delimiter used prior to the given subitem (e.g. {","}). * `separator`: The separator to display before the item. Always set on subitems, and set on top-level items if `disallow_custom_separators` is not given. Controlled by `special_separators` (for top-level items) and `subitem_separator_map` (for subitems). * `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified, {show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is called. Use {no_show_decorations = true} to prevent this. ]==] function export.parse_list_with_inline_modifiers_and_separate_params(data) validate_argument_related_fields(data) local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods if raw_args then local params = data.params local termarg_spec = params[termarg] if termarg_spec == true or not termarg_spec.list then internal_error("Term spec in `data.params` must have `list` set", termarg_spec) end if termarg_spec == true or not (termarg_spec.allow_holes or termarg_spec.disallow_holes) then internal_error("Term spec in `data.params` must have either `allow_holes` or `disallow_holes` set", termarg_spec) end export.augment_params_with_modifiers(params, param_mods) args = process_params(raw_args, params) else args = data.processed_args end local process_args_before_parsing = data.process_args_before_parsing if process_args_before_parsing then process_args_before_parsing(args) end -- Find the maximum index among any of the list parameters. local term_args = args[termarg] -- As a special case, the term args might not have a `maxindex` field because they might have -- been declared with `disallow_holes = true`, so fall back to the actual length of the list -- using the table_len function, since # can be unpredictable with arbitrary tables. local maxmaxindex = term_args.maxindex or table_len(term_args) for _, v in pairs(args) do if type(v) == "table" and v.maxindex and v.maxindex > maxmaxindex then maxmaxindex = v.maxindex end end local special_separators = data.special_separators or export.default_special_separators local items, lang_cache, use_semicolon = {}, data.lang_cache or {} local lang = fetch_argument(args, data.lang) if lang then lang_cache[lang:getCode()] = lang end local sc = fetch_argument(args, data.sc) local term_dest = data.term_dest or "term" -- FIXME: this is vulnerable to abusive inputs like 1000000=. local itemno = 0 for i = 1, maxmaxindex do local term = term_args[i] if data.disallow_custom_separators or not special_separators[term] then itemno = itemno + 1 -- Compute whether any of the separate indexed params exist for this index. local any_param_at_index for param_mod in pairs(param_mods) do local argval = args[param_mod] -- Careful with argument values that may be `false`. if argval then argval = argval[itemno] end if not argval_missing(argval) then any_param_at_index = true break end end if data.stop_when and data.stop_when{ term = term, -- FIXME, we should just pass in `any_param_at_index` directly. any_param_at_index = term ~= nil or any_param_at_index, orig_index = i, itemno = itemno, stored_itemno = #items + 1, } then break end -- If any of the params used for formatting this term is present, create a term and add it to the list. if not data.dont_skip_items and term == nil and not any_param_at_index then track("skipped-term", data.track_module) else if not term then track("missing-term", data.track_module) end local termobj = { itemno = itemno, orig_index = i, } if not data.disallow_custom_separators then termobj.separator = i == 1 and "" or special_separators[term_args[i - 1]] end -- Add 1 because first term index starts at 2. local paramname = termarg + i - 1 if term then local function generate_obj(term, parse_err) return export.generate_obj_maybe_parsing_lang_prefix { term = term, termobj = data.splitchar and {} or termobj, term_dest = term_dest, paramname = paramname, parse_lang_prefix = data.parse_lang_prefix, parse_err = parse_err, allow_bad_lang_prefix = data.allow_bad_lang_prefix, allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes, lang_cache = lang_cache, } end parse_inline_modifiers(term, { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, splitchar = data.splitchar, preserve_splitchar = true, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, outer_container = data.splitchar and termobj or nil, pre_normalize_modifiers = data.pre_normalize_modifiers, }) end -- FIXME: Make into an error, then remove after a month. if data.no_show_qualifiers then track("no_show_qualifiers") end local term_data = { args = args, param_mods = param_mods, itemno = itemno, termobj = termobj, term_dest = term_dest, has_subitems = not not data.splitchar, lang = lang, -- As a special case, if the caller defined a scN= separate param, set it on all subitems if there -- are multiple, falling back to the overall sc= param. sc = args.sc and args.sc[itemno] or sc, subitem_param_handling = data.subitem_param_handling, subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map, allow_conflicting_inline_mods_and_separate_params = data.allow_conflicting_inline_mods_and_separate_params, postprocess_termobj = postprocess_termobj, disallow_custom_separators = data.disallow_custom_separators, use_semicolon = use_semicolon, no_show_decorations = data.no_show_decorations or data.no_show_qualifiers, } copy_separate_params_to_termobj_and_postprocess(term_data) use_semicolon = term_data.use_semicolon insert(items, termobj) end end end if not data.disallow_custom_separators then -- Set the default separator of all those items for which a separator wasn't explicitly given to the default -- separator, defaulting to comma + space; but if any items have embedded commas, set the separator to -- semicolon + space. for _, item in ipairs(items) do if not item.separator then item.separator = use_semicolon and "; " or data.default_separator or ", " end end end return items, args, {use_semicolon = use_semicolon} end --[==[ Parse a single term that may have properties specified through inline modifiers or separate parameters. This differs from `parse_list_with_inline_modifiers_and_separate_params()` in that the latter is for parsing a list of terms, each of which may have properties specified through inline modifiers or separate parameters. Both functions optionally support having multiple subitems in a single term. This function is used e.g. for form-of templates ({{tl|inflection of}}/{{tl|infl of}}, {{tl|form of}}, and specific templates such as {{tl|alt form}}/{{tl|alternative form of}}, {{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others); for etymology templates ({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates like {{tl|ellipsis}}, {{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and for other templates with an argument structure similar to {{tl|l}} or {{tl|m}}. In these templates there is a term specified using a numeric parameter and associated separate parameters specifying term properties such as {{para|t}} for the gloss or {{para|tr}} for manual transliteration. All such properties can also be specified through inline modifiers attached directly to each term (`<t:...>`, `<tr:...>`, etc.). Normally it is an error if both an inline modifier and separate parameter for the same value are given, but this can be overridden (in which case inline modifiers take precedence over separate parameters when both occur). Some notable properties of this function: # Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]], containing only the term list itself along with any other parameters that are '''not''' term properties (such as a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is ''augmented'' with parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the latter case where raw argument processing is done by the caller, they must build the partial `params` structure; augment it themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the processed arguments. In both cases, the return value of this function contains two values, an object specifying the term and all properties; and the processed arguments structure, so that the non-term-property arguments can be processed as appropriate. # Optionally, the term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality is used, for example, by form-of templates. When this feature is invoked, the format of the term object changes; instead of directly being an object describing the term and its properties, it is an object with a `terms` field containing a list of per-subitem objects along with other top-level fields describing per-term properties. By default, if there are separate parameters specified along with multiple subitems, an error occurs, but this is controllable; currently, you can request that the parameters be assigned to the first or last subitem. `data` is an object containing several possible fields. 1. Fields that are required or recommended (usually related to argument processing): * `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from {frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw arguments. * `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both) of `raw_args` and `processed_args` must be set. * `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified manually. * `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters, '''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically "augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the "augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by [[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be correctly handled.) * `termarg` ('''required'''): The argument containing the item with attached inline modifiers to be parsed. Usually a numeric value such as {1} or {2}. * `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used internally to track pages containing template invocations with certain properties. * `lang` ('''recommended'''): The language object for the language of the item or subitems, or the name of the argument to fetch the object from. It is not strictly necessary to specify this, as this function only initializes items based on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used for certain purposes: *# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language prefix). *# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same language code may appear in several subitems). The value of `lang` can be any of the following: * If a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching. * Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed to the function as its only argument. * Otherwise, it is used directly. * `sc` ('''recommended'''): The script object for the item or subitems, or the name of the argument to fetch the object from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not strictly necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned items if not otherwise set (e.g. by the {{para|sc}} parameter or `<sc:...>` inline modifier). The most common value is {"sc"}. * `make_separate_g_into_list`: Set this to {true} if separate gender parameters exist are are specified using {{para|g}}, {{para|g2}}, etc. instead of using a single comma-separated {{para|g}} field. 2. Other argument-related fields: * `adjust_params_before_arg_processing`: An optional function to further adjust the `params` structure prior to calling `process()` in [[Module:parameters]]. This should be used when there are mismatches between the format of a given property as an inline modifier and the corresponding property as a separate parameter (as with the {{para|g}} parameter and {{cd|<g:...>}} modifier, but this particular case is handled by the `make_separate_g_into_list` field). * `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed arguments. It should make modifications in-place. * `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language prefixes are stripped off. Defaults to {"term"}. * `pre_normalize_modifiers`: As in `parse_inline_modifiers()`. * `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence. * `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set, causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem. If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations. 3. Fields related to language prefixes: * `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in [[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the `termlangs` field, and the first specified object is stored in the `lang` field). * `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above. * `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an error and will tend to remain unfixed. 4. Fields related to multiple subitems in the term: * `splitchar`: A Lua pattern. If specified, the user-specified argument can consist of multiple delimiter-separated subitems, each of which may be followed by inline modifiers. In this case, the first returned value is no longer an object describing the item, but instead an object with a `terms` field, whose value is a list describing the subitems (whose format is the same as the normal format of the item when `splitchar` is not specified). Each subitem object will have a `delimiter` field holding the actual delimiter occurring before the subitem, which is useful in the case where `splitchar` matches multiple possible characters. In this case, it is possible to specify that a given modifier can only occur after the last subitem and effectively modifies the whole collection of subitems by setting `overall = true` on the modifier. In this case, the modifier's value will be stored in the top-level object (the object with the `terms` field specifying the subitems). Note that splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). In addition, the algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers or inside of square brackets, which do not trigger splitting (whether or not contained within protected sequences). * `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in [[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified (see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that comma+whitespace sequences won't be split. * `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems. The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored into all subitems. * `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator` field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar` are recognized. Delimiters not in this map display as-is. Two values are returned, an object describing the item (or subitems) and the processed `args` structure. In the returned item, there will be one field set for each specified property (either through inline modifiers or separate parameters). If subitems are not allowed, the item directly has fields set on it for the specified properties. If subitems ''are'' allowed, the item contains a `terms` field, which is a list of subitem objects, each of which has fields set on it for the specified properties of that subitem. In addition, the following fields may be set on the item or each subitem: * `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given. * `termlang`: If there is a language prefix, the corresponding language object is stored here (only if `parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set). * `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are set, the list of corresponding language objects is stored here. * `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified; (c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value. * `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b) `sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value. * `delimiter`: If subitems are allowed, this specifies the delimiter used prior to the given subitem (e.g. {","}). * `separator`: If subitems are allowed, this specifies the displayed form of the delimiter to be shown before a given subitem. The mapping from user-specified delimiters to displayed separators is handled by `subitem_separator_map`; see above. The first subitem always has a blank string in the `separator` field. * `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified, {show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is called. Use {no_show_decorations = true} to prevent this. ]==] function export.parse_term_with_inline_modifiers_and_separate_params(data) validate_argument_related_fields(data) local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods if raw_args then local params = data.params local termarg_spec = params[termarg] if type(termarg_spec) == "table" and termarg_spec.list then internal_error("Term spec in `data.params` must not have `list` set", termarg_spec) end export.augment_params_with_modifiers(params, param_mods, "always") if data.make_separate_g_into_list then -- HACK: g= is a list for compatibility, but sublist as an inline parameter. params.g = {list = true, item_dest = "genders", type = "genders", flatten = true} end local adjust_params_before_arg_processing = data.adjust_params_before_arg_processing if adjust_params_before_arg_processing then adjust_params_before_arg_processing(params) end args = process_params(raw_args, params) else args = data.processed_args end local process_args_before_parsing = data.process_args_before_parsing if process_args_before_parsing then process_args_before_parsing(args) end local term, lang_cache = args[termarg], data.lang_cache local lang = fetch_argument(args, data.lang) if lang and lang_cache then lang_cache[lang:getCode()] = lang end local sc = fetch_argument(args, data.sc) local term_dest = data.term_dest or "term" if not term then track("missing-term", data.track_module) end local termobj, splitchar = {}, data.splitchar if term then local function generate_obj(term, parse_err) return export.generate_obj_maybe_parsing_lang_prefix { term = term, termobj = splitchar and {} or termobj, term_dest = term_dest, paramname = termarg, parse_lang_prefix = data.parse_lang_prefix, parse_err = parse_err, allow_bad_lang_prefix = data.allow_bad_lang_prefix, allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes, lang_cache = lang_cache, } end parse_inline_modifiers(term, { paramname = termarg, param_mods = param_mods, generate_obj = generate_obj, splitchar = splitchar, preserve_splitchar = true, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, outer_container = splitchar and termobj or nil, pre_normalize_modifiers = data.pre_normalize_modifiers, }) end -- FIXME: Make into an error, then remove after a month. if data.no_show_qualifiers then track("no_show_qualifiers") end copy_separate_params_to_termobj_and_postprocess { args = args, param_mods = param_mods, termobj = termobj, has_subitems = not not splitchar, lang = lang, sc = sc, subitem_param_handling = data.subitem_param_handling, subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map, allow_conflicting_inline_mods_and_separate_params = data.allow_conflicting_inline_mods_and_separate_params, no_show_decorations = data.no_show_decorations or data.no_show_qualifiers, } if splitchar and termobj.terms[2] then track("parse-term-multiple-subitems", data.track_module) track("parse-term-multiple-subitems") end return termobj, args end return export 1kp3k2x2ejn7z1nhgqsp5b8fxr2qs7w and 0 8060 54873 2026-09-27T21:22:11Z ~2026-51979-14 7217 Gesceop tramet þe hafaþ '=={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ˈænd ===Wordstǣr=== Of {{inh|en|ang|and}} ===Fégung=== # [[and]] [[Flocc: Nīwenglisc fégung]]' 54873 wikitext text/x-wiki =={{sprǣc|en}}== ===Rihtstefn=== * IPA: /ˈænd ===Wordstǣr=== Of {{inh|en|ang|and}} ===Fégung=== # [[and]] [[Flocc: Nīwenglisc fégung]] 5ogya3j4nn79odg007uyim2i6i9r8n6 Bysen:inherited 10 8061 54874 2026-09-27T21:23:53Z Deadend0914 7211 Gesceop tramet þe hafaþ '<includeonly>{{#invoke:etymology/templates|inherited}}</includeonly><!--' 54874 wikitext text/x-wiki <includeonly>{{#invoke:etymology/templates|inherited}}</includeonly><!-- bq7ejryulb4otc87jkojcpd6hknr5sk Bysen:inh 10 8062 54875 2026-09-27T21:23:58Z Deadend0914 7211 Redirected page to [[Bysen:inherited]] 54875 wikitext text/x-wiki #REDIRECT [[Bysen:inherited]] nihfwodm2xnif41h8wseemi0bdrolfc Module:etymology/templates 828 8063 54876 2026-09-27T21:24:38Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local require_when_needed = require("Module:require when needed") local get_current_L2 = require_when_needed("Module:pages", "get_current_L2") local get_lang_by_name = require_when_needed("Module:languages", "getByCanonicalName") local is_content_page = require_when_needed("Module:pages", "is_content_page") local process_params = require_when_needed("Module:parameters", "process") local trim = mw.text.trim local lower = mw.ustring.lower local...' 54876 Scribunto text/plain local export = {} local require_when_needed = require("Module:require when needed") local get_current_L2 = require_when_needed("Module:pages", "get_current_L2") local get_lang_by_name = require_when_needed("Module:languages", "getByCanonicalName") local is_content_page = require_when_needed("Module:pages", "is_content_page") local process_params = require_when_needed("Module:parameters", "process") local trim = mw.text.trim local lower = mw.ustring.lower local etymology_module = "Module:etymology" local headword_data_module = "Module:headword/data" local etymology_specialized_module = "Module:etymology/specialized" local parameter_utilities_module = "Module:parameter utilities" -- For testing local force_cat = false local allowed_conjs = {"and", "or", ",", "/", "~", ";"} -- Sinitic lects (Mandarin, Cantonese, Hokkien, etc.) are full languages, but Chinese entries sit -- under a single ==Chinese== L2, while romanization entries (pinyin, jyutping, pe̍h-ōe-jī) have lect -- L2s, and content is often shared between them. Contact languages (Chinese-based creoles and mixed -- languages) have their own L2s and are excluded. local function is_sinitic(lang) return lang:inFamily("zhx") and not lang:inFamily("qfa-cnt") end local content_page local function is_content_page_cached() if content_page == nil then content_page = is_content_page(mw.title.getCurrentTitle()) end return content_page end -- Throw an error if `lang` (the language of the entry) doesn't match -- the L2 header that the template is invoked under. local function check_lang_matches_L2(lang, nocat) if nocat or not lang or lang:getCode() == "und" or (lang.hasType and lang:hasType("family")) then return end local headword_data = mw.loadData(headword_data_module) if headword_data.large_pages[headword_data.pagename] then return end if not is_content_page_cached() then return end local current_L2 = get_current_L2() if not current_L2 then return end local full_name = lang:getFullName() if full_name == current_L2 then return end -- Accept any Sinitic language under any Sinitic L2. if is_sinitic(lang) then local L2_lang = get_lang_by_name(current_L2) if L2_lang and is_sinitic(L2_lang) then return end end local lang_desc = lang:getCode() .. " (" .. lang:getCanonicalName() .. ")" if lang:getFullCode() ~= lang:getCode() then lang_desc = lang_desc .. ", an etymology-only language whose full language is " .. lang:getFullCode() .. " (" .. full_name .. ")" end error("Language '" .. lang_desc .. "' does not match the L2 header (" .. current_L2 .. ").") end local function parse_etym_args(parent_args, base_params, has_dest_lang) local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "q", "l", "ref"}}, } local sourcearg, termarg if has_dest_lang then sourcearg, termarg = 2, 3 else sourcearg, termarg = 1, 2 end local terms, args = m_param_utils.parse_term_with_inline_modifiers_and_separate_params { params = base_params, param_mods = param_mods, raw_args = parent_args, termarg = termarg, track_module = "etymology", lang = function(args) return args[sourcearg][#args[sourcearg]] end, sc = "sc", -- Don't do this, doesn't seem to make sense. -- parse_lang_prefix = true, make_separate_g_into_list = true, splitchar = ",", subitem_param_handling = "last", } -- If term param 3= is empty, there will be no terms in terms.terms. To facilitate further code and for -- compatibility,, insert one. It will display as <small>[Term?]</small>. if not terms.terms[1] then terms.terms[1] = { lang = args[sourcearg][#args[sourcearg]], sc = args.sc, } end return terms.terms, args end function export.parse_2_lang_args(parent_args, has_text, no_family) local boolean = {type = "boolean"} local params = { [1] = { required = true, type = "language", default = "und" }, [2] = { required = true, sublist = true, type = "language", family = not no_family, default = "und" }, [3] = true, [4] = {alias_of = "alt"}, [5] = {alias_of = "t"}, ["senseid"] = true, ["nocat"] = boolean, ["sort"] = true, ["sourceconj"] = true, ["conj"] = {set = allowed_conjs, default = ","}, } if has_text then params["notext"] = boolean params["nocap"] = boolean end local terms, args = parse_etym_args(parent_args, params, "has dest lang") check_lang_matches_L2(args[1], args.nocat) return terms, args end -- Implementation of deprecated {{etyl}}. Provided to make histories more legible. function export.etyl(frame) local params = { [1] = {required = true, type = "language", default = "und"}, [2] = {type = "language", default = "en"}, ["sort"] = {}, } -- Empty language means English, but "-" means no language. Yes, confusing... local args = frame:getParent().args if args[2] and trim(args[2]) == "-" then params[2] = nil args = process_params({ [1] = args[1], ["sort"] = args.sort }, params) else args = process_params(args, params) end check_lang_matches_L2(args[2]) return require(etymology_module).format_source { lang = args[2], source = args[1], sort_key = args.sort, force_cat = force_cat, } end -- Implementation of {{derived}}/{{der}}. function export.derived(frame) local parent_args = frame:getParent().args local terms, args = export.parse_2_lang_args(parent_args) return require(etymology_module).format_derived { lang = args[1], sources = args[2], terms = terms, sort_key = args.sort, nocat = args.nocat, sourceconj = args.sourceconj, conj = args.conj, template_name = "derived", force_cat = force_cat, } end -- Implementation of {{borrowed}}/{{bor}}. function export.borrowed(frame) local parent_args = frame:getParent().args local terms, args = export.parse_2_lang_args(parent_args) return require(etymology_module).format_borrowed { lang = args[1], sources = args[2], terms = terms, sort_key = args.sort, nocat = args.nocat, sourceconj = args.sourceconj, conj = args.conj, force_cat = force_cat, } end function export.inherited(frame) local parent_args = frame:getParent().args local terms, args = export.parse_2_lang_args(parent_args) local sources = args[2] if sources[2] then -- Because this doesn't really make sense. error("[[Template:inherited]] doesn't support multiple comma-separated sources") end return require(etymology_module).format_inherited { lang = args[1], terms = terms, sort_key = args.sort, nocat = args.nocat, conj = args.conj, force_cat = force_cat, } end function export.cognate(frame) local params = { [1] = { required = true, sublist = true, type = "language", family = true, default = "und" }, [2] = true, [3] = {alias_of = "alt"}, [4] = {alias_of = "t"}, sourceconj = true, ["conj"] = {set = allowed_conjs, default = ","}, sort = true, } local parent_args = frame:getParent().args local terms, args = parse_etym_args(parent_args, params, false) return require(etymology_module).format_cognate { sources = args[1], terms = terms, sort_key = args.sort, sourceconj = args.sourceconj, conj = args.conj, force_cat = force_cat, } end function export.noncognate(frame) return export.cognate(frame) end -- Supports various specialized types of borrowings, according to `frame.args.bortype`: -- "learned" = {{lbor}}/{{learned borrowing}} -- "semi-learned" = {{slbor}}/{{semi-learned borrowing}} -- "orthographic" = {{obor}}/{{orthographic borrowing}} -- "unadapted" = {{ubor}}/{{unadapted borrowing}} -- "calque" = {{cal}}/{{calque}} -- "partial-calque" = {{pcal}}/{{partial calque}} -- "semantic-loan" = {{sl}}/{{semantic loan}} -- "transliteration" = {{translit}}/{{transliteration}} -- "phono-semantic-matching" = {{psm}}/{{phono-semantic matching}} function export.specialized_borrowing(frame) local parent_args = frame:getParent().args local terms, args = export.parse_2_lang_args(parent_args, "has text") local m_etymology_specialized = require(etymology_specialized_module) return m_etymology_specialized.specialized_borrowing { bortype = frame.args.bortype, lang = args[1], sources = args[2], terms = terms, sort_key = args.sort, nocap = args.nocap, notext = args.notext, nocat = args.nocat, sourceconj = args.sourceconj, conj = args.conj, senseid = args.senseid, force_cat = force_cat, } end -- Implementation of miscellaneous templates such as {{abbrev}}, {{back-formation}}, {{clipping}}, {{ellipsis}}, -- {{rebracketing}} and {{reduplication}} that have a single associated term. function export.misc_variant(frame) local iparams = { ["ignore-params"] = true, text = {required = true}, oftext = true, cat = {list = true}, -- allow and compress holes conj = true, } local iargs = process_params(frame.args, iparams) local boolean = {type = "boolean"} local params = { [1] = {required = true, type = "language", default = "und"}, [2] = true, [3] = {alias_of = "alt"}, [4] = {alias_of = "t"}, nocap = boolean, -- should be processed in the template itself notext = boolean, nocat = boolean, conj = {set = allowed_conjs}, sort = true, } -- |ignore-params= parameter to module invocation specifies -- additional parameter names to allow in template invocation, separated by -- commas. They must consist of ASCII letters or numbers or hyphens. local ignore_params = iargs["ignore-params"] if ignore_params then ignore_params = trim(ignore_params) if not ignore_params:match("^[%w%-,]+$") then error("Invalid characters in |ignore-params=: " .. ignore_params:gsub("[%w%-,]+", "")) end for param in ignore_params:gmatch("[%w%-]+") do if params[param] then error("Duplicate param |" .. param .. " in |ignore-params=: already specified in params") end params[param] = true end end local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "q", "l", "ref"}}, } local parent_args = frame:getParent().args local terms, args = m_param_utils.parse_term_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, track_module = "etymology", -- Don't set lang here as we want to know whether there was a lang prefix or not. sc = "sc", parse_lang_prefix = true, allow_multiple_lang_prefixes = true, make_separate_g_into_list = true, splitchar = ",", subitem_param_handling = "last", } check_lang_matches_L2(args[1], args.nocat) return require(etymology_module).format_misc_variant { lang = args[1], notext = args.notext, text = iargs.text, oftext = iargs.oftext, terms = terms.terms, sort_key = args.sort, conj = args.conj or iargs.conj or "and", nocat = args.nocat, cats = iargs.cat, force_cat = force_cat, } end -- Implementation of miscellaneous templates such as {{doublet}} that can take multiple terms. Doesn't handle {{blend}} -- or {{univerbation}}, which display + signs between elements and use compound_like in [[Module:affix/templates]]. function export.misc_variant_multiple_terms(frame) local iparams = { text = {required = true}, oftext = true, cat = {list = true}, -- allow and compress holes conj = true, } local iargs = process_params(frame.args, iparams) local boolean = {type = "boolean"} local params = { [1] = {required = true, type = "language", template_default = "und"}, [2] = {list = true, allow_holes = true}, nocap = boolean, -- should be processed in the template itself notext = boolean, nocat = boolean, conj = {set = allowed_conjs}, sort = true, } local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { -- We want to require an index for all params. {default = true, require_index = true}, {group = {"link", "q", "l", "ref"}}, } local parent_args = frame:getParent().args local terms, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, parse_lang_prefix = true, allow_multiple_lang_prefixes = true, track_module = "etymology-templates-doublet", disallow_custom_separators = true, -- For compatibility, we need to not skip completely unspecified items. It is common, for example, to do -- {{suffix|lang||foo}} to generate "+ -foo". dont_skip_items = true, -- Don't set lang here as we want to know whether there was a lang prefix or not. sc = "sc.default", } check_lang_matches_L2(args[1], args.nocat) return require(etymology_module).format_misc_variant { lang = args[1], notext = args.notext, text = iargs.text, oftext = iargs.oftext, terms = terms, sort_key = args.sort, conj = args.conj or iargs.conj or "and", nocat = args.nocat, cats = iargs.cat, force_cat = force_cat, } end -- Implementation of miscellaneous templates such as {{unknown}} that have no associated terms. do local function get_args(frame) local boolean = {type = "boolean"} local params = { [1] = {required = true, type = "language", default = "und"}, ["title"] = true, ["nocap"] = boolean, -- should be processed in the template itself ["notext"] = boolean, ["nocat"] = boolean, ["sort"] = true, } if frame.args.title2_alias then params[2] = {alias_of = "title"} end local args = process_params(frame:getParent().args, params) check_lang_matches_L2(args[1], args.nocat) return args end function export.misc_variant_no_term(frame) local args = get_args(frame) return require(etymology_module).format_misc_variant_no_term { lang = args[1], notext = args.notext, title = args.title or frame.args.text, nocat = args.nocat, cat = frame.args.cat, sort_key = args.sort, force_cat = force_cat, } end -- This function works similarly to misc_variant_no_term(), but with some automatic linking to the glossary in -- `title`. function export.onomatopoeia(frame) local args = get_args(frame) local title = args.title if title and (lower(title) == "imitative" or lower(title) == "imitation") then title = "[[Appendix:Glossary#imitative|" .. title .. "]]" end return require(etymology_module).format_misc_variant_no_term { lang = args[1], notext = args.notext, title = title or frame.args.text, nocat = args.nocat, cat = frame.args.cat, sort_key = args.sort, force_cat = force_cat, } end end return export 62uclsg0tnc97nmhtj6qaxzsgsp4laq Module:etymology 828 8064 54877 2026-09-27T21:25:05Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} -- For testing local force_cat = false local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local insert = table.insert local new_title = mw.title.new local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local...' 54877 Scribunto text/plain local export = {} -- For testing local force_cat = false local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local insert = table.insert local new_title = mw.title.new local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_language_data_module_name(...) get_language_data_module_name = require(languages_module).getDataModuleName return get_language_data_module_name(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function serial_comma_join(...) serial_comma_join = require(table_module).serialCommaJoin return serial_comma_join(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function track(page, code) local tracking_page = "etymology/" .. page debug_track(tracking_page) if code then debug_track(tracking_page .. "/" .. code) end end local function join_segs(segs, conj) if not segs[2] then return segs[1] elseif conj == "and" or conj == "or" then return serial_comma_join(segs, {conj = conj}) end local sep if conj == "," or conj == ";" then sep = conj .. " " elseif conj == "/" then sep = "/" elseif conj == "~" then sep = " ~ " elseif conj then error(("Internal error: Unrecognized conjunction \"%s\""):format(conj)) else error(("Internal error: No value supplied for conjunction"):format(conj)) end return concat(segs, sep) end -- Returns true if `lang` is the same as `source`, or a variety of it. local function lang_is_source(lang, source) return lang:getCode() == source:getCode() or lang:hasParent(source) end --[==[ Format one or more links as specified in `termobjs`, a list of term objects of the format accepted by `full_link()` in [[Module:links]], including decorations (qualifiers, labels and references). `conj` is used to join multiple terms and must be specified if there is more than one term. `template_name` is the template name used in debug tracking and must be specified. Optional `sourcetext` is text to prepend to the concatenated terms, separated by a space if the concatenated terms are non-empty (which is always the case unless there is a single term with the value "-"). If `decorations_on_outside` is given, any decorations specified in the first term go on the outside of (i.e before) `sourcetext`; otherwise they will end up on the inside. ]==] function export.format_links(termobjs, conj, template_name, sourcetext, decorations_on_outside) if not template_name then error("Internal error: Must specify `template_name` to format_links()") end for i, termobj in ipairs(termobjs) do if termobj.lang:hasType("family") or termobj.lang:getFamilyCode() == "qfa-sub" then if termobj.term and termobj.term ~= "-" then debug_track(template_name .. "/family-with-term") end termobj.term = "-" end if termobj.term == "-" then --[=[ [[Special:WhatLinksHere/Wiktionary:Tracking/cognate/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/derived/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/borrowed/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/calque/no-term]] ]=] debug_track(template_name .. "/no-term") termobjs[i] = i == 1 and sourcetext or "" else if i == 1 and decorations_on_outside and sourcetext then termobj.pretext = sourcetext .. " " sourcetext = nil end termobjs[i] = (i == 1 and sourcetext and sourcetext .. " " or "") .. full_link(termobj, "term") end end return join_segs(termobjs, conj) end function export.get_display_and_cat_name(source, raw) local display, cat_name if source:getCode() == "und" then display = "undetermined" cat_name = "other languages" elseif source:getCode() == "mul" then display = raw and "translingual" or "[[w:Translingualism|translingual]]" cat_name = "Translingual" elseif source:getCode() == "mul-tax" then display = raw and "taxonomic name" or "[[w:Biological nomenclature|taxonomic name]]" cat_name = "taxonomic names" else display = raw and source:getCanonicalName() or source:makeWikipediaLink() cat_name = source:getDisplayForm() end return display, cat_name end function export.insert_source_cat_get_display(data) local categories, lang, source = data.categories, data.lang, data.source local display, cat_name = export.get_display_and_cat_name(source, data.raw) if lang and not data.nocat then -- Add the category, but only if there is a current language if not categories then categories = {} end local langname = lang:getFullName() -- If `lang` is an etym-only language, we need to check both it and its parent full language against `source`. -- Otherwise if e.g. `lang` is Medieval Latin and `source` is Latin, we'll end up wrongly constructing a -- category 'Latin terms derived from Latin'. insert(categories, langname .. ( lang_is_source(lang, source) and " terms borrowed back into " .. cat_name or " " .. (data.borrowing_type or "terms derived") .. " from " .. cat_name )) end return display, categories end function export.format_source(data) local lang, sort_key = data.lang, data.sort_key -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/sortkey]] if sort_key then track("sortkey") end local display, categories = export.insert_source_cat_get_display(data) if lang and not data.nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories categories = format_categories(categories, lang, sort_key, nil, data.force_cat or force_cat) else categories = "" end return "<span class=\"etyl\">" .. display .. categories .. "</span>" end --[==[ Format sources for etymology templates such as {{tl|bor}}, {{tl|der}}, {{tl|inh}} and {{tl|cog}}. There may potentially be more than one source language (except currently {{tl|inh}}, which doesn't support it because it doesn't really make sense). In that case, all but the last source language is linked to the first term, but only if there is such a term and this linking makes sense, i.e. either (1) the term page exists after stripping diacritics according to the source language in question, or (2) the result of stripping diacritics according to the source language in question results in a different page from the same process applied with the last source language. For example, {{m|ru|соля́нка}} will link to [[солянка]] but {{m|en|соля́нка}} will link to [[соля́нка]] with an accent, and since they are different pages, the use of English as a non-final source with term 'соля́нка' will link to [[соля́нка]] even though it doesn't exist, on the assumption that it is merely a redlink that might exist. If none of the above criteria apply, a non-final source language will be linked to the Wikipedia entry for the language, just as final source languages always are. `data` contains the following fields: * `lang`: The destination language object into which the terms were borrowed, inherited or otherwise derived. Used for categorization and can be nil, as with {{tl|cog}}. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see above. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_sources(data) local lang, sources, terms, borrowing_type, sort_key, categories, nocat = data.lang, data.sources, data.terms, data.borrowing_type, data.sort_key, data.categories, data.nocat local term1, sources_n, source_segs = terms[1], #sources, {} local final_link_page local term1_term, term1_sc = term1.term, term1.sc if sources_n > 1 and term1_term and term1_term ~= "-" then final_link_page = get_link_page(term1_term, sources[sources_n], term1_sc) end for i, source in ipairs(sources) do local seg, display_term if i < sources_n and term1_term and term1_term ~= "-" then local link_page = get_link_page(term1_term, source, term1_sc) display_term = (link_page ~= final_link_page) or (link_page and not not new_title(link_page):getContent()) end -- TODO: if the display forms or transliterations are different, display the terms separately. if display_term then local display, this_cats = export.insert_source_cat_get_display{ lang = lang, source = source, borrowing_type = borrowing_type, raw = true, categories = categories, nocat = nocat, } seg = language_link { lang = source, term = term1_term, alt = display, tr = "-", } if lang and not nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories this_cats = format_categories(this_cats, lang, sort_key, nil, data.force_cat or force_cat) else this_cats = "" end seg = "<span class=\"etyl\">" .. seg .. this_cats .. "</span>" else seg = export.format_source{ lang = lang, source = source, borrowing_type = borrowing_type, sort_key = sort_key, categories = categories, nocat = nocat, } end insert(source_segs, seg) end return join_segs(source_segs, data.sourceconj or "and") end -- Internal implementation of {{cognate}}/{{cog}} template. function export.format_cognate(data) return export.format_derived { sources = data.sources, terms = data.terms, sort_key = data.sort_key, sourceconj = data.sourceconj, conj = data.conj, template_name = "cognate", force_cat = data.force_cat, } end --[==[ Internal implementation of {{derived}}/{{der}} template. This is called externally from [[Module:affix]], [[Module:affixusex]] and [[Module:see]] and needs to support decorations (qualifiers, labels and references) on the outside of the sources for use by those modules. `data` contains the following fields: * `lang`: The destination language object into which the terms were derived. Used for categorization and can be nil, as with {{tl|cog}}; in this case, no categories are added. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see `format_sources()`. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `conj`: Conjunction used to separate multiple terms. '''Required'''. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are as for `conj` above. * `decorations_on_outside`: If specified, any decorations (qualifiers, labels or references) in the first term in `terms` will be displayed on the outside of (before) the source language(s) in `sources`. Normally this should be specified if there is only one term possible in `terms`. * `template_name`: Name of the template invoking this function. Must be specified. Only used for tracking pages. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_derived(data) local terms = data.terms local sourcetext = export.format_sources(data) return export.format_links(terms, data.conj, data.template_name, sourcetext, data.decorations_on_outside) end function export.insert_borrowed_cat(categories, lang, source) if lang_is_source(lang, source) then return end -- If both are the same, we want e.g. [[:Category:English terms borrowed back into English]] not -- [[:Category:English terms borrowed from English]]; the former is inserted automatically by format_source(). -- The second parameter here doesn't matter as it only affects `display`, which we don't use. insert(categories, lang:getFullName() .. " terms borrowed from " .. select(2, export.get_display_and_cat_name(source, "raw"))) end -- Internal implementation of {{borrowed}}/{{bor}} template. function export.format_borrowed(data) local categories = {} if not data.nocat then local lang = data.lang for _, source in ipairs(data.sources) do export.insert_borrowed_cat(categories, lang, source) end end data = shallow_copy(data) data.categories = categories return export.format_links(data.terms, data.conj, "borrowed", export.format_sources(data)) end do -- Generate the non-ancestor error message. local function show_language(lang) local retval = ("%s (%s)"):format(lang:makeCategoryLink(), lang:getCode()) if lang:hasType("etymology-only") then retval = retval .. (" (an etymology-only language whose regular parent is %s)"):format( show_language(lang:getParent())) end return retval end -- Check that `lang` has `otherlang` (which may be an etymology-only language) as an ancestor. Throw an error if -- not. When `lang` is a family, verifies that `otherlang` is a language in that family. function export.check_ancestor(lang, otherlang) -- When `lang` is a family, verify `otherlang` is in that family or in its parent family. if lang.hasType and lang:hasType("family") then local family_code = lang:getCode() local function in_family_code(fcode, other) if not fcode or fcode == "" then return false end if other.inFamily and other:inFamily(fcode) then return true end if other.getFamilyCode and other:getFamilyCode() == fcode then return true end return false end local in_family = in_family_code(family_code, otherlang) if not in_family then local parent_code if lang.getParent then local parent_family = lang:getParent() if parent_family and parent_family.getCode then parent_code = parent_family:getCode() end end if not parent_code and family_code:find("-", 1, true) then parent_code = family_code:match("^(.+)-[^-]+$") end if parent_code then in_family = in_family_code(parent_code, otherlang) end end if not in_family then local other_display = (otherlang.getCanonicalName and otherlang:getCanonicalName()) or (otherlang.getCode and otherlang:getCode()) or tostring(otherlang) local fam_display = (lang.getCanonicalName and lang:getCanonicalName()) or family_code error(("%s is not in family %s; inherited ancestor under a family must be a language in that family or its parent family.") :format(other_display, fam_display)) end return end -- FIXME: I don't know if this function works correctly with etym-only languages in `lang`. I have fixed up -- the module link code appropriately (June 2024) but the remaining logic is untouched. if lang:hasAncestor(otherlang) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/variety]] -- Track inheritance from varieties of Latin that shouldn't have any descendants (everything except Old Latin, Classical Latin and Vulgar Latin). if otherlang:getFullCode() == "la" then otherlang = otherlang:getCode() if not (otherlang == "itc-ola" or otherlang == "la-cla" or otherlang == "la-vul") then track("bad ancestor", otherlang) end end return end local ancestors = lang:getAncestors() local postscript local etym_module_link = lang:hasType("etymology-only") and "[[Module:etymology languages/data]] or " or "" local module_link = "[[" .. get_language_data_module_name(lang:getFullCode()) .. "]]" if not ancestors[1] then postscript = show_language(lang) .. " has no ancestors." else local ancestor_list = {} for _, ancestor in ipairs(ancestors) do insert(ancestor_list, show_language(ancestor)) end postscript = ("The ancestor%s of %s %s %s."):format( ancestors[2] and "s" or "", lang:getCanonicalName(), ancestors[2] and "are" or "is", concat(ancestor_list, " and ")) end error(("%s is not set as an ancestor of %s in %s%s. %s") :format(show_language(otherlang), show_language(lang), etym_module_link, module_link, postscript)) end end -- Internal implementation of {{inherited}}/{{inh}} template. function export.format_inherited(data) local lang, terms, nocat = data.lang, data.terms, data.nocat local source = terms[1].lang local categories = {} if not nocat then insert(categories, lang:getFullName() .. " terms inherited from " .. source:getCanonicalName()) end export.check_ancestor(lang, source) data = shallow_copy(data) data.categories = categories data.source = source return export.format_links(terms, data.conj, "inherited", export.format_source(data)) end -- Internal implementation of "misc variant" templates such as {{abbrev}}, {{clipping}}, {{reduplication}} and the like. function export.format_misc_variant(data) local lang, notext, terms, cats, parts = data.lang, data.notext, data.terms, data.cats, {} if not notext then insert(parts, data.text) end if terms[1] then if not notext then -- FIXME: If term is given as '-', we should consider displaying just "Clipping" not "Clipping of". insert(parts, " " .. (data.oftext or "of")) end local termparts = {} -- Make links out of all the parts. for _, termobj in ipairs(terms) do local result if termobj.lang then result = export.format_derived { lang = lang, terms = {termobj}, sources = termobj.termlangs or {termobj.lang}, template_name = "misc_variant", decorations_on_outside = true, force_cat = data.force_cat, } else termobj.lang = lang result = export.format_links({termobj}, nil, "misc_variant") end table.insert(termparts, result) end local linktext = join_segs(termparts, data.conj) if not notext and linktext ~= "" then insert(parts, " ") end insert(parts, linktext) end local categories = {} if not data.nocat and cats then for _, cat in ipairs(cats) do insert(categories, lang:getFullName() .. " " .. cat) end end if categories[1] then insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end -- Implementation of miscellaneous templates such as {{unknown}} and {{onomatopoeia}} that have no associated terms. function export.format_misc_variant_no_term(data) local parts = {} if not data.notext then insert(parts, data.title) end if not data.nocat and data.cat then local lang, categories = data.lang, {} insert(categories, lang:getFullName() .. " " .. data.cat) insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end return export bz4oy4imn1osztk04jf87q163ctpr9d Module:etymology/specialized 828 8065 54878 2026-09-27T21:26:30Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local m_str_utils = require("Module:string utilities") local en_utilities_module = "Module:en-utilities" local etymology_module = "Module:etymology" local gsub = m_str_utils.gsub local insert = table.insert local pluralize = require(en_utilities_module).pluralize local upper = m_str_utils.upper -- This function handles all the messiness of different types of specialized borrowings. It should insert any -- borrowing-type-specific categories in...' 54878 Scribunto text/plain local export = {} local m_str_utils = require("Module:string utilities") local en_utilities_module = "Module:en-utilities" local etymology_module = "Module:etymology" local gsub = m_str_utils.gsub local insert = table.insert local pluralize = require(en_utilities_module).pluralize local upper = m_str_utils.upper -- This function handles all the messiness of different types of specialized borrowings. It should insert any -- borrowing-type-specific categories into `categories` unless `nocat` is given, and return the text to display -- before the source + term (or "" for no text). local function get_specialized_borrowing_text_insert_cats(data) local bortype, categories, lang, terms, source, nocap, nocat, senseid = data.bortype, data.categories, data.lang, data.terms, data.source, data.nocap, data.nocat, data.senseid local function inscat(cat) if not nocat then local display, sourcedisp = require(etymology_module).get_display_and_cat_name(source, "raw") if cat:find("DISPLAY") then cat = cat:gsub("DISPLAY", display) elseif cat:find("SOURCE") then cat = cat:gsub("SOURCE", sourcedisp) else cat = cat .. " " .. sourcedisp end insert(categories, lang:getFullName() .. " " .. cat) end end -- `text` is the display text for the borrowing type, which gets converted -- into a link. -- `appendix` is a the glossary anchor, which defaults to `text` -- `prep` is the preposition between the borrowing type and the language -- name (e.g. "of", "from") -- `pos` is the part of speech for the borrowing type ("noun" or -- "adjective"; defaults to "noun") -- `plural` is the plural form of the borrowing type; if not specified, -- the pluralize function is used local text, appendix, prep, pos, plural if bortype == "calque" then text, prep = "calque", "of" inscat("terms calqued from") elseif bortype == "partial-calque" then text, prep = "partial calque", "of" inscat("terms partially calqued from") elseif bortype == "semantic-loan" then text, prep = "semantic loan", "from" inscat("semantic loans from") elseif bortype == "transliteration" then text, prep = "transliteration", "of" inscat("terms borrowed from") inscat("transliterations of DISPLAY terms") elseif bortype == "phono-semantic-matching" then text, prep = "phono-semantic matching", "of" inscat("phono-semantic matchings from") else local langcode = lang:getCode() local lang_is_source = langcode == source:getCode() if lang_is_source then -- Track, because this shouldn't be happening. A language can only have itself as a source further up the chain after a borrowing, which is always "derived". require("Module:debug/track"){ "etymology/specialized/self-as-source", "etymology/specialized/self-as-source/" .. langcode } inscat("terms borrowed back into") else inscat("terms borrowed from") if bortype ~= "borrowing" then inscat(bortype .. " borrowings from") end end if bortype == "borrowing" then text, appendix, prep, pos = "borrowed", "loanword", "from", "adjective" elseif ( bortype == "learned" or bortype == "semi-learned" or bortype == "orthographic" or bortype == "unadapted" ) then text, prep = bortype .. " borrowing", "from" elseif bortype == "adapted" then text, prep = bortype .. " borrowing", "of" else error("Internal error: Unrecognized bortype: " .. bortype) end end -- If the term is suppressed, the preposition should always be "from": -- "Calque of Chinese 中國". -- "Calque from Chinese" (not "Calque of Chinese"). if terms[1].term == "-" then prep = "from" end appendix = "Appendix:Glossary#" .. (appendix or text) if senseid then local senseids, output = mw.text.split(senseid, '!!'), {} for i, id in ipairs(senseids) do -- FIXME: This should be done via a function. insert(output, mw.getCurrentFrame():preprocess('{{senseno|' .. lang:getCode() .. '|' .. id .. (i == 1 and not nocap and "|uc=1" or "") .. '}}')) end local link if senseid:find('!!') then link, text = "are", pos == "adjective" and text or plural or pluralize(text) else link = pos == "adjective" and "is" or "is a" end text = mw.text.listToText(output) .. " " .. link .. " " .. '[[' .. appendix .. '|' .. text .. ']]' else text = "[[" .. appendix .. "|" .. (nocap and text or gsub(text, "^.", upper)) .. "]]" end return text .. " " .. prep .. " " end function export.specialized_borrowing(data) local lang, sources, terms = data.lang, data.sources, data.terms local categories = {} local text for _, source in ipairs(sources) do text = get_specialized_borrowing_text_insert_cats { bortype = data.bortype, categories = categories, lang = lang, terms = terms, source = source, nocap = data.nocap, nocat = data.nocat, senseid = data.senseid, } end text = data.notext and "" or text local sourcetext = require(etymology_module).format_sources { lang = lang, sources = sources, terms = terms, sort_key = data.sort_key, categories = categories, nocat = data.nocat, sourceconj = data.sourceconj, } return text .. require(etymology_module).format_links(terms, data.conj, "etymology/specialized", sourcetext) end return export bvgt9bjaugs1kua7m5g9pj20a0vngnq Module:parse utilities 828 8066 54879 2026-09-27T21:29:46Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local fun_is_callable_module = "Module:fun/isCallable" local languages_module = "Module:languages" local parameters_module = "Module:parameters" local string_char_module = "Module:string/char" local string_utilities_module = "Module:string utilities" local table_insert_if_not_module = "Module:table/insertIfNot" local assert = assert local concat = table.concat local dump = mw.dumpObject local error = error local insert = table.insert local ipa...' 54879 Scribunto text/plain local export = {} local fun_is_callable_module = "Module:fun/isCallable" local languages_module = "Module:languages" local parameters_module = "Module:parameters" local string_char_module = "Module:string/char" local string_utilities_module = "Module:string utilities" local table_insert_if_not_module = "Module:table/insertIfNot" local assert = assert local concat = table.concat local dump = mw.dumpObject local error = error local insert = table.insert local ipairs = ipairs local list_to_text = mw.text.listToText local pairs = pairs local require = require local sort = table.sort local type = type local ugsub = mw.ustring.gsub local function convert_val(...) convert_val = require(parameters_module).convert_val return convert_val(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function insert_if_not(...) insert_if_not = require(table_insert_if_not_module) return insert_if_not(...) end local function is_callable(...) is_callable = require(fun_is_callable_module) return is_callable(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function u(...) u = require(string_char_module) return u(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end --[==[ intro: In order to understand the following parsing code, you need to understand how inflected text specs work. They are intended to work with inflected text where individual words to be inflected may be followed by inflection specs in angle brackets. The format of the text inside of the angle brackets is up to the individual language and part-of-speech specific implementation. A real-world example is as follows: `<nowiki>[[медичний|меди́чна]]<+> [[сестра́]]<*,*#.pr></nowiki>`. This is the inflection of the Ukrainian multiword expression {{m|uk|меди́чна сестра́||nurse|lit=medical sister}}, consisting of two words: the adjective {{m|uk|меди́чна||medical|pos=feminine singular}} and the noun {{m|uk|сестра́||sister}}. The specs in angle brackets follow each word to be inflected; for example, `<+>` means that the preceding word should be declined as an adjective. The code below works in terms of balanced expressions, which are bounded by delimiters such as `< >` or `[ ]`. The intention is to allow separators such as spaces to be embedded inside of delimiters; such embedded separators will not be parsed as separators. For example, Ukrainian noun specs allow footnotes in brackets to be inserted inside of angle brackets; something like `меди́чна<+> сестра́<pr.[this is a footnote]>` is legal, as is `<nowiki>[[медичний|меди́чна]]<+> [[сестра́]]<pr.[this is an <i>italicized footnote</i>]></nowiki>`, and the parsing code should not be confused by the embedded brackets, spaces or angle brackets. The parsing is done by two functions, which work in close concert: {parse_balanced_segment_run()} and {split_alternating_runs()}. To illustrate, consider the following: {parse_balanced_segment_run("foo<M.proper noun> bar<F>", "<", ">")} =<br /> { {"foo", "<M.proper noun>", " bar", "<F>", ""}} then {split_alternating_runs({"foo", "<M.proper noun>", " bar", "<F>", ""}, " ")} =<br /> { {{"foo", "<M.proper noun>", ""}, {"bar", "<F>", ""}}} Here, we start out with a typical inflected text spec `foo<M.proper noun> bar<F>`, call {parse_balanced_segment_run()} on it, and call {split_alternating_runs()} on the result. The output of {parse_balanced_segment_run()} is a list where even-numbered segments are bounded by the bracket-like characters passed into the function, and odd-numbered segments consist of the surrounding text. {split_alternating_runs()} is called on this, and splits '''only''' the odd-numbered segments, grouping all segments between the specified character. Note that the inner lists output by {split_alternating_runs()} are themselves in the same format as the output of {parse_balanced_segment_run()}, with bracket-bounded text in the even-numbered segments. Hence, such lists can be passed again to {split_alternating_runs()}. ]==] --[==[ Parse a string containing matched instances of parens, brackets or the like. Return a list of strings, alternating between textual runs not containing the open/close characters and runs beginning and ending with the open/close characters. For example, {parse_balanced_segment_run("foo(x(1)), bar(2)", "(", ")") = {"foo", "(x(1))", ", bar", "(2)", ""}} ]==] function export.parse_balanced_segment_run(segment_run, open, close) return split(segment_run, "(%b" .. open .. close .. ")") end -- The following is an equivalent, older implementation that does not use %b (written before I was aware of %b). --[=[ function export.parse_balanced_segment_run(segment_run, open, close) local break_on_open_close = split(segment_run, "([%" .. open .. "%" .. close .. "])") local text_and_specs = {} local level = 0 local seg_group = {} for i, seg in ipairs(break_on_open_close) do if i % 2 == 0 then if seg == open then insert(seg_group, seg) level = level + 1 else assert(seg == close) insert(seg_group, seg) level = level - 1 if level < 0 then error("Unmatched " .. close .. " sign: '" .. segment_run .. "'") elseif level == 0 then insert(text_and_specs, concat(seg_group)) seg_group = {} end end elseif level > 0 then insert(seg_group, seg) else insert(text_and_specs, seg) end end if level > 0 then error("Unmatched " .. open .. " sign: '" .. segment_run .. "'") end return text_and_specs end ]=] --[==[ Like parse_balanced_segment_run() but accepts multiple sets of delimiters. For example, {parse_multi_delimiter_balanced_segment_run("foo[bar(baz[bat])], quux<glorp>", {{"[", "]"}, {"(", ")"}, {"<", ">"}}) = {"foo", "[bar(baz[bat])]", ", quux", "<glorp>", ""}}. Each element in the list of delimiter pairs is a string specifying an equivalence class of possible delimiter characters. You can use this, for example, to allow either "[" or "&amp;#91;" to be treated equivalently, with either one closed by either "]" or "&amp;#93;". To do this, first replace "&amp;#91;" and "&amp;#93;" with single Unicode characters such as U+FFF0 and U+FFF1, and then specify a two-character string containing "[" and U+FFF0 as the opening delimiter, and a two-character string containing "]" and U+FFF1 as the corresponding closing delimiter. If `no_error_on_unmatched` is given and an error is found during parsing, a string is returned containing the error message instead of throwing an error. ]==] function export.parse_multi_delimiter_balanced_segment_run(segment_run, delimiter_pairs, no_error_on_unmatched) local escaped_delimiter_pairs = {} local open_to_close_map = {} local open_close_items = {} local open_items = {} for _, open_close in ipairs(delimiter_pairs) do local open, close = open_close[1], open_close[2] open = open:gsub("([%[%]%%%%-])", "%%%1") close = close:gsub("([%[%]%%%%-])", "%%%1") insert(open_close_items, open) insert(open_close_items, close) insert(open_items, open) open = "[" .. open .. "]" close = "[" .. close .. "]" open_to_close_map[open] = close insert(escaped_delimiter_pairs, {open, close}) end local open_close_pattern = "([" .. concat(open_close_items) .. "])" local open_pattern = "([" .. concat(open_items) .. "])" local break_on_open_close = split(segment_run, open_close_pattern) local text_and_specs = {} local level = 0 local seg_group = {} local open_at_level_zero for i, seg in ipairs(break_on_open_close) do if i % 2 == 0 then insert(seg_group, seg) if level == 0 then if not umatch(seg, open_pattern) then local errmsg = "Unmatched close sign " .. seg .. ": '" .. segment_run .. "'" if no_error_on_unmatched then return errmsg else error(errmsg) end end assert(open_at_level_zero == nil) for _, open_close in ipairs(escaped_delimiter_pairs) do local open = open_close[1] if umatch(seg, open) then open_at_level_zero = open break end end if open_at_level_zero == nil then error(("Internal error: Segment %s didn't match any open regex"):format(seg)) end level = level + 1 elseif umatch(seg, open_at_level_zero) then level = level + 1 elseif umatch(seg, open_to_close_map[open_at_level_zero]) then level = level - 1 assert(level >= 0) if level == 0 then insert(text_and_specs, concat(seg_group)) seg_group = {} open_at_level_zero = nil end end elseif level > 0 then insert(seg_group, seg) else insert(text_and_specs, seg) end end if level > 0 then local errmsg = "Unmatched open sign " .. open_at_level_zero .. ": '" .. segment_run .. "'" if no_error_on_unmatched then return errmsg else error(errmsg) end end return text_and_specs end --[==[ Check whether a term contains top-level HTML. We want to distinguish inline modifiers from HTML. We assume an inline modifier is either a boolean modifier like `<bor>` or a prefix modifier like `<tr:Miryem>`. All other things inside of angle brackets, e.g. `<nowiki><span class="foo"></nowiki>`, `<nowiki></span></nowiki>`, `<nowiki><br/></nowiki>`, etc., should be flagged as HTML (typically caused by wrapping an argument in {{tl|m|...}}, {{tl|af|...}} or similar, but sometimes specified directly, e.g. `<nowiki><sup>6</sup></nowiki>`). By default, we assume the tag in an inline modifier contains either letters, numbers, hyphens or underscore (but not spaces), and must either stand alone or be followed by a colon, leading to a default HTML-checking pattern of {"<[%w_%-]*[^%w_%-:>]"}. But this can be modified; e.g. [[Module:tl-pronunciation]] allows modifiers of the form `<<var>pos</var>^<var>defn</var>>` or `<<var>pos</var>,<var>pos</var>,<var>pos</var>^<var>defn</var>>`, and would need to use its own HTML pattern. It's important we restrict the check for HTML to top-level to allow for generated HTML inside of e.g. qualifier tags, such as `<nowiki>foo<q:similar to {{m|fr|bar}}></nowiki>`. ]==] function export.term_contains_top_level_html(term, html_pattern) html_pattern = html_pattern or "<[%w_%-]*[^%w_%-:>]" -- If no HTML anywhere, the answer is no. if not term:find(html_pattern) then return false end -- Otherwise, we have to call parse_balanced_segment_run() and check alternate runs at top level. local runs = export.parse_balanced_segment_run(term, "<", ">") for i = 2, #runs, 2 do if runs[i]:find("^" .. html_pattern) then return true end end return false end --[==[ Check whether a term appears to have already been passed through `full_link()`. Passing it again will mangle it in various ways; at best it will have unnecessary lang/script wrapping, which might do nothing but might result in overly large fonts or other issues. We also check for uses of {{tl|ja-r/args}}, {{tl|ryu-r/args}} or {{tl|ko-l/args}}, which will be manged by `full_link()`. If this check succeeds, use the text raw instead of passing through `full_link()`. ]==] function export.term_already_linked(term) return term:find("<span") or term:find("{{ja%-r|") or term:find("{{ryu%-r|") or term:find("{{ko%-l|") end --[==[ Split a list of alternating textual runs of the format returned by `parse_balanced_segment_run` on `splitchar`. This only splits the odd-numbered textual runs (the portions between the balanced open/close characters). The return value is a list of lists, where each list contains an odd number of elements, where the even-numbered elements of the sublists are the original balanced textual run portions. For example, if we do {parse_balanced_segment_run("foo<M.proper noun> bar<F>", "<", ">") = {"foo", "<M.proper noun>", " bar", "<F>", ""}} then {split_alternating_runs({"foo", "<M.proper noun>", " bar", "<F>", ""}, " ") = {{"foo", "<M.proper noun>", ""}, {"bar", "<F>", ""}}} Note that we did not touch the text "<M.proper noun>" even though it contains a space in it, because it is an even-numbered element of the input list. This is intentional and allows for embedded separators inside of brackets/parens/etc. Note also that the inner lists in the return value are of the same form as the input list (i.e. they consist of alternating textual runs where the even-numbered segments are balanced runs), and can in turn be passed to split_alternating_runs(). If `preserve_splitchar` is passed in, the split character is included in the output, as follows: {split_alternating_runs({"foo", "<M.proper noun>", " bar", "<F>", ""}, " ", true) = {{"foo", "<M.proper noun>", ""}, {" "}, {"bar", "<F>", ""}}} Consider what happens if the original string has multiple spaces between brackets, and multiple sets of brackets without spaces between them. {parse_balanced_segment_run("foo[dated][low colloquial] baz-bat quux xyzzy[archaic]", "[", "]") = {"foo", "[dated]", "", "[low colloquial]", " baz-bat quux xyzzy", "[archaic]", ""}} then {split_alternating_runs({"foo", "[dated]", "", "[low colloquial]", " baz-bat quux xyzzy", "[archaic]", ""}, "[ %-]") = {{"foo", "[dated]", "", "[low colloquial]", ""}, {"baz"}, {"bat"}, {"quux"}, {"xyzzy", "[archaic]", ""}}} If `preserve_splitchar` is passed in, the split character is included in the output, as follows: {split_alternating_runs({"foo", "[dated]", "", "[low colloquial]", " baz bat quux xyzzy", "[archaic]", ""}, "[ %-]", true) = {{"foo", "[dated]", "", "[low colloquial]", ""}, {" "}, {"baz"}, {"-"}, {"bat"}, {" "}, {"quux"}, {" "}, {"xyzzy", "[archaic]", ""}}} As can be seen, the even-numbered elements in the outer list are one-element lists consisting of the separator text. ]==] function export.split_alternating_runs(segment_runs, splitchar, preserve_splitchar) local grouped_runs = {} local run = {} for i, seg in ipairs(segment_runs) do if i % 2 == 0 then insert(run, seg) else local parts = split(seg, preserve_splitchar and "(" .. splitchar .. ")" or splitchar) insert(run, parts[1]) for j=2,#parts do insert(grouped_runs, run) run = {parts[j]} end end end if #run > 0 then insert(grouped_runs, run) end return grouped_runs end --[==[ After calling `parse_multi_delimiter_balanced_segment_run()`, rejoin delimiter-bounded textual runs (i.e. textual runs surrounded by certain matched delimiters) with the runs on either side. This can be used when some of the matched delimiters are specified only in order to ensure that delimiters inside of other delimiters aren't parsed. As an example, [[Module:object usage]] calls {m_parse_utilities.parse_multi_delimiter_balanced_segment_run(object, {{"[", "]"}, {"(", ")"}, {"<", ">"}})} but the actual syntax of {{tl|+obj}} only uses parens and angle brackets as delimiters. Square brackets are included so that internal links are treated as units (i.e. parens and angle brackets occurring inside of them aren't parsed), but beyond that we don't treat square brackets as delimiters, so we want to rejoin square-bracket-delimited textual runs with adjacent runs before further parsing. There are two primary workflows when using this function: # If you only care about balanced delimiters occurring inside of other balanced delimiters (e.g. in the above example with [[Module:object usage]], you can call `rejoin_delimited_runs()` directly after `parse_multi_delimiter_balanced_segment_run()`. # However, if you care about single delimiters such as commas and slashes occurring inside of balanced delimiters (e.g. if you allow multiple comma-separated terms, e.g. of which can have associated inline modifiers, and you don't want commas inside of internal links to be treated as delimiters), you need to call `rejoin_delimited_runs()` ''after'' calling `split_alternating_runs()`. This is used, for example, in `parse_inline_modifiers()` for exactly this reason, when a `splitchar` is provided. `data` is an object of properties. Currently there are two: `runs` (the output of calling `parse_multi_delimiter_balanced_segment_run()`, i.e. a list of textual runs, where even-numbered elements begin and end with a matched delimiter and odd-numbered elements are surrounding text) and `delimiter_pattern` (a Lua pattern matching delimited textual runs that we want to rejoin with the surrounding text). `delimiter_pattern` should normally be anchored at the beginning; e.g. {"^%["} would be the correct pattern to use when rejoining square-bracket-delimited textual runs, as described above. ]==] function export.rejoin_delimited_runs(data) local joined_runs = {} local i = 1 while i <= #data.runs do local run = data.runs[i] if i % 2 == 0 and run:find(data.delimiter_pattern) then joined_runs[#joined_runs] = joined_runs[#joined_runs] .. run .. data.runs[i + 1] i = i + 2 else insert(joined_runs, run) i = i + 1 end end return joined_runs end function export.strip_spaces(text) return (ugsub(text, "^%s*(.-)%s*$", "%1")) end --[==[ Apply an arbitrary function `frob` to the "raw-text" segments in a split run set (the output of split_alternating_runs()). We leave alone stuff within balanced delimiters (footnotes, inflection specs and the like), as well as splitchars themselves if present. `preserve_splitchar` indicates whether splitchars are present in the split run set. `frob` is a function of one argument (the string to frob) and should return one argument (the frobbed string). We operate by only frobbing odd-numbered segments, and only in odd-numbered runs if preserve_splitchar is given. ]==] function export.frob_raw_text_alternating_runs(split_run_set, frob, preserve_splitchar) for i, run in ipairs(split_run_set) do if not preserve_splitchar or i % 2 == 1 then for j, segment in ipairs(run) do if j % 2 == 1 then run[j] = frob(segment) end end end end end --[==[ Like split_alternating_runs() but applies an arbitrary function `frob` to "raw-text" segments in the result (i.e. not stuff within balanced delimiters such as footnotes and inflection specs, and not splitchars if present). `frob` is a function of one argument (the string to frob) and should return one argument (the frobbed string). ]==] function export.split_alternating_runs_and_frob_raw_text(run, splitchar, frob, preserve_splitchar) local split_runs = export.split_alternating_runs(run, splitchar, preserve_splitchar) export.frob_raw_text_alternating_runs(split_runs, frob, preserve_splitchar) return split_runs end --[==[ FIXME: Older entry point. Call `split_alternating_runs_and_frob_raw_text()` in [[Module:parse utilities]] directly. Like `split_alternating_runs()` but strips spaces from both ends of the odd-numbered elements (only in odd-numbered runs if `preserve_splitchar` is given). Effectively we leave alone the footnotes and splitchars themselves, but otherwise strip extraneous spaces. Spaces in the middle of an element are also left alone. ]==] function export.split_alternating_runs_and_strip_spaces(segment_runs, splitchar, preserve_splitchar) return export.split_alternating_runs_and_frob_raw_text(segment_runs, splitchar, export.strip_spaces, preserve_splitchar) end --[==[ Split the non-modifier parts of an alternating run (after parse_balanced_segment_run() is called) on a Lua pattern, but not on certain sequences involving characters in that pattern (e.g. comma+whitespace). `splitchar` is the pattern to split on; `preserve_splitchar` indicates whether to preserve the delimiter and is the same as in split_alternating_runs(). `escape_fun` is called beforehand on each run of raw text and should return two values: the escaped run and whether unescaping is needed. If any call to `escape_fun` indicates that unescaping is needed, `unescape_fun` will be called on each run of raw text after splitting on `splitchar`. The return value of this function is as in split_alternating_runs(). ]==] function export.split_alternating_runs_escaping(run, splitchar, preserve_splitchar, escape_fun, unescape_fun) -- First replace comma with a temporary character in comma+whitespace sequences. local need_unescape = false for i in ipairs(run) do if i % 2 == 1 and escape_fun then local this_need_unescape run[i], this_need_unescape = escape_fun(run[i]) need_unescape = need_unescape or this_need_unescape end end if need_unescape then return export.split_alternating_runs_and_frob_raw_text(run, splitchar, unescape_fun, preserve_splitchar) else return export.split_alternating_runs(run, splitchar, preserve_splitchar) end end --[==[ Replace comma with a temporary char in comma + whitespace. ]==] function export.escape_comma_whitespace(run, tempcomma) tempcomma = tempcomma or u(0xFFF0) local escaped = false if run:find("\\,") then -- FIXME: we should probably convert literal \\ to \ to allow people to put a backslash before a comma that -- should be passed through; but maybe it's enough to use an HTML escape for the comma or backslash. run = (run:gsub("\\,", tempcomma)) -- discard backslash before comma, doing its duty to protect the comma escaped = true end if run:find(",%s") then run = (run:gsub(",(%s)", tempcomma .. "%1")) escaped = true end return run, escaped end --[==[ Undo the replacement of comma with a temporary char. ]==] function export.unescape_comma_whitespace(run, tempcomma) tempcomma = tempcomma or u(0xFFF0) return (run:gsub(tempcomma, ",")) end --[==[ Split the non-modifier parts of an alternating run (after parse_balanced_segment_run() is called) on comma, but not on comma+whitespace. See `split_on_comma()` above for more information and the meaning of `tempcomma`. ]==] function export.split_alternating_runs_on_comma(run, tempcomma) tempcomma = tempcomma or u(0xFFF0) -- Replace comma with a temporary char in comma + whitespace. local function escape_comma_whitespace(seg) return export.escape_comma_whitespace(seg, tempcomma) end -- Undo replacement of comma with a temporary char in comma + whitespace. local function unescape_comma_whitespace(seg) return export.unescape_comma_whitespace(seg, tempcomma) end return export.split_alternating_runs_escaping(run, ",", false, escape_comma_whitespace, unescape_comma_whitespace) end --[==[ Split text on a Lua pattern, but not on certain sequences involving characters in that pattern (e.g. comma+whitespace). `splitchar` is the pattern to split on; `preserve_splitchar` indicates whether to preserve the delimiter between split segments. `escape_fun` is called beforehand on the text and should return two values: the escaped run and whether unescaping is needed. If the call to `escape_fun` indicates that unescaping is needed, `unescape_fun` will be called on each run of text after splitting on `splitchar`. The return value of this a list of runs, interspersed with delimiters if `preserve_splitchar` is specified. ]==] function export.split_escaping(text, splitchar, preserve_splitchar, escape_fun, unescape_fun) if not umatch(text, splitchar) then return {text} end -- If there are square or angle brackets, we don't want to split on delimiters inside of them. To effect this, we -- use parse_multi_delimiter_balanced_segment_run() to parse balanced brackets, then do delimiter splitting on the -- non-bracketed portions of text using split_alternating_runs_escaping(), and concatenate back to a list of -- strings. When calling parse_multi_delimiter_balanced_segment_run(), we make sure not to throw an error on -- unbalanced brackets; in that case, we fall through to the code below that handles the case without brackets. if text:find("[%[<]") then local runs = export.parse_multi_delimiter_balanced_segment_run(text, {{"[", "]"}, {"<", ">"}}, "no error on unmatched") if type(runs) ~= "string" then local split_runs = export.split_alternating_runs_escaping(runs, splitchar, preserve_splitchar, escape_fun, unescape_fun) for i = 1, #split_runs do split_runs[i] = concat(split_runs[i]) end return split_runs end end -- First escape sequences we don't want to count for splitting. local need_unescape if escape_fun then text, need_unescape = escape_fun(text) end local parts = split(text, preserve_splitchar and "(" .. splitchar .. ")" or splitchar) if need_unescape then for i = 1, #parts, (preserve_splitchar and 2 or 1) do parts[i] = unescape_fun(parts[i]) end end return parts end --[==[ Split text on comma, but not on comma+whitespace. This is similar to `mw.text.split(text, ",")` but will not split on commas directly followed by whitespace, to handle embedded commas in terms (which are almost always followed by a space). `tempcomma` is the Unicode character to temporarily use when doing the splitting; normally U+FFF0, but you can specify a different character if you use U+FFF0 for some internal purpose. ]==] function export.split_on_comma(text, tempcomma) -- Don't do anything if no comma. Note that split_escaping() has a similar check at the beginning, so if there's a -- comma we effectively do this check twice, but this is worth it to optimize for the common no-comma case. if not text:find(",") then return {text} end tempcomma = tempcomma or u(0xFFF0) -- Replace comma with a temporary char in comma + whitespace. local function escape_comma_whitespace(run) return export.escape_comma_whitespace(run, tempcomma) end -- Undo replacement of comma with a temporary char in comma + whitespace. local function unescape_comma_whitespace(run) return export.unescape_comma_whitespace(run, tempcomma) end return export.split_escaping(text, ",", false, escape_comma_whitespace, unescape_comma_whitespace) end --[==[ Ensure that Wikicode (template calls, bracketed links, HTML, bold/italics, etc.) displays literally in error messages by inserting a Unicode word-joiner symbol after all characters that may trigger Wikicode interpretation. Replacing with equivalent HTML escapes doesn't work because they are displayed literally. I could not get this to work using <nowiki>...</nowiki> (those tags display literally), using using {{#tag:nowiki|...}} (same thing) or using mw.getCurrentFrame():extensionTag("nowiki", ...) (everything gets converted to a strip marker `UNIQ--nowiki-00000000-QINU` or similar). FIXME: This is a massive hack; there must be a better way. ]==] function export.escape_wikicode(term) term = term:gsub("([%[<'{])", "%1" .. u(0x2060)) return term end function export.make_parse_err(arg_gloss) return function(msg, stack_frames_to_ignore) error(export.escape_wikicode(("%s: %s"):format(msg, arg_gloss)), stack_frames_to_ignore) end end -- Parse a term that may include a link '[[LINK]]' or a two-part link '[[LINK|DISPLAY]]'. FIXME: Doesn't currently -- handle embedded links like '[[FOO]] [[BAR]]' or [[FOO|BAR]] [[BAZ]]' or '[[FOO]]s'; if they are detected, it returns -- the term unchanged and `nil` for the display form. local function parse_bracketed_term(term, parse_err) local inside = term:match("^%[%[(.*)%]%]$") if inside then if inside:find("%[%[") or inside:find("%]%]") then -- embedded links, e.g. '[[FOO]] [[BAR]]'; FIXME: we should process them properly return term, nil end local parts = split(inside, "|") if #parts > 2 then parse_err("Saw more than two parts inside a bracketed link") end return parts[1], parts[2] end return term, nil end --[==[ Parse a term that may have a language code (or possibly multiple plus-separated language codes, if `data.allow_multiple` is given) preceding it (e.g. {la:minūtia} or {grc:[[σκῶρ|σκατός]]} or {nan-hbl+hak:[[毋]][[知]]}). Return five arguments: # the original prefixed term; in the case of a Wikipedia or Wikisource prefix followed by a two-part link, it is a two-part link with the Wikipedia/Wikisource prefix moved inside the link; in the case of a Wikipedia or Wikisource prefix followed by a redundant one-part link, the brackets are removed; # the language object corresponding to the language code (possibly a family object if `data.allow_family` is given), or a list of such objects if `data.allow_multiple` is given; # the link if the unprefixed term is of the form <code>[[<var>link</var>|<var>display</var>]]</code> or of the form <code>[[<var>link</var>]]</code>, otherwise the full unprefixed term; # the display part if the term is of the form <code>[[<var>link</var>|<var>display</var>]]</code> or has a Wikipedia or Wikisource prefix (in which case the part minus the prefix and any following language code will be returned, with redundant brackets stripped), else {nil}; # {true} if the term has a Wikipedia/Wikisource prefix, else {false}. Etymology-only languages are always allowed. This function also correctly handles Wikipedia prefixes (e.g. {w:Abatemarco} or {w:it:Colle Val d'Elsa} or {lw:ru:Филарет}) and Wikisource prefixes (e.g. {s:Twelve O'Clock} or {s:[[Walden/Chapter XVIII|Walden]]} or {s:fr:Perceval ou le conte du Graal} or {s:ro:[[Domnul Vucea|Mr. Vucea]]} or {ls:ko:이상적 부인} or {ls:ko:[[조선 독립의 서#一. 槪論|조선 독립의 서]]}) and converts them into two-part links, with the display form not including the Wikipedia or Wikisource prefix unless it was explicitly specified using a two-part link as in {lw:ru:[[Филарет (Дроздов)|Митрополи́т Филаре́т]]} or {ls:ko:[[조선 독립의 서#一. 槪論|조선 독립의 서]]}. The difference between {w:} ("Wikipedia") and {lw:} ("Wikipedia link") is that the latter requires a language code and returns the corresponding language object; same for the difference between {s:} ("Wikisource") and {ls:} ("Wikisource link"). NOTE: Embedded links are not correctly handled currently. If an embedded link is detected, the whole term is returned as the link part (third argument), and the display part is nil. If you construct your own link from the link and display parts, you must check for this. The calling convention is to pass in a single argument `data` containing the following fields: * `term`: The term to parse. * `parse_err`: An optional function of one or two arguments to display an error. (The second argument to the function is the number of stack frames to ignore when calling error(); if you declare your error function with only one argument, things will still work fine.) * `paramname`: If `parse_err` is omitted, this should be a string naming a parameter to display in the error message, along with the term in question, and will be used to generate a `parse_err` function using `make_parse_err()`. (If `paramname` is omitted, just the term itself appears in the error message.) * `allow_multiple`: Allow multiple plus-separated language codes, e.g. {nan-hbl+hak:[[毋]][[知]]}. See above. * `allow_family`: Allow family objects to appear in place of language codes. * `allow_bad`: Don't throw an error on invalid language code prefixes; instead, include the prefix and colon as part of the term. Note that if a prefix doesn't look like a language code (e.g. if it's a number), the code won't even try to parse it as a language code, regardless of the `allow_bad` setting, but will always include it in the term. * `lang_cache`: A table mapping language codes to language objects, where invalid language codes are indicated by the value `false`. If this field is specified, the cache will be consulted before calling `getByCode()` in [[Module:languages]], and the result cached. If not specified, no cache will be used. ]==] function export.parse_term_with_lang(data) local term = data.term local parse_err = data.parse_err or data.paramname and export.make_parse_err(("%s=%s"):format(data.paramname, term)) or export.make_parse_err(term) -- Parse off an initial language code (e.g. 'la:minūtia' or 'grc:[[σκῶρ|σκατός]]'). First check for Wikipedia -- prefixes ('w:Abatemarco' or 'w:it:Colle Val d'Elsa' or 'lw:zh:邹衡') and Wikisource prefixes -- ('s:ro:[[Domnul Vucea|Mr. Vucea]]' or 'ls:ko:이상적 부인'). Wikipedia/Wikisource language codes follow a similar -- format to Wiktionary language codes (see below). Here and below we don't parse if there's a space after the -- colon (happens e.g. if the user uses {{desc|...}} inside of {{col}}, grrr ...). local termlang, foreign_wiki, actual_term = term:match("^(l?[ws]):([a-z][a-z][a-z-]*):([^ ].*)$") if not termlang then termlang, actual_term = term:match("^([ws]):([^ ].*)$") end if termlang then local wiki_links = termlang:find("^l") local base_wiki_prefix = termlang:find("w$") and "w:" or "s:" local wiki_prefix = base_wiki_prefix .. (foreign_wiki and foreign_wiki .. ":" or "") local link, display = parse_bracketed_term(actual_term, parse_err) if link:find("%[%[") or display and display:find("%[%[") then -- FIXME, this should be handlable with the right parsing code parse_err("Cannot have embedded brackets following a Wikipedia (w:... or lw:...) link; expand the term to a fully bracketed term w:[[LINK|DISPLAY]] or similar") end local lang = wiki_links and get_lang(foreign_wiki, parse_err, "allow etym") or nil local prefixed_link = wiki_prefix .. link if display then return ("[[%s|%s]]"):format(prefixed_link, display), lang, prefixed_link, display, true else -- Return the link minus any language codes as the fourth term (display form). Previously we returned `actual_term` -- but this causes problems with redundant Wikipedia links of the form `w:[[Dragon Ball Z]]`. Don't generate a -- two-part link so you can specify a display form in 3=. Note that the fourth and fifth params are currently only -- used in [[Module:quote]]. return prefixed_link, lang, prefixed_link, link, true end end -- Wiktionary language codes are in one of the following formats, where 'x' is a lowercase letter and 'X' an -- uppercase letter: -- xx -- xxx -- xxx-xxx -- xxx-xxx-xxx (esp. for protolanguages) -- xx-xxx (for etymology-only languages) -- xx-xxx-xxx (maybe? for etymology-only languages) -- xx-XX (for etymology-only languages, where XX is a country code, e.g. en-US) -- xxx-XX (for etymology-only languages, where XX is a country code) -- xx-xxx-XX (for etymology-only languages, where XX is a country code) -- xxx-xxx-XX (for etymology-only langauges, where XX is a country code, e.g. nan-hbl-PH) -- Things like xxx-x+ (e.g. cmn-pinyin, cmn-tongyong) -- VL., LL., etc. -- -- We check the for nonstandard Latin etymology language codes separately, and otherwise make only the following -- assumptions: -- (1) There are one to three hyphen-separated components. -- (2) The last component can consist of two uppercase ASCII letters; otherwise, all components contain only -- lowercase ASCII letters. -- (3) Each component must have at least two letters. -- (4) The first component must have two or three letters. local function is_possible_lang_code(code) -- Special hack for Latin variants, which can have nonstandard etym codes, e.g. VL., LL. if code:find("^[A-Z]L%.$") then return true end return code:find("^([a-z][a-z][a-z]?)$") or code:find("^[a-z][a-z][a-z]?%-[A-Z][A-Z]$") or code:find("^[a-z][a-z][a-z]?%-[a-z][a-z]+$") or code:find("^[a-z][a-z][a-z]?%-[a-z][a-z]+%-[A-Z][A-Z]$") or code:find("^[a-z][a-z][a-z]?%-[a-z][a-z]+%-[a-z][a-z]+$") end local function get_by_code(code, allow_bad) local lang if data.lang_cache then lang = data.lang_cache[code] end if lang == nil then lang = get_lang(code, not allow_bad and parse_err or nil, "allow etym", data.allow_family) if data.lang_cache then data.lang_cache[code] = lang or false end end return lang or nil end if data.allow_multiple then local termlang_spec termlang_spec, actual_term = term:match("^([a-zA-Z.,+-]+):([^ ].*)$") if termlang_spec then termlang = split(termlang_spec, "[,+]") local all_possible_code = true for _, code in ipairs(termlang) do if not is_possible_lang_code(code) then all_possible_code = false break end end if all_possible_code then local saw_nil = false for i, code in ipairs(termlang) do termlang[i] = get_by_code(code, data.allow_bad) if not termlang[i] then saw_nil = true end end if saw_nil then termlang = nil else term = actual_term end else termlang = nil end end else termlang, actual_term = term:match("^([a-zA-Z.-]+):([^ ].*)$") if termlang then if is_possible_lang_code(termlang) then termlang = get_by_code(termlang, data.allow_bad) if termlang then term = actual_term end else termlang = nil end end end local link, display = parse_bracketed_term(term, parse_err) return term, termlang, link, display, false end --[==[ Maybe parse any language prefix off of a given term and store the term and language(s) into a new or existing object. This function is useful for implementing a `generate_obj` handler of `parse_inline_modifiers` that can support terms with prefixed language(s). This wraps `parse_term_with_lang` and has the same handling of language prefixes as that function. The calling convention is to pass in a single argument `data` containing the following fields (NOTE: you must set `parse_lang_prefix` to get language-prefix-parsing behavior): * `term`: The term to parse. This is the only required parameter. * `parse_lang_prefix`: This must be specified in order for language prefixes to be recognized and parsed off. * `termobj`: The existing object to store results into. If unspecified, a new object is created. * `term_dest`: The field in `termobj` into which the term itself (minus any language prefix) is stored. If unspecified, defaults to {"term"}. * `parse_err`: An optional function of one or two arguments to display an error. (The second argument to the function is the number of stack frames to ignore when calling error(); if you declare your error function with only one argument, things will still work fine.) * `paramname`: If `parse_err` is omitted, this should be a string naming a parameter to display in the error message, along with the term in question, and will be used to generate a `parse_err` function using `make_parse_err()`. (If `paramname` is omitted, just the term itself appears in the error message.) * `allow_multiple_lang_prefixes`: Allow multiple plus-separated language codes, e.g. {nan-hbl+hak:[[毋]][[知]]}. See `parse_term_with_lang` for more information. * `allow_bad_lang_prefix`: Don't throw an error on invalid language code prefixes; instead, include the prefix and colon as part of the term. Note that if a prefix doesn't look like a language code (e.g. if it's a number), the code won't even try to parse it as a language code, regardless of this setting, but will always include it in the term. * `allow_family_as_lang_prefix`: Allow family objects to appear in place of language codes. * `lang_cache`: A table mapping language codes to language objects, where invalid language codes are indicated by the value `false`. If this field is specified, the cache will be consulted before calling `getByCode()` in [[Module:languages]], and the result cached. If not specified, no cache will be used. The return value is the object in `data.termobj` (if non-{nil}) or a newly-created object (otherwise), with the parsed-off term stored in the field named by the `data.term_dest` property (normally `.term`). If `data.allow_multiple_lang_prefixes` was not given and a language prefix was parsed off, the corresponding language object is stored into both `lang` and `termlang` (the storage into `termlang` is so that the presence of a language prefix can specifically be determined in the event that `lang` is already set in an existing termobj). If `data.allow_multiple_lang_prefixes` was given and one or more language prefixes were parsed off, the first such language object is stored into `lang`, and the list of all language objects stored into `termlangs. ]==] function export.generate_obj_maybe_parsing_lang_prefix(data) local term = data.term local term_dest = data.term_dest or "term" local termobj = data.termobj or {} if data.parse_lang_prefix and term:find(":", nil, true) then local actual_term, termlangs = export.parse_term_with_lang { term = term, parse_err = data.parse_err, paramname = data.paramname, allow_bad = data.allow_bad_lang_prefix, allow_multiple = data.allow_multiple_lang_prefixes, allow_family = data.allow_family_as_lang_prefix, lang_cache = data.lang_cache, } termobj[term_dest] = actual_term ~= "" and actual_term or nil if termlangs then -- If we couldn't parse a language code, don't overwrite an existing setting in `lang` -- that may have originated from a separate |langN= param. if data.allow_multiple_lang_prefixes then termobj.termlangs = termlangs termobj.lang = termlangs and termlangs[1] or nil else termobj.termlang = termlangs termobj.lang = termlangs end end else termobj[term_dest] = term ~= "" and term or nil end return termobj end --[==[ Parse a term that may have inline modifiers attached (e.g. {rifiuti<q:plural-only>} or {rinfusa<t:bulk cargo><lit:resupplying><qq:more common in the plural {{m|it|rinfuse}}>}). * `arg` is the term to parse. * `props` is an object holding further properties controlling how to parse the term (only `param_mods` and `generate_obj` are required): ** `paramname` is the name of the parameter where `arg` comes from, or nil if this isn't available (it is used only in error messages). ** `param_mods` is a table describing the allowed inline modifiers (see below). ** `generate_obj` is a function of one or two arguments that should parse the argument minus the inline modifiers and return a corresponding parsed object (into which the inline modifiers will be rewritten). If declared with one argument, that will be the raw value to parse; if declared with two arguments, the second argument will be the `parse_err` function (see below). ** `parse_err` is an optional function of one argument (an error message) and should display the error message, along with any desired contextual text (e.g. the argument name and value that triggered the error). If omitted, a default function will be generated which displays the error along with the original value of `arg` (passed through {escape_wikicode()} above to ensure that Wikicode (such as links) is displayed literally). ** `splitchar` is a Lua pattern. If specified, `arg` can consist of multiple delimiter-separated terms, each of which may be followed by inline modifiers, and the return value will be a list of parsed objects instead of a single object. Note that splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). The algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers, which do not trigger splitting (whether or not contained within protected sequences). ** `outer_container`, if specified, is used when multiple delimiter-separated terms are possible, and is the object into which the list of per-term objects is stored (into the `terms` field) and into which any modifiers that are given the `overall` property (see below) will be stored. If given, this value will be returned as the value of {parse_inline_modifiers()}. If `outer_container` is not given, {parse_inline_modifiers()} will return the list of per-term objects directly, and no modifier may have an `overall` property. ** `preserve_splitchar`, if specified, causes the actual delimiter matched by `splitchar` to be returned in the parsed object describing the element that comes after the delimiter. The delimiter is stored in a key whose name is controlled by `delimiter_key`, which defaults to "delimiter". ** `delimiter_key` controls the key into which the actual delimiter is written when `preserve_splitchar` is used. See above. ** `escape_fun` and `unescape_fun` are as in split_escaping() and split_alternating_runs_escaping() above and control the protected sequences that won't be split. By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that comma+whitespace sequences won't be split. Set to `false` to disable escaping/unescaping. ** `pre_normalize_modifiers`, if specified, is a function of one argument, which can be used to "normalize" modifiers prior to further parsing. This is used, for example, in [[Module:tl-pronunciation]] to convert modifiers of the form `<noun^expectation; hope>` to `<t:noun^expectation; hope>`, so they can be processed as standard modifiers. It is also used in [[Module:ar-verb]] to convert footnotes of the form `[rare]` to `<footnote:[rare]>`, to allow for mixing bracketed footnotes and inline modifiers when overriding verbal nouns and such. It could similarly be used to handle boolean modifiers like `<slb>` in {{tl|desc}} and convert them to a standard form `<slb:1>`. It runs just before parsing out the modifier prefix and value, and is passed an object containing fields `modtext` (the un-normalized modifier text, including surrounding angle brackets, or in some cases, text surrounded by other delimiters such as square brackets, if `parse_inline_modifiers_from_segments()` is being called and the caller did their own parsing of balanced segment runs) and `parse_err` (the passed-in or autogenerated function to signal an error during parsing; a function of one argument, a message, which throws an error displaying that message). It should return a single value, the normalized value of `modtext`, including surrounding angle brackets. `param_mods` is a table describing allowed modifiers. The keys of the table are modifier prefixes and the values are tables describing how to parse and store the associated modifier values. Here is a typical example, for an item that takes the standard modifiers associated with `full_link()` in [[Module:links]], as well as left and right qualifiers and labels: { local param_mods = { alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = {}, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "g". `sublist = true` automatically splits on comma (optionally -- with surrounding whitespace). item_dest = "genders", sublist = true, }, pos = {}, lit = {}, id = {}, sc = { -- Automatically parse as a script code and convert to a script object. type = "script", }, -- Qualifiers and labels q = { type = "qualifier", }, qq = { type = "qualifier", }, l = { type = "labels", }, ll = { type = "labels", }, } } In the table values: * `item_dest` specifies the destination key to store the object into (if not the same as the modifier key itself). * `type`, `set`, `sublist` and `convert` have the same meaning as in [[Module:parameters]] and are used for converting the object from the string form given by the user into the form needed for further processing. Note that `type` makes use of additional properties that may be specified. Specifically, if {type = "language"}, the properties `family` and `method` are also examined, and if {type = "family"} or {type = "script"}, the property `method` is examined. * `store` describes how to store the converted modifier value into the parsed object. If omitted, the converted value is simply written into the parsed object under the appropriate key; but an error is generated if the key already has a value. (This means that multiple occurrences of a given modifier are allowed if `store` is given, but not otherwise.) `store` can be one of the following: ** {"insert"}: the converted value is appended to the key's value using {insert()}; if the key has no value, it is first converted to an empty list; ** {"insertIfNot"}: is similar but appends the value using {insertIfNot()} in [[Module:table]]; ** {"insert-flattened"}, the converted value is assumed to be a list and the objects are appended one-by-one into the key's existing value using {insert()}; ** {"insertIfNot-flattened"} is similar but appends using {insertIfNot()} in [[Module:table]]; (WARNING: When using {"insert-flattened"} and {"insertIfNot-flattened"}, if there is no existing value for the key, the converted value is just stored directly. This means that future appends will side-effect that value, so make sure that the return value of the conversion function for this key generates a fresh list each time.) ** a function of one argument, an object with the following properties: *** `dest`: the object to write the value into; *** `key`: the field where the value should be written; *** `converted`: the (converted) value to write; *** `raw_val`: the raw, user-specified value (a string); *** `parse_err`: a function of one argument (an error string), which signals an error, and includes extra context in the message about the modifier in question, the angle-bracket spec that includes the modifier in it, the overall value, and (if `paramname` was given) the parameter holding the overall value. * `overall` only applies if `splitchar` is given. In this case, the modifier applies to the entire argument rather than to an individual term in the argument, and must occur after the last item separated by `splitchar`, instead of being allowed to occur after any of them. The modifier will be stored into the outer container object, which must exist (i.e. `outer_container` must have been given). The return value of {parse_inline_modifiers()} depends on whether `splitchar` and `outer_container` have been given. If neither is given, the return value is the object returned by `generate_obj`. If `splitchar` but not `outer_container` is given, the return value is a list of per-term objects, each of which is generated by `generate_obj`. If both `splitchar` and `outer_container` are given, the return value is the value of `outer_container` and the per-term objects are stored into the `terms` field of this object. ]==] function export.parse_inline_modifiers(arg, props) local segments local function rejoin_bracket_delimited_runs(segments) return export.rejoin_delimited_runs { runs = segments, delimiter_pattern = "^%[.*%]$", } end local rejoin_square_brackets_after_split = false -- The following is an optimization. If we see a square bracket (normally a double square bracket internal link -- [[...]]), we want to not treat delimiter characters inside (either <...> balanced delimiters or separators such -- as commas) as delimiters. But this requires a more sophisticated and slower algorithm, and most of the time it -- isn't needed because there are no square brackets. So we check for a square bracket and fall back to a simpler -- algorithm otherwise (which, since it involves only a single balanced delimiter, can use the built-in %b() Lua -- pattern syntax, which AFAIK is implemented in C). if arg:find("%[") then segments = export.parse_multi_delimiter_balanced_segment_run(arg, {{"[", "]"}, {"<", ">"}}) if not props.splitchar then segments = rejoin_bracket_delimited_runs(segments) else rejoin_square_brackets_after_split = true end else segments = export.parse_balanced_segment_run(arg, "<", ">") end local function verify_no_overall() for _, mod_props in pairs(props.param_mods) do if mod_props.overall then error("Internal caller error: Can't specify `overall` for a modifier in `param_mods` unless `outer_container` property is given") end end end if not props.splitchar then if props.outer_container then error("Internal caller error: Can't specify `outer_container` property unless `splitchar` is given") end verify_no_overall() return export.parse_inline_modifiers_from_segments { group = segments, group_index = nil, separated_groups = nil, arg = arg, props = props, } else local terms = {} if props.outer_container then props.outer_container.terms = terms else verify_no_overall() end local escape_fun = props.escape_fun if escape_fun == nil then escape_fun = export.escape_comma_whitespace end local unescape_fun = props.unescape_fun if unescape_fun == nil then unescape_fun = export.unescape_comma_whitespace end local separated_groups = export.split_alternating_runs_escaping(segments, props.splitchar, props.preserve_splitchar, escape_fun, unescape_fun) for j = 1, #separated_groups, (props.preserve_splitchar and 2 or 1) do if rejoin_square_brackets_after_split then separated_groups[j] = rejoin_bracket_delimited_runs(separated_groups[j]) end local parsed = export.parse_inline_modifiers_from_segments { group = separated_groups[j], group_index = j, separated_groups = separated_groups, arg = arg, props = props, } if props.preserve_splitchar and j > 1 then parsed[props.delimiter_key or "delimiter"] = separated_groups[j - 1][1] end insert(terms, parsed) end if props.outer_container then return props.outer_container else return terms end end end --[==[ Parse a single term that may have inline modifiers attached. This is a helper function of {parse_inline_modifiers()} but is exported separately in case the caller needs to make their own call to {parse_balanced_segment_run()} (as in [[Module:quote]], which splits on several matched delimiters simultaneously). It takes only a single argument, `data`, which is an object with the following fields: * `group`: A list of segments as output by {parse_balanced_segment_run()} (see the overall comment at the top of [[Module:parse utilities]]), or one of the lists returned by calling {split_alternating_runs()}. * `separated_groups`: The list of groups (each of which is of the form of `group`) describing all the terms in the argument parsed by {parse_inline_modifiers()}, or {nil} if this isn't applicable (i.e. multiple terms aren't allowed in the argument). Currently used only the check the number of groups in the list against `group_index`. * `group_index`: The index into `separated_groups` where `group` can be found, or {nil} if not applicable (see below). * `arg`: The original user-specified argument being parsed; used only for error messages and only when `props.parse_err` is not specified. * `props`: The `props` argument to {parse_inline_modifiers()}. The return value is the object created by `generate_obj`, with properties filled in describing the modifiers of the term in question. Note that `props.outer_container` and the `overall` setting of the `props.param_mods` structure are respected, but `props.splitchar` is ignored because the splitting happens in the caller. Specifically, if there are any modifiers with the `overall` setting, `props.separated_groups` and `props.group_index` must be given so that the function is able to determine if the modifier is indeed attached to the last term, and `props.outer_container` must be given because that is where such modifiers are stored. Otherwise, none of these settings need be given. ]==] function export.parse_inline_modifiers_from_segments(data) local props = data.props local group = data.group local function get_valid_prefixes() local valid_prefixes = {} for param_mod, mod_props in pairs(props.param_mods) do if not mod_props.deprecated then insert(valid_prefixes, param_mod) end end sort(valid_prefixes) return valid_prefixes end local function get_arg_gloss() if props.paramname then return ("%s=%s"):format(props.paramname, data.arg) else return data.arg end end local parse_err = props.parse_err or export.make_parse_err(get_arg_gloss()) local term_obj = props.generate_obj(group[1], parse_err) for k = 2, #group - 1, 2 do if group[k + 1] ~= "" then parse_err("Extraneous text '" .. group[k + 1] .. "' after modifier") end local group_k = group[k] if props.pre_normalize_modifiers then -- FIXME: For some use cases, we might have to pass more information. group_k = props.pre_normalize_modifiers { modtext = group_k, parse_err = parse_err } end local modtext = group_k:match("^<(.*)>$") if not modtext then parse_err("Internal error: Modifier '" .. group_k .. "' isn't surrounded by angle brackets") end local prefix, val = modtext:match("^([a-zA-Z0-9+_-]+):(.*)$") if not prefix then local valid_prefixes = get_valid_prefixes() for i, valid_prefix in ipairs(valid_prefixes) do valid_prefixes[i] = "'" .. valid_prefix .. ":'" end parse_err(("Modifier %s%s lacks a prefix, should begin with one of %s"):format( group_k, group_k ~= group[k] and (" (normalized from %s)"):format(group[k]) or "", list_to_text(valid_prefixes))) end local prefix_parse_err if props.parse_err then prefix_parse_err = function(msg, stack_frames_to_ignore) props.parse_err(("%s: modifier prefix '%s' in %s"):format(msg, prefix, group[k]), stack_frames_to_ignore) end else prefix_parse_err = export.make_parse_err(("modifier prefix '%s' in %s in %s"):format( prefix, group[k], get_arg_gloss())) end if props.param_mods[prefix] then local mod_props = props.param_mods[prefix] if mod_props.replaced_by == false then prefix_parse_err( ("Prefix has been removed and is no longer valid%s%s"):format( mod_props.reason and ", " .. mod_props.reason or "", mod_props.instead and "; instead, " .. mod_props.instead or "") ) elseif mod_props.replaced_by then prefix_parse_err( ("Prefix has been replaced by '%s'%s"):format( mod_props.replaced_by, mod_props.reason and ", " .. mod_props.reason or "") ) end local key = mod_props.item_dest or prefix local dest if mod_props.overall then if not data.separated_groups then prefix_parse_err("Internal error: `data.separated_groups` not given when `overall` is seen") end if not props.outer_container then -- This should have been caught earlier during validation in parse_inline_modifiers(). prefix_parse_err("Internal error: `props.outer_container` not given when `overall` is seen") end if data.group_index ~= #data.separated_groups then prefix_parse_err("Prefix should occur after the last comma-separated term") end dest = props.outer_container else dest = term_obj end local converted = val if mod_props.type or mod_props.set or mod_props.sublist or mod_props.convert then -- WARNING: Here as an optimization we embed some knowledge of convert_val() in [[Module:parameters]], -- specifically that if none of `type`, `set`, `sublist` and `convert` are set, the conversion is an -- identity operation and can be skipped. (convert_val() also makes use of the fields `method` and -- `family`, but only if `type` is set to certain values such as "language", "family" or "script", and -- makes use of the field `required`, but only if `set` is set.) If this becomes problematic, consider -- removing the optimization. converted = convert_val(converted, prefix_parse_err, mod_props) end local store = props.param_mods[prefix].store if not store then if dest[key] then prefix_parse_err("Prefix occurs twice") end dest[key] = converted elseif store == "insert" then if not dest[key] then dest[key] = {converted} else insert(dest[key], converted) end elseif store == "insertIfNot" then if not dest[key] then dest[key] = {converted} else insert_if_not(dest[key], converted) end elseif store == "insert-flattened" then if not dest[key] then dest[key] = converted else for _, obj in ipairs(converted) do insert(dest[key], obj) end end elseif store == "insertIfNot-flattened" then if not dest[key] then dest[key] = converted else for _, obj in ipairs(converted) do insert_if_not(dest[key], obj) end end elseif type(store) == "string" then prefix_parse_err(("Internal caller error: Unrecognized value '%s' for `store` property"):format(store)) elseif not is_callable(store) then prefix_parse_err(("Internal caller error: Unrecognized type for `store` property %s"):format(dump(store))) else store{ dest = dest, key = key, converted = converted, raw = val, parse_err = prefix_parse_err } end else local valid_prefixes = get_valid_prefixes() for i, valid_prefix in ipairs(valid_prefixes) do valid_prefixes[i] = "'" .. valid_prefix .. "'" end prefix_parse_err("Unrecognized prefix, should be one of " .. list_to_text(valid_prefixes)) end end return term_obj end return export 7m7efxuys6lveo94nvwye55dw2c55on Module:etymology/templates/descendant 828 8067 54881 2026-09-27T21:41:39Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local descendants_tree_module = "Module:descendants tree" local etymology_style_css = "Module:etymology/style.css" local labels_module = "Module:labels" local languages_module = "Module:languages" local links_module = "Module:links" local parameter_utilities_module = "Module:parameter utilities" local scripts_module = "Module:scripts" local table_mod...' 54881 Scribunto text/plain local export = {} local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local descendants_tree_module = "Module:descendants tree" local etymology_style_css = "Module:etymology/style.css" local labels_module = "Module:labels" local languages_module = "Module:languages" local links_module = "Module:links" local parameter_utilities_module = "Module:parameter utilities" local scripts_module = "Module:scripts" local table_module = "Module:table" local table_module_list_to_set = "Module:table/listToSet" local template_styles_module = "Module:TemplateStyles" local concat = table.concat local insert = table.insert local list_to_set = require(table_module_list_to_set) local error_on_no_descendants = false local function track(page) return require(debug_track_module)("descendant/" .. page) end local function ine(arg) if arg == "" then return nil else return arg end end local function add_tooltip(text, tooltip) return '<span class="desc-arr" title="' .. tooltip .. '">' .. text .. '</span>' end -- Boolean params indicating whether a descendant term (or all terms) are particular sorts of borrowings. local bortypes = {"inh", "bor", "lbor", "slb", "obor", "translit", "der", "clq", "pclq", "sml", "unc"} local bortype_set = list_to_set(bortypes) -- Aliases of clq=. local calque_aliases = {"cal", "calq", "calque"} local calque_alias_set = list_to_set(calque_aliases) -- Aliases of pclq=. local partial_calque_aliases = {"pcal", "pcalq", "pcalque"} local partial_calque_alias_set = list_to_set(partial_calque_aliases) local semi_learned_borrowing_aliases = {"slbor"} local semi_learned_borrowing_alias_set = list_to_set(semi_learned_borrowing_aliases) --- Return a function of one argument `field` (a param name), which fetches `args`[`field`].default if index == 0, else --- `container`[`field`]. local function get_val(container, args, index) return function(field) if index == 0 then return args[field].default else return container[field] end end end local function get_arrow(container, args, index) local val = get_val(container, args, index) local arrow if val("bor") then arrow = add_tooltip("→", "borrowed") elseif val("lbor") then arrow = add_tooltip("→", "learned borrowing") elseif val("slb") then arrow = add_tooltip("→", "semi-learned borrowing") elseif val("obor") then arrow = add_tooltip("→", "orthographic borrowing") elseif val("translit") then arrow = add_tooltip("→", "transliteration") elseif val("clq") then arrow = add_tooltip("→", "calque") elseif val("pclq") then arrow = add_tooltip("→", "partial calque") elseif val("sml") then arrow = add_tooltip("→", "semantic loan") elseif val("inh") or (val("unc") and not val("der")) then arrow = add_tooltip(">", "inherited") else arrow = "" end -- allow der=1 in conjunction with bor=1 to indicate e.g. English "pars recta" -- derived and borrowed from Latin "pars". if val("der") then arrow = arrow .. add_tooltip("⇒", "reshaped by analogy or addition of morphemes") end if val("unc") then arrow = arrow .. add_tooltip("?", "uncertain") end if arrow ~= "" then arrow = arrow .. " " end return arrow end -- Return the pre-decoration text for the `index`th term, or the overall pre-decoration text if index == 0. local function get_pre_decorations(container, args, index) if index > 0 then -- per term decorations are handled at the subitem level, by full_link(). return nil, nil end local val = get_val(container, args, index) return val("l"), val("q") end -- Return the post-decoration text for the `index`th term, or the overall post-decoration text if index == 0. local function get_post_decorations(container, args, index, lang) local val = get_val(container, args, index) local boolean_labels = {} if val("inh") then insert(boolean_labels, "inherited") end if val("lbor") then insert(boolean_labels, "learned") end if val("slb") then insert(boolean_labels, "semi-learned") end if val("translit") then insert(boolean_labels, "transliteration") end if val("clq") then insert(boolean_labels, "calque") end if val("pclq") then insert(boolean_labels, "partial calque") end if val("sml") then insert(boolean_labels, "semantic loan") end if index > 0 then -- per term decorations are handled at the subitem level, by full_link(). return boolean_labels else local quals, dash_labels quals = val("qq") if val("ll") then local labels = require(labels_module).show_labels { lang = lang, labels = val("ll"), nocat = true, open = false, close = false, no_track_already_seen = true, ok_to_destructively_modify = true, -- doesn't apply to `labels` } if labels ~= "" then dash_labels = " &mdash; " .. labels end end return boolean_labels, quals, dash_labels end end local function desc_or_desc_tree(frame, desc_tree) local params local boolean = {type = "boolean"} if desc_tree then params = { [1] = {required = true, type = "language", family = true, default = "gem-pro"}, [2] = {required = true, list = true, allow_holes = true, default = "*fuhsaz"}, notext = boolean, noalts = boolean, noparent = boolean, } else params = { [1] = {required = true, type = "language", family = true, default = "en"}, [2] = {list = true, allow_holes = true, template_default = "word"}, alts = boolean, } end -- Add other single params. params.sclang = boolean params.sclb = {replaced_by = "sclang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"} params.nolang = boolean params.nolb = {replaced_by = "nolang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"} local parent_args if frame.args[1] then parent_args = frame.args else parent_args = frame:getParent().args end -- Error to catch most uses of old-style parameters. if ine(parent_args[4]) and not ine(parent_args[3]) and not ine(parent_args.tr2) and not ine(parent_args.ts2) and not ine(parent_args.t2) and not ine(parent_args.gloss2) and not ine(parent_args.g2) and not ine(parent_args.alt2) then error("You specified a term in 4= and not one in 3=. You probably meant to use t= to specify a gloss instead. " .. "If you intended to specify two terms, put the second term in 3=.") end if not ine(parent_args[3]) and not ine(parent_args.alt2) and not ine(parent_args.tr2) and not ine(parent_args.ts2) and ine(parent_args.g2) then error("You specified a gender in g2= but no term in 3=. You were probably trying to specify two genders for " .. "a single term. To do that, put both genders in g=, comma-separated.") end local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l", "q"}}, {param = "lb", replaced_by = false, instead = "use 'l' for left labels or 'll' for right labels"}, {param = bortypes, type = "boolean", overall = true, separate_no_index = true}, {param = calque_aliases, alias_of = "clq"}, {param = partial_calque_aliases, alias_of = "pclq"}, {param = semi_learned_borrowing_aliases, alias_of = "slb"}, } local groups, args, globalprops = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, -- Need some work to support this. -- parse_lang_prefix = true, track_module = "descendant", -- Due to allowing families as langs and substituting 'und', it's easier to do this later. -- lang = function() ... end sc = "sc.default", splitchar = "[,~]", subitem_separator_map = {[","] = " / ", ["~"] = " ~ "}, pre_normalize_modifiers = function(data) local modtext = data.modtext modtext = modtext:match("^<(.*)>$") if not modtext then error(("Internal error: Passed-in modifier isn't surrounded by angle brackets: %s"):format( data.modtext)) end if bortype_set[modtext] or calque_alias_set[modtext] or partial_calque_alias_set[modtext] or semi_learned_borrowing_alias_set[modtext] then modtext = modtext .. ":1" end return "<" .. modtext .. ">" end, } local lang = args[1] local namespace = mw.title.getCurrentTitle().nsText if (namespace == "" or namespace == "Reconstruction") and ( lang:hasType("appendix-constructed") and not lang:hasType("regular")) then error("Terms in appendix-only constructed languages may not be given as descendants.") end local fetch_alt_forms = desc_tree and not args.noalts or not desc_tree and args.alts local m_desctree if desc_tree or fetch_alt_forms then m_desctree = require(descendants_tree_module) end if lang:getCode() ~= lang:getFullCode() then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/etymological]] track("etymological") track("etymological/" .. lang:getCode()) end local is_family = lang:hasType("family") local proxy_lang if is_family then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/family]] track("family") track("family/" .. lang:getCode()) proxy_lang = require(languages_module).getByCode("und") else proxy_lang = lang end local langname if is_family then -- The display form for families includes the word "languages", which we probably don't want to -- display. langname = lang:getCanonicalName() else langname = lang:getDisplayForm() end local langtag if args.sclang then local sc_to_use = args.sc.default if not sc_to_use then local first_termobj = groups[1] and groups[1].terms[1] if not first_termobj then error("sclang= given but no term exists to display the script name of") end sc_to_use = first_termobj.sc if not sc_to_use then local first_term = first_termobj.term or first_termobj.alt if not first_term then error("sclang= given but first specified item no term or display form to display the script name of") end if first_termobj.lang then sc_to_use = first_termobj.lang:findBestScript(first_term) elseif is_family then sc_to_use = require(scripts_module).findBestScriptWithoutLang(first_term, "none is last resort") else sc_to_use = lang:findBestScript(first_term) end end end langtag = sc_to_use:getDisplayForm(lang) else langtag = langname end local terms_for_descendant_trees = {} -- Keep track of descendants whose descendant tree we fetch. Don't fetch the same descendant tree twice (which -- can happen especially with Arabic-script terms with the same unvocalized spelling but differing vocalization). -- This happens e.g. with Ottoman Turkish [[پورتقال]], which has {{desctree|fa-cls|پُرْتُقَال|پُرْتِقَال|bor=1}}, with -- two terms that have the same unvocalized spelling. local terms_and_ids_fetched = {} local descendant_terms_seen = {} local parts = {} for i, group in ipairs(groups) do local group_parts = {} local terms_for_alt_forms = {} for _, item in ipairs(group.terms) do local link = "" item.lang = item.lang or proxy_lang item.track_sc = true -- Construct a link out of `item`. Also add the term to the list of descendant trees and/or alternative -- forms to fetch, if the page+ID combination hasn't already been seen. if item.term ~= "-" then -- including term == nil link = require(links_module).full_link(item, nil, true) if item.term and (desc_tree or fetch_alt_forms) then local m_links = require(links_module) -- Fetches information under entry. If term is of type A//B, it checks A. local entry_name = m_links.get_link_page(m_links.remove_links(mw.ustring.gsub(item.term, "//.+$", "")), lang, item.sc) -- NOTE: We use the term and ID as the key, but not the language. This is OK currently because -- all terms have the same language; but if we ever add support for a term-specific language, -- we need to fix this. local term_and_id = item.id and entry_name .. "!!!" .. item.id or entry_name if not terms_and_ids_fetched[term_and_id] then terms_and_ids_fetched[term_and_id] = true local term_for_fetching = { lang = lang, entry_name = entry_name, id = item.id } if desc_tree then if is_family then error("No support currently (and probably ever) for fetching a descendant tree when a family code instead of language code is given") end if error_on_no_descendants then require(table_module).insertIfNot(descendant_terms_seen, { term = item.term, id = item.id }) end table.insert(terms_for_descendant_trees, term_for_fetching) end if fetch_alt_forms then if is_family then error("No support currently (and probably ever) for fetching alternative forms when a family code instead of language code is given") end -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/alts]] track("alts") table.insert(terms_for_alt_forms, term_for_fetching) end end end elseif item.tr or item.ts or item.gloss or item.genders then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term]] track("no term") item.term = nil item.show_decorations = true link = require(links_module).full_link(item, nil, true) link = link :gsub("<small>%[Term%?%]</small> ", "") :gsub("<small>%[Term%?%]</small>&nbsp;", "") :gsub("%[%[Category:[^%[%]]+ term requests%]%]", "") else -- display no link at all -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term or annotations]] track("no term or annotations") end if link ~= "" then insert(group_parts, item.separator) insert(group_parts, link) end end if group_parts[1] then for _, altterm in ipairs(terms_for_alt_forms) do local altform = m_desctree.get_alternative_forms(altterm.lang, altterm.entry_name, altterm.id, globalprops.use_semicolon and "; " or ", ") if altform ~= "" then insert(group_parts, globalprops.use_semicolon and "; " or ", ") insert(group_parts, altform) end end local group_link = concat(group_parts) insert(parts, group.separator) if not args.notext then insert(parts, get_arrow(group, args, i)) end -- no pre-qualifiers/labels and no post-qualifiers/dash-labels local post_boolean_labels = get_post_decorations(group, args, i, proxy_lang) if post_boolean_labels and post_boolean_labels[1] then group_link = require(decorations_module).format_decorations { lang = proxy_lang, text = group_link, ll = post_boolean_labels, } end insert(parts, group_link) end end local descendant_trees = {} for _, descterm in ipairs(terms_for_descendant_trees) do -- When I ([[User:Benwing2]]) first implemented this in Nov 2020, I had `maxmaxindex > 1` as the last argument. -- Since then, [[User:Fytcha]] changed the last param to `true`. local descendant_tree = m_desctree.get_descendants(descterm.lang, descterm.entry_name, descterm.id, true) if descendant_tree and descendant_tree ~= "" then insert(descendant_trees, descendant_tree) end end if error_on_no_descendants and desc_tree and not descendant_trees[1] then local function format_term_seen(term_seen) if term_seen.id then return ("[[%s]] with ID '%s'"):format(term_seen.term, term_seen.id) else return ("[[%s]]"):format(term_seen.term) end end if #descendant_terms_seen == 0 then error("[[Template:desctree]] invoked but no terms to retrieve descendants from") elseif #descendant_terms_seen == 1 then error(("No Descendants section was found in the entry %s under the header for %s"):format( format_term_seen(descendant_terms_seen[1]), lang:getFullName())) else for i, term_seen in ipairs(descendant_terms_seen) do descendant_terms_seen[i] = format_term_seen(term_seen) end error(("No Descendants section was found in any of the entries %s under the header for %s"):format( concat(descendant_terms_seen, ", "), lang:getFullName())) end end local descendants = concat(descendant_trees) if args.noparent then return descendants end local initial_labels, initial_quals = get_pre_decorations(nil, args, 0) local final_boolean_labels, final_quals, final_dash_labels = get_post_decorations(nil, args, 0, proxy_lang) local all_linktext = concat(parts) if initial_labels and initial_labels[1] or initial_quals and initial_quals[1] or final_boolean_labels and final_boolean_labels[1] or final_quals and final_quals[1] then all_linktext = require(decorations_module).format_decorations { lang = proxy_lang, text = all_linktext, l = initial_labels, q = initial_quals, ll = final_boolean_labels, qq = final_quals, } end if final_dash_labels then all_linktext = all_linktext .. final_dash_labels end all_linktext = all_linktext .. descendants if args.notext then return all_linktext end local initial_arrow = get_arrow(nil, args, 0) if args.nolang then return initial_arrow .. all_linktext else return concat { initial_arrow, langtag, ":", all_linktext ~= "" and " " or "", all_linktext } end end function export.descendant(frame) return desc_or_desc_tree(frame, false) .. require(template_styles_module)(etymology_style_css) end function export.descendants_tree(frame) return desc_or_desc_tree(frame, true) end return export eu10g3uzxrhe1n2f4cxiwv3j90yvwqw Module:parameters/finalizeSet 828 8068 54883 2026-09-27T21:47:26Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local parameters_track_module = "Module:parameters/track" local dump = mw.dumpObject local error = error local format = string.format local pairs = pairs local tostring = tostring local type = type local function track(...) track = require(parameters_track_module) return track(...) end local type_err = 'expected set members to be of type "string" or "number", but saw %s' --[==[ -- Takes `t`, a list or key map which defines a set, and returns a key map for t...' 54883 Scribunto text/plain local parameters_track_module = "Module:parameters/track" local dump = mw.dumpObject local error = error local format = string.format local pairs = pairs local tostring = tostring local type = type local function track(...) track = require(parameters_track_module) return track(...) end local type_err = 'expected set members to be of type "string" or "number", but saw %s' --[==[ -- Takes `t`, a list or key map which defines a set, and returns a key map for the set (which might be the original input value). In addition to distinguishing lists from key maps, this function performs various validation checks: * Sets may only contain strings and numbers (i.e. the two valid data types for Scribunto parameters). * List inputs must be contiguous arrays, with no other keys in use (even if they are non-numbers). * Key maps are treated as having two different kinds of key: ** If a key's value is boolean {true}, it is a standard key. ** If a key's value is anything else, it is an alias of the specified key (which must not be an alias itself). For instance, if the key {"foo"} is set to {"bar"}, then {"foo"} is an alias of the key {"bar"}, which must be a non-alias key that is also in the key map.]==] return function(t, name) -- Iterates over `t` using pairs(), doing separate key map and list parses simultaneously. If the key map parse succeeds, `t` will simply be returned, but if the list parse succeeds, a key map of the values in the list will be returned instead. -- Lists and key maps are mutually exclusive, as they can be distinguished by the presence of boolean `true` as a value: -- (1) If `t` is a list, then `true` is an invalid value, because sets may only contain strings and numbers. -- (2) If `t` is a key map, then it must contain at least one `true` as a value, because any keys which do not have `true` as their value are (by definition) aliases, and aliases must be pointed to a non-alias key (i.e. a key which has the value `true`), which is only possible if one or more keys are set to `true`. -- (Formally, an empty table could be either, but it's more efficient to treat it as a key map.) local i, new_map, list_err, map_err, duplicate_in_list = 0 -- The two parses are separated into blocks, which can be skipped on any -- further iterations if that parse fails. for k, v in pairs(t) do local k_type, v_type = type(k) -- The while-blocks make it possible to use `break` as a substitute for -- `goto`, and both unconditionally terminate after one iteration. while not list_err do -- Catch holes using the same method as [[Module:table/isArray]], -- but also check for non-number keys. i = i + 1 if k_type ~= "number" or t[i] == nil then list_err = "input list is not contiguous" break end -- If `t` is a list then `v` is a set member, so it must be a string -- or number. v_type = type(v) if not (v_type == "string" or v_type == "number") then list_err = format( type_err, v_type == "boolean" and tostring(v) or v_type ) -- Populate `new_map` with valid keys. elseif not new_map then new_map = {[v] = true} -- If new_map[v] is already set, then `t` has duplicates. This -- should be tracked, but only if `t` does turn out to be a list, so -- flag it with `duplicate_in_list` for now. elseif new_map[v] then duplicate_in_list = true else new_map[v] = true end break end while not map_err do -- If `t` is a list then `k` is a set member, so it must be a string -- or number. if not (k_type == "string" or k_type == "number") then map_err = format( type_err, k_type == "boolean" and tostring(k) or k_type ) break -- If `v` is true, `k` is a non-alias key. elseif v == true then break -- Filter out invalid self-aliases. elseif k == v then map_err = format( "set member %s cannot be an alias of itself", dump(k) ) break elseif not v_type then v_type = type(v) end -- Only possible to be an alias of another key, which must be a -- string or number. if not (v_type == "string" or v_type == "number") then map_err = format( 'expected set key %s have the value true or a value of type "string" or "number", but saw %s', dump(k), v_type == "boolean" and tostring(v) or v_type ) break end -- Must be an alias of a non-alias, so the value for the key `v` -- must be `true`. local main = t[v] if main ~= true then map_err = format( "set member %s is specified as an alias of %s, %s", dump(k), dump(v), main == nil and "which is not in the set" or "which is also an alias" ) end break end -- If both parses have failed, throw an error with the two messages. if list_err and map_err then error(format( "Internal error: the `set` spec%s cannot be parsed as either a list or a key map:\nlist parse: %s\nmap parse: %s", name and format(" for parameter %s", dump(name)) or "", map_err, list_err )) end end -- If `t` can be parsed as a key map, just return it. if not map_err then return t -- Otherwise, `t` is a list, so track duplicate entries if it has been -- flagged. elseif duplicate_in_list then track("duplicate entry in set list") end -- Return the new key map. return new_map end aoylq9j5kt5vr096lzl6bv7hjufnzka Module:parse interface 828 8069 54884 2026-09-27T21:47:51Z Deadend0914 7211 Gesceop tramet þe hafaþ 'local export = {} local string_utilities_module = "Module:string utilities" local parse_utilities_module = "Module:parse utilities" local table_module = "Module:table" --[=[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target funct...' 54884 Scribunto text/plain local export = {} local string_utilities_module = "Module:string utilities" local parse_utilities_module = "Module:parse utilities" local table_module = "Module:table" --[=[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls. ]=] local function rfind(...) rfind = require(string_utilities_module).find return rfind(...) end local function rsplit(...) rsplit = require(string_utilities_module).split return rsplit(...) end local function split_on_comma(...) split_on_comma = require(parse_utilities_module).split_on_comma return split_on_comma(...) end local function split_escaping(...) split_escaping = require(parse_utilities_module).split_escaping return split_escaping(...) end local function parse_inline_modifiers(...) parse_inline_modifiers = require(parse_utilities_module).parse_inline_modifiers return parse_inline_modifiers(...) end local function parse_term_with_lang(...) parse_term_with_lang = require(parse_utilities_module).parse_term_with_lang return parse_term_with_lang(...) end local function term_contains_top_level_html(...) term_contains_top_level_html = require(parse_utilities_module).term_contains_top_level_html return term_contains_top_level_html(...) end local function escape_comma_whitespace(...) escape_comma_whitespace = require(parse_utilities_module).escape_comma_whitespace return escape_comma_whitespace(...) end local function unescape_comma_whitespace(...) unescape_comma_whitespace = require(parse_utilities_module).unescape_comma_whitespace return unescape_comma_whitespace(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function decode_entities(...) -- FIXME: Why are we doing this? It was added to [[Module:form of/templates]] in -- https://en.wiktionary.org/w/index.php?title=Module:form_of/templates&diff=prev&oldid=81900806 on 2024-09-24 -- by [[User:Theknightwho]] with the comment "Optimisations + decode HTML entities.". -- -- NOTE: We could add a check for & in the term before calling decode_entities(), but in practice, -- [[Module:string utilities]] is essentially always loaded so there's little point. str_decode_entities = require(string_utilities_module).decode_entities return str_decode_entities(...) end --[==[ This is an almost drop-in replacement for split_on_comma() in [[Module:parse utilities]], with optimizations to avoid loading and running the while algorithm in [[Module:parse utilities]] except when necessary. ]==] function export.split_on_comma(val) if val:find(",%s") or (val:find(",") and val:find("[\\%[<]")) then -- Comma after whitespace not split; nor are backslash-escaped commas or commas inside of square or -- angle brackets. If we see any of these, use the more sophisticated algorithm in -- [[Module:parse utilities]]. Otherwise it's safe to just split on commas directly. This optimization -- avoids loading [[Module:parse utilities]] unnecessarily. return split_on_comma(val) else return rsplit(val, ",") end end --[==[ This is similar to parse_term_with_lang() in [[Module:parse utilities]], but if there is no colon + non-space in the term, it will be returned directly and not parsed into link/display format. If you need the link/display arguments even in the absence of a language prefix, call [[Module:parse utilities]] directly. ]==] function export.parse_term_with_lang(data) if data.term:find(":[^ ]") then return parse_term_with_lang(data) else return data.term, nil, nil, nil end end --[==[ This is an almost drop-in replacement for parse_inline_modifiers() in [[Module:parse utilities]] except that # it won't attempt to parse inline modifiers if it detects that top-level HTML is present (but it will still split on `splitchar` if given, unless it detects the presence of the {{tl|,}} template); # it has a default for `generate_obj` that simply sets `lang` and `term` after calling `decode_entities()` on the term (FIXME: this was inherited from code added to [[Module:form of/templates]] by [[User:Theknightwho]]; I don't know why it is necessary); # it has a lot of optimizations to avoid loading [[Module:parse utilities]] in simple cases where there are no `<` signs and (when `splitchar` is given) either there are no delimiters present at all or no characters present that will make a simple split on `splitchar` invalid. Generally you should use this in preference to either calling parse_inline_modifiers() directly in [[Module:parse utilities]] or rolling your own front-end function. ]==] function export.parse_inline_modifiers(val, props) local paramname, lang, splitchar = props.paramname, props.lang, props.splitchar local preserve_splitchar, escape_fun, unescape_fun = props.preserve_splitchar, props.escape_fun, props.unescape_fun local outer_container = props.outer_container local generate_obj = props.generate_obj or function(term) return {lang = lang, term = decode_entities(term)} end local delimiter_key = props.delimiter_key or "delimiter" -- Check for inline modifier, e.g. מרים<tr:Miryem>. But exclude HTML entry with <span ...>, <i ...>, <br/> or -- similar in it, caused by wrapping an argument in {{l|...}}, {{af|...}} or similar. Basically, all tags of -- the sort we parse here should consist of a less-than sign, plus letters, plus a colon, e.g. <tr:...>, so if -- we see a tag on the outer level that isn't in this format, we don't try to parse it. The restriction to the -- outer level is to allow generated HTML inside of e.g. qualifier tags, such as foo<q:similar to {{m|fr|bar}}>. if val:find("<") and not term_contains_top_level_html(val) then if not props.generate_obj then props = shallow_copy(props) props.generate_obj = generate_obj end return parse_inline_modifiers(val, props) end if not splitchar then return generate_obj(val) end local retval if splitchar == "," and escape_fun == nil and unescape_fun == nil then if val:find(",</") then -- This happens when there's an embedded {{,}} template, as in [[MMR]], [[TMA]], [[DEI]], where an -- initialism expands to multiple terms; easiest not to try and parse the lemma spec as multiple lemmas. retval = {val} else retval = export.split_on_comma(val) end for i, split in ipairs(retval) do retval[i] = generate_obj(split) if preserve_splitchar and i > 1 then retval[delimiter_key] = "," end end elseif rfind(val, splitchar) then if val:find(",</") then -- This happens when there's an embedded {{,}} template, as in [[MMR]], [[TMA]], [[DEI]], where an -- initialism expands to multiple terms; easiest not to try and parse the lemma spec as multiple lemmas. retval = {val} elseif escape_fun or unescape_fun or val:find(",%s") or val:find("[\\%[<]") then local defaulted_escape_fun, defaulted_unescape_fun if escape_fun == nil then defaulted_escape_fun = escape_comma_whitespace end if unescape_fun == nil then defaulted_unescape_fun = unescape_comma_whitespace end retval = split_escaping(val, splitchar, preserve_splitchar, defaulted_escape_fun, defaulted_unescape_fun) elseif preserve_splitchar then retval = rsplit(val, "(" .. splitchar .. ")") else retval = rsplit(val, splitchar) end if preserve_splitchar then local new_retval = {} for j = 1, #retval, 2 do local obj = generate_obj(retval[j]) if j > 1 then obj[delimiter_key] = retval[j - 1] end table.insert(new_retval, obj) end retval = new_retval else for i, split in ipairs(retval) do retval[i] = generate_obj(split) end end else retval = {generate_obj(val)} end if outer_container then outer_container.terms = retval return outer_container end return retval end return export bq5sys0ifgirhdamppe6s3lw9xawfoo Brūcend:Deadend0914 2 8070 54886 2026-09-27T22:58:01Z Deadend0914 7211 Gesceop tramet þe hafaþ '=Standardization!= ==Parts of Speach== * (a lot of these are already established, this is just a list) *Noun - Nama ** Masculine - Werlic ** Feminine - Wīflic ** Neuter - Nāhwæðer *Adjective - Tōgeīecendlic (note: this is also an adjective) *Adverb - Bīword *Case - Cāsus ** Accusative - Wrēgendlic ** Dative - Forgifendlic ** Genitive - Āgniendlic ** Imperative - Bebēodendlic ** Instrumental - Tōllic ** Nominative - Nemniendlic ** Subjunctive - Underþ...' 54886 wikitext text/x-wiki =Standardization!= ==Parts of Speach== * (a lot of these are already established, this is just a list) *Noun - Nama ** Masculine - Werlic ** Feminine - Wīflic ** Neuter - Nāhwæðer *Adjective - Tōgeīecendlic (note: this is also an adjective) *Adverb - Bīword *Case - Cāsus ** Accusative - Wrēgendlic ** Dative - Forgifendlic ** Genitive - Āgniendlic ** Imperative - Bebēodendlic ** Instrumental - Tōllic ** Nominative - Nemniendlic ** Subjunctive - Underþēodendlic *Conjunction - Fēgung *Declension(s) - Declīnung(a) *Gender - Cynn *Inflection - Gebīegednes *Infinitive (or Indefinite) - Ungeendigendlic *Interjection - Betwuxālegednes *Participle - Dǣlnimend *Person - Hād *Plural - Manigfealdlic (mnf.) *Preposition - Foresetednes *Pronoun - Bīnama *Singular - Ānfealdlic (anf.) *Time - Tīd ** Past - Forþgewiten ** Perfect - Fullfremed ** Present - Andweardnes *Verb - Word ** Strong - Strang ** Weak - Unstrang =Templates= ==Noun Example== *(script used is sprǣc|language-code surrounded by brackets ({}) =={{sprǣc|ang}}== ===Āwendednessa=== * Here, you put alternate forms of words. Optional ===Rihtstefn=== * Here, you put the pronunciation of the word * IPA: /whatever-word/ ===Wordstǣr=== * Here, you put where the word came from. I'm still working on the script but it will be inh|current-language|language-from|word * For now, use this format below * Of (language, use the dative) ===(Gender if applicable) Nama=== * Still working on the script here, for now use 3 ' around the word! '''example''' * For definitions, use a #, and for examples, use #: # It will look like this #: This is an example! ====Declīnung==== * Here, you put the inflection table! The script is depending on the specific noun's conjugation! I'll make a page for that soon! ====Wendunga==== * Here, you put translations! Only for Englisc words! * Categories go further down, but it doesn't really matter where, because they don't appear on the actual page anyway. We use Flocc:(whatever category name) inside square brackets "[]" 4fqfig9mjiyuimx9fcybv7zaeetuq3e 54887 54886 2026-09-27T23:03:37Z Deadend0914 7211 /* Templates */ 54887 wikitext text/x-wiki =Standardization!= ==Parts of Speach== * (a lot of these are already established, this is just a list) *Noun - Nama ** Masculine - Werlic ** Feminine - Wīflic ** Neuter - Nāhwæðer *Adjective - Tōgeīecendlic (note: this is also an adjective) *Adverb - Bīword *Case - Cāsus ** Accusative - Wrēgendlic ** Dative - Forgifendlic ** Genitive - Āgniendlic ** Imperative - Bebēodendlic ** Instrumental - Tōllic ** Nominative - Nemniendlic ** Subjunctive - Underþēodendlic *Conjunction - Fēgung *Declension(s) - Declīnung(a) *Gender - Cynn *Inflection - Gebīegednes *Infinitive (or Indefinite) - Ungeendigendlic *Interjection - Betwuxālegednes *Participle - Dǣlnimend *Person - Hād *Plural - Manigfealdlic (mnf.) *Preposition - Foresetednes *Pronoun - Bīnama *Singular - Ānfealdlic (anf.) *Time - Tīd ** Past - Forþgewiten ** Perfect - Fullfremed ** Present - Andweardnes *Verb - Word ** Strong - Strang ** Weak - Unstrang =Templates= ==Noun Example (use 2 "=")== *(script used is sprǣc|language-code surrounded by brackets ({}). =={{sprǣc|ang}}== ===Āwendednessa (use 3 "=")=== * Here, you put alternate forms of words. Optional. ===Rihtstefn (use 3 "=")=== * Here, you put the pronunciation of the word. * IPA: /whatever-word/ ===Wordstǣr (use 3 "=")=== * Here, you put where the word came from. I'm still working on the script but it will be inh|current-language|language-from|word * For now, use this format below * Of (language, use the dative) ====Gesibbword==== * Same language Cognates ====Sweostorword==== * Non-same language Cognates) ====Cildru==== * Descendants ===(Gender if applicable) Nama (use 3 "=")=== * Still working on the script here, for now use 3 ' around the word! '''example''' * For definitions, use a #, and for examples, use #: # It will look like this #: This is an example! ====Declīnung (use 4 "=")==== * Here, you put the inflection table! The script is depending on the specific noun's conjugation! I'll make a page for that soon! ====Wendunga (use 4 "=")==== * Here, you put translations! Only for Englisc words! * Categories go further down, but it doesn't really matter where, because they don't appear on the actual page anyway. We use Flocc:(whatever category name) inside square brackets "[]" rwy0uqrqbhrgigp2te4q1fq3utwpwch