Wikikamus mswiktionary https://ms.wiktionary.org/wiki/Wikikamus:Laman_Utama MediaWiki 1.47.0-wmf.20 case-sensitive Media Khas Perbincangan Pengguna Perbincangan pengguna Wikikamus Perbincangan Wikikamus Fail Perbincangan fail MediaWiki Perbincangan MediaWiki Templat Perbincangan templat Bantuan Perbincangan bantuan Kategori Perbincangan kategori Lampiran Perbincangan lampiran Rima Perbincangan rima Tesaurus Perbincangan tesaurus Indeks Perbincangan indeks Petikan Perbincangan petikan Rekonstruksi Perbincangan rekonstruksi Padanan isyarat Perbincangan padanan isyarat Konkordans Perbincangan konkordans TimedText TimedText talk Modul Perbincangan modul Acara Perbincangan acara Modul:languages 828 8666 375364 375316 2026-09-22T04:33:42Z Hakimi97 2668 Membatalkan semakan [[Special:Diff/375316|375316]] oleh [[Special:Contributions/Hakimi97|Hakimi97]] ([[User talk:Hakimi97|bincang]]) 375364 Scribunto text/plain --[==[ intro: This module implements fetching of language-specific information and processing text in a given language. ===Types of languages=== There are two types of languages: full languages and etymology-only languages. The essential difference is that only full languages appear in L2 headings in vocabulary entries, and hence categories like [[:Category:French nouns]] exist only for full languages. Etymology-only languages have either a full language or another etymology-only language as their parent (in the parent-child inheritance sense), and for etymology-only languages with another etymology-only language as their parent, a full language can always be derived by following the parent links upwards. For example, "Canadian French", code `fr-CA`, is an etymology-only language whose parent is the full language "French", code `fr`. An example of an etymology-only language with another etymology-only parent is "Northumbrian Old English", code `ang-nor`, which has "Anglian Old English", code `ang-ang` as its parent; this is an etymology-only language whose parent is "Old English", code `ang`, which is a full language. (This is because Northumbrian Old English is considered a variety of Anglian Old English.) Sometimes the parent is the "Undetermined" language, code `und`; this is the case, for example, for "substrate" languages such as "Pre-Greek", code `qsb-grc`, and "the BMAC substrate", code `qsb-bma`. It is important to distinguish language ''parents'' from language ''ancestors''. The parent-child relationship is one of containment, i.e. if X is a child of Y, X is considered a variety of Y. On the other hand, the ancestor-descendant relationship is one of descent in time. For example, "Classical Latin", code `la-cla`, and "Late Latin", code `la-lat`, are both etymology-only languages with "Latin", code `la`, as their parents, because both of the former are varieties of Latin. However, Late Latin does *NOT* have Classical Latin as its parent because Late Latin is *not* a variety of Classical Latin; rather, it is a descendant. There is in fact a separate `ancestors` field that is used to express the ancestor-descendant relationship, and Late Latin's ancestor is given as Classical Latin. It is also important to note that sometimes an etymology-only language is actually the conceptual ancestor of its parent language. This happens, for example, with "Old Italian" (code `roa-oit`), which is an etymology-only variant of full language "Italian" (code `it`), and with "Old Latin" (code `itc-ola`), which is an etymology-only variant of Latin. In both cases, the full language has the etymology-only variant listed as an ancestor. This allows a Latin term to inherit from Old Latin using the {{tl|inh}} template (where in this template, "inheritance" refers to ancestral inheritance, i.e. inheritance in time, rather than in the parent-child sense); likewise for Italian and Old Italian. Full languages come in three subtypes: * {regular}: This indicates a full language that is attested according to [[WT:CFI]] and therefore permitted in the main namespace. There may also be reconstructed terms for the language, which are placed in the {Reconstruction} namespace and must be prefixed with * to indicate a reconstruction. Most full languages are natural (not constructed) languages, but a few constructed languages (e.g. Esperanto and Volapük, among others) are also allowed in the mainspace and considered regular languages. * {reconstructed}: This language is not attested according to [[WT:CFI]], and therefore is allowed only in the {Reconstruction} namespace. All terms in this language are reconstructed, and must be prefixed with *. Languages such as Proto-Indo-European and Proto-Germanic are in this category. * {appendix-constructed}: This language is attested but does not meet the additional requirements set out for constructed languages ([[WT:CFI#Constructed languages]]). Its entries must therefore be in the Appendix namespace, but they are not reconstructed and therefore should not have * prefixed in links. Most constructed languages are of this subtype. Both full languages and etymology-only languages have a {Language} object associated with them, which is fetched using the {getByCode} function in [[Modul:languages]] to convert a language code to a {Language} object. Depending on the options supplied to this function, etymology-only languages may or may not be accepted, and family codes may be accepted (returning a {Family} object as described in [[Modul:families]]). There are also separate {getByCanonicalName} functions in [[Modul:languages]] and [[Modul:etymology languages]] to convert a language's canonical name to a {Language} object (depending on whether the canonical name refers to a full or etymology-only language). ===Textual representations=== Textual strings belonging to a given language come in several different ''text variants'': # The ''input text'' is what the user supplies in wikitext, in the parameters to {{tl|m}}, {{tl|l}}, {{tl|ux}}, {{tl|t}}, {{tl|lang}} and the like. # The ''corrected input text'' is the input text with some corrections and/or normalizations applied, such as bad-character replacements for certain languages, like replacing `l` or `1` to [[palochka]] in some languages written in Cyrillic. (FIXME: This currently goes under the name ''display text'' but that will be repurposed below. Also, [[User:Surjection]] suggests renaming this to ''normalized input text'', but "normalized" is used in a different sense in [[Modul:usex]].) # The ''display text'' is the text in the form as it will be displayed to the user. This is what appears in headwords, in usexes, in displayed internal links, etc. This can include accent marks that are removed to form the stripped display text (see below), as well as embedded bracketed links that are variously processed further. The display text is generated from the corrected input text by applying language-specific transformations; for most languages, there will be no such transformations. The general reason for having a difference between input and display text is to allow for extra information in the input text that is not displayed to the user but is sent to the transliteration module. Note that having different display and input text is only supported currently through special-casing but will be generalized. Examples of transformations are: (1) Removing the {{cd|^}} that is used in certain East Asian (and possibly other unicameral) languages to indicate capitalization of the transliteration (which is currently special-cased); (2) for Korean, removing or otherwise processing hyphens (which is currently special-cased); (3) for Arabic, removing a ''sukūn'' diacritic placed over a ''tāʔ marbūṭa'' (like this: ةْ) to indicate that the ''tāʔ marbūṭa'' is pronounced and transliterated as /t/ instead of being silent [NOTE, NOT IMPLEMENTED YET]; (4) for Thai and Khmer, converting space-separated words to bracketed words and resolving respelling substitutions such as `[กรีน/กฺรีน]`, which indicate how to transliterate given words [NOTE, NOT IMPLEMENTED YET except in language-specific templates like {{tl|th-usex}}]. ## The ''right-resolved display text'' is the result of removing brackets around one-part embedded links and resolving two-part embedded links into their right-hand components (i.e. converting two-part links into the displayed form). The process of right-resolution is what happens when you call {{cd|remove_links()}} in [[Modul:links]] on some text. When applied to the display text, it produces exactly what the user sees, without any link markup. # The ''stripped display text'' is the result of applying diacritic-stripping to the display text. ## The ''left-resolved stripped display text'' [NEED BETTER NAME] is the result of applying left-resolution to the stripped display text, i.e. similar to right-resolution but resolving two-part embedded links into their left-hand components (i.e. the linked-to page). If the display text refers to a single page, the resulting of applying diacritic stripping and left-resolution produces the ''logical pagename''. # The ''physical pagename text'' is the result of converting the stripped display text into physical page links. If the stripped display text contains embedded links, the left side of those links is converted into physical page links; otherwise, the entire text is considered a pagename and converted in the same fashion. The conversion does three things: (1) converts characters not allowed in pagenames into their "unsupported title" representation, e.g. {{cd|Unsupported titles/`gt`}} in place of the logical name {{cd|>}}; (2) handles certain special-cased unsupported-title logical pagenames, such as {{cd|Unsupported titles/Space}} in place of {{cd|[space]}} and {{cd|Unsupported titles/Ancient Greek dish}} in place of a very long Greek name for a gourmet dish as found in Aristophanes; (3) converts "mammoth" pagenames such as [[a]] into their appropriate split component, e.g. [[a/languages A to L]]. # The ''source translit text'' is the text as supplied to the language-specific {{cd|transliterate()}} method. The form of the source translit text may need to be language-specific, e.g Thai and Khmer will need the corrected input text, whereas other languages may need to work off the display text. [FIXME: It's still unclear to me how embedded bracketed links are handled in the existing code.] In general, embedded links need to be right-resolved (see above), but when this happens is unclear to me [FIXME]. Some languages have a chop-up-and-paste-together scheme that sends parts of the text through the transliterate mechanism, and for others (those listed with "cont" in {{cd|substitution}} in [[Modul:languages/data]]) they receive the full input text, but preprocessed in certain ways. (The wisdom of this is still unclear to me.) # The ''transliterated text'' (or ''transliteration'') is the result of transliterating the source translit text. Unlike for all the other text variants except the transcribed text, it is always in the Latin script. # The ''transcribed text'' (or ''transcription'') is the result of transcribing the source translit text, where "transcription" here means a close approximation to the phonetic form of the language in languages (e.g. Akkadian, Sumerian, Ancient Egyptian, maybe Tibetan) that have a wide difference between the written letters and spoken form. Unlike for all the other text variants other than the transliterated text, it is always in the Latin script. Currently, the transcribed text is always supplied manually be the user; there is no such thing as a {{cd|transcribe()}} method on language objects. # The ''sort key'' is the text used in sort keys for determining the placing of pages in categories they belong to. The sort key is generated from the pagename or a specified ''sort base'' by lowercasing, doing language-specific transformations and then uppercasing the result. If the sort base is supplied and is generated from input text, it needs to be converted to display text, have embedded links removed through right-resolution and have diacritic-stripping applied. # There are other text variants that occur in usexes (specifically, there are normalized variants of several of the above text variants), but we can skip them for now. The following methods exist on {Language} objects to convert between different text variants: # {correctInputText} (currently called {makeDisplayText}): This converts input text to corrected input text. # {stripDiacritics}: This converts to stripped display text. [FIXME: This needs some rethinking. In particular, {stripDiacritics} is sometimes called on input text, corrected input text or display text (in various paths inside of [[Modul:links]], and, in the case of input text, usually from other modules). We need to make sure we don't try to convert input text to display text twice, but at the same time we need to support calling it directly on input text since so many modules do this. This means we need to add a parameter indicating whether the passed-in text is input, corrected input, or display text; if the former two, we call {correctInputText} ourselves.] # {logicalToPhysical}: This converts logical pagenames to physical pagenames. # {transliterate}: This appears to convert input text with embedded brackets removed into a transliteration. [FIXME: This needs some rethinking. In particular, it calls {processDisplayText} on its input, which won't work for Thai and Khmer, so we may need language-specific flags indicating whether to pass the input text directly to the language transliterate method. In addition, I'm not sure how embedded links are handled in the existing translit code; a lot of callers remove the links themselves before calling {transliterate()}, which I assume is wrong.] # {makeSortKey}: This converts display text (?) to a sort key. [FIXME: Clarify this.] ]==] local export = {} local debug_track_module = "Modul:debug/track" local etymology_languages_data_module = "Modul:etymology languages/data" local families_module = "Modul:families" local headword_page_module = "Modul:headword/page" local json_module = "Modul:JSON" local language_like_module = "Modul:language-like" local languages_data_module = "Modul:languages/data" local languages_data_patterns_module = "Modul:languages/data/patterns" local links_data_module = "Modul:links/data" local load_module = "Modul:load" local scripts_module = "Modul:scripts" local scripts_data_module = "Modul:scripts/data" local string_encode_entities_module = "Modul:string/encode entities" local string_pattern_escape_module = "Modul:string/patternEscape" local string_replacement_escape_module = "Modul:string/replacementEscape" local string_utilities_module = "Modul:string utilities" local table_module = "Modul:table" local utilities_module = "Modul:utilities" local wikimedia_languages_module = "Modul:wikimedia languages" local mw = mw local string = string local table = table local char = string.char local concat = table.concat local find = string.find local floor = math.floor local get_by_code -- Defined below. local get_data_module_name -- Defined below. local get_extra_data_module_name -- Defined below. local getmetatable = getmetatable local gmatch = string.gmatch local gsub = string.gsub local insert = table.insert local ipairs = ipairs local is_known_language_tag = mw.language.isKnownLanguageTag local make_object -- Defined below. local match = string.match local next = next local pairs = pairs local remove = table.remove local require = require local select = select local setmetatable = setmetatable local sub = string.sub local type = type local unstrip = mw.text.unstrip -- Loaded as needed by findBestScript. local Hans_chars local Hant_chars local function check_object(...) check_object = require(utilities_module).check_object return check_object(...) end local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function decode_entities(...) decode_entities = require(string_utilities_module).decode_entities return decode_entities(...) end local function decode_uri(...) decode_uri = require(string_utilities_module).decode_uri return decode_uri(...) end local function deep_copy(...) deep_copy = require(table_module).deepCopy return deep_copy(...) end local function encode_entities(...) encode_entities = require(string_encode_entities_module) return encode_entities(...) end local function get_L2_sort_key(...) get_L2_sort_key = require(headword_page_module).get_L2_sort_key return get_L2_sort_key(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function find_best_script_without_lang(...) find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang return find_best_script_without_lang(...) end local function get_family(...) get_family = require(families_module).getByCode return get_family(...) end local function get_plaintext(...) get_plaintext = require(utilities_module).get_plaintext return get_plaintext(...) end local function get_wikimedia_lang(...) get_wikimedia_lang = require(wikimedia_languages_module).getByCode return get_wikimedia_lang(...) end local function keys_to_list(...) keys_to_list = require(table_module).keysToList return keys_to_list(...) end local function list_to_set(...) list_to_set = require(table_module).listToSet return list_to_set(...) end local function load_data(...) load_data = require(load_module).load_data return load_data(...) end local function make_family_object(...) make_family_object = require(families_module).makeObject return make_family_object(...) end local function pattern_escape(...) pattern_escape = require(string_pattern_escape_module) return pattern_escape(...) end local function replacement_escape(...) replacement_escape = require(string_replacement_escape_module) return replacement_escape(...) end local function safe_require(...) safe_require = require(load_module).safe_require return safe_require(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function to_json(...) to_json = require(json_module).toJSON return to_json(...) end local function u(...) u = require(string_utilities_module).char return u(...) end local function ugsub(...) ugsub = require(string_utilities_module).gsub return ugsub(...) end local function ulen(...) ulen = require(string_utilities_module).len return ulen(...) end local function ulower(...) ulower = require(string_utilities_module).lower return ulower(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end local function uupper(...) uupper = require(string_utilities_module).upper return uupper(...) end local function track(page) debug_track("languages/" .. page) return true end local function normalize_code(code) return load_data(languages_data_module).aliases[code] or code end local function check_inputs(self, check, default, ...) local n = select("#", ...) if n == 0 then return false end local ret = check(self, (...)) if ret ~= nil then return ret elseif n > 1 then local inputs = {...} for i = 2, n do ret = check(self, inputs[i]) if ret ~= nil then return ret end end end return default end local function make_link(self, target, display) local prefix, main if self:getFamilyCode() == "qfa-sub" then prefix, main = display:match("^(the )(.*)") if not prefix then prefix, main = display:match("^(a )(.*)") end end return (prefix or "") .. "[[" .. target .. "|" .. (main or display) .. "]]" end -- Convert risky characters to HTML entities, which minimizes interference once returned (e.g. for "sms:a", "<!-- -->" etc.). local function escape_risky_characters(text) -- Spacing characters in isolation generally need to be escaped in order to be properly processed by the MediaWiki -- software. if umatch(text, "^%s*$") then return encode_entities(text, text) end return encode_entities(text, "!#%&*+/:;<=>?@[\\]_{|}") end -- Temporarily convert various formatting characters to PUA to prevent them from being disrupted by the substitution process. local function doTempSubstitutions(text, subbedChars, keepCarets, noTrim) -- Clone so that we don't insert any extra patterns into the table in package.loaded. For some reason, using require -- seems to keep memory use down; probably because the table is always cloned. local patterns = shallow_copy(require(languages_data_patterns_module)) if keepCarets then insert(patterns, "((\\+)%^)") insert(patterns, "((%^))") end -- Ensure any whitespace at the beginning and end is temp substituted, to prevent it from being accidentally -- trimmed. We only want to trim any final spaces added during the substitution process (e.g. by a module), which -- means we only do this during the first round of temp substitutions. if not noTrim then insert(patterns, "^([\128-\191\244]*(%s+))") insert(patterns, "((%s+)[\128-\191\244]*)$") end -- Pre-substitution, of "[[" and "]]", which makes pattern matching more accurate. text = gsub(text, "%f[%[]%[%[", "\1"):gsub("%f[%]]%]%]", "\2") local i = #subbedChars for _, pattern in ipairs(patterns) do -- Patterns ending in \0 stand are for things like "[[" or "]]"), so the inserted PUA are treated as breaks -- between terms by modules that scrape info from pages. local term_divider pattern = gsub(pattern, "%z$", function(divider) term_divider = divider == "\0" return "" end) text = gsub(text, pattern, function(...) local m = {...} local m1New = m[1] for k = 2, #m do local n = i + k - 1 subbedChars[n] = m[k] local byte2 = floor(n / 4096) % 64 + (term_divider and 128 or 136) local byte3 = floor(n / 64) % 64 + 128 local byte4 = n % 64 + 128 m1New = gsub(m1New, pattern_escape(m[k]), "\244" .. char(byte2) .. char(byte3) .. char(byte4), 1) end i = i + #m - 1 return m1New end) end text = gsub(text, "\1", "%[%["):gsub("\2", "%]%]") return text, subbedChars end -- Reinsert any formatting that was temporarily substituted. local function undoTempSubstitutions(text, subbedChars) for i = 1, #subbedChars do local byte2 = floor(i / 4096) % 64 + 128 local byte3 = floor(i / 64) % 64 + 128 local byte4 = i % 64 + 128 text = gsub(text, "\244[" .. char(byte2) .. char(byte2+8) .. "]" .. char(byte3) .. char(byte4), replacement_escape(subbedChars[i])) end text = gsub(text, "\1", "%[%["):gsub("\2", "%]%]") return text end -- Check if the raw text is an unsupported title, and if so return that. Otherwise, remove HTML entities. We do the -- pre-conversion to avoid loading the unsupported title list unnecessarily. local function checkNoEntities(text) local textNoEnc = decode_entities(text) if textNoEnc ~= text and load_data(links_data_module).unsupported_titles[text] then return text else return textNoEnc end end -- If no script object is provided (or if it's invalid or None), get one. local function checkScript(text, self, sc) if not check_object("script", true, sc) or sc:getCode() == "None" then return self:findBestScript(text) end return sc end local function normalize(text, sc) text = sc:fixDiscouragedSequences(text) return sc:toFixedNFD(text) end --[=[ Subfunction of iterateSectionSubstitutions(). Process an individual chunk of text according to the specifications in `substitution_data`. The input parameters are all as in the documentation of iterateSectionSubstitutions() except for `recursed`, which is set to true if we called ourselves recursively to process a script-specific setting or script-wide fallback. Returns two values: the processed text and the actual substitution data used to do the substitutions (same as the `actual_substitution_data` return value to iterateSectionSubstitutions()). ]=] local function doSubstitutions(self, text, sc, substitution_data, data_field, function_name, recursed) -- BE CAREFUL in this function because the value at any level can be `false`, which causes no processing to be done -- and blocks any further fallback processing. local actual_substitution_data = substitution_data -- If there are language-specific substitutes given in the data module, use those. if type(substitution_data) == "table" then -- If a script is specified, run this function with the script-specific data before continuing. local sc_code = sc:getCode() local has_substitution_data = false if substitution_data[sc_code] ~= nil then has_substitution_data = true if substitution_data[sc_code] then text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data[sc_code], data_field, function_name, true) end -- Hant, Hans and Hani are usually treated the same, so add a special case to avoid having to specify each one -- separately. elseif sc_code:match("^Han") and substitution_data.Hani ~= nil then has_substitution_data = true if substitution_data.Hani then text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data.Hani, data_field, function_name, true) end -- Substitution data with key 1 in the outer table may be given as a fallback. elseif substitution_data[1] ~= nil then has_substitution_data = true if substitution_data[1] then text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data[1], data_field, function_name, true) end end -- Iterate over all strings in the "from" subtable, and gsub with the corresponding string in "to". We work with -- the NFD decomposed forms, as this simplifies many substitutions. if substitution_data.from then has_substitution_data = true for i, from in ipairs(substitution_data.from) do -- Normalize each loop, to ensure multi-stage substitutions work correctly. text = sc:toFixedNFD(text) text = ugsub(text, sc:toFixedNFD(from), substitution_data.to[i] or "") end end if substitution_data.remove_diacritics then has_substitution_data = true text = sc:toFixedNFD(text) -- Convert exceptions to PUA. local remove_exceptions, substitutes = substitution_data.remove_exceptions if remove_exceptions then substitutes = {} local i = 0 for _, exception in ipairs(remove_exceptions) do exception = sc:toFixedNFD(exception) text = ugsub(text, exception, function(m) i = i + 1 local subst = u(0x80000 + i) substitutes[subst] = m return subst end) end end -- Strip diacritics. text = ugsub(text, "[" .. substitution_data.remove_diacritics .. "]", "") -- Convert exceptions back. if remove_exceptions then text = text:gsub("\242[\128-\191]*", substitutes) end end if not has_substitution_data and sc._data[data_field] then -- If language-specific sort key (etc.) is nil, fall back to script-wide sort key (etc.). text, actual_substitution_data = doSubstitutions(self, text, sc, sc._data[data_field], data_field, function_name, true) end elseif type(substitution_data) == "string" then -- If there is a dedicated function module, use that. local module = safe_require("Modul:" .. substitution_data) if module then -- TODO: translit functions should take objects, not codes. -- TODO: translit functions should be called with form NFD. if function_name == "tr" then if not module[function_name] then error(("Internal error: Module [[%s]] has no function named 'tr'"):format(substitution_data)) end text = module[function_name](text, self._code, sc:getCode()) elseif function_name == "stripDiacritics" then -- FIXME, get rid of this arm after renaming makeEntryName -> stripDiacritics. if module[function_name] then text = module[function_name](sc:toFixedNFD(text), self, sc) elseif module.makeEntryName then text = module.makeEntryName(sc:toFixedNFD(text), self, sc) else error(("Internal error: Module [[%s]] has no function named 'stripDiacritics' or 'makeEntryName'" ):format(substitution_data)) end else if not module[function_name] then error(("Internal error: Module [[%s]] has no function named '%s'"):format( substitution_data, function_name)) end text = module[function_name](sc:toFixedNFD(text), self, sc) end else error("Substitution data '" .. substitution_data .. "' does not match an existing module.") end elseif substitution_data == nil and sc._data[data_field] then -- If language-specific sort key (etc.) is nil, fall back to script-wide sort key (etc.). text, actual_substitution_data = doSubstitutions(self, text, sc, sc._data[data_field], data_field, function_name, true) end -- Don't normalize to NFC if this is the inner loop or if a module returned nil. if recursed or not text then return text, actual_substitution_data end -- Fix any discouraged sequences created during the substitution process, and normalize into the final form. return sc:toFixedNFC(sc:fixDiscouragedSequences(text)), actual_substitution_data end --[=[ Split the text into sections, based on the presence of temporarily substituted formatting characters, then iterate over each section to apply substitutions (e.g. transliteration or diacritic stripping). This avoids putting PUA (Private Use Area) characters through language-specific modules, which may be unequipped for them. This function is passed the following values: * `self` (the Language object); * `text` (the text to process); * `sc` (the script of the text, which must be specified; callers should call checkScript() as needed to autodetect the script of the text if not given explicitly by the user); * `subbedChars` (an array of the same length as the text, indicating which characters have been substituted and by what, or {nil} if no substitutions are to happen); * `keepCarets` (DOCUMENT ME); * `substitution_data` (the data indicating which substitutions to apply, taken directly from `data_field` in the language's data structure in a submodule of [[Modul:languages/data]]); * `data_field` (the field from which `substitution_data` was fetched, such as {"sort_key"} or {"strip_diacritics"}); * `function_name` (the name of the function to call to do the substitution, in case `substitution_data` specifies a module to do the substitution); * `notrim` (don't trim whitespace at the edges of `text`; set when computing the sort key, because whitespace at the beginning of a sort key is significant and causes the resulting page to be sorted at the beginning of the category it's in). Return three values: # the processed text; # the value of `subbedChars` that was passed in, possibly modified with additional character substitutions; will be {nil} if {nil} was passed in; # the actual substitution data that was used to apply substitutions to `text`; this may be different from the value of `substitution_data` passed in if that value recursively specified script-specific substitutions or if no substitution data could be found in the language-specific data (e.g. {nil} was passed in or a structure was passed in that had no setting for the script given in `sc`), but a script-wide fallback value was set; currently it is only used by {makeSortKey()}. ]=] local function iterateSectionSubstitutions(self, text, sc, subbedChars, keepCarets, substitution_data, data_field, function_name, notrim) local sections -- See [[Modul:languages/data]]. if not find(text, "\244") or load_data(languages_data_module).substitution[self._code] == "cont" then sections = {text} else sections = split(text, "\244[\128-\143][\128-\191]*", true) end local actual_substitution_data for _, section in ipairs(sections) do -- Don't bother processing empty strings or whitespace (which may also not be handled well by dedicated -- modules). if gsub(section, "%s+", "") ~= "" then local sub, this_actual_substitution_data = doSubstitutions(self, section, sc, substitution_data, data_field, function_name) actual_substitution_data = this_actual_substitution_data -- Second round of temporary substitutions, in case any formatting was added by the main substitution -- process. However, don't do this if the section contains formatting already (as it would have had to have -- been escaped to reach this stage, and therefore should be given as raw text). if sub and subbedChars then local noSub for _, pattern in ipairs(require(languages_data_patterns_module)) do if match(section, pattern .. "%z?") then noSub = true end end if not noSub then sub, subbedChars = doTempSubstitutions(sub, subbedChars, keepCarets, true) end end if not sub then text = sub break end text = sub and gsub(text, pattern_escape(section), replacement_escape(sub), 1) or text end end if not notrim then -- Trim, unless there are only spacing characters, while ignoring any final formatting characters. -- Do not trim sort keys because spaces at the beginning are significant. text = text and text:gsub("^([\128-\191\244]*)%s+(%S)", "%1%2"):gsub("(%S)%s+([\128-\191\244]*)$", "%1%2") or nil end return text, subbedChars, actual_substitution_data end -- Process carets (and any escapes). Default to simple removal, if no pattern/replacement is given. local function processCarets(text, pattern, repl) local rep repeat text, rep = gsub(text, "\\\\(\\*^)", "\3%1") until rep == 0 return (text:gsub("\\^", "\4") :gsub(pattern or "%^", repl or "") :gsub("\3", "\\") :gsub("\4", "^")) end -- Remove carets if they are used to capitalize parts of transliterations (unless they have been escaped). local function removeCarets(text, sc) if not sc:hasCapitalization() and sc:isTransliterated() and text:find("^", 1, true) then return processCarets(text) else return text end end local Language = {} --[==[ Return the language code of the language. Example: {fr"} for French. ]==] function Language:getCode() return self._code end --[==[ Return the canonical name of the language. This is the name used to represent that language on Wiktionary, and is guaranteed to be unique to that language alone. Example: {"French"} for French. ]==] function Language:getCanonicalName() local name = self._name if name == nil then name = self._data[1] self._name = name end return name end --[==[ Return the display form of the language. The display form of a language, family or script is the form it takes when appearing as the <code><var>source</var></code> in categories such as <code>English terms derived from <var>source</var></code> or <code>English given names from <var>source</var></code>, and is also the displayed text in `##makeCategoryLink()` links. For full and etymology-only languages, this is the same as the canonical name, but for families, it reads <code>"<var>name</var> languages"</code> (e.g. {"Indo-Iranian languages"}), and for scripts, it reads <code>"<var>name</var> script"</code> (e.g. {"Arabic script"}). ]==] function Language:getDisplayForm() local form = self._displayForm if form == nil then form = self:getCanonicalName() -- Add article and " substrate" to substrates that lack them. if self:getFamilyCode() == "qfa-sub" then if not (sub(form, 1, 4) == "the " or sub(form, 1, 2) == "a ") then form = "a " .. form end if not match(form, " [Ss]ubstrate") then form = form .. " substrate" end end self._displayForm = form end return form end --[==[ Return the value which should be used in the HTML `lang=` attribute for tagged text in the language. ]==] function Language:getHTMLAttribute(sc, region) local code = self._code if not find(code, "-", 1, true) then return code .. "-" .. sc:getCode() .. (region and "-" .. region or "") end local parent = self:getParent() region = region or match(code, "%f[%u][%u-]+%f[%U]") if parent then return parent:getHTMLAttribute(sc, region) end -- TODO: ISO family codes can also be used. return "mis-" .. sc:getCode() .. (region and "-" .. region or "") end --[==[ Return a list of the aliases that the language is known by, excluding the canonical name. Aliases are synonyms for the language in question. The names are not guaranteed to be unique, in that sometimes more than one language is known by the same name. Example: { {"High German", "New High German", "Deutsch"}} for {{lnl|de}}. ]==] function Language:getAliases() self:loadInExtraData() return require(language_like_module).getAliases(self) end --[==[ Return a list of the known subvarieties of a given language, excluding subvarieties that have been given explicit etymology-only language codes. The names are not guaranteed to be unique, in that sometimes a given name refers to a subvariety of more than one language. Example: { {"Southern Aymara", "Central Aymara"}} for {{lnl|ay}}. Note that the returned value can have nested tables in it, when a subvariety goes by more than one name. Example: { {"North Azerbaijani", "South Azerbaijani", {"Afshar", "Afshari", "Afshar Azerbaijani", "Afchar"}, {"Qashqa'i", "Qashqai", "Kashkay"}, "Sonqor"}} for {{lnl|az}}. Here, for example, Afshar, Afshari, Afshar Azerbaijani and Afchar all refer to the same subvariety, whose preferred name is Afshar (the one listed first). To avoid a return value with nested tables in it, specify a non-{nil} value for the `flatten` parameter; in that case, the return value would be { {"North Azerbaijani", "South Azerbaijani", "Afshar", "Afshari", "Afshar Azerbaijani", "Afchar", "Qashqa'i", "Qashqai", "Kashkay", "Sonqor"}}. ]==] function Language:getVarieties(flatten) self:loadInExtraData() return require(language_like_module).getVarieties(self, flatten) end --[==[ Return a table of the "other names" that the language is known by, which are listed in the `other_names` field in the language's extra-data module. It should be noted that the `other_names` field itself is deprecated, and entries listed there should eventually be moved to either `aliases` or `varieties`, or removed if they refer to larger entities which the language in question is actually a part of. ]==] function Language:getOtherNames() -- To be eventually removed, once there are no more uses of the `other_names` field. self:loadInExtraData() return require(language_like_module).getOtherNames(self) end --[==[ Return a combined table of the canonical name, aliases, varieties and other names of a given language. ]==] function Language:getAllNames() self:loadInExtraData() return require(language_like_module).getAllNames(self) end --[==[ Return a table of types as a set (with the types as keys). The possible types are * {language}: This is a language, either full or etymology-only. * {full}: This is a "full" (not etymology-only) language, i.e. the union of {regular}, {reconstructed} and {appendix-constructed}. Note that the types {full} and {etymology-only} also exist for families, so if you want to check specifically for a full language and you have an object that might be a family, you should use {hasType("language", "full")} and not simply {hasType("full")}. * {etymology-only}: This is an etymology-only (not full) language, whose parent is another etymology-only language or a full language. Note that the types {full} and {etymology-only} also exist for families, so if you want to check specifically for an etymology-only language and you have an object that might be a family, you should use {hasType("language", "etymology-only")} and not simply {hasType("etymology-only")}. * {regular}: This indicates a full language that is attested according to [[WT:CFI]] and therefore permitted in the main namespace. There may also be reconstructed terms for the language, which are placed in the {Reconstruction} namespace and must be prefixed with `*` to indicate a reconstruction. Most full languages are natural (not constructed) languages, but a few constructed languages (e.g. Esperanto and Volapük, among others) are also allowed in the mainspace and considered regular languages. * {reconstructed}: This language is not attested according to [[WT:CFI]], and therefore is allowed only in the {Reconstruction} namespace. All terms in this language are reconstructed, and must be prefixed with `*`. Languages such as Proto-Indo-European and Proto-Germanic are in this category. '''Exception:''' Reconstructed languages can have entries in the mainspace using the ''anti-asterisk'' feature. This requires that both the headword for the mainspace term (as specified using {{para|head}}) and links to the term are prefixed with the anti-asterisk indicator `!!`. * {appendix-constructed}: This language is attested but does not meet the additional requirements set out for constructed languages ([[WT:CFI#Constructed languages]]). Its entries must therefore be in the Appendix namespace, but they are not reconstructed and therefore should not have * prefixed in links. ]==] function Language:getTypes() local types = self._types if types == nil then types = {language = true} if self:getFullCode() == self._code then types.full = true else types["etymology-only"] = true end for t in gmatch(self._data.type, "[^,]+") do types[t] = true end self._types = types end return types end --[==[ Given a list of types as strings, return {true} if the language has all of them. ]==] function Language:hasType(...) Language.hasType = require(language_like_module).hasType return self:hasType(...) end --[==[ Return a list containing `WikimediaLanguage` objects (see [[Modul:wikimedia languages]]), which represent languages and their codes as they are used in Wikimedia projects for interwiki linking and such. More than one object may be returned, as a single Wiktionary language may correspond to multiple Wikimedia languages. For example, Wiktionary's single code `sh` (Serbo-Croatian) maps to four Wikimedia codes: `sh` (Serbo-Croatian), `bs` (Bosnian), `hr` (Croatian) and `sr` (Serbian). The code for the Wikimedia language is retrieved from the `wikimedia_codes` property in the data modules. If that property is not present, the code of the current language is used. If none of the available codes is actually a valid Wikimedia code, an empty list is returned. ]==] function Language:getWikimediaLanguages() local wm_langs = self._wikimediaLanguageObjects if wm_langs == nil then local codes = self:getWikimediaLanguageCodes() wm_langs = {} for i = 1, #codes do wm_langs[i] = get_wikimedia_lang(codes[i]) end self._wikimediaLanguageObjects = wm_langs end return wm_langs end function Language:getWikimediaLanguageCodes() local wm_langs = self._wikimediaLanguageCodes if wm_langs == nil then wm_langs = self._data.wikimedia_codes if wm_langs then wm_langs = split(wm_langs, ",", true, true) else local code = self._code if is_known_language_tag(code) then wm_langs = {code} else -- Inherit, but only if no codes are specified in the data *and* -- the language code isn't a valid Wikimedia language code. local parent = self:getParent() wm_langs = parent and parent:getWikimediaLanguageCodes() or {} end end self._wikimediaLanguageCodes = wm_langs end return wm_langs end --[==[ Return the name of the Wikipedia article for the language. `project` specifies the language and project to retrieve the article from, defaulting to {"enwiki"} for the English Wikipedia. Normally if specified it should be the project code for a specific-language Wikipedia e.g. "zhwiki" for the Chinese Wikipedia, but it can be any project, including non-Wikipedia ones. If the project is the English Wikipedia and the property {wikipedia_article} is present in the data module it will be used first. In all other cases, a sitelink will be generated from {:getWikidataItem} (if set). The resulting value (or lack of value) is cached so that subsequent calls are fast. If no value could be determined, and `noCategoryFallback` is {false}, `##Language:getCategoryName` is used as fallback; otherwise, {nil} is returned. Note that if `noCategoryFallback` is {nil} or omitted, it defaults to {false} if the project is the English Wikipedia, otherwise to {true}. In other words, under normal circumstances, if the English Wikipedia article couldn't be retrieved, the return value will fall back to a link to the language's category, but this won't normally happen for any other project. ]==] function Language:getWikipediaArticle(noCategoryFallback, project) Language.getWikipediaArticle = require(language_like_module).getWikipediaArticle return self:getWikipediaArticle(noCategoryFallback, project) end function Language:makeWikipediaLink() return make_link(self, "w:" .. self:getWikipediaArticle(), self:getCanonicalName()) end --[==[ Return the name of the Wikimedia Commons category page for the language. ]==] function Language:getCommonsCategory() Language.getCommonsCategory = require(language_like_module).getCommonsCategory return self:getCommonsCategory() end --[==[ Return the Wikidata item ID for the language or {nil}. This corresponds to the the second field in the data modules. ]==] function Language:getWikidataItem() Language.getWikidataItem = require(language_like_module).getWikidataItem return self:getWikidataItem() end --[==[ Return a list of `Script` objects for all scripts that the language is written in. See [[Modul:scripts]]. ]==] function Language:getScripts() local scripts = self._scriptObjects if scripts == nil then local codes = self:getScriptCodes() if codes[1] == "All" then scripts = load_data(scripts_data_module) else scripts = {} for i = 1, #codes do scripts[i] = get_script(codes[i]) end end self._scriptObjects = scripts end return scripts end --[==[ Return the list of script codes in the language's data file. ]==] function Language:getScriptCodes() local scripts = self._scriptCodes if scripts == nil then scripts = self._data[4] if scripts then local codes, n = {}, 0 for code in gmatch(scripts, "[^,]+") do n = n + 1 -- Special handling of "Hants", which represents "Hani", "Hant" and "Hans" collectively. if code == "Hants" then codes[n] = "Hani" codes[n + 1] = "Hant" codes[n + 2] = "Hans" n = n + 2 else codes[n] = code end end scripts = codes else scripts = {"None"} end self._scriptCodes = scripts end return scripts end --[==[ Given some text, iterate through the scripts of a given language trying to find the script that best matches the text. Return a {Script} object representing the script. If no match is found at all, it returns the {None} script object. ]==] function Language:findBestScript(text, forceDetect) if not text or text == "" or text == "-" then return get_script("None") end -- Differs from table returned by getScriptCodes, as Hants is not normalized into its constituents. local codes = self._bestScriptCodes if codes == nil then codes = self._data[4] codes = codes and split(codes, ",", true, true) or {"None"} self._bestScriptCodes = codes end local first_sc = codes[1] if first_sc == "All" then return find_best_script_without_lang(text) end local codes_len = #codes if not (forceDetect or first_sc == "Hants" or codes_len > 1) then first_sc = get_script(first_sc) local charset = first_sc.characters return charset and umatch(text, "[" .. charset .. "]") and first_sc or get_script("None") end -- Remove all formatting characters. text = get_plaintext(text) -- Remove all spaces and any ASCII punctuation. Some non-ASCII punctuation is script-specific, so can't be removed. text = ugsub(text, "[%s!\"#%%&'()*,%-./:;?@[\\%]_{}]+", "") if #text == 0 then return get_script("None") end -- Try to match every script against the text, -- and return the one with the most matching characters. local bestcount, bestscript, length = 0 for i = 1, codes_len do local sc = codes[i] -- Special case for "Hants", which is a special code that represents whichever of "Hant" or "Hans" best matches, -- or "Hani" if they match equally. This avoids having to list all three. In addition, "Hants" will be treated -- as the best match if there is at least one matching character, under the assumption that a Han script is -- desirable in terms that contain a mix of Han and other scripts (not counting those which use Jpan or Kore). if sc == "Hants" then local Hani = get_script("Hani") if not Hant_chars then Hant_chars = load_data("Modul:zh/data/ts") Hans_chars = load_data("Modul:zh/data/st") end local t, s, found = 0, 0 -- This is faster than using mw.ustring.gmatch directly. for ch in gmatch((ugsub(text, "[" .. Hani.characters .. "]", "\255%0")), "\255(.[\128-\191]*)") do found = true if Hant_chars[ch] then t = t + 1 if Hans_chars[ch] then s = s + 1 end elseif Hans_chars[ch] then s = s + 1 else t, s = t + 1, s + 1 end end if found then if t == s then return Hani end return get_script(t > s and "Hant" or "Hans") end else sc = get_script(sc) if not length then length = ulen(text) end -- Count characters by removing everything in the script's charset and comparing to the original length. local charset = sc.characters local count = charset and length - ulen((ugsub(text, "[" .. charset .. "]+", ""))) or 0 if count >= length then return sc elseif count > bestcount then bestcount = count bestscript = sc end end end -- Return best matching script, or otherwise None. return bestscript or get_script("None") end --[==[ Return a `Family` object for the language family that the language belongs to. See [[Modul:families]]. ]==] function Language:getFamily() local family = self._familyObject if family == nil then family = self:getFamilyCode() -- If the value is nil, it's cached as false. family = family and get_family(family) or false self._familyObject = family end return family or nil end --[==[ Return the family code in the language's data file. ]==] function Language:getFamilyCode() local family = self._familyCode if family == nil then -- If the value is nil, it's cached as false. family = self._data[3] or false self._familyCode = family end return family or nil end function Language:getFamilyName() local family = self._familyName if family == nil then family = self:getFamily() -- If the value is nil, it's cached as false. family = family and family:getCanonicalName() or false self._familyName = family end return family or nil end do local function check_family(self, family) if type(family) == "table" then family = family:getCode() end if self:getFamilyCode() == family then return true end local self_family = self:getFamily() if self_family:inFamily(family) then return true -- If the family isn't a real family (e.g. creoles) check any ancestors. elseif self_family:inFamily("qfa-not") then local ancestors = self:getAncestors() for _, ancestor in ipairs(ancestors) do if ancestor:inFamily(family) then return true end end end end --[==[ Check whether the language belongs to `family` (which can be a family code or object). A list of objects can be given in place of `family`; in that case, return true if the language belongs to any of the specified families. Note that some languages (in particular, certain creoles) can have multiple immediate ancestors potentially belonging to different families; in that case, return true if the language belongs to any of the specified families. ]==] function Language:inFamily(...) if self:getFamilyCode() == nil then return false end return check_inputs(self, check_family, false, ...) end end function Language:getParent() local parent = self._parentObject if parent == nil then parent = self:getParentCode() -- If the value is nil, it's cached as false. parent = parent and get_by_code(parent, nil, true, true) or false self._parentObject = parent end return parent or nil end function Language:getParentCode() local parent = self._parentCode if parent == nil then -- If the value is nil, it's cached as false. parent = self._data.parent or false self._parentCode = parent end return parent or nil end function Language:getParentName() local parent = self._parentName if parent == nil then parent = self:getParent() -- If the value is nil, it's cached as false. parent = parent and parent:getCanonicalName() or false self._parentName = parent end return parent or nil end function Language:getParentChain() local chain = self._parentChain if chain == nil then chain = {} local parent, n = self:getParent(), 0 while parent do n = n + 1 chain[n] = parent parent = parent:getParent() end self._parentChain = chain end return chain end do local function check_lang(self, lang) for _, parent in ipairs(self:getParentChain()) do if (type(lang) == "string" and lang or lang:getCode()) == parent:getCode() then return true end end end function Language:hasParent(...) return check_inputs(self, check_lang, false, ...) end end --[==[ If the language is etymology-only, iterate through parents until a full language or family is found, and the corresponding object is returned. If the language is a full language, return that language. ]==] function Language:getFull() local full = self._fullObject if full == nil then full = self:getFullCode() full = full == self._code and self or get_by_code(full) self._fullObject = full end return full end --[==[ If the language is an etymology-only language, iterate through parents until a full language or family is found, and the corresponding code is returned. If the language is a full language, return that language's code. ]==] function Language:getFullCode() return self._fullCode or self._code end --[==[ If the language is an etymology-only language, iterate through parents until a full language or family is found, and the corresponding canonical name is returned. If the language is a full language, return that language's canonical name. ]==] function Language:getFullName() local full = self._fullName if full == nil then full = self:getFull():getCanonicalName() self._fullName = full end return full end --[==[ Return a list of {Language} objects for all languages that this language is directly descended from. Generally this is only a single language, but creoles, pidgins and mixed languages can have multiple ancestors. ]==] function Language:getAncestors() local ancestors = self._ancestorObjects if ancestors == nil then ancestors = {} local ancestor_codes = self:getAncestorCodes() if #ancestor_codes > 0 then for _, ancestor in ipairs(ancestor_codes) do insert(ancestors, get_by_code(ancestor, nil, true)) end else local fam = self:getFamily() local protoLang = fam and fam:getProtoLanguage() or nil -- For the cases where the current language is the proto-language -- of its family, or an etymology-only language that is ancestral to that -- proto-language, we need to step up a level higher right from the -- start. if protoLang and ( protoLang:getCode() == self._code or (self:hasType("etymology-only") and protoLang:hasAncestor(self)) ) then fam = fam:getFamily() protoLang = fam and fam:getProtoLanguage() or nil end while not protoLang and not (not fam or fam:getCode() == "qfa-not") do fam = fam:getFamily() protoLang = fam and fam:getProtoLanguage() or nil end insert(ancestors, protoLang) end self._ancestorObjects = ancestors end return ancestors end do -- Avoid a language being its own ancestor via class inheritance. We only need to check for this if the language has -- inherited an ancestor table from its parent, because we never want to drop ancestors that have been explicitly -- set in the data. Recursively iterate over ancestors until we either find self or run out. If self is found, -- return true. local function check_ancestor(self, lang) local codes = lang:getAncestorCodes() if not codes then return nil end for i = 1, #codes do local code = codes[i] if code == self._code then return true end local anc = get_by_code(code, nil, true) if check_ancestor(self, anc) then return true end end end --[==[ Return a list of `Language` codes for all languages that this language is directly descended from. Generally this is only a single language, but creoles, pidgins and mixed languages can have multiple ancestors. ]==] function Language:getAncestorCodes() if self._ancestorCodes then return self._ancestorCodes end local data = self._data local codes = data.ancestors if codes == nil then codes = {} self._ancestorCodes = codes return codes end codes = split(codes, ",", true, true) self._ancestorCodes = codes -- If there are no codes or the ancestors weren't inherited data, there's nothing left to check. if #codes == 0 or self:getData(false, "raw").ancestors ~= nil then return codes end local i = 1 while i <= #codes do if check_ancestor(self, self) then remove(codes, i) else i = i + 1 end end return codes end end --[==[ Given a list of language objects or codes, return true if at least one of them is an ancestor. This includes any etymology-only children of that ancestor. If the language's ancestor(s) are etymology-only languages, it will also return true for those language parent(s) (e.g. if Vulgar Latin is the ancestor, it will also return true for its parent, Latin). However, a parent is excluded from this if the ancestor is also ancestral to that parent (e.g. if Classical Persian is the ancestor, Persian would return {false}, because Classical Persian is also ancestral to Persian). ]==] function Language:hasAncestor(...) local function iterateOverAncestorTree(node, func, parent_check) local ancestors = node:getAncestors() local ancestorsParents = {} for _, ancestor in ipairs(ancestors) do -- When checking the parents of the other language, and the ancestor is also a parent, skip to the next -- ancestor, so that we exclude any etymology-only children of that parent that are not directly related -- (see below). local ret = (parent_check or not node:hasParent(ancestor)) and func(ancestor) or iterateOverAncestorTree(ancestor, func, parent_check) if ret then return ret end end -- Check the parents of any ancestors. We don't do this if checking the parents of the other language, so that -- we exclude any etymology-only children of those parents that are not directly related (e.g. if the ancestor -- is Vulgar Latin and we are checking New Latin, we want it to return false because they are on different -- ancestral branches. As such, if we're already checking the parent of New Latin (Latin) we don't want to -- compare it to the parent of the ancestor (Latin), as this would be a false positive; it should be one or the -- other). if not parent_check then return nil end for _, ancestor in ipairs(ancestors) do local ancestorParents = ancestor:getParentChain() for _, ancestorParent in ipairs(ancestorParents) do if ancestorParent:getCode() == self._code or ancestorParent:hasAncestor(ancestor) then break else insert(ancestorsParents, ancestorParent) end end end for _, ancestorParent in ipairs(ancestorsParents) do local ret = func(ancestorParent) if ret then return ret end end end local function do_iteration(otherlang, parent_check) -- otherlang can't be self if (type(otherlang) == "string" and otherlang or otherlang:getCode()) == self._code then return false end repeat if iterateOverAncestorTree( self, function(ancestor) return ancestor:getCode() == (type(otherlang) == "string" and otherlang or otherlang:getCode()) end, parent_check ) then return true elseif type(otherlang) == "string" then otherlang = get_by_code(otherlang, nil, true) end otherlang = otherlang:getParent() parent_check = false until not otherlang end local parent_check = true for _, otherlang in ipairs{...} do local ret = do_iteration(otherlang, parent_check) if ret then return true end end return false end do local function construct_node(lang, memo) local branch, ancestors = {lang = lang:getCode()} memo[lang:getCode()] = branch for _, ancestor in ipairs(lang:getAncestors()) do if ancestors == nil then ancestors = {} end insert(ancestors, memo[ancestor:getCode()] or construct_node(ancestor, memo)) end branch.ancestors = ancestors return branch end function Language:getAncestorChain() local chain = self._ancestorChain if chain == nil then chain = construct_node(self, {}) self._ancestorChain = chain end return chain end end function Language:getAncestorChainOld() local chain = self._ancestorChain if chain == nil then chain = {} local step = self while true do local ancestors = step:getAncestors() step = #ancestors == 1 and ancestors[1] or nil if not step then break end insert(chain, step) end self._ancestorChain = chain end return chain end local function fetch_descendants(self, fmt) local descendants, family = {}, self:getFamily() -- Iterate over all three datasets. for _, data in ipairs{ require("Modul:languages/code to canonical name"), require("Modul:etymology languages/code to canonical name"), require("Modul:families/code to canonical name"), } do for code in pairs(data) do local lang = get_by_code(code, nil, true, true) if not lang then error(("Internal error: code '%s' can't be fetched"):format(code)) end -- Test for a descendant. Earlier tests weed out most candidates, while the more intensive tests are only used sparingly. if ( code ~= self._code and -- Not self. lang:inFamily(family) and -- In the same family. ( family:getProtoLanguageCode() == self._code or -- Self is the protolanguage. self:hasDescendant(lang) or -- Full hasDescendant check. (lang:getFullCode() == self._code and not self:hasAncestor(lang)) -- Etymology-only child which isn't an ancestor. ) ) then if fmt == "object" then insert(descendants, lang) elseif fmt == "code" then insert(descendants, code) elseif fmt == "name" then insert(descendants, lang:getCanonicalName()) end end end end return descendants end function Language:getDescendants() local descendants = self._descendantObjects if descendants == nil then descendants = fetch_descendants(self, "object") self._descendantObjects = descendants end return descendants end function Language:getDescendantCodes() local descendants = self._descendantCodes if descendants == nil then descendants = fetch_descendants(self, "code") self._descendantCodes = descendants end return descendants end function Language:getDescendantNames() local descendants = self._descendantNames if descendants == nil then descendants = fetch_descendants(self, "name") self._descendantNames = descendants end return descendants end do local function check_lang(self, lang) if type(lang) == "string" then lang = get_by_code(lang, nil, true) end if lang:hasAncestor(self) then return true end end function Language:hasDescendant(...) return check_inputs(self, check_lang, false, ...) end end local function fetch_children(self, fmt) local m_etym_data = require(etymology_languages_data_module) local self_code, children = self._code, {} for code, lang in pairs(m_etym_data) do local _lang = lang repeat local parent = _lang.parent if parent == self_code then if fmt == "object" then insert(children, get_by_code(code, nil, true)) elseif fmt == "code" then insert(children, code) elseif fmt == "name" then insert(children, lang[1]) end break end _lang = m_etym_data[parent] until not _lang end return children end function Language:getChildren() local children = self._childObjects if children == nil then children = fetch_children(self, "object") self._childObjects = children end return children end function Language:getChildrenCodes() local children = self._childCodes if children == nil then children = fetch_children(self, "code") self._childCodes = children end return children end function Language:getChildrenNames() local children = self._childNames if children == nil then children = fetch_children(self, "name") self._childNames = children end return children end function Language:hasChild(...) local lang = ... if not lang then return false elseif type(lang) == "string" then lang = get_by_code(lang, nil, true) end if lang:hasParent(self) then return true end return self:hasChild(select(2, ...)) end --[==[ Return the name of the main category of that language. Example: {French language"} for French, whose category is at [[:Category:French language]]. Unless optional argument `nocap` is given, the language name at the beginning of the returned value will be capitalized. This capitalization is correct for category names, but not if the language name is lowercase and the returned value of this function is used in the middle of a sentence. ]==] function Language:getCategoryName(nocap) local name = self._categoryName if name == nil then name = self:getCanonicalName() -- If a substrate, omit any leading article. if self:getFamilyCode() == "qfa-sub" then name = name:gsub("^sebuah ", ""):gsub("^suatu ", "") end -- Only add " language" if a full language. if self:hasType("full") then -- Unless the canonical name already ends with "language", "lect" or their derivatives, add " language". if not (match(name, "^[Bb]ahasa") or match(name, "^[Ll]ek")) then name = "Bahasa " .. name end end self._categoryName = name end if nocap then return name end return mw.getContentLanguage():ucfirst(name) end --[==[ Create a link to the category; the link text is the canonical name. ]==] function Language:makeCategoryLink() return make_link(self, ":Kategori:" .. self:getCategoryName(), self:getDisplayForm()) end function Language:getStandardCharacters(sc) local standard_chars = self._data.standard_chars if type(standard_chars) ~= "table" then return standard_chars elseif sc and type(sc) ~= "string" then check_object("script", nil, sc) sc = sc:getCode() end if (not sc) or sc == "None" then local scripts = {} for _, script in pairs(standard_chars) do insert(scripts, script) end return concat(scripts) end if standard_chars[sc] then return standard_chars[sc] .. (standard_chars[1] or "") end end --[==[ Strip diacritics from display text `text` (in a language-specific fashion), which is in the script `sc`. If `sc` is omitted or {nil}, the script is autodetected. This also strips certain punctuation characters from the end and (in the case of Spanish upside-down question mark and exclamation points) from the beginning; strips any whitespace at the end of the text or between the text and final stripped punctuation characters; and applies some language-specific Unicode normalizations to replace discouraged characters with their prescribed alternatives. Return the stripped text. ]==] function Language:stripDiacritics(text, sc) if (not text) or text == "" then return text end sc = checkScript(text, self, sc) text = normalize(text, sc) -- FIXME, rename makeEntryName to stripDiacritics and get rid of second and third return values -- everywhere local _ text, _, _ = iterateSectionSubstitutions(self, text, sc, nil, nil, self._data.strip_diacritics or self._data.entry_name, "strip_diacritics", "stripDiacritics") text = umatch(text, "^[¿¡]?(.-[^%s%p].-)%s*[؟?!;՛՜ ՞ ՟?!︖︕।॥။၊་།]?$") or text return text end --[==[ Convert a ''logical'' pagename (the pagename as it appears to the user, after diacritics and punctuation have been stripped) to a ''physical'' pagename (the pagename as it appears in the MediaWiki database). Reasons for a difference between the two are (a) unsupported titles such as `[ ]` (with square brackets in them), `#` (pound/hash sign) and `¯\_(ツ)_/¯` (with underscores), as well as overly long titles of various sorts; (b) "mammoth" pages that are split into parts (e.g. `a`, which is split into physical pagenames `a/languages A to L` and `a/languages M to Z`). For almost all purposes, you should work with logical and not physical pagenames. But there are certain use cases that require physical pagenames, such as checking the existence of a page or retrieving a page's contents. `pagename` is the logical pagename to be converted. `is_reconstructed_or_appendix` indicates whether the page is in the `Reconstruction` or `Appendix` namespaces. If it is omitted or has the value {nil}, the pagename is checked for an initial asterisk, and if found, the page is assumed to be a `Reconstruction` page. Setting a value of `false` or `true` to `is_reconstructed_or_appendix` disables this check and allows for mainspace pagenames that begin with an asterisk. ]==] function Language:logicalToPhysical(pagename, is_reconstructed_or_appendix) -- FIXME: This probably shouldn't happen but it happens when makeEntryName() receives nil. if pagename == nil then track("nil-passed-to-logicalToPhysical") return nil end local initial_asterisk if is_reconstructed_or_appendix == nil then local pagename_minus_initial_asterisk initial_asterisk, pagename_minus_initial_asterisk = pagename:match("^(%*)(.*)$") if pagename_minus_initial_asterisk then is_reconstructed_or_appendix = true pagename = pagename_minus_initial_asterisk elseif self:hasType("appendix-constructed") then is_reconstructed_or_appendix = true end end if not is_reconstructed_or_appendix then -- Check if the pagename is a listed unsupported title. local unsupportedTitles = load_data(links_data_module).unsupported_titles if unsupportedTitles[pagename] then return "Tajuk tidak disokong/" .. unsupportedTitles[pagename] end end -- Set `unsupported` as true if certain conditions are met. local unsupported -- Check if there's an unsupported character. \239\191\189 is the replacement character U+FFFD, which can't be typed -- directly here due to an abuse filter. Unix-style dot-slash notation is also unsupported, as it is used for -- relative paths in links, as are 3 or more consecutive tildes. Note: match is faster with magic -- characters/charsets; find is faster with plaintext. if ( match(pagename, "[#<>%[%]_{|}]") or find(pagename, "\239\191\189") or match(pagename, "%f[^%z/]%.%.?%f[%z/]") or find(pagename, "~~~") ) then unsupported = true -- If it looks like an interwiki link. elseif find(pagename, ":") then local prefix = gsub(pagename, "^:*(.-):.*", ulower) if ( load_data("Modul:data/namespaces")[prefix] or load_data("Modul:data/interwikis")[prefix] ) then unsupported = true end end -- Escape unsupported characters so they can be used in titles. ` is used as a delimiter for this, so a raw use of -- it in an unsupported title is also escaped here to prevent interference; this is only done with unsupported -- titles, though, so inclusion won't in itself mean a title is treated as unsupported (which is why it's excluded -- from the earlier test). if unsupported then -- FIXME: This conversion needs to be different for reconstructed pages with unsupported characters. There -- aren't any currently, but if there ever are, we need to fix this e.g. to put them in something like -- Reconstruction:Proto-Indo-European/Unsupported titles/`lowbar``num`. local unsupported_characters = load_data(links_data_module).unsupported_characters pagename = pagename:gsub("[#<>%[%]_`{|}\239]\191?\189?", unsupported_characters) :gsub("%f[^%z/]%.%.?%f[%z/]", function(m) return (gsub(m, "%.", "`period`")) end) :gsub("~~~+", function(m) return (gsub(m, "~", "`tilde`")) end) pagename = "Tajuk tidak disokong/" .. pagename elseif not is_reconstructed_or_appendix then -- Check if this is a mammoth page. If so, which subpage should we link to? local m_links_data = load_data(links_data_module) local mammoth_page_type = m_links_data.mammoth_pages[pagename] if mammoth_page_type then local canonical_name = self:getFullName() if canonical_name ~= "Rentas bahasa" and canonical_name ~= "Bahasa Melayu" then local this_subpage local L2_sort_key = get_L2_sort_key(canonical_name) for _, subpage_spec in ipairs(m_links_data.mammoth_page_subpage_types[mammoth_page_type]) do -- unpack() fails utterly on data loaded using mw.loadData() even if offsets are given local subpage, pattern = subpage_spec[1], subpage_spec[2] if pattern == true or L2_sort_key:match(pattern) then this_subpage = subpage break end end if not this_subpage then error(("Internal error: Bad data in mammoth_page_subpage_pages in [[Modul:links/data]] for mammoth page %s, type %s; last entry didn't have 'true' in it"):format( pagename, mammoth_page_type)) end pagename = pagename .. "/" .. this_subpage end end end return (initial_asterisk or "") .. pagename end --[==[ Strip the diacritics from a display pagename and convert the resulting logical pagename into a physical pagename. This allows you, for example, to retrieve the contents of the page or check its existence. WARNING: This is deprecated and will be going away. It is a simple composition of `##Language:stripDiacritics` and `##Language:logicalToPhysical`; most callers only want the former, and if you need both, call them both yourself. `text` and `sc` are as in `##Language:stripDiacritics`, and `is_reconstructed_or_appendix` is as in `##Language:logicalToPhysical`. ]==] function Language:makeEntryName(text, sc, is_reconstructed_or_appendix) track("makeEntryName called") return self:logicalToPhysical(self:stripDiacritics(text, sc), is_reconstructed_or_appendix) end --[==[ Generate term alternants (e.g. simplified versions of traditional Chinese terms) using a language-specific method, and return them as a table. If the language provides no method for generating alternants, return a table containing only the input term. ]==] function Language:generateAlternants(text, sc) local generate_alternants = self._data.generate_alternants if generate_alternants == nil then return {text} end sc = checkScript(text, self, sc) return require("Modul:" .. generate_alternants).generateAlternants(text, self, sc) end --[==[ Create a sort key for the given stripped text, following the rules appropriate for the language. This removes diacritical marks from the stripped text if they are not considered significant for sorting, and may perform some other changes. Any initial hyphen is also removed, and anything in parentheses is removed as well. The `sort_key` setting for each language in the data modules defines the replacements made by this function, or it gives the name of the module that takes the stripped text and returns a sortkey. ]==] function Language:makeSortKey(text, sc) if (not text) or text == "" then return text end if match(text, "<[^<>]+>") then track("track HTML tag") end -- Remove directionality/control characters, bold, italics, soft hyphens, strip markers and HTML tags. -- FIXME: Partly duplicated with remove_formatting() in [[Modul:links]]. text = ugsub(text, "[\194\173\216\156\226\128\142\226\128\143\226\128\170-\226\128\174\226\129\166-\226\129\175]", "") text = text:gsub("('*)'''(.-'*)'''", "%1%2"):gsub("('*)''(.-'*)''", "%1%2") text = gsub(unstrip(text), "<[^<>]+>", "") text = decode_uri(text, "PATH") text = checkNoEntities(text) -- Remove initial hyphens and * unless the term only consists of spacing + punctuation characters. text = ugsub(text, "^([􀀀-􏿽]*)[-־ـ᠊*]+([􀀀-􏿽]*)(.*[^%s%p].*)", "%1%2%3") sc = checkScript(text, self, sc) text = normalize(text, sc) text = removeCarets(text, sc) -- For languages with dotted dotless i, ensure that "İ" is sorted as "i", and "I" is sorted as "ı". if self:hasDottedDotlessI() then text = gsub(text, "I\204\135", "i") -- decomposed "İ" :gsub("I", "ı") text = sc:toFixedNFD(text) end -- Convert to lowercase, make the sortkey, then convert to uppercase. Where the language has dotted dotless i, it is -- usually not necessary to convert "i" to "İ" and "ı" to "I" first, because "I" will always be interpreted as -- conventional "I" (not dotless "İ") by any sorting algorithms, which will have been taken into account by the -- sortkey substitutions themselves. However, if no sortkey substitutions have been specified, then conversion is -- necessary so as to prevent "i" and "ı" both being sorted as "I". -- -- An exception is made for scripts that (sometimes) sort by scraping page content, as that means they are sensitive -- to changes in capitalization (as it changes the target page). if not sc:sortByScraping() then text = ulower(text) end local actual_substitution_data, _ -- Don't trim whitespace here because it's significant at the beginning of a sort key or sort base. text, _, actual_substitution_data = iterateSectionSubstitutions(self, text, sc, nil, nil, self._data.sort_key, "sort_key", "makeSortKey", "notrim") if not sc:sortByScraping() then if self:hasDottedDotlessI() and not actual_substitution_data then text = text:gsub("ı", "I"):gsub("i", "İ") text = sc:toFixedNFC(text) end text = uupper(text) end -- Remove parentheses, as long as they are either preceded or followed by something. text = gsub(text, "(.)[()]+", "%1"):gsub("[()]+(.)", "%1") text = escape_risky_characters(text) return text end --[==[ Create the form used as as a basis for display text and transliteration. FIXME: Rename to correctInputText(). ]==] local function processDisplayText(text, self, sc, keepCarets, keepPrefixes) local subbedChars = {} text, subbedChars = doTempSubstitutions(text, subbedChars, keepCarets) text = decode_uri(text, "PATH") text = checkNoEntities(text) sc = checkScript(text, self, sc) text = normalize(text, sc) text, subbedChars = iterateSectionSubstitutions(self, text, sc, subbedChars, keepCarets, self._data.display_text, "display_text", "makeDisplayText") text = removeCarets(text, sc) -- Remove any interwiki link prefixes (unless they have been escaped or this has been disabled). if find(text, ":") and not keepPrefixes then local rep repeat text, rep = gsub(text, "\\\\(\\*:)", "\3%1") until rep == 0 text = gsub(text, "\\:", "\4") while true do local prefix = gsub(text, "^(.-):.+", function(m1) return (gsub(m1, "\244[\128-\191]*", "")) end) -- Check if the prefix is an interwiki, though ignore capitalised Wiktionary:, which is a namespace. if not prefix or prefix == text or prefix == "Wikikamus" or not (load_data("Modul:data/interwikis")[ulower(prefix)] or prefix == "") then break end text = gsub(text, "^(.-):(.*)", function(m1, m2) local ret = {} for subbedChar in gmatch(m1, "\244[\128-\191]*") do insert(ret, subbedChar) end return concat(ret) .. m2 end) end text = gsub(text, "\3", "\\"):gsub("\4", ":") end return text, subbedChars end --[==[ Make the display text (i.e. what is displayed on the page). ]==] function Language:makeDisplayText(text, sc, keepPrefixes) if not text or text == "" then return text end local subbedChars text, subbedChars = processDisplayText(text, self, sc, nil, keepPrefixes) text = escape_risky_characters(text) return undoTempSubstitutions(text, subbedChars) end --[==[ Transliterate the text from the given script into Latin script (see [[Wiktionary:Transliteration and romanization]]). The language must have the `translit` property for this to work; if it is not present, {nil} is returned. The `sc` parameter is handled by the transliteration module, and how it is handled is specific to that module. Some transliteration modules may tolerate {nil} as the script, others require it to be one of the possible scripts that the module can transliterate, and will throw an error if it's not one of them. For this reason, the `sc` parameter should always be provided when writing non-language-specific code. The `module_override` parameter is used to override the default module that is used to provide the transliteration. This is useful in cases where you need to demonstrate a particular module in use, but there is no default module yet, or you want to demonstrate an alternative version of a transliteration module before making it official. It should not be used in real modules or templates, only for testing. All uses of this parameter are tracked by [[Wiktionary:Tracking/languages/module_override]]. '''Known bugs''': * This function assumes {tr(s1) .. tr(s2) == tr(s1 .. s2)}. When this assertion fails, wikitext markups like <nowiki>'''</nowiki> can cause wrong transliterations. * HTML entities like `&amp;apos;`, often used to escape wikitext markups, do not work. ]==] function Language:transliterate(text, sc, module_override) -- If there is no text, or the language doesn't have transliteration data and there's no override, return nil. if not text or text == "" or text == "-" then return text end -- If the script is not transliteratable (and no override is given), return nil. sc = checkScript(text, self, sc) if not (sc:isTransliterated() or module_override) then -- temporary tracking to see if/when this gets triggered track("non-transliterable") track("non-transliterable/" .. self._code) track("non-transliterable/" .. sc:getCode()) track("non-transliterable/" .. sc:getCode() .. "/" .. self._code) return nil end -- Remove any strip markers. text = unstrip(text) -- Do not process the formatting into PUA characters for certain languages. local processed = load_data(languages_data_module).substitution[self._code] ~= "none" -- Get the display text with the keepCarets flag set. local subbedChars if processed then text, subbedChars = processDisplayText(text, self, sc, true) end -- Transliterate (using the module override if applicable). text, subbedChars = iterateSectionSubstitutions(self, text, sc, subbedChars, true, module_override or self._data.translit, "translit", "tr") if not text then return nil end -- Incomplete transliterations return nil. local charset = sc.characters if charset and umatch(text, "[" .. charset .. "]") then -- Remove any characters in Latin, which includes Latin characters also included in other scripts (as these are -- false positives), as well as any PUA substitutions. Anything remaining should only be script code "None" -- (e.g. numerals). local check_text = ugsub(text, "[" .. get_script("Latn").characters .. "􀀀-􏿽]+", "") -- Set none_is_last_resort_only flag, so that any non-None chars will cause a script other than "None" to be -- returned. if find_best_script_without_lang(check_text, true):getCode() ~= "None" then return nil end end if processed then text = escape_risky_characters(text) text = undoTempSubstitutions(text, subbedChars) end -- If the script does not use capitalization, then capitalize any letters of the transliteration which are -- immediately preceded by a caret (and remove the caret). if text and not sc:hasCapitalization() and text:find("^", 1, true) then text = processCarets(text, "%^([\128-\191\244]*%*?)([^\128-\191\244][\128-\191]*)", function(m1, m2) return m1 .. uupper(m2) end) end -- Track module overrides. if module_override ~= nil then track("module_override") end return text end do local function handle_language_spec(self, spec, sc) local ret = self["_" .. spec] if ret == nil then ret = self._data[spec] if type(ret) == "string" then ret = list_to_set(split(ret, ",", true, true)) end self["_" .. spec] = ret end if type(ret) == "table" then ret = ret[sc:getCode()] end return not not ret end function Language:overrideManualTranslit(sc) return handle_language_spec(self, "override_translit", sc) end function Language:link_tr(sc) return handle_language_spec(self, "link_tr", sc) end end --[==[ Return {true} if the language has a transliteration module, or {false} if it doesn't. ]==] function Language:hasTranslit() return not not self._data.translit end --[==[ Return {true} if the language uses the letters I/ı and İ/i, or {false} if it doesn't.]==] function Language:hasDottedDotlessI() return not not self._data.dotted_dotless_i end function Language:toJSON(opts) local strip_diacritics, strip_diacritics_patterns, strip_diacritics_remove_diacritics = self._data.strip_diacritics if strip_diacritics then if strip_diacritics.from then strip_diacritics_patterns = {} for i, from in ipairs(strip_diacritics.from) do insert(strip_diacritics_patterns, {from = from, to = strip_diacritics.to[i] or ""}) end end strip_diacritics_remove_diacritics = strip_diacritics.remove_diacritics end -- mainCode should only end up non-nil if dontCanonicalizeAliases is passed to make_object(). -- props should either contain zero-argument functions to compute the value, or the value itself. local props = { ancestors = function() return self:getAncestorCodes() end, canonicalName = function() return self:getCanonicalName() end, categoryName = function() return self:getCategoryName("nocap") end, code = self._code, mainCode = self._mainCode, parent = function() return self:getParentCode() end, full = function() return self:getFullCode() end, stripDiacriticsPatterns = strip_diacritics_patterns, stripDiacriticsRemoveDiacritics = strip_diacritics_remove_diacritics, family = function() return self:getFamilyCode() end, aliases = function() return self:getAliases() end, varieties = function() return self:getVarieties() end, otherNames = function() return self:getOtherNames() end, scripts = function() return self:getScriptCodes() end, type = function() return keys_to_list(self:getTypes()) end, wikimediaLanguages = function() return self:getWikimediaLanguageCodes() end, wikidataItem = function() return self:getWikidataItem() end, wikipediaArticle = function() return self:getWikipediaArticle(true) end, } local ret = {} for prop, val in pairs(props) do if not opts.skip_fields or not opts.skip_fields[prop] then if type(val) == "function" then ret[prop] = val() else ret[prop] = val end end end -- Use `deep_copy` when returning a table, so that there are no editing restrictions imposed by `mw.loadData`. return opts and opts.lua_table and deep_copy(ret) or to_json(ret, opts) end function export.getDataModuleName(code) local letter = match(code, "^(%l)%l%l?$") return "Modul:" .. ( letter == nil and "languages/data/exceptional" or #code == 2 and "languages/data/2" or "languages/data/3/" .. letter ) end get_data_module_name = export.getDataModuleName function export.getExtraDataModuleName(code) return get_data_module_name(code) .. "/extra" end get_extra_data_module_name = export.getExtraDataModuleName do local function make_stack(data) local key_types = { [2] = "unique", aliases = "unique", otherNames = "unique", type = "append", varieties = "unique", wikipedia_article = "unique", wikimedia_codes = "unique" } local function __index(self, k) local stack, key_type = getmetatable(self), key_types[k] -- Data that isn't inherited from the parent. if key_type == "unique" then local v = stack[stack[make_stack]][k] if v == nil then local layer = stack[0] if layer then -- Could be false if there's no extra data. v = layer[k] end end return v -- Data that is appended by each generation. elseif key_type == "append" then local parts, offset, n = {}, 0, stack[make_stack] for i = 1, n do local part = stack[i][k] if part == nil then offset = offset + 1 else parts[i - offset] = part end end return offset ~= n and concat(parts, ",") or nil end local n = stack[make_stack] while true do local layer = stack[n] if not layer then -- Could be false if there's no extra data. return nil end local v = layer[k] if v ~= nil then return v end n = n - 1 end end local function __newindex() error("table is read-only") end local function __pairs(self) -- Iterate down the stack, caching keys to avoid duplicate returns. local stack, seen = getmetatable(self), {} local n = stack[make_stack] local iter, state, k, v = pairs(stack[n]) return function() repeat repeat k = iter(state, k) if k == nil then n = n - 1 local layer = stack[n] if not layer then -- Could be false if there's no extra data. return nil end iter, state, k = pairs(layer) end until not (k == nil or seen[k]) -- Get the value via a lookup, as the one returned by the -- iterator will be the raw value from the current layer, -- which may not be the one __index will return for that -- key. Also memoize the key in `seen` (even if the lookup -- returns nil) so that it doesn't get looked up again. -- TODO: store values in `self`, avoiding the need to create -- the `seen` table. The iterator will need to iterate over -- `self` with `next` first to find these on future loops. v, seen[k] = self[k], true until v ~= nil return k, v end end local __ipairs = require(table_module).indexIpairs function make_stack(data) local stack = { data, [make_stack] = 1, -- stores the length and acts as a sentinel to confirm a given metatable is a stack. __index = __index, __newindex = __newindex, __pairs = __pairs, __ipairs = __ipairs, } stack.__metatable = stack return setmetatable({}, stack), stack end return make_stack(data) end local function get_stack(data) local stack = getmetatable(data) return stack and type(stack) == "table" and stack[make_stack] and stack or nil end --[==[ <span style="color: var(--wikt-palette-red,#BA0000)">This function is not for use in entries or other content pages.</span> Return a blob of data about the language. The format of this blob is undocumented, and perhaps unstable; it's intended for things like the module's own unit-tests, which are "close friends" with the module and will be kept up-to-date as the format changes. If `extra` is set, any extra data in the relevant `/extra` module will be included. (Note that it will be included anyway if it has already been loaded into the language object.) If `raw` is set, then the returned data will not contain any data inherited from parent objects. -- Do NOT use these methods! -- All uses should be pre-approved on the talk page! ]==] function Language:getData(extra, raw) if extra then self:loadInExtraData() end local data = self._data -- If raw is not set, just return the data. if not raw then return data end local stack = get_stack(data) -- If there isn't a stack or its length is 1, return the data. Extra data (if any) will be included, as it's -- stored at key 0 and doesn't affect the reported length. if stack == nil then return data end local n = stack[make_stack] if n == 1 then return data end extra = stack[0] -- If there isn't any extra data, return the top layer of the stack. if extra == nil then return stack[n] end -- If there is, return a new stack which has the top layer at key 1 and the extra data at key 0. data, stack = make_stack(stack[n]) stack[0] = extra return data end function Language:loadInExtraData() -- Only full languages have extra data. if not self:hasType("language", "full") then return end local data = self._data -- If there's no stack, create one. local stack = get_stack(self._data) if stack == nil then data, stack = make_stack(data) -- If already loaded, return. elseif stack[0] ~= nil then return end self._data = data -- Load extra data from the relevant module and add it to the stack at key 0, so that the __index and __pairs -- metamethods will pick it up, since they iterate down the stack until they run out of layers. local code = self._code local modulename = get_extra_data_module_name(code) -- No data cached as false. stack[0] = modulename and load_data(modulename)[code] or false end --[==[ Return the name of the module containing the language's data, e.g. [[Modul:languages/data/3/k]] for three-letter codes beginning with `k`. ]==] function Language:getDataModuleName() local name = self._dataModuleName if name == nil then name = self:hasType("etymology-only") and etymology_languages_data_module or get_data_module_name(self._mainCode or self._code) self._dataModuleName = name end return name end --[==[ Return the name of the module containing the language's extra data, e.g. [[Modul:languages/data/3/k/extra]] for three-letter codes beginning with `k`. ]==] function Language:getExtraDataModuleName() local name = self._extraDataModuleName if name == nil then name = not self:hasType("etymology-only") and get_extra_data_module_name(self._mainCode or self._code) or false self._extraDataModuleName = name end return name or nil end function export.makeObject(code, data, dontCanonicalizeAliases) local data_type = type(data) if data_type ~= "table" then error(("bad argument #2 to 'makeObject' (table expected, got %s)"):format(data_type)) end -- Convert any aliases. local input_code = code code = normalize_code(code) input_code = dontCanonicalizeAliases and input_code or code local parent if data.parent then parent = get_by_code(data.parent, nil, true, true) else parent = Language end parent.__index = parent local lang = {_code = input_code} -- This can only happen if dontCanonicalizeAliases is passed to make_object(). if code ~= input_code then lang._mainCode = code end local parent_data = parent._data if parent_data == nil then -- Full code is the same as the code. lang._fullCode = parent._code or code else -- Copy full code. lang._fullCode = parent._fullCode local stack = get_stack(parent_data) if stack == nil then parent_data, stack = make_stack(parent_data) end -- Insert the input data as the new top layer of the stack. local n = stack[make_stack] + 1 data, stack[n], stack[make_stack] = parent_data, data, n end lang._data = data return setmetatable(lang, parent) end make_object = export.makeObject end --[==[ Find the language whose code matches the one provided. If it exists, return a `Language` object representing the language. Otherwise, return {nil}, unless `paramForError` is given, in which case an error is generated. If `paramForError` is {true}, a generic error message mentioning the bad code is generated; otherwise `paramForError` should be a string or number specifying the parameter that the code came from, and this parameter will be mentioned in the error message along with the bad code. If `allowEtymLang` is specified, etymology-only language codes are allowed and looked up along with full language codes. If `allowFamily` is specified, language family codes are allowed and looked up along with normal language codes. ]==] function export.getByCode(code, paramForError, allowEtymLang, allowFamily) -- [[ Track uses of paramForError, ultimately so it can be removed, as error-handling should be done by -- [[Modul:parameters]], not here. ]] -- I (Benwing) disagree; that is an idealistic goal unlikely to be achievable -- realistically. if paramForError ~= nil then track("paramForError") end if type(code) ~= "string" then local typ if not code then typ = "nil" elseif check_object("language", true, code) then typ = "sebuah objek bahasa" elseif check_object("family", true, code) then typ = "sebuah objek keluarga" else typ = "sebuah " .. type(code) end error("The function getByCode expects a string as its first argument, but received " .. typ .. ".") end local m_data = load_data(languages_data_module) if m_data.aliases[code] or m_data.track[code] then track(code) end local norm_code = normalize_code(code) -- Get the data, checking for etymology-only languages if allowEtymLang is set. local data = load_data(get_data_module_name(norm_code))[norm_code] or allowEtymLang and load_data(etymology_languages_data_module)[norm_code] -- If no data was found and allowFamily is set, check the family data. If the main family data was found, make the -- object with [[Modul:families]] instead, as family objects have different methods. However, if it's an -- etymology-only family, use make_object in this module (which handles object inheritance), and the -- family-specific methods will be inherited from the parent object. if data == nil and allowFamily then data = load_data("Modul:families/data")[norm_code] if data ~= nil then if data.parent == nil then return make_family_object(norm_code, data) elseif not allowEtymLang then data = nil end end end local retval = code and data and make_object(code, data) if not retval and paramForError then require("Modul:languages/errorGetBy").code(code, paramForError, allowEtymLang, allowFamily) end return retval end get_by_code = export.getByCode --[==[ Find the language whose canonical name (the name used to represent that language on Wiktionary) or other name matches the one provided. If it exists, return a `Language` object representing the language. Otherwise, return {nil}, unless `paramForError` is given, in which case an error is generated. If `allowEtymLang` is specified, etymology-only language codes are allowed and looked up along with full language codes. If `allowFamily` is specified, language family codes are allowed and looked up along with normal language codes. The canonical name of languages should always be unique (it is an error for two languages on Wiktionary to share the same canonical name), so this is guaranteed to give at most one result. This function is powered by [[Modul:languages/canonical names]], which contains a pre-generated mapping of full-language canonical names to codes. It is generated by going through the [[:Category:Language data modules]] for full languages. When `allowEtymLang` is specified for the above function, [[Modul:etymology languages/canonical names]] may also be used, and when `allowFamily` is specified for the above function, [[Modul:families/canonical names]] may also be used. ]==] function export.getByCanonicalName(name, errorIfInvalid, allowEtymLang, allowFamily) local byName = load_data("Modul:languages/canonical names") local code = byName and byName[name] if not code and allowEtymLang then byName = load_data("Modul:etymology languages/canonical names") code = byName and byName[name] or byName[gsub(name, "^[Ss]ubstratum", "")] or byName[gsub(name, "^sebuah ", "")] or byName[gsub(name, "^sebuah [Ss]ubstratum ", ""):gsub("", "")] or -- For etymology families like "ira-pro". -- FIXME: This is not ideal, as it allows " languages" to be appended to any etymology-only language, too. byName[match(name, "^[Bb]ahasa%-bahasa (.*)$")] end if not code and allowFamily then byName = load_data("Modul:families/canonical names") code = byName[name] or byName[match(name, "^^[Bb]ahasa%-bahasa (.*)$")] end local retval = code and get_by_code(code, errorIfInvalid, allowEtymLang, allowFamily) if not retval and errorIfInvalid then require("Modul:languages/errorGetBy").canonicalName(name, allowEtymLang, allowFamily) end return retval end --[==[ Used by [[Modul:languages/data/2]] (et al.) and [[Modul:etymology languages/data]], [[Modul:families/data]], [[Modul:scripts/data]] and [[Modul:writing systems/data]] to finalize the data into the format that is actually returned. ]==] function export.finalizeData(data, main_type, variety) local fields = {"type"} if main_type == "language" then insert(fields, 4) -- script codes insert(fields, "ancestors") insert(fields, "link_tr") insert(fields, "override_translit") insert(fields, "wikimedia_codes") elseif main_type == "script" then insert(fields, 3) -- writing system codes end -- Families and writing systems have no extra fields to process. local fields_len = #fields for _, entity in next, data do if variety then -- Move parent from 3 to "parent" and family from "family" to 3. These are different for the sake of -- convenience, since very few varieties have the family specified, whereas all of them have a parent. entity.parent, entity[3] = entity[3], entity.family entity.family = nil -- Give the type "regular" iff not a variety and no other types are assigned. elseif not (entity.type or entity.parent) then entity.type = "regular" end for i = 1, fields_len do local key = fields[i] local field = entity[key] if field and type(field) == "string" then entity[key] = gsub(field, "%s*,%s*", ",") end end end return data end --[==[ For backwards compatibility only; modules should require the error themselves. ]==] function export.err(lang_code, param, code_desc, template_tag, not_real_lang) return require("Modul:languages/error")(lang_code, param, code_desc, template_tag, not_real_lang) end return export dbicznkmimzt0t4bckchxj0kvhhiklz Modul:headword 828 9757 375353 369309 2026-09-22T03:13:28Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714451|92714451]]) 375353 Scribunto text/plain local export = {} -- Named constants for all modules used, to make it easier to swap out sandbox versions. local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local gender_and_number_module = "Module:gender and number" local headword_data_module = "Module:headword/data" local headword_page_module = "Module:headword/page" local links_module = "Module:links" local load_module = "Module:load" local pages_module = "Module:pages" local palindromes_module = "Module:palindromes" local scripts_module = "Module:scripts" local scripts_data_module = "Module:scripts/data" local script_utilities_module = "Module:script utilities" local script_utilities_data_module = "Module:script utilities/data" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local dump = mw.dumpObject local insert = table.insert local ipairs = ipairs local max = math.max local new_title = mw.title.new local pairs = pairs local require = require local toNFC = mw.ustring.toNFC local toNFD = mw.ustring.toNFD local type = type local ufind = mw.ustring.find local ugmatch = mw.ustring.gmatch local ugsub = mw.ustring.gsub local umatch = mw.ustring.match --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==] local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function contains(...) contains = require(table_module).contains return contains(...) end local function encode_entities(...) encode_entities = require(string_utilities_module).encode_entities return encode_entities(...) end local function extend(...) extend = require(table_module).extend return extend(...) end local function find_best_script_without_lang(...) find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang return find_best_script_without_lang(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function format_genders(...) format_genders = require(gender_and_number_module).format_genders return format_genders(...) end local function format_decorations(...) format_decorations = require(decorations_module).format_decorations return format_decorations(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_current_L2(...) get_current_L2 = require(pages_module).get_current_L2 return get_current_L2(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function is_palindrome(...) is_palindrome = require(palindromes_module).is_palindrome return is_palindrome(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function load_data(...) load_data = require(load_module).load_data return load_data(...) end local function pattern_escape(...) pattern_escape = require(string_utilities_module).pattern_escape return pattern_escape(...) end local function pluralize(...) pluralize = require(en_utilities_module).pluralize return pluralize(...) end local function process_page(...) process_page = require(headword_page_module).process_page return process_page(...) end local function remove_links(...) remove_links = require(links_module).remove_links return remove_links(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function tag_text(...) tag_text = require(script_utilities_module).tag_text return tag_text(...) end local function tag_transcription(...) tag_transcription = require(script_utilities_module).tag_transcription return tag_transcription(...) end local function tag_translit(...) tag_translit = require(script_utilities_module).tag_translit return tag_translit(...) end local function trim(...) trim = require(string_utilities_module).trim return trim(...) end local function ulen(...) ulen = require(string_utilities_module).len return ulen(...) end local function ucfirst(...) ucfirst = require(string_utilities_module).ucfirst return ucfirst(...) end --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local m_data local function get_data() m_data = load_data(headword_data_module) return m_data end local script_data local function get_script_data() script_data = load_data(scripts_data_module) return script_data end local script_utilities_data local function get_script_utilities_data() script_utilities_data = load_data(script_utilities_data_module) return script_utilities_data end -- If set to true, categories always appear, even in non-mainspace pages local test_force_categories = false -- Add a tracking category to track entries with certain (unusually undesirable) properties. `track_id` is an identifier -- for the particular property being tracked and goes into the tracking page. Specifically, this adds a link in the -- page text to [[Wiktionary:Tracking/headword/TRACK_ID]], meaning you can find all entries with the `track_id` property -- by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID]]. -- -- If `lang` (a language object) is given, an additional tracking page [[Wiktionary:Tracking/headword/TRACK_ID/CODE]] is -- linked to where CODE is the language code of `lang`, and you can find all entries in the combination of `track_id` -- and `lang` by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID/CODE]]. This makes it possible to -- isolate only the entries with a specific tracking property that are in a given language. Note that if `lang` -- references at etymology-only language, both that language's code and its full parent's code are tracked. local function track(track_id, lang) local tracking_page = "headword/" .. track_id if lang and lang:hasType("etymology-only") then debug_track{tracking_page, tracking_page .. "/" .. lang:getCode(), tracking_page .. "/" .. lang:getFullCode()} elseif lang then debug_track{tracking_page, tracking_page .. "/" .. lang:getCode()} else debug_track(tracking_page) end return true end local function text_in_script(text, script_code) local sc = get_script(script_code) if not sc then error("Internal error: Bad script code " .. script_code) end local characters = sc.characters local out if characters then text = ugsub(text, "%W", "") out = ufind(text, "[" .. characters .. "]") end if out then return true else return false end end local spacingPunctuation = "[%s%p]+" --[[ List of punctuation or spacing characters that are found inside of words. Used to exclude characters from the regex above. ]] local wordPunc = "-#%%&@־׳״'.·*’་•:᠊" local notWordPunc = "[^" .. wordPunc .. "]+" --[=[ Format a term (either a head term or an inflection term) along with any decorations (left or right qualifiers, labels, references or customized separator). `part` is the object specifying the term (and `lang` the language of the term), which should optionally contain: * left qualifiers in `q`, an array of strings; * right qualifiers in `qq`, an array of strings; * left labels in `l`, an array of strings; * right labels in `ll`, an array of strings; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; * a separator in `separator`, defaulting to " <i>or</i> " if this is not the first term (j > 1), otherwise "". `formatted` is the formatted version of the term itself, and `j` is the index of the term. ]=] local function format_term_with_decorations(lang, part, formatted, j) local function part_non_empty(field) local list = part[field] if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `part.%s`=%s, should be \"table\""):format(field, dump(list))) end return list[1] end if part_non_empty("q") or part_non_empty("qq") or part_non_empty("l") or part_non_empty("ll") or part_non_empty("refs") then formatted = format_decorations { lang = lang, text = formatted, q = part.q, qq = part.qq, l = part.l, ll = part.ll, refs = part.refs, } end local separator = part.separator or j > 1 and " <i>atau</i> " -- use "" to request no separator if separator then formatted = separator .. formatted end return formatted end --[==[Return true if the given head is multiword according to the algorithm used in full_headword().]==] function export.head_is_multiword(head) for possibleWordBreak in ugmatch(head, spacingPunctuation) do if umatch(possibleWordBreak, notWordPunc) then return true end end return false end do local function workaround_to_exclude_chars(s) return (ugsub(s, notWordPunc, "\2%1\1")) end --[==[ Add appropriate links to `head`, correctly handling multiword terms. This is intended for multiword terms but can be used for any term if you want links added to single-word terms as well. If you want to only add links to multiword terms, first check that the term is multiword using `head_is_multiword`. If `default` is specified, this will escape colons so that they don't get interpreted as interwiki links. This should generally only be used when `head` is an actual pagename or is taken from a {{para|pagename}} parameter, not when taken from a {{para|head}} parameter. ]==] function export.add_multiword_links(head, default) head = "\1" .. ugsub(head, spacingPunctuation, workaround_to_exclude_chars) .. "\2" if default then head = head :gsub("(\1[^\2]*)\\([:#][^\2]*\2)", "%1\\\\%2") :gsub("(\1[^\2]*)([:#][^\2]*\2)", "%1\\%2") end --Escape any remaining square brackets to stop them breaking links (e.g. "[citation needed]"). head = encode_entities(head, "[]", true, true) --[=[ use this when workaround is no longer needed: head = "[[" .. ugsub(head, WORDBREAKCHARS, "]]%1[[") .. "]]" Remove any empty links, which could have been created above at the beginning or end of the string. ]=] return (head :gsub("\1\2", "") :gsub("[\1\2]", {["\1"] = "[[", ["\2"] = "]]"})) end end local function non_categorizable(full_raw_pagename) return full_raw_pagename:find("^Lampiran:Gerak isyarat/") or -- Unsupported titles with descriptive names. (full_raw_pagename:find("^Tajuk tidak disokong/") and not full_raw_pagename:find("`")) end local function tag_text_and_add_decorations(data, head, formatted, j) -- Add language and script wrapper. formatted = tag_text(formatted, data.lang, head.sc, "head", nil, j == 1 and data.id or nil) -- Add decorations (qualifiers, labels, references and separator). return format_term_with_decorations(data.lang, head, formatted, j) end -- Format a headword with transliterations. local function format_headword(data) -- Are there non-empty transliterations? local has_translits = false local has_manual_translits = false ------ Format the headwords. ------ local head_parts = {} local unique_head_parts = {} local has_multiple_heads = not not data.heads[2] for j, head in ipairs(data.heads) do if head.tr or head.ts then has_translits = true end if head.tr and head.tr_manual or head.ts then has_manual_translits = true end local formatted -- Apply processing to the headword, for formatting links and such. if head.term:find("[[", nil, true) and head.sc:getCode() ~= "Image" then formatted = language_link{term = head.term, lang = data.lang} else formatted = data.lang:makeDisplayText(head.term, head.sc, true) end local head_part = tag_text_and_add_decorations(data, head, formatted, j) insert(head_parts, head_part) -- If multiple heads, try to determine whether all heads display the same. To do this we need to effectively -- rerun the text tagging and addition of decorations, using 1 for all indices. if has_multiple_heads then local unique_head_part if j == 1 then unique_head_part = head_part else unique_head_part = tag_text_and_add_decorations(data, head, formatted, 1) end unique_head_parts[unique_head_part] = true end end local set_size = 0 if has_multiple_heads then for _ in pairs(unique_head_parts) do set_size = set_size + 1 end end if set_size == 1 then head_parts = head_parts[1] else head_parts = concat(head_parts) end if has_manual_translits then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr/LANGCODE]] track("manual-tr", data.lang) end ------ Format the transliterations and transcriptions. ------ local translits_formatted if has_translits then local translit_parts = {} for _, head in ipairs(data.heads) do if head.tr or head.ts then local this_parts = {} if head.tr then insert(this_parts, tag_translit(head.tr, data.lang:getCode(), "head", nil, head.tr_manual)) if head.ts then insert(this_parts, " ") end end if head.ts then insert(this_parts, "/" .. tag_transcription(head.ts, data.lang:getCode(), "head") .. "/") end insert(translit_parts, concat(this_parts)) end end translits_formatted = " (" .. concat(translit_parts, " <i>atau</i> ") .. ")" local function format_transliteration_page_link(langname) local transliteration_pagename = "Transliterasi bahasa " .. langname local transliteration_page = new_title(transliteration_pagename, "Wikikamus") if transliteration_page and transliteration_page:getContent() then return ("[[Wikikamus:%s|•]]"):format(transliteration_pagename) end return nil end local translit_page_link = format_transliteration_page_link(data.lang:getCanonicalName()) -- If data.lang is an etymology-only language and we didn't find a translation page for it, fall back to the -- full parent. if not translit_page_link and data.lang:hasType("etymology-only") then translit_page_link = format_transliteration_page_link(data.lang:getFullName()) end if translit_page_link then translits_formatted = " " .. translit_page_link .. translits_formatted end else translits_formatted = "" end ------ Paste heads and transliterations/transcriptions. ------ local lemma_gloss if data.gloss then lemma_gloss = ' <span class="ib-content qualifier-content">' .. data.gloss .. '</span>' else lemma_gloss = "" end return head_parts .. translits_formatted .. lemma_gloss end local function format_headword_genders(data, is_varform_only) local retval = "" if data.genders and data.genders[1] then if data.gloss then retval = "," end local pos_for_cat if not data.nogendercat and not is_varform_only then local no_gender_cat = (m_data or get_data()).no_gender_cat if not (no_gender_cat[data.lang:getCode()] or no_gender_cat[data.lang:getFullCode()]) then pos_for_cat = (m_data or get_data()).pos_for_gender_number_cat[data.pos_category:gsub("^reconstructed ", "")] end end local text, cats = format_genders(data.genders, data.lang, pos_for_cat) if cats then extend(data.categories, cats) end retval = retval .. "&nbsp;" .. text end return retval end -- Forward reference local format_inflections local function format_inflection_parts(data, parts) for j, part in ipairs(parts) do if type(part) ~= "table" then part = {term = part} end local partaccel = part.accel local face = part.face or "bold" if face ~= "bold" and face ~= "plain" and face ~= "hypothetical" then error("The face `" .. face .. "` " .. ( (script_utilities_data or get_script_utilities_data()).faces[face] and "should not be used for non-headword terms on the headword line." or "is invalid." )) end -- Here the final part 'or data.nolinkinfl' allows to have 'nolinkinfl=true' -- right into the 'data' table to disable inflection links of the entire headword -- when inflected forms aren't entry-worthy, e.g.: in Vulgar Latin local nolinkinfl = part.face == "hypothetical" or (part.nolink and track("nolink") or part.nolinkinfl) or ( data.nolink and track("nolink") or data.nolinkinfl) local formatted if part.label then -- FIXME: There should be a better way of italicizing a label. As is, this isn't customizable. formatted = "<i>" .. part.label .. "</i>" else -- Convert the term into a full link. Don't show a transliteration here unless enable_auto_translit is -- requested, either at the `parts` level (i.e. per inflection) or at the `data.inflections` level (i.e. -- specified for all inflections). This is controllable in {{head}} using autotrinfl=1 for all inflections, -- or fNautotr=1 for an individual inflection (remember that a single inflection may be associated with -- multiple terms). The reason for doing this is to avoid clutter in headword lines by default in languages -- where the script is relatively straightforward to read by learners (e.g. Greek, Russian), but allow it -- to be enabled in languages with more complex scripts (e.g. Arabic). -- -- FIXME: With nested inflections, should we also respect `enable_auto_translit` at the top level of the -- nested inflections structure? local tr = part.tr or not (parts.enable_auto_translit or data.inflections.enable_auto_translit) and "-" or nil local postprocess_annotations if part.inflections then postprocess_annotations = function(infldata) insert(infldata.annotations, format_inflections(data, part.inflections)) end end formatted = full_link( { term = not nolinkinfl and part.term or nil, alt = part.alt or (nolinkinfl and part.term or nil), lang = part.lang or data.lang, sc = part.sc or parts.sc or nil, gloss = part.gloss, pos = part.pos, lit = part.lit, id = part.id, genders = part.genders, tr = tr, ts = part.ts, accel = partaccel or parts.accel, postprocess_annotations = postprocess_annotations, }, face ) end parts[j] = format_term_with_decorations(part.lang or data.lang, part, formatted, j) end local parts_output if parts[1] then parts_output = (parts.label and " " or "") .. concat(parts) elseif parts.request then parts_output = " <small>[sila nyatakan]</small>" insert(data.categories, "Permohonan fleksi dalam entri bahasa " .. data.lang:getFullName()) else parts_output = "" end local parts_label = parts.label and ("<i>" .. parts.label .. "</i>") or "" return format_term_with_decorations(data.lang, parts, parts_label .. parts_output, 1) end -- Format the inflections following the headword or nested after a given inflection. Declared local above. function format_inflections(data, inflections) if inflections and inflections[1] then -- Format each inflection individually. for key, infl in ipairs(inflections) do inflections[key] = format_inflection_parts(data, infl) end return concat(inflections, ", ") else return "" end end -- Format the top-level inflections following the headword. Currently this just adds parens around the -- formatted comma-separated inflections in `data.inflections`. local function format_top_level_inflections(data) local result = format_inflections(data, data.inflections) if result ~= "" then return " (" .. result .. ")" else return result end end -- Forward reference local check_red_link_inflections -- Check a single inflection (which consists of a label and zero or more terms, each possibly with nested inflections) -- for red links. If so, insert a red-link category based on `plpos` (the plural part of speech to insert in the -- category), stop further processing, and return true. If no red links found, return false. local function check_red_link_inflection_parts(data, parts, plpos) for _, part in ipairs(parts) do if type(part) ~= "table" then part = {term = part} end local term = part.term if term and not term:find("%[%[") then local stripped_physical_term = get_link_page(term, data.lang, part.sc or parts.sc or nil) if stripped_physical_term then local title = mw.title.new(stripped_physical_term) if title and not title:getContent() then insert(data.categories, data.lang:getFullName() .. " " .. plpos .. " with red links in their headword lines") return true end end end if part.inflections then if check_red_link_inflections(data, part.inflections, plpos) then return true end end end return false end -- Check a set of inflections (each of which describes a single inflection of the term, such as feminine or plural, and -- consists of a label and zero or more terms, each possibly with nested inflections) for red links. If so, insert a -- red-link category based on `plpos` (the plural part of speech to insert in the category), stop further processing, -- and return true. If no red links found, return false. function check_red_link_inflections(data, inflections, plpos) if inflections and inflections[1] then -- Check each inflection individually. for key, infl in ipairs(inflections) do if check_red_link_inflection_parts(data, infl, plpos) then return true end end end return false end -- Check the top-level inflections in `data.inflections`, along with any nested inflections, for red links. If so, -- insert a red-link category based on `plpos` (the plural part of speech to insert in the category), stop further -- processing, and return true. If no red links found, return false. local function check_red_link_inflections_top_level(data, plpos) return check_red_link_inflections(data, data.inflections, plpos) end --[==[ Returns the plural form of `pos`, a raw part of speech input, which could be singular or plural. Irregular plural POS are taken into account (e.g. "kanji" pluralizes to "kanji"). ]==] function export.pluralize_pos(pos) -- Make the plural form of the part of speech return (m_data or get_data()).irregular_plurals[pos] or pos:sub(-1) == "s" and pos or pluralize(pos) end --[==[ Return "lemma" if the given POS is a lemma, "non-lemma form" if a non-lemma form, or nil if unknown. The POS passed in must be in its plural form ("nouns", "prefixes", etc.). If you have a POS in its singular form, call {export.pluralize_pos()} above to pluralize it in a smart fashion that knows when to add "-s" and when to add "-es", and also takes into account any irregular plurals. If `best_guess` is given and the POS is in neither the lemma nor non-lemma list, guess based on whether it ends in " forms"; otherwise, return nil. ]==] function export.pos_lemma_or_nonlemma(plpos, best_guess) local m_headword_data = m_data or get_data() local isLemma = m_headword_data.lemmas -- Is it a lemma category? if isLemma[plpos] then return "Lema" end local plpos_no_recon = plpos:gsub("^reconstructed ", "") if isLemma[plpos_no_recon] then return "Lema" end -- Is it a nonlemma category? local isNonLemma = m_headword_data.nonlemmas if isNonLemma[plpos] or isNonLemma[plpos_no_recon] then return "Bentuk bukan lema" end local plpos_no_mut = plpos:gsub("^mutated ", "") if isLemma[plpos_no_mut] or isNonLemma[plpos_no_mut] then return "Bentuk bukan lema" elseif best_guess then return plpos:find("^Bentuk ") and "Bentuk bukan lema" or "Lema" else return nil end end --[==[ Canonicalize a part of speech as specified in 2= in {{tl|head}}. This checks for POS aliases and non-lemma form aliases ending in 'f', and then pluralizes if the POS term does not have an invariable plural. ]==] function export.canonicalize_pos(pos) -- FIXME: Temporary code to throw an error for alias 'pre' (= preposition) that will go away. if pos == "pre" then -- Don't throw error on 'pref' as it's an alias for "prefix". error("POS 'pre' for 'preposition' no longer allowed as it's too ambiguous; use 'prep'") end -- Likewise for pro = pronoun. if pos == "pro" or pos == "prof" then error("POS 'pro' for 'pronoun' no longer allowed as it's too ambiguous; use 'pron'") end local m_headword_data = m_data or get_data() if m_headword_data.pos_aliases[pos] then pos = m_headword_data.pos_aliases[pos] elseif pos:sub(-1) == "f" then pos = pos:sub(1, -2) pos = "Bentuk " .. (m_headword_data.pos_aliases[pos] or pos) end return export.pluralize_pos(pos) end -- Find and return the maximum index in the array `data[element]` (which may have gaps in it), and initialize it to a -- zero-length array if unspecified. Check to make sure all keys are numeric (other than "maxindex", which is set by -- [[Module:parameters]] for list parameters), all values are strings, and unless `allow_blank_string` is given, -- no blank (zero-length) strings are present. local function init_and_find_maximum_index(data, element, allow_blank_string) local maxind = 0 if not data[element] then data[element] = {} end local typ = type(data[element]) if typ ~= "table" then error(("Internal error: In full_headword(), `data.%s` must be an array but is a %s"):format(element, typ)) end for k, v in pairs(data[element]) do if k ~= "maxindex" then if type(k) ~= "number" then error(("Internal error: Unrecognized non-numeric key '%s' in `data.%s`"):format(k, element)) end if k > maxind then maxind = k end if v then if type(v) ~= "string" then error(("Internal error: For key '%s' in `data.%s`, value should be a string but is a %s"):format(k, element, type(v))) end if not allow_blank_string and v == "" then error(("Internal error: For key '%s' in `data.%s`, blank string not allowed; use 'false' for the default"):format(k, element)) end end end end return maxind end --[==[ -- Add the page to various maintenance categories for the language and the -- whole page. These are placed in the headword somewhat arbitrarily, but -- mainly because headword templates are mandatory for entries (meaning that -- in theory it provides full coverage). -- -- This is provided as an external entry point so that modules which transclude -- information from other entries (such as {{tl|ja-see}}) can take advantage -- of this feature as well, because they are used in place of a conventional -- headword template.]==] do -- Handle any manual sortkeys that have been specified in raw categories -- by tracking if they are the same or different from the automatically- -- generated sortkey, so that we can track them in maintenance -- categories. local function handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) sortkey = sortkey or lang:makeSortKey(page.pagename) -- If there are raw categories with no sortkey, then they will be -- sorted based on the default MediaWiki sortkey, so we check against -- that. if tbl == true then if page.raw_defaultsort ~= sortkey then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik") end return end local redundant, different for k in pairs(tbl) do if k == sortkey then redundant = true else different = true end end if redundant then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih lewah") end if different then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik") end return sortkey end function export.maintenance_cats(page, lang, lang_cats, page_cats) extend(page_cats, page.cats) lang = lang:getFull() -- since we are just generating categories local canonical = lang:getCanonicalName() local tbl = page.wikitext_topic_cat[lang:getCode()] local sortkey = nil if tbl then sortkey = handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori topik yang menggunakan penanda mentah") end tbl = page.wikitext_langname_cat[canonical] if tbl then handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori nama bahasa yang menggunakan penanda mentah") end if get_current_L2() ~= "Bahasa " .. canonical then insert(lang_cats, "Entri bahasa " .. canonical .. " dengan pengepala bahasa tidak betul") -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header/LANGCODE]] track("pengepala bahasa tidak betul", lang) end end end --[==[This is the primary external entry point. {{lua|full_headword(data)}} This is used by {{temp|head}} and various language-specific headword templates (e.g. {{temp|ru-adj}} for Russian adjectives, {{temp|de-noun}} for German nouns, etc.) to display an entire headword line. See [[#Further explanations for full_headword()]] ]==] function export.full_headword(data) -- Prevent data from being destructively modified. data = shallow_copy(data) ------------ 1. Basic checks for old-style (multi-arg) calling convention. ------------ if data.getCanonicalName then error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) of properties, not a language object") end if not data.lang or type(data.lang) ~= "table" or not data.lang.getCode then error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) and `data.lang` must be a language object") end if data.id and type(data.id) ~= "string" then error("Internal error: The id in the data table should be a string.") end ------------ 2. Initialize pagename etc. ------------ local langcode = data.lang:getCode() local full_langcode = data.lang:getFullCode() local langname = data.lang:getCanonicalName() local full_langname = data.lang:getFullName() local raw_pagename = data.pagename local page local m_headword_data = m_data or get_data() if raw_pagename and raw_pagename ~= m_headword_data.pagename then -- for testing, doc pages, etc. -- data.pagename is often set on documentation and test pages through the pagename= parameter of various -- templates, to emulate running on that page. Having a large number of such test templates on a single -- page often leads to timeouts, because we fetch and parse the contents of each page in turn. However, -- we don't really need to do that and can function fine without fetching and parsing the contents of a -- given page, so turn off content fetching/parsing (and also setting the DEFAULTSORT key through a parser -- function, which is *slooooow*) in certain namespaces where test and documentation templates are likely to -- be found and where actual content does not live (User, Template, Module). local actual_namespace = m_headword_data.page.namespace local no_fetch_content = actual_namespace == "User" or actual_namespace == "Template" or actual_namespace == "Module" page = process_page(raw_pagename, no_fetch_content) else page = m_headword_data.page end local namespace = page.namespace if data.altform then -- Temporary tracking for use of old altform= track("altform", data.lang) end local is_varform_only = data.var and data.var ~= "both" local is_varform_both = data.var == "both" ------------ 3. Initialize `data.heads` table; if old-style, convert to new-style. ------------ if type(data.heads) == "table" and type(data.heads[1]) == "table" then -- new-style if data.translits or data.transcriptions then error("Internal error: In full_headword(), if `data.heads` is new-style (array of head objects), `data.translits` and `data.transcriptions` cannot be given") end else -- convert old-style `heads`, `translits` and `transcriptions` to new-style local maxind = max( init_and_find_maximum_index(data, "heads"), init_and_find_maximum_index(data, "translits", true), init_and_find_maximum_index(data, "transcriptions", true) ) for i = 1, maxind do data.heads[i] = { term = data.heads[i], tr = data.translits[i], ts = data.transcriptions[i], } end end -- Make sure there's at least one head. if not data.heads[1] then data.heads[1] = {} end ------------ 4. Initialize and validate `data.categories` and `data.whole_page_categories`, and determine `pos_category` if not given, and add basic categories. ------------ init_and_find_maximum_index(data, "categories") init_and_find_maximum_index(data, "whole_page_categories") local pos_category_already_present = false if data.categories[1] then local escaped_langname = pattern_escape(full_langname) local matches_lang_pattern = "^" .. escaped_langname .. " " for _, cat in ipairs(data.categories) do -- Does the category begin with the language name? If not, tag it with a tracking category. if not cat:find(matches_lang_pattern) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category/LANGCODE]] track("no lang category", data.lang) end end -- If `pos_category` not given, try to infer it from the first specified category. If this doesn't work, we -- throw an error below. if not data.pos_category and data.categories[1]:find(matches_lang_pattern) then data.pos_category = data.categories[1]:gsub(matches_lang_pattern, "") -- Optimization to avoid inserting category already present. pos_category_already_present = true end end if not data.pos_category then error("Internal error: `data.pos_category` not specified and could not be inferred from the categories given in " .. "`data.categories`. Either specify the plural part of speech in `data.pos_category` " .. "(e.g. \"proper nouns\") or ensure that the first category in `data.categories` is formed from the " .. "language's canonical name plus the plural part of speech (e.g. \"Norwegian Bokmål proper nouns\")." ) end -- Insert a category at the beginning for the part of speech unless it's already present or `data.noposcat` given. if not pos_category_already_present and not data.noposcat and not is_varform_only then local pos_category = ucfirst(data.pos_category) .. " bahasa " .. full_langname -- FIXME: [[User:Theknightwho]] Why is this special case here? Please add an explanatory comment. if pos_category ~= "Aksara Han rentas bahasa" then insert(data.categories, 1, pos_category) end end -- Try to determine whether the part of speech refers to a lemma or a non-lemma form; if we can figure this out, -- add an appropriate category. local postype = export.pos_lemma_or_nonlemma(data.pos_category) if not postype then -- We don't know what this category is, so tag it with a tracking category. -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/LANGCODE]] track("unrecognized pos", data.lang) -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS/LANGCODE]] track("unrecognized pos/pos/" .. data.pos_category, data.lang) elseif not data.noposcat and not is_varform_only then insert(data.categories, 1, ucfirst(postype) .. " bahasa " .. full_langname) end -- Categorize variant forms into 'variant lemmas' or 'variant non-lemma forms'. Originally proposed in -- [[Wiktionary:Beer parlour/2024/June#Decluttering the altform mess]] as 'alternative forms'; renamed in -- [[Wiktionary:Beer parlour/2026/July#Renaming the "alternative forms" categories]]. if (is_varform_only or is_varform_both) and postype then insert(data.categories, 1, postype .. " kelaianan bahasa " .. full_langname) end ------------ 5. Create a default headword, and add links to multiword page names. ------------ -- Determine if this is an "anti-asterisk" term, i.e. an attested term in a language that must normally be -- reconstructed. local is_anti_asterisk = data.heads[1].term and data.heads[1].term:find("^!!") local lang_reconstructed = data.lang:hasType("reconstructed") if is_anti_asterisk then if not lang_reconstructed then error("Anti-asterisk feature (head= beginning with !!) can only be used with reconstructed languages") end lang_reconstructed = false end -- Determine if term is reconstructed local is_reconstructed = namespace == "Rekonstruksi" or lang_reconstructed -- Create a default headword based on the pagename, which is determined in -- advance by the data module so that it only needs to be done once. local default_head = page.pagename -- Add links to multi-word page names when appropriate if not (is_reconstructed or data.nolinkhead) then local no_links = m_headword_data.no_multiword_links if not (no_links[langcode] or no_links[full_langcode]) and export.head_is_multiword(default_head) then default_head = export.add_multiword_links(default_head, true) end end if is_reconstructed then default_head = "*" .. default_head end ------------ 6. Check the namespace against the language type. ------------ if namespace == "" then if lang_reconstructed then error("Entries in " .. langname .. " must be placed in the Rekonstruksi: namespace") elseif data.lang:hasType("appendix-constructed") then error("Entries in " .. langname .. " must be placed in the Lampiran: namespace") end elseif namespace == "Petikan" or namespace == "Tesaurus" then error("Headword templates should not be used in the " .. namespace .. ": namespace.") end ------------ 7. Fill in missing values in `data.heads`. ------------ -- True if any script among the headword scripts has spaces in it. local any_script_has_spaces = false -- True if any term has a redundant head= param. local has_redundant_head_param = false for _, head in ipairs(data.heads) do ------ 7a. If missing head, replace with default head. if not head.term then head.term = default_head elseif head.term == default_head then has_redundant_head_param = true elseif is_anti_asterisk and head.term == "!!" then -- If explicit head=!! is given, it's an anti-asterisk term and we fill in the default head. head.term = "!!" .. default_head elseif head.term:find("^[!?]$") then -- If explicit head= just consists of ! or ?, add it to the end of the default head. head.term = default_head .. head.term end head.term_no_initial_bang_bang = is_anti_asterisk and head.term:sub(3) or head.term if is_reconstructed then local head_term = head.term if head_term:find("%[%[") then head_term = remove_links(head_term) end if head_term:sub(1, 1) ~= "*" then error("The headword '" .. head_term .. "' must begin with '*' to indicate that it is reconstructed.") end end ------ 7b. Try to detect the script(s) if not provided. If a per-head script is provided, that takes precedence, ------ otherwise fall back to the overall script if given. If neither given, autodetect the script. local auto_sc = data.lang:findBestScript(head.term) if ( auto_sc:getCode() == "None" and find_best_script_without_lang(head.term):getCode() ~= "None" ) then insert(data.categories, "Perkataan bahasa " .. full_langname .. " dalam bentuk tulisan tidak piawai") end if not (head.sc or data.sc) then -- No script code given, so use autodetected script. head.sc = auto_sc else if not head.sc then -- Overall script code given. head.sc = data.sc end -- Track uses of sc parameter. if head.sc:getCode() == auto_sc:getCode() then track("redundant script code", data.lang) if not data.no_script_code_cat then insert(data.categories, "Perkataan dengan kod tulisan lewah bahasa " .. full_langname ) end else track("non-redundant manual script code", data.lang) if not data.no_script_code_cat then insert(data.categories, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. full_langname ) end end end -- If using a discouraged character sequence, add to maintenance category. if head.sc:hasNormalizationFixes() == true then local composed_head = toNFC(head.term) if head.sc:fixDiscouragedSequences(composed_head) ~= composed_head then insert(data.whole_page_categories, "Laman menggunakan jujukan aksara tidak digalakkan") end end any_script_has_spaces = any_script_has_spaces or head.sc:hasSpaces() ------ 7c. Create automatic transliterations for any non-Latin headwords without manual translit given ------ (provided automatic translit is available, e.g. not in Persian or Hebrew). -- Make transliterations head.tr_manual = nil -- Try to generate a transliteration if necessary if head.tr == "-" then head.tr = nil else local notranslit = m_headword_data.notranslit if not (notranslit[langcode] or notranslit[full_langcode]) and head.sc:isTransliterated() then head.tr_manual = not not head.tr local text = head.term_no_initial_bang_bang if not data.lang:link_tr(head.sc) then text = remove_links(text) end local automated_tr = data.lang:transliterate(text, head.sc) if automated_tr then local manual_tr = head.tr if manual_tr then if remove_links(manual_tr) == remove_links(automated_tr) then insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi lewah") else insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi manual tidak lewah") end end if not manual_tr then head.tr = automated_tr end end -- There is still no transliteration? -- Add the entry to a cleanup category. if not head.tr then head.tr = "<small>transliterasi diperlukan</small>" -- FIXME: No current support for 'Request for transliteration of Classical Persian terms' or similar. -- Consider adding this support in [[Module:category tree/poscatboiler/data/entry maintenance]]. insert(data.categories, "Permintaan transliterasi perkataan bahasa " .. full_langname) else -- Otherwise, trim it. head.tr = trim(head.tr) end end end -- Link to the transliteration entry for languages that require this. if head.tr and data.lang:link_tr(head.sc) then head.tr = full_link{ term = head.tr, lang = data.lang, sc = get_script("Latn"), tr = "-" } end end ------------ 8. Maybe tag the title with the appropriate script code, using the `display_title` mechanism. ------------ -- Assumes that the scripts in "toBeTagged" will never occur in the Reconstruction namespace. -- (FIXME: Don't make assumptions like this, and if you need to do so, throw an error if the assumption is violated.) -- Avoid tagging ASCII as Hani even when it is tagged as Hani in the headword, as in [[check]]. The check for ASCII -- might need to be expanded to a check for any Latin characters and whitespace or punctuation. local display_title -- Where there are multiple headwords, use the script for the first. This assumes the first headword is similar to -- the pagename, and that headwords that are in different scripts from the pagename aren't first. This seems to be -- about the best we can do (alternatively we could potentially do script detection on the pagename). local dt_script = data.heads[1].sc local dt_script_code = dt_script:getCode() local page_non_ascii = namespace == "" and not page.pagename:find("^[%z\1-\127]+$") local unsupported_pagename, unsupported = page.full_raw_pagename:gsub("^Tajuk tidak disokong/", "") if unsupported == 1 and page.unsupported_titles[unsupported_pagename] then display_title = 'Tajuk tidak disokong/<span class="' .. dt_script_code .. '">' .. page.unsupported_titles[unsupported_pagename] .. '</span>' elseif page_non_ascii and m_headword_data.toBeTagged[dt_script_code] or (dt_script_code == "Jpan" and (text_in_script(page.pagename, "Hira") or text_in_script(page.pagename, "Kana"))) or (dt_script_code == "Kore" and text_in_script(page.pagename, "Hang")) then display_title = '<span class="' .. dt_script_code .. '">' .. page.full_raw_pagename .. '</span>' -- Keep Han entries region-neutral in the display title. elseif page_non_ascii and (dt_script_code == "Hant" or dt_script_code == "Hans") then display_title = '<span class="Hani">' .. page.full_raw_pagename .. '</span>' elseif namespace == "Rekonstruksi" then local matched display_title, matched = ugsub( page.full_raw_pagename, "^(Rekonstruksi:[^/]+/)(.+)$", function(before, term) return before .. tag_text(term, data.lang, dt_script) end ) if matched == 0 then display_title = nil end end -- FIXME: Generalize this. -- If the current language uses Aran (Nastaliq), e.g. Urdu, and there's more than one language on the page, don't -- set the display title because we don't want Nastaliq for terms that also exist in other languages that don't -- display in Nastaliq (e.g. Arabic or Persian). Because the word "Urdu" occurs near the end of the alphabet, Urdu -- fonts tend to override the fonts of other languages. FIXME: This is checking for more than one language on the -- page but instead needs to check if there are any languages using scripts other than Aran. if dt_script_code == "Aran" and page.L2_list.n > 1 then display_title = nil end if display_title then mw.getCurrentFrame():callParserFunction( "DISPLAYTITLE", display_title ) end ------------ 9. Insert additional categories. ------------ if data.force_cat_output then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/force cat output]] track("force cat output") end if has_redundant_head_param then if not data.no_redundant_head_cat then -- This is not the right way to go about this; too many exceptions and problems due to language-specific headword -- handling customization. If we want this, it should be opt-in by a given language passing in the default headword. -- insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan parameter pengepala lewah") end end -- If the first head is multiword (after removing links), maybe insert into "LANG multiword terms". if not data.nomultiwordcat and not is_varform_only and any_script_has_spaces and postype == "lemma" then local no_multiword_cat = m_headword_data.no_multiword_cat if not (no_multiword_cat[langcode] or no_multiword_cat[full_langcode]) then -- Check for spaces or hyphens, but exclude prefixes and suffixes. -- Use the pagename, not the head= value, because the latter may have extra -- junk in it, e.g. superscripted text that throws off the algorithm. local no_hyphen = m_headword_data.hyphen_not_multiword_sep -- Exclude hyphens if the data module states that they should for this language. local checkpattern = (no_hyphen[langcode] or no_hyphen[full_langcode]) and ".[%s፡]." or ".[%s%-፡]." local is_multiword = umatch(page.pagename, checkpattern) if is_multiword and not non_categorizable(page.full_raw_pagename) then insert(data.categories, "Perkataan berbilang kata bahasa " .. full_langname) elseif not is_multiword then local long_word_threshold = m_headword_data.long_word_thresholds[langcode] or m_headword_data.long_word_thresholds[full_langcode] if long_word_threshold and ulen(page.pagename) >= long_word_threshold then insert(data.categories, "Perkataan panjang bahasa " .. full_langname) end end end end -- Determine whether to insert a category 'LANGNAME POS in SCRIPT'. If there are multiple heads, we may need to check -- each head, as the heads may (theoretically) have different scripts. local default_sccat = m_headword_data.default_sccat if data.sccat or not is_varform_only and (default_sccat[langcode] or langcode ~= full_langcode and default_sccat[full_langcode]) then local function needs_sccat(sccat_entry, sc) if sccat_entry == true or not sccat_entry then return sccat_entry end if type(sccat_entry) == "table" then local in_list = contains(sccat_entry, sc:getCode()) if sccat_entry[1] == "not" then in_list = not in_list end return in_list end return nil end for _, head in ipairs(data.heads) do -- First check the `sccat` specified at the {{head}} level. local this_needs_sccat = needs_sccat(data.sccat, head.sc) -- If that wasn't given, check the default sccat at the language level for the lang code. if this_needs_sccat == nil and not is_varform_only then this_needs_sccat = needs_sccat(default_sccat[langcode], head.sc) end -- If that wasn't found and the lang code is an etym code, check the default sccat at the parent language level. if this_needs_sccat == nil and not is_varform_only and langcode ~= full_langcode then this_needs_sccat = needs_sccat(default_sccat[full_langcode], head.sc) end if this_needs_sccat then insert(data.categories, ucfirst(data.pos_category) .. " bahasa " .. full_langname .. " dalam " .. head.sc:getDisplayForm(data.lang)) end end end -- Reconstructed terms often use weird combinations of scripts and realistically aren't spelled so much as notated. if namespace ~= "Rekonstruksi" and not is_varform_only then -- Map from languages to a string containing the characters to ignore when considering whether a term has -- multiple written scripts in it. Typically these are Greek or Cyrillic letters used for their phonetic -- values. local characters_to_ignore = { ["aaq"] = "αάὰ", -- Penobscot (Algonquian) ["acy"] = "δθ", -- Cypriot Arabic ["aez"] = "β", -- Aeka (Trans-New Guinea) ["anc"] = "γ", -- Ngas (Chadic/Afroasiatic) ["aou"] = "χ", -- A'ou (Kra-Dai) ["art-blk"] = "ч", -- Bolak (conlang) ["awg"] = "β", -- Anguthimri (Pama-Nyungan) ["az"] = "ь", -- Azerbaijani (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["ba"] = "ь", -- Bashkir (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["bhp"] = "β", -- Bima (Austronesian) ["bjz"] = "β", -- Baruga (Trans-New Guinea) ["byk"] = "θ", -- Biao (Kra-Dai) ["cdy"] = "θ", -- Chadong (Kra-Dai) ["chp"] = "θ", -- Chipewyan (Athabaskan) ["cjh"] = "χ", -- Upper Chehalis (Salishan) ["clm"] = "χ", -- Klallam (Salishan) ["col"] = "χ", -- Colombia-Wenatchi (Salishan) ["coo"] = "χθ", -- Comox (Salishan) ["crx"] = "θ", -- Carrier (Athabaskan) ["ets"] = "θ", -- Yekhee (Edoid/Niger-Congo) ["ett"] = "χ", -- Etruscan (isolate; in romanizations) ["fla"] = "χ", -- Montana Salish (Salishan) ["grt"] = "་", -- Garo (South Asian Sino-Tibetan) ["gmw-gts"] = "χ", -- Gottscheerish (Bavarian variant spoken in Slovenia) ["hur"] = "χθ", -- Halkomelem (Salishan) ["itc-psa"] = "f", -- Pre-Samnite (Italic; normally written in Greek) ["izh"] = "ь", -- Ingrian (Finnic) ["kic"] = "θ", -- Kickapoo (Algonquian) ["kk"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["ky"] = "ь", -- Kyrgyz (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["lil"] = "χ", -- Lillooet (Salishan) ["lsi"] = "ꓹ", -- Lashi (Lolo-Burmese/Sino-Tibetan; represents a glottal stop) ["mhz"] = "β", -- Mor (Austronesian) ["mqn"] = "β", -- Moronene (Austronesian) ["neg"]= "ӡā", -- Negidal (Tungusic; normally in Cyrillic) ["oka"] = "χ", -- Okanagan (Salishan) ["ole"] = "θ", -- Olekha (Sino-Tibetan) ["oui"] = "γβ", -- Old Uyghur (Turkic; FIXME: others? E.g. Greek delta (δ)?) ["pox"] = "χ", -- Polabian (West Slavic) ["rif"] = "ε", -- Tarifit (Berber) ["rom"] = "Θθ", -- Romani (Indic: International Standard; two different thetas???) ["rpn"] = "β", -- Repanbitip (Austronesian) ["sah"] = "ь", -- Yakut (Turkic; 1929 - 1939 Latin spelling) ["sit-jap"] = "χ", -- Japhug (Sino-Tibetan) ["sjw"] = "θ", -- Shawnee (Algonquian) ["squ"] = "χ", -- Squamish (Salishan) ["str"] = "χθ", -- Saanich (Salishan) ["teh"] = "χ", -- Tehuelche (Chonan; spoken in Argentina) ["tep"] = "η", -- Tepecano (Uto-Aztecan) ["thp"] = "χ", -- Thompson (Salishan) ["tk"] = "ь", -- Turkmen (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["tt"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["twa"] = "χ", -- Twana (Salishan) ["wbl"] = "ы", -- Wakhi (Iranian) ["xbc"] = "ϸ", -- Bactrian (Iranian; represents š; normally written in Greek) ["yha"] = "θ", -- Baha (Kra-Dai) ["za"] = "зч", -- Zhuang (Tai/Kra-Dai); 1957-1982 alphabet used two Cyrillic letters (as well as some others like -- ƃ, ƅ, ƨ, ɯ and ɵ that look like Cyrillic or Greek but are actually Latin) ["zlw-slv"] = "χђћ", -- Slovincian (West Slavic; FIXME: χ is Greek, the other two are Cyrillic, but I'm not sure -- the currect characters are being chosen in the entry names) ["zng"] = "θ", -- Mang (Mon-Khmer) ["ztp"] = "θ", -- Loxicha Zapotec (Zapotecan) } -- Determine how many real scripts are found in the pagename, where we exclude symbols and such. We exclude -- scripts whose `character_category` is false as well as Zmth (mathematical notation symbols), which has a -- category of "Mathematical notation symbols". When counting scripts, we need to elide language-specific -- variants because e.g. Beng and as-Beng have slightly different characters but we don't want to consider them -- two different scripts (e.g. [[এৰ]] has two characters which are detected respectively as Beng and as-Beng). local seen_scripts = {} local num_seen_scripts = 0 local num_loops = 0 local canon_pagename = page.pagename local ch_to_ignore = characters_to_ignore[full_langcode] if ch_to_ignore then canon_pagename = ugsub(canon_pagename, "[" .. ch_to_ignore .. "]", "") end while true do if canon_pagename == "" or num_seen_scripts >= 2 or num_loops >= 10 then break end -- Make sure we don't get into a loop checking the same script over and over again; happens with e.g. [[ᠪᡳ]] num_loops = num_loops + 1 local pagename_script = find_best_script_without_lang(canon_pagename, "None only as last resort") local script_chars = pagename_script.characters if not script_chars then -- we are stuck; this happens with None break end local script_code = pagename_script:getCode() local replaced canon_pagename, replaced = ugsub(canon_pagename, "[" .. script_chars .. "]", "") if ( replaced and script_code ~= "Zmth" and (script_data or get_script_data())[script_code] and script_data[script_code].character_category ~= false ) then script_code = script_code:gsub("^.-%-", "") if not seen_scripts[script_code] then seen_scripts[script_code] = true num_seen_scripts = num_seen_scripts + 1 end end end if num_seen_scripts > 1 then insert(data.categories, "Perkataan bahasa " .. full_langname .. " dieja dalam berbilang tulisan") end end -- Categorise for unusual characters. Takes into account combining characters, so that we can categorise for characters with diacritics that aren't encoded as atomic characters (e.g. U̠). These can be in two formats: single combining characters (i.e. character + diacritic(s)) or double combining characters (i.e. character + diacritic(s) + character). Each can have any number of diacritics. local standard = data.lang:getStandardCharacters() if not is_varform_only and standard and not non_categorizable(page.full_raw_pagename) then local function char_category(char) local specials = { ["#"] = "number sign", ["("] = "parentheses", [")"] = "parentheses", ["<"] = "angle brackets", [">"] = "angle brackets", ["["] = "square brackets", ["]"] = "square brackets", ["_"] = "underscore", ["{"] = "braces", ["|"] = "vertical line", ["}"] = "braces", ["ß"] = "ẞ", ["\205\133"] = "", -- this is UTF-8 for U+0345 ( ͅ) ["\239\191\189"] = "replacement character", } char = toNFD(char) :gsub(".[\128-\191]*", function(m) local new_m = specials[m] new_m = new_m or m:uupper() return new_m end) return toNFC(char) end if full_langcode ~= "hi" and full_langcode ~= "lo" then local standard_chars_scripts = {} for _, head in ipairs(data.heads) do standard_chars_scripts[head.sc:getCode()] = true end -- Iterate over the scripts, in case there is more than one (as they can have different sets of standard characters). for code in pairs(standard_chars_scripts) do local sc_standard = data.lang:getStandardCharacters(code) if sc_standard then if page.pagename_len > 1 then local explode_standard = {} local function explode(char) explode_standard[char] = true return "" end local sc_standard_exploded = ugsub(sc_standard, page.comb_chars.combined_double, explode) -- The following is correct; it relies on side-effecing the explode_standard[] table. ugsub(sc_standard_exploded, page.comb_chars.combined_single, explode):gsub(".[\128-\191]*", explode) local num_cat_inserted for char in pairs(page.explode_pagename) do if not explode_standard[char] then if char:find("[0-9]") then if not num_cat_inserted then insert(data.categories, "Perkataan dieja dengan nombor bahasa " .. full_langname) num_cat_inserted = true end elseif ufind(char, page.emoji_pattern) then insert(data.categories, "Perkataan dieja dengan emoji bahasa " .. full_langname) else local upper = char_category(char) if not explode_standard[upper] then char = upper end insert(data.categories, "Perkataan dieja dengan " .. char .. " bahasa " .. full_langname) end end end end -- If a diacritic doesn't appear in any of the standard characters, also categorise for it generally. sc_standard = toNFD(sc_standard) for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_single) do if not umatch(sc_standard, diacritic) then insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname) end end for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_double) do if not umatch(sc_standard, diacritic) then insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname) end end end end -- Ancient Greek, Hindi and Lao handled the old way for now, as their standard chars still need to be converted to the new format (because there are a lot of them). elseif ulen(page.pagename) ~= 1 then for character in ugmatch(page.pagename, "([^" .. standard .. "])") do local upper = char_category(character) if not umatch(upper, "[" .. standard .. "]") then character = upper end insert(data.categories, "Perkataan dieja dengan " .. character .. " bahasa " .. full_langname) end end end if not is_varform_only and data.heads[1].sc:isSystem("alphabet") then local pagename, i = page.pagename:ulower(), 2 while umatch(pagename, "(%a)" .. ("%1"):rep(i)) do i = i + 1 insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan " .. i .. " contoh huruf yang sama berturut-turut") end end -- Categorise for palindromes if not is_varform_only and not data.nopalindromecat and namespace ~= "Rekonstruksi" and ulen(page.pagename) > 2 -- FIXME: Use of first script here seems hacky. What is the clean way of doing this in the presence of -- multiple scripts? and is_palindrome(page.pagename, data.lang, data.heads[1].sc) then insert(data.categories, "Palindrom bahasa " .. full_langname) end if namespace == "" and not lang_reconstructed then for _, head in ipairs(data.heads) do if page.full_raw_pagename ~= get_link_page(remove_links(head.term), data.lang, head.sc) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch/LANGCODE]] track("pagename spelling mismatch", data.lang) break end end end -- Add red link category if called for and we're not a "large" page, where such checks are disabled. if data.checkredlinks and not m_headword_data.large_pages[m_headword_data.pagename] then local plposcat = type(data.checkredlinks) == "string" and data.checkredlinks or data.pos_category check_red_link_inflections_top_level(data, plposcat) end -- Add to various maintenance categories. export.maintenance_cats(page, data.lang, data.categories, data.whole_page_categories) ------------ 10. Format and return headwords, genders, inflections and categories. ------------ -- Format and return all the gathered information. This may add more categories (e.g. gender/number categories), -- so make sure we do it before evaluating `data.categories`. local text = '<span class="headword-line">' .. format_headword(data) .. format_headword_genders(data, is_varform_only) .. format_top_level_inflections(data) .. '</span>' -- Language-specific categories. local cats = format_categories( data.categories, data.lang, data.sort_key, page.encoded_pagename, data.force_cat_output or test_force_categories, data.heads[1].sc ) -- Language-agnostic categories. local whole_page_cats = format_categories( data.whole_page_categories, nil, "-" ) return text .. cats .. whole_page_cats end return export dsnbdd62bmi4jy62atkjng0gw1w5e6o 375385 375353 2026-09-22T07:10:05Z Hakimi97 2668 Betulkan ejaan 375385 Scribunto text/plain local export = {} -- Named constants for all modules used, to make it easier to swap out sandbox versions. local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local gender_and_number_module = "Module:gender and number" local headword_data_module = "Module:headword/data" local headword_page_module = "Module:headword/page" local links_module = "Module:links" local load_module = "Module:load" local pages_module = "Module:pages" local palindromes_module = "Module:palindromes" local scripts_module = "Module:scripts" local scripts_data_module = "Module:scripts/data" local script_utilities_module = "Module:script utilities" local script_utilities_data_module = "Module:script utilities/data" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local dump = mw.dumpObject local insert = table.insert local ipairs = ipairs local max = math.max local new_title = mw.title.new local pairs = pairs local require = require local toNFC = mw.ustring.toNFC local toNFD = mw.ustring.toNFD local type = type local ufind = mw.ustring.find local ugmatch = mw.ustring.gmatch local ugsub = mw.ustring.gsub local umatch = mw.ustring.match --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==] local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function contains(...) contains = require(table_module).contains return contains(...) end local function encode_entities(...) encode_entities = require(string_utilities_module).encode_entities return encode_entities(...) end local function extend(...) extend = require(table_module).extend return extend(...) end local function find_best_script_without_lang(...) find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang return find_best_script_without_lang(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function format_genders(...) format_genders = require(gender_and_number_module).format_genders return format_genders(...) end local function format_decorations(...) format_decorations = require(decorations_module).format_decorations return format_decorations(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_current_L2(...) get_current_L2 = require(pages_module).get_current_L2 return get_current_L2(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function is_palindrome(...) is_palindrome = require(palindromes_module).is_palindrome return is_palindrome(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function load_data(...) load_data = require(load_module).load_data return load_data(...) end local function pattern_escape(...) pattern_escape = require(string_utilities_module).pattern_escape return pattern_escape(...) end local function pluralize(...) pluralize = require(en_utilities_module).pluralize return pluralize(...) end local function process_page(...) process_page = require(headword_page_module).process_page return process_page(...) end local function remove_links(...) remove_links = require(links_module).remove_links return remove_links(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function tag_text(...) tag_text = require(script_utilities_module).tag_text return tag_text(...) end local function tag_transcription(...) tag_transcription = require(script_utilities_module).tag_transcription return tag_transcription(...) end local function tag_translit(...) tag_translit = require(script_utilities_module).tag_translit return tag_translit(...) end local function trim(...) trim = require(string_utilities_module).trim return trim(...) end local function ulen(...) ulen = require(string_utilities_module).len return ulen(...) end local function ucfirst(...) ucfirst = require(string_utilities_module).ucfirst return ucfirst(...) end --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local m_data local function get_data() m_data = load_data(headword_data_module) return m_data end local script_data local function get_script_data() script_data = load_data(scripts_data_module) return script_data end local script_utilities_data local function get_script_utilities_data() script_utilities_data = load_data(script_utilities_data_module) return script_utilities_data end -- If set to true, categories always appear, even in non-mainspace pages local test_force_categories = false -- Add a tracking category to track entries with certain (unusually undesirable) properties. `track_id` is an identifier -- for the particular property being tracked and goes into the tracking page. Specifically, this adds a link in the -- page text to [[Wiktionary:Tracking/headword/TRACK_ID]], meaning you can find all entries with the `track_id` property -- by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID]]. -- -- If `lang` (a language object) is given, an additional tracking page [[Wiktionary:Tracking/headword/TRACK_ID/CODE]] is -- linked to where CODE is the language code of `lang`, and you can find all entries in the combination of `track_id` -- and `lang` by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID/CODE]]. This makes it possible to -- isolate only the entries with a specific tracking property that are in a given language. Note that if `lang` -- references at etymology-only language, both that language's code and its full parent's code are tracked. local function track(track_id, lang) local tracking_page = "headword/" .. track_id if lang and lang:hasType("etymology-only") then debug_track{tracking_page, tracking_page .. "/" .. lang:getCode(), tracking_page .. "/" .. lang:getFullCode()} elseif lang then debug_track{tracking_page, tracking_page .. "/" .. lang:getCode()} else debug_track(tracking_page) end return true end local function text_in_script(text, script_code) local sc = get_script(script_code) if not sc then error("Internal error: Bad script code " .. script_code) end local characters = sc.characters local out if characters then text = ugsub(text, "%W", "") out = ufind(text, "[" .. characters .. "]") end if out then return true else return false end end local spacingPunctuation = "[%s%p]+" --[[ List of punctuation or spacing characters that are found inside of words. Used to exclude characters from the regex above. ]] local wordPunc = "-#%%&@־׳״'.·*’་•:᠊" local notWordPunc = "[^" .. wordPunc .. "]+" --[=[ Format a term (either a head term or an inflection term) along with any decorations (left or right qualifiers, labels, references or customized separator). `part` is the object specifying the term (and `lang` the language of the term), which should optionally contain: * left qualifiers in `q`, an array of strings; * right qualifiers in `qq`, an array of strings; * left labels in `l`, an array of strings; * right labels in `ll`, an array of strings; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; * a separator in `separator`, defaulting to " <i>or</i> " if this is not the first term (j > 1), otherwise "". `formatted` is the formatted version of the term itself, and `j` is the index of the term. ]=] local function format_term_with_decorations(lang, part, formatted, j) local function part_non_empty(field) local list = part[field] if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `part.%s`=%s, should be \"table\""):format(field, dump(list))) end return list[1] end if part_non_empty("q") or part_non_empty("qq") or part_non_empty("l") or part_non_empty("ll") or part_non_empty("refs") then formatted = format_decorations { lang = lang, text = formatted, q = part.q, qq = part.qq, l = part.l, ll = part.ll, refs = part.refs, } end local separator = part.separator or j > 1 and " <i>atau</i> " -- use "" to request no separator if separator then formatted = separator .. formatted end return formatted end --[==[Return true if the given head is multiword according to the algorithm used in full_headword().]==] function export.head_is_multiword(head) for possibleWordBreak in ugmatch(head, spacingPunctuation) do if umatch(possibleWordBreak, notWordPunc) then return true end end return false end do local function workaround_to_exclude_chars(s) return (ugsub(s, notWordPunc, "\2%1\1")) end --[==[ Add appropriate links to `head`, correctly handling multiword terms. This is intended for multiword terms but can be used for any term if you want links added to single-word terms as well. If you want to only add links to multiword terms, first check that the term is multiword using `head_is_multiword`. If `default` is specified, this will escape colons so that they don't get interpreted as interwiki links. This should generally only be used when `head` is an actual pagename or is taken from a {{para|pagename}} parameter, not when taken from a {{para|head}} parameter. ]==] function export.add_multiword_links(head, default) head = "\1" .. ugsub(head, spacingPunctuation, workaround_to_exclude_chars) .. "\2" if default then head = head :gsub("(\1[^\2]*)\\([:#][^\2]*\2)", "%1\\\\%2") :gsub("(\1[^\2]*)([:#][^\2]*\2)", "%1\\%2") end --Escape any remaining square brackets to stop them breaking links (e.g. "[citation needed]"). head = encode_entities(head, "[]", true, true) --[=[ use this when workaround is no longer needed: head = "[[" .. ugsub(head, WORDBREAKCHARS, "]]%1[[") .. "]]" Remove any empty links, which could have been created above at the beginning or end of the string. ]=] return (head :gsub("\1\2", "") :gsub("[\1\2]", {["\1"] = "[[", ["\2"] = "]]"})) end end local function non_categorizable(full_raw_pagename) return full_raw_pagename:find("^Lampiran:Gerak isyarat/") or -- Unsupported titles with descriptive names. (full_raw_pagename:find("^Tajuk tidak disokong/") and not full_raw_pagename:find("`")) end local function tag_text_and_add_decorations(data, head, formatted, j) -- Add language and script wrapper. formatted = tag_text(formatted, data.lang, head.sc, "head", nil, j == 1 and data.id or nil) -- Add decorations (qualifiers, labels, references and separator). return format_term_with_decorations(data.lang, head, formatted, j) end -- Format a headword with transliterations. local function format_headword(data) -- Are there non-empty transliterations? local has_translits = false local has_manual_translits = false ------ Format the headwords. ------ local head_parts = {} local unique_head_parts = {} local has_multiple_heads = not not data.heads[2] for j, head in ipairs(data.heads) do if head.tr or head.ts then has_translits = true end if head.tr and head.tr_manual or head.ts then has_manual_translits = true end local formatted -- Apply processing to the headword, for formatting links and such. if head.term:find("[[", nil, true) and head.sc:getCode() ~= "Image" then formatted = language_link{term = head.term, lang = data.lang} else formatted = data.lang:makeDisplayText(head.term, head.sc, true) end local head_part = tag_text_and_add_decorations(data, head, formatted, j) insert(head_parts, head_part) -- If multiple heads, try to determine whether all heads display the same. To do this we need to effectively -- rerun the text tagging and addition of decorations, using 1 for all indices. if has_multiple_heads then local unique_head_part if j == 1 then unique_head_part = head_part else unique_head_part = tag_text_and_add_decorations(data, head, formatted, 1) end unique_head_parts[unique_head_part] = true end end local set_size = 0 if has_multiple_heads then for _ in pairs(unique_head_parts) do set_size = set_size + 1 end end if set_size == 1 then head_parts = head_parts[1] else head_parts = concat(head_parts) end if has_manual_translits then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr/LANGCODE]] track("manual-tr", data.lang) end ------ Format the transliterations and transcriptions. ------ local translits_formatted if has_translits then local translit_parts = {} for _, head in ipairs(data.heads) do if head.tr or head.ts then local this_parts = {} if head.tr then insert(this_parts, tag_translit(head.tr, data.lang:getCode(), "head", nil, head.tr_manual)) if head.ts then insert(this_parts, " ") end end if head.ts then insert(this_parts, "/" .. tag_transcription(head.ts, data.lang:getCode(), "head") .. "/") end insert(translit_parts, concat(this_parts)) end end translits_formatted = " (" .. concat(translit_parts, " <i>atau</i> ") .. ")" local function format_transliteration_page_link(langname) local transliteration_pagename = "Transliterasi bahasa " .. langname local transliteration_page = new_title(transliteration_pagename, "Wikikamus") if transliteration_page and transliteration_page:getContent() then return ("[[Wikikamus:%s|•]]"):format(transliteration_pagename) end return nil end local translit_page_link = format_transliteration_page_link(data.lang:getCanonicalName()) -- If data.lang is an etymology-only language and we didn't find a translation page for it, fall back to the -- full parent. if not translit_page_link and data.lang:hasType("etymology-only") then translit_page_link = format_transliteration_page_link(data.lang:getFullName()) end if translit_page_link then translits_formatted = " " .. translit_page_link .. translits_formatted end else translits_formatted = "" end ------ Paste heads and transliterations/transcriptions. ------ local lemma_gloss if data.gloss then lemma_gloss = ' <span class="ib-content qualifier-content">' .. data.gloss .. '</span>' else lemma_gloss = "" end return head_parts .. translits_formatted .. lemma_gloss end local function format_headword_genders(data, is_varform_only) local retval = "" if data.genders and data.genders[1] then if data.gloss then retval = "," end local pos_for_cat if not data.nogendercat and not is_varform_only then local no_gender_cat = (m_data or get_data()).no_gender_cat if not (no_gender_cat[data.lang:getCode()] or no_gender_cat[data.lang:getFullCode()]) then pos_for_cat = (m_data or get_data()).pos_for_gender_number_cat[data.pos_category:gsub("^reconstructed ", "")] end end local text, cats = format_genders(data.genders, data.lang, pos_for_cat) if cats then extend(data.categories, cats) end retval = retval .. "&nbsp;" .. text end return retval end -- Forward reference local format_inflections local function format_inflection_parts(data, parts) for j, part in ipairs(parts) do if type(part) ~= "table" then part = {term = part} end local partaccel = part.accel local face = part.face or "bold" if face ~= "bold" and face ~= "plain" and face ~= "hypothetical" then error("The face `" .. face .. "` " .. ( (script_utilities_data or get_script_utilities_data()).faces[face] and "should not be used for non-headword terms on the headword line." or "is invalid." )) end -- Here the final part 'or data.nolinkinfl' allows to have 'nolinkinfl=true' -- right into the 'data' table to disable inflection links of the entire headword -- when inflected forms aren't entry-worthy, e.g.: in Vulgar Latin local nolinkinfl = part.face == "hypothetical" or (part.nolink and track("nolink") or part.nolinkinfl) or ( data.nolink and track("nolink") or data.nolinkinfl) local formatted if part.label then -- FIXME: There should be a better way of italicizing a label. As is, this isn't customizable. formatted = "<i>" .. part.label .. "</i>" else -- Convert the term into a full link. Don't show a transliteration here unless enable_auto_translit is -- requested, either at the `parts` level (i.e. per inflection) or at the `data.inflections` level (i.e. -- specified for all inflections). This is controllable in {{head}} using autotrinfl=1 for all inflections, -- or fNautotr=1 for an individual inflection (remember that a single inflection may be associated with -- multiple terms). The reason for doing this is to avoid clutter in headword lines by default in languages -- where the script is relatively straightforward to read by learners (e.g. Greek, Russian), but allow it -- to be enabled in languages with more complex scripts (e.g. Arabic). -- -- FIXME: With nested inflections, should we also respect `enable_auto_translit` at the top level of the -- nested inflections structure? local tr = part.tr or not (parts.enable_auto_translit or data.inflections.enable_auto_translit) and "-" or nil local postprocess_annotations if part.inflections then postprocess_annotations = function(infldata) insert(infldata.annotations, format_inflections(data, part.inflections)) end end formatted = full_link( { term = not nolinkinfl and part.term or nil, alt = part.alt or (nolinkinfl and part.term or nil), lang = part.lang or data.lang, sc = part.sc or parts.sc or nil, gloss = part.gloss, pos = part.pos, lit = part.lit, id = part.id, genders = part.genders, tr = tr, ts = part.ts, accel = partaccel or parts.accel, postprocess_annotations = postprocess_annotations, }, face ) end parts[j] = format_term_with_decorations(part.lang or data.lang, part, formatted, j) end local parts_output if parts[1] then parts_output = (parts.label and " " or "") .. concat(parts) elseif parts.request then parts_output = " <small>[sila nyatakan]</small>" insert(data.categories, "Permohonan fleksi dalam entri bahasa " .. data.lang:getFullName()) else parts_output = "" end local parts_label = parts.label and ("<i>" .. parts.label .. "</i>") or "" return format_term_with_decorations(data.lang, parts, parts_label .. parts_output, 1) end -- Format the inflections following the headword or nested after a given inflection. Declared local above. function format_inflections(data, inflections) if inflections and inflections[1] then -- Format each inflection individually. for key, infl in ipairs(inflections) do inflections[key] = format_inflection_parts(data, infl) end return concat(inflections, ", ") else return "" end end -- Format the top-level inflections following the headword. Currently this just adds parens around the -- formatted comma-separated inflections in `data.inflections`. local function format_top_level_inflections(data) local result = format_inflections(data, data.inflections) if result ~= "" then return " (" .. result .. ")" else return result end end -- Forward reference local check_red_link_inflections -- Check a single inflection (which consists of a label and zero or more terms, each possibly with nested inflections) -- for red links. If so, insert a red-link category based on `plpos` (the plural part of speech to insert in the -- category), stop further processing, and return true. If no red links found, return false. local function check_red_link_inflection_parts(data, parts, plpos) for _, part in ipairs(parts) do if type(part) ~= "table" then part = {term = part} end local term = part.term if term and not term:find("%[%[") then local stripped_physical_term = get_link_page(term, data.lang, part.sc or parts.sc or nil) if stripped_physical_term then local title = mw.title.new(stripped_physical_term) if title and not title:getContent() then insert(data.categories, data.lang:getFullName() .. " " .. plpos .. " with red links in their headword lines") return true end end end if part.inflections then if check_red_link_inflections(data, part.inflections, plpos) then return true end end end return false end -- Check a set of inflections (each of which describes a single inflection of the term, such as feminine or plural, and -- consists of a label and zero or more terms, each possibly with nested inflections) for red links. If so, insert a -- red-link category based on `plpos` (the plural part of speech to insert in the category), stop further processing, -- and return true. If no red links found, return false. function check_red_link_inflections(data, inflections, plpos) if inflections and inflections[1] then -- Check each inflection individually. for key, infl in ipairs(inflections) do if check_red_link_inflection_parts(data, infl, plpos) then return true end end end return false end -- Check the top-level inflections in `data.inflections`, along with any nested inflections, for red links. If so, -- insert a red-link category based on `plpos` (the plural part of speech to insert in the category), stop further -- processing, and return true. If no red links found, return false. local function check_red_link_inflections_top_level(data, plpos) return check_red_link_inflections(data, data.inflections, plpos) end --[==[ Returns the plural form of `pos`, a raw part of speech input, which could be singular or plural. Irregular plural POS are taken into account (e.g. "kanji" pluralizes to "kanji"). ]==] function export.pluralize_pos(pos) -- Make the plural form of the part of speech return (m_data or get_data()).irregular_plurals[pos] or pos:sub(-1) == "s" and pos or pluralize(pos) end --[==[ Return "lemma" if the given POS is a lemma, "non-lemma form" if a non-lemma form, or nil if unknown. The POS passed in must be in its plural form ("nouns", "prefixes", etc.). If you have a POS in its singular form, call {export.pluralize_pos()} above to pluralize it in a smart fashion that knows when to add "-s" and when to add "-es", and also takes into account any irregular plurals. If `best_guess` is given and the POS is in neither the lemma nor non-lemma list, guess based on whether it ends in " forms"; otherwise, return nil. ]==] function export.pos_lemma_or_nonlemma(plpos, best_guess) local m_headword_data = m_data or get_data() local isLemma = m_headword_data.lemmas -- Is it a lemma category? if isLemma[plpos] then return "Lema" end local plpos_no_recon = plpos:gsub("^reconstructed ", "") if isLemma[plpos_no_recon] then return "Lema" end -- Is it a nonlemma category? local isNonLemma = m_headword_data.nonlemmas if isNonLemma[plpos] or isNonLemma[plpos_no_recon] then return "Bentuk bukan lema" end local plpos_no_mut = plpos:gsub("^mutated ", "") if isLemma[plpos_no_mut] or isNonLemma[plpos_no_mut] then return "Bentuk bukan lema" elseif best_guess then return plpos:find("^Bentuk ") and "Bentuk bukan lema" or "Lema" else return nil end end --[==[ Canonicalize a part of speech as specified in 2= in {{tl|head}}. This checks for POS aliases and non-lemma form aliases ending in 'f', and then pluralizes if the POS term does not have an invariable plural. ]==] function export.canonicalize_pos(pos) -- FIXME: Temporary code to throw an error for alias 'pre' (= preposition) that will go away. if pos == "pre" then -- Don't throw error on 'pref' as it's an alias for "prefix". error("POS 'pre' for 'preposition' no longer allowed as it's too ambiguous; use 'prep'") end -- Likewise for pro = pronoun. if pos == "pro" or pos == "prof" then error("POS 'pro' for 'pronoun' no longer allowed as it's too ambiguous; use 'pron'") end local m_headword_data = m_data or get_data() if m_headword_data.pos_aliases[pos] then pos = m_headword_data.pos_aliases[pos] elseif pos:sub(-1) == "f" then pos = pos:sub(1, -2) pos = "Bentuk " .. (m_headword_data.pos_aliases[pos] or pos) end return export.pluralize_pos(pos) end -- Find and return the maximum index in the array `data[element]` (which may have gaps in it), and initialize it to a -- zero-length array if unspecified. Check to make sure all keys are numeric (other than "maxindex", which is set by -- [[Module:parameters]] for list parameters), all values are strings, and unless `allow_blank_string` is given, -- no blank (zero-length) strings are present. local function init_and_find_maximum_index(data, element, allow_blank_string) local maxind = 0 if not data[element] then data[element] = {} end local typ = type(data[element]) if typ ~= "table" then error(("Internal error: In full_headword(), `data.%s` must be an array but is a %s"):format(element, typ)) end for k, v in pairs(data[element]) do if k ~= "maxindex" then if type(k) ~= "number" then error(("Internal error: Unrecognized non-numeric key '%s' in `data.%s`"):format(k, element)) end if k > maxind then maxind = k end if v then if type(v) ~= "string" then error(("Internal error: For key '%s' in `data.%s`, value should be a string but is a %s"):format(k, element, type(v))) end if not allow_blank_string and v == "" then error(("Internal error: For key '%s' in `data.%s`, blank string not allowed; use 'false' for the default"):format(k, element)) end end end end return maxind end --[==[ -- Add the page to various maintenance categories for the language and the -- whole page. These are placed in the headword somewhat arbitrarily, but -- mainly because headword templates are mandatory for entries (meaning that -- in theory it provides full coverage). -- -- This is provided as an external entry point so that modules which transclude -- information from other entries (such as {{tl|ja-see}}) can take advantage -- of this feature as well, because they are used in place of a conventional -- headword template.]==] do -- Handle any manual sortkeys that have been specified in raw categories -- by tracking if they are the same or different from the automatically- -- generated sortkey, so that we can track them in maintenance -- categories. local function handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) sortkey = sortkey or lang:makeSortKey(page.pagename) -- If there are raw categories with no sortkey, then they will be -- sorted based on the default MediaWiki sortkey, so we check against -- that. if tbl == true then if page.raw_defaultsort ~= sortkey then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik") end return end local redundant, different for k in pairs(tbl) do if k == sortkey then redundant = true else different = true end end if redundant then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih lewah") end if different then insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik") end return sortkey end function export.maintenance_cats(page, lang, lang_cats, page_cats) extend(page_cats, page.cats) lang = lang:getFull() -- since we are just generating categories local canonical = lang:getCanonicalName() local tbl = page.wikitext_topic_cat[lang:getCode()] local sortkey = nil if tbl then sortkey = handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori topik yang menggunakan penanda mentah") end tbl = page.wikitext_langname_cat[canonical] if tbl then handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats) insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori nama bahasa yang menggunakan penanda mentah") end if get_current_L2() ~= "Bahasa " .. canonical then insert(lang_cats, "Entri bahasa " .. canonical .. " dengan pengepala bahasa tidak betul") -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header/LANGCODE]] track("pengepala bahasa tidak betul", lang) end end end --[==[This is the primary external entry point. {{lua|full_headword(data)}} This is used by {{temp|head}} and various language-specific headword templates (e.g. {{temp|ru-adj}} for Russian adjectives, {{temp|de-noun}} for German nouns, etc.) to display an entire headword line. See [[#Further explanations for full_headword()]] ]==] function export.full_headword(data) -- Prevent data from being destructively modified. data = shallow_copy(data) ------------ 1. Basic checks for old-style (multi-arg) calling convention. ------------ if data.getCanonicalName then error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) of properties, not a language object") end if not data.lang or type(data.lang) ~= "table" or not data.lang.getCode then error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) and `data.lang` must be a language object") end if data.id and type(data.id) ~= "string" then error("Internal error: The id in the data table should be a string.") end ------------ 2. Initialize pagename etc. ------------ local langcode = data.lang:getCode() local full_langcode = data.lang:getFullCode() local langname = data.lang:getCanonicalName() local full_langname = data.lang:getFullName() local raw_pagename = data.pagename local page local m_headword_data = m_data or get_data() if raw_pagename and raw_pagename ~= m_headword_data.pagename then -- for testing, doc pages, etc. -- data.pagename is often set on documentation and test pages through the pagename= parameter of various -- templates, to emulate running on that page. Having a large number of such test templates on a single -- page often leads to timeouts, because we fetch and parse the contents of each page in turn. However, -- we don't really need to do that and can function fine without fetching and parsing the contents of a -- given page, so turn off content fetching/parsing (and also setting the DEFAULTSORT key through a parser -- function, which is *slooooow*) in certain namespaces where test and documentation templates are likely to -- be found and where actual content does not live (User, Template, Module). local actual_namespace = m_headword_data.page.namespace local no_fetch_content = actual_namespace == "User" or actual_namespace == "Template" or actual_namespace == "Module" page = process_page(raw_pagename, no_fetch_content) else page = m_headword_data.page end local namespace = page.namespace if data.altform then -- Temporary tracking for use of old altform= track("altform", data.lang) end local is_varform_only = data.var and data.var ~= "both" local is_varform_both = data.var == "both" ------------ 3. Initialize `data.heads` table; if old-style, convert to new-style. ------------ if type(data.heads) == "table" and type(data.heads[1]) == "table" then -- new-style if data.translits or data.transcriptions then error("Internal error: In full_headword(), if `data.heads` is new-style (array of head objects), `data.translits` and `data.transcriptions` cannot be given") end else -- convert old-style `heads`, `translits` and `transcriptions` to new-style local maxind = max( init_and_find_maximum_index(data, "heads"), init_and_find_maximum_index(data, "translits", true), init_and_find_maximum_index(data, "transcriptions", true) ) for i = 1, maxind do data.heads[i] = { term = data.heads[i], tr = data.translits[i], ts = data.transcriptions[i], } end end -- Make sure there's at least one head. if not data.heads[1] then data.heads[1] = {} end ------------ 4. Initialize and validate `data.categories` and `data.whole_page_categories`, and determine `pos_category` if not given, and add basic categories. ------------ init_and_find_maximum_index(data, "categories") init_and_find_maximum_index(data, "whole_page_categories") local pos_category_already_present = false if data.categories[1] then local escaped_langname = pattern_escape(full_langname) local matches_lang_pattern = "^" .. escaped_langname .. " " for _, cat in ipairs(data.categories) do -- Does the category begin with the language name? If not, tag it with a tracking category. if not cat:find(matches_lang_pattern) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category/LANGCODE]] track("no lang category", data.lang) end end -- If `pos_category` not given, try to infer it from the first specified category. If this doesn't work, we -- throw an error below. if not data.pos_category and data.categories[1]:find(matches_lang_pattern) then data.pos_category = data.categories[1]:gsub(matches_lang_pattern, "") -- Optimization to avoid inserting category already present. pos_category_already_present = true end end if not data.pos_category then error("Internal error: `data.pos_category` not specified and could not be inferred from the categories given in " .. "`data.categories`. Either specify the plural part of speech in `data.pos_category` " .. "(e.g. \"proper nouns\") or ensure that the first category in `data.categories` is formed from the " .. "language's canonical name plus the plural part of speech (e.g. \"Norwegian Bokmål proper nouns\")." ) end -- Insert a category at the beginning for the part of speech unless it's already present or `data.noposcat` given. if not pos_category_already_present and not data.noposcat and not is_varform_only then local pos_category = ucfirst(data.pos_category) .. " bahasa " .. full_langname -- FIXME: [[User:Theknightwho]] Why is this special case here? Please add an explanatory comment. if pos_category ~= "Aksara Han rentas bahasa" then insert(data.categories, 1, pos_category) end end -- Try to determine whether the part of speech refers to a lemma or a non-lemma form; if we can figure this out, -- add an appropriate category. local postype = export.pos_lemma_or_nonlemma(data.pos_category) if not postype then -- We don't know what this category is, so tag it with a tracking category. -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/LANGCODE]] track("unrecognized pos", data.lang) -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS/LANGCODE]] track("unrecognized pos/pos/" .. data.pos_category, data.lang) elseif not data.noposcat and not is_varform_only then insert(data.categories, 1, ucfirst(postype) .. " bahasa " .. full_langname) end -- Categorize variant forms into 'variant lemmas' or 'variant non-lemma forms'. Originally proposed in -- [[Wiktionary:Beer parlour/2024/June#Decluttering the altform mess]] as 'alternative forms'; renamed in -- [[Wiktionary:Beer parlour/2026/July#Renaming the "alternative forms" categories]]. if (is_varform_only or is_varform_both) and postype then insert(data.categories, 1, postype .. " kelainan bahasa " .. full_langname) end ------------ 5. Create a default headword, and add links to multiword page names. ------------ -- Determine if this is an "anti-asterisk" term, i.e. an attested term in a language that must normally be -- reconstructed. local is_anti_asterisk = data.heads[1].term and data.heads[1].term:find("^!!") local lang_reconstructed = data.lang:hasType("reconstructed") if is_anti_asterisk then if not lang_reconstructed then error("Anti-asterisk feature (head= beginning with !!) can only be used with reconstructed languages") end lang_reconstructed = false end -- Determine if term is reconstructed local is_reconstructed = namespace == "Rekonstruksi" or lang_reconstructed -- Create a default headword based on the pagename, which is determined in -- advance by the data module so that it only needs to be done once. local default_head = page.pagename -- Add links to multi-word page names when appropriate if not (is_reconstructed or data.nolinkhead) then local no_links = m_headword_data.no_multiword_links if not (no_links[langcode] or no_links[full_langcode]) and export.head_is_multiword(default_head) then default_head = export.add_multiword_links(default_head, true) end end if is_reconstructed then default_head = "*" .. default_head end ------------ 6. Check the namespace against the language type. ------------ if namespace == "" then if lang_reconstructed then error("Entries in " .. langname .. " must be placed in the Rekonstruksi: namespace") elseif data.lang:hasType("appendix-constructed") then error("Entries in " .. langname .. " must be placed in the Lampiran: namespace") end elseif namespace == "Petikan" or namespace == "Tesaurus" then error("Headword templates should not be used in the " .. namespace .. ": namespace.") end ------------ 7. Fill in missing values in `data.heads`. ------------ -- True if any script among the headword scripts has spaces in it. local any_script_has_spaces = false -- True if any term has a redundant head= param. local has_redundant_head_param = false for _, head in ipairs(data.heads) do ------ 7a. If missing head, replace with default head. if not head.term then head.term = default_head elseif head.term == default_head then has_redundant_head_param = true elseif is_anti_asterisk and head.term == "!!" then -- If explicit head=!! is given, it's an anti-asterisk term and we fill in the default head. head.term = "!!" .. default_head elseif head.term:find("^[!?]$") then -- If explicit head= just consists of ! or ?, add it to the end of the default head. head.term = default_head .. head.term end head.term_no_initial_bang_bang = is_anti_asterisk and head.term:sub(3) or head.term if is_reconstructed then local head_term = head.term if head_term:find("%[%[") then head_term = remove_links(head_term) end if head_term:sub(1, 1) ~= "*" then error("The headword '" .. head_term .. "' must begin with '*' to indicate that it is reconstructed.") end end ------ 7b. Try to detect the script(s) if not provided. If a per-head script is provided, that takes precedence, ------ otherwise fall back to the overall script if given. If neither given, autodetect the script. local auto_sc = data.lang:findBestScript(head.term) if ( auto_sc:getCode() == "None" and find_best_script_without_lang(head.term):getCode() ~= "None" ) then insert(data.categories, "Perkataan bahasa " .. full_langname .. " dalam bentuk tulisan tidak piawai") end if not (head.sc or data.sc) then -- No script code given, so use autodetected script. head.sc = auto_sc else if not head.sc then -- Overall script code given. head.sc = data.sc end -- Track uses of sc parameter. if head.sc:getCode() == auto_sc:getCode() then track("redundant script code", data.lang) if not data.no_script_code_cat then insert(data.categories, "Perkataan dengan kod tulisan lewah bahasa " .. full_langname ) end else track("non-redundant manual script code", data.lang) if not data.no_script_code_cat then insert(data.categories, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. full_langname ) end end end -- If using a discouraged character sequence, add to maintenance category. if head.sc:hasNormalizationFixes() == true then local composed_head = toNFC(head.term) if head.sc:fixDiscouragedSequences(composed_head) ~= composed_head then insert(data.whole_page_categories, "Laman menggunakan jujukan aksara tidak digalakkan") end end any_script_has_spaces = any_script_has_spaces or head.sc:hasSpaces() ------ 7c. Create automatic transliterations for any non-Latin headwords without manual translit given ------ (provided automatic translit is available, e.g. not in Persian or Hebrew). -- Make transliterations head.tr_manual = nil -- Try to generate a transliteration if necessary if head.tr == "-" then head.tr = nil else local notranslit = m_headword_data.notranslit if not (notranslit[langcode] or notranslit[full_langcode]) and head.sc:isTransliterated() then head.tr_manual = not not head.tr local text = head.term_no_initial_bang_bang if not data.lang:link_tr(head.sc) then text = remove_links(text) end local automated_tr = data.lang:transliterate(text, head.sc) if automated_tr then local manual_tr = head.tr if manual_tr then if remove_links(manual_tr) == remove_links(automated_tr) then insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi lewah") else insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi manual tidak lewah") end end if not manual_tr then head.tr = automated_tr end end -- There is still no transliteration? -- Add the entry to a cleanup category. if not head.tr then head.tr = "<small>transliterasi diperlukan</small>" -- FIXME: No current support for 'Request for transliteration of Classical Persian terms' or similar. -- Consider adding this support in [[Module:category tree/poscatboiler/data/entry maintenance]]. insert(data.categories, "Permintaan transliterasi perkataan bahasa " .. full_langname) else -- Otherwise, trim it. head.tr = trim(head.tr) end end end -- Link to the transliteration entry for languages that require this. if head.tr and data.lang:link_tr(head.sc) then head.tr = full_link{ term = head.tr, lang = data.lang, sc = get_script("Latn"), tr = "-" } end end ------------ 8. Maybe tag the title with the appropriate script code, using the `display_title` mechanism. ------------ -- Assumes that the scripts in "toBeTagged" will never occur in the Reconstruction namespace. -- (FIXME: Don't make assumptions like this, and if you need to do so, throw an error if the assumption is violated.) -- Avoid tagging ASCII as Hani even when it is tagged as Hani in the headword, as in [[check]]. The check for ASCII -- might need to be expanded to a check for any Latin characters and whitespace or punctuation. local display_title -- Where there are multiple headwords, use the script for the first. This assumes the first headword is similar to -- the pagename, and that headwords that are in different scripts from the pagename aren't first. This seems to be -- about the best we can do (alternatively we could potentially do script detection on the pagename). local dt_script = data.heads[1].sc local dt_script_code = dt_script:getCode() local page_non_ascii = namespace == "" and not page.pagename:find("^[%z\1-\127]+$") local unsupported_pagename, unsupported = page.full_raw_pagename:gsub("^Tajuk tidak disokong/", "") if unsupported == 1 and page.unsupported_titles[unsupported_pagename] then display_title = 'Tajuk tidak disokong/<span class="' .. dt_script_code .. '">' .. page.unsupported_titles[unsupported_pagename] .. '</span>' elseif page_non_ascii and m_headword_data.toBeTagged[dt_script_code] or (dt_script_code == "Jpan" and (text_in_script(page.pagename, "Hira") or text_in_script(page.pagename, "Kana"))) or (dt_script_code == "Kore" and text_in_script(page.pagename, "Hang")) then display_title = '<span class="' .. dt_script_code .. '">' .. page.full_raw_pagename .. '</span>' -- Keep Han entries region-neutral in the display title. elseif page_non_ascii and (dt_script_code == "Hant" or dt_script_code == "Hans") then display_title = '<span class="Hani">' .. page.full_raw_pagename .. '</span>' elseif namespace == "Rekonstruksi" then local matched display_title, matched = ugsub( page.full_raw_pagename, "^(Rekonstruksi:[^/]+/)(.+)$", function(before, term) return before .. tag_text(term, data.lang, dt_script) end ) if matched == 0 then display_title = nil end end -- FIXME: Generalize this. -- If the current language uses Aran (Nastaliq), e.g. Urdu, and there's more than one language on the page, don't -- set the display title because we don't want Nastaliq for terms that also exist in other languages that don't -- display in Nastaliq (e.g. Arabic or Persian). Because the word "Urdu" occurs near the end of the alphabet, Urdu -- fonts tend to override the fonts of other languages. FIXME: This is checking for more than one language on the -- page but instead needs to check if there are any languages using scripts other than Aran. if dt_script_code == "Aran" and page.L2_list.n > 1 then display_title = nil end if display_title then mw.getCurrentFrame():callParserFunction( "DISPLAYTITLE", display_title ) end ------------ 9. Insert additional categories. ------------ if data.force_cat_output then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/force cat output]] track("force cat output") end if has_redundant_head_param then if not data.no_redundant_head_cat then -- This is not the right way to go about this; too many exceptions and problems due to language-specific headword -- handling customization. If we want this, it should be opt-in by a given language passing in the default headword. -- insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan parameter pengepala lewah") end end -- If the first head is multiword (after removing links), maybe insert into "LANG multiword terms". if not data.nomultiwordcat and not is_varform_only and any_script_has_spaces and postype == "lemma" then local no_multiword_cat = m_headword_data.no_multiword_cat if not (no_multiword_cat[langcode] or no_multiword_cat[full_langcode]) then -- Check for spaces or hyphens, but exclude prefixes and suffixes. -- Use the pagename, not the head= value, because the latter may have extra -- junk in it, e.g. superscripted text that throws off the algorithm. local no_hyphen = m_headword_data.hyphen_not_multiword_sep -- Exclude hyphens if the data module states that they should for this language. local checkpattern = (no_hyphen[langcode] or no_hyphen[full_langcode]) and ".[%s፡]." or ".[%s%-፡]." local is_multiword = umatch(page.pagename, checkpattern) if is_multiword and not non_categorizable(page.full_raw_pagename) then insert(data.categories, "Perkataan berbilang kata bahasa " .. full_langname) elseif not is_multiword then local long_word_threshold = m_headword_data.long_word_thresholds[langcode] or m_headword_data.long_word_thresholds[full_langcode] if long_word_threshold and ulen(page.pagename) >= long_word_threshold then insert(data.categories, "Perkataan panjang bahasa " .. full_langname) end end end end -- Determine whether to insert a category 'LANGNAME POS in SCRIPT'. If there are multiple heads, we may need to check -- each head, as the heads may (theoretically) have different scripts. local default_sccat = m_headword_data.default_sccat if data.sccat or not is_varform_only and (default_sccat[langcode] or langcode ~= full_langcode and default_sccat[full_langcode]) then local function needs_sccat(sccat_entry, sc) if sccat_entry == true or not sccat_entry then return sccat_entry end if type(sccat_entry) == "table" then local in_list = contains(sccat_entry, sc:getCode()) if sccat_entry[1] == "not" then in_list = not in_list end return in_list end return nil end for _, head in ipairs(data.heads) do -- First check the `sccat` specified at the {{head}} level. local this_needs_sccat = needs_sccat(data.sccat, head.sc) -- If that wasn't given, check the default sccat at the language level for the lang code. if this_needs_sccat == nil and not is_varform_only then this_needs_sccat = needs_sccat(default_sccat[langcode], head.sc) end -- If that wasn't found and the lang code is an etym code, check the default sccat at the parent language level. if this_needs_sccat == nil and not is_varform_only and langcode ~= full_langcode then this_needs_sccat = needs_sccat(default_sccat[full_langcode], head.sc) end if this_needs_sccat then insert(data.categories, ucfirst(data.pos_category) .. " bahasa " .. full_langname .. " dalam " .. head.sc:getDisplayForm(data.lang)) end end end -- Reconstructed terms often use weird combinations of scripts and realistically aren't spelled so much as notated. if namespace ~= "Rekonstruksi" and not is_varform_only then -- Map from languages to a string containing the characters to ignore when considering whether a term has -- multiple written scripts in it. Typically these are Greek or Cyrillic letters used for their phonetic -- values. local characters_to_ignore = { ["aaq"] = "αάὰ", -- Penobscot (Algonquian) ["acy"] = "δθ", -- Cypriot Arabic ["aez"] = "β", -- Aeka (Trans-New Guinea) ["anc"] = "γ", -- Ngas (Chadic/Afroasiatic) ["aou"] = "χ", -- A'ou (Kra-Dai) ["art-blk"] = "ч", -- Bolak (conlang) ["awg"] = "β", -- Anguthimri (Pama-Nyungan) ["az"] = "ь", -- Azerbaijani (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["ba"] = "ь", -- Bashkir (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["bhp"] = "β", -- Bima (Austronesian) ["bjz"] = "β", -- Baruga (Trans-New Guinea) ["byk"] = "θ", -- Biao (Kra-Dai) ["cdy"] = "θ", -- Chadong (Kra-Dai) ["chp"] = "θ", -- Chipewyan (Athabaskan) ["cjh"] = "χ", -- Upper Chehalis (Salishan) ["clm"] = "χ", -- Klallam (Salishan) ["col"] = "χ", -- Colombia-Wenatchi (Salishan) ["coo"] = "χθ", -- Comox (Salishan) ["crx"] = "θ", -- Carrier (Athabaskan) ["ets"] = "θ", -- Yekhee (Edoid/Niger-Congo) ["ett"] = "χ", -- Etruscan (isolate; in romanizations) ["fla"] = "χ", -- Montana Salish (Salishan) ["grt"] = "་", -- Garo (South Asian Sino-Tibetan) ["gmw-gts"] = "χ", -- Gottscheerish (Bavarian variant spoken in Slovenia) ["hur"] = "χθ", -- Halkomelem (Salishan) ["itc-psa"] = "f", -- Pre-Samnite (Italic; normally written in Greek) ["izh"] = "ь", -- Ingrian (Finnic) ["kic"] = "θ", -- Kickapoo (Algonquian) ["kk"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["ky"] = "ь", -- Kyrgyz (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["lil"] = "χ", -- Lillooet (Salishan) ["lsi"] = "ꓹ", -- Lashi (Lolo-Burmese/Sino-Tibetan; represents a glottal stop) ["mhz"] = "β", -- Mor (Austronesian) ["mqn"] = "β", -- Moronene (Austronesian) ["neg"]= "ӡā", -- Negidal (Tungusic; normally in Cyrillic) ["oka"] = "χ", -- Okanagan (Salishan) ["ole"] = "θ", -- Olekha (Sino-Tibetan) ["oui"] = "γβ", -- Old Uyghur (Turkic; FIXME: others? E.g. Greek delta (δ)?) ["pox"] = "χ", -- Polabian (West Slavic) ["rif"] = "ε", -- Tarifit (Berber) ["rom"] = "Θθ", -- Romani (Indic: International Standard; two different thetas???) ["rpn"] = "β", -- Repanbitip (Austronesian) ["sah"] = "ь", -- Yakut (Turkic; 1929 - 1939 Latin spelling) ["sit-jap"] = "χ", -- Japhug (Sino-Tibetan) ["sjw"] = "θ", -- Shawnee (Algonquian) ["squ"] = "χ", -- Squamish (Salishan) ["str"] = "χθ", -- Saanich (Salishan) ["teh"] = "χ", -- Tehuelche (Chonan; spoken in Argentina) ["tep"] = "η", -- Tepecano (Uto-Aztecan) ["thp"] = "χ", -- Thompson (Salishan) ["tk"] = "ь", -- Turkmen (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["tt"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938) ["twa"] = "χ", -- Twana (Salishan) ["wbl"] = "ы", -- Wakhi (Iranian) ["xbc"] = "ϸ", -- Bactrian (Iranian; represents š; normally written in Greek) ["yha"] = "θ", -- Baha (Kra-Dai) ["za"] = "зч", -- Zhuang (Tai/Kra-Dai); 1957-1982 alphabet used two Cyrillic letters (as well as some others like -- ƃ, ƅ, ƨ, ɯ and ɵ that look like Cyrillic or Greek but are actually Latin) ["zlw-slv"] = "χђћ", -- Slovincian (West Slavic; FIXME: χ is Greek, the other two are Cyrillic, but I'm not sure -- the currect characters are being chosen in the entry names) ["zng"] = "θ", -- Mang (Mon-Khmer) ["ztp"] = "θ", -- Loxicha Zapotec (Zapotecan) } -- Determine how many real scripts are found in the pagename, where we exclude symbols and such. We exclude -- scripts whose `character_category` is false as well as Zmth (mathematical notation symbols), which has a -- category of "Mathematical notation symbols". When counting scripts, we need to elide language-specific -- variants because e.g. Beng and as-Beng have slightly different characters but we don't want to consider them -- two different scripts (e.g. [[এৰ]] has two characters which are detected respectively as Beng and as-Beng). local seen_scripts = {} local num_seen_scripts = 0 local num_loops = 0 local canon_pagename = page.pagename local ch_to_ignore = characters_to_ignore[full_langcode] if ch_to_ignore then canon_pagename = ugsub(canon_pagename, "[" .. ch_to_ignore .. "]", "") end while true do if canon_pagename == "" or num_seen_scripts >= 2 or num_loops >= 10 then break end -- Make sure we don't get into a loop checking the same script over and over again; happens with e.g. [[ᠪᡳ]] num_loops = num_loops + 1 local pagename_script = find_best_script_without_lang(canon_pagename, "None only as last resort") local script_chars = pagename_script.characters if not script_chars then -- we are stuck; this happens with None break end local script_code = pagename_script:getCode() local replaced canon_pagename, replaced = ugsub(canon_pagename, "[" .. script_chars .. "]", "") if ( replaced and script_code ~= "Zmth" and (script_data or get_script_data())[script_code] and script_data[script_code].character_category ~= false ) then script_code = script_code:gsub("^.-%-", "") if not seen_scripts[script_code] then seen_scripts[script_code] = true num_seen_scripts = num_seen_scripts + 1 end end end if num_seen_scripts > 1 then insert(data.categories, "Perkataan bahasa " .. full_langname .. " dieja dalam berbilang tulisan") end end -- Categorise for unusual characters. Takes into account combining characters, so that we can categorise for characters with diacritics that aren't encoded as atomic characters (e.g. U̠). These can be in two formats: single combining characters (i.e. character + diacritic(s)) or double combining characters (i.e. character + diacritic(s) + character). Each can have any number of diacritics. local standard = data.lang:getStandardCharacters() if not is_varform_only and standard and not non_categorizable(page.full_raw_pagename) then local function char_category(char) local specials = { ["#"] = "number sign", ["("] = "parentheses", [")"] = "parentheses", ["<"] = "angle brackets", [">"] = "angle brackets", ["["] = "square brackets", ["]"] = "square brackets", ["_"] = "underscore", ["{"] = "braces", ["|"] = "vertical line", ["}"] = "braces", ["ß"] = "ẞ", ["\205\133"] = "", -- this is UTF-8 for U+0345 ( ͅ) ["\239\191\189"] = "replacement character", } char = toNFD(char) :gsub(".[\128-\191]*", function(m) local new_m = specials[m] new_m = new_m or m:uupper() return new_m end) return toNFC(char) end if full_langcode ~= "hi" and full_langcode ~= "lo" then local standard_chars_scripts = {} for _, head in ipairs(data.heads) do standard_chars_scripts[head.sc:getCode()] = true end -- Iterate over the scripts, in case there is more than one (as they can have different sets of standard characters). for code in pairs(standard_chars_scripts) do local sc_standard = data.lang:getStandardCharacters(code) if sc_standard then if page.pagename_len > 1 then local explode_standard = {} local function explode(char) explode_standard[char] = true return "" end local sc_standard_exploded = ugsub(sc_standard, page.comb_chars.combined_double, explode) -- The following is correct; it relies on side-effecing the explode_standard[] table. ugsub(sc_standard_exploded, page.comb_chars.combined_single, explode):gsub(".[\128-\191]*", explode) local num_cat_inserted for char in pairs(page.explode_pagename) do if not explode_standard[char] then if char:find("[0-9]") then if not num_cat_inserted then insert(data.categories, "Perkataan dieja dengan nombor bahasa " .. full_langname) num_cat_inserted = true end elseif ufind(char, page.emoji_pattern) then insert(data.categories, "Perkataan dieja dengan emoji bahasa " .. full_langname) else local upper = char_category(char) if not explode_standard[upper] then char = upper end insert(data.categories, "Perkataan dieja dengan " .. char .. " bahasa " .. full_langname) end end end end -- If a diacritic doesn't appear in any of the standard characters, also categorise for it generally. sc_standard = toNFD(sc_standard) for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_single) do if not umatch(sc_standard, diacritic) then insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname) end end for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_double) do if not umatch(sc_standard, diacritic) then insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname) end end end end -- Ancient Greek, Hindi and Lao handled the old way for now, as their standard chars still need to be converted to the new format (because there are a lot of them). elseif ulen(page.pagename) ~= 1 then for character in ugmatch(page.pagename, "([^" .. standard .. "])") do local upper = char_category(character) if not umatch(upper, "[" .. standard .. "]") then character = upper end insert(data.categories, "Perkataan dieja dengan " .. character .. " bahasa " .. full_langname) end end end if not is_varform_only and data.heads[1].sc:isSystem("alphabet") then local pagename, i = page.pagename:ulower(), 2 while umatch(pagename, "(%a)" .. ("%1"):rep(i)) do i = i + 1 insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan " .. i .. " contoh huruf yang sama berturut-turut") end end -- Categorise for palindromes if not is_varform_only and not data.nopalindromecat and namespace ~= "Rekonstruksi" and ulen(page.pagename) > 2 -- FIXME: Use of first script here seems hacky. What is the clean way of doing this in the presence of -- multiple scripts? and is_palindrome(page.pagename, data.lang, data.heads[1].sc) then insert(data.categories, "Palindrom bahasa " .. full_langname) end if namespace == "" and not lang_reconstructed then for _, head in ipairs(data.heads) do if page.full_raw_pagename ~= get_link_page(remove_links(head.term), data.lang, head.sc) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch]] -- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch/LANGCODE]] track("pagename spelling mismatch", data.lang) break end end end -- Add red link category if called for and we're not a "large" page, where such checks are disabled. if data.checkredlinks and not m_headword_data.large_pages[m_headword_data.pagename] then local plposcat = type(data.checkredlinks) == "string" and data.checkredlinks or data.pos_category check_red_link_inflections_top_level(data, plposcat) end -- Add to various maintenance categories. export.maintenance_cats(page, data.lang, data.categories, data.whole_page_categories) ------------ 10. Format and return headwords, genders, inflections and categories. ------------ -- Format and return all the gathered information. This may add more categories (e.g. gender/number categories), -- so make sure we do it before evaluating `data.categories`. local text = '<span class="headword-line">' .. format_headword(data) .. format_headword_genders(data, is_varform_only) .. format_top_level_inflections(data) .. '</span>' -- Language-specific categories. local cats = format_categories( data.categories, data.lang, data.sort_key, page.encoded_pagename, data.force_cat_output or test_force_categories, data.heads[1].sc ) -- Language-agnostic categories. local whole_page_cats = format_categories( data.whole_page_categories, nil, "-" ) return text .. cats .. whole_page_cats end return export phbggqxeqxeecc32oqns9mjq2wituwm Modul:links 828 9771 375349 281073 2026-09-22T03:12:32Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92723705|92723705]]) 375349 Scribunto text/plain local export = {} --[=[ [[Unsupported titles]], pages with high memory usage, extraction modules and part-of-speech names are listed at [[Module:links/data]]. Other modules used: [[Module:script utilities]] [[Module:scripts]] [[Module:languages]] and its submodules [[Module:gender and number]] [[Module:debug/track]] ]=] local anchors_module = "Module:anchors" local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local form_of_module = "Module:form of" local gender_and_number_module = "Module:gender and number" local languages_module = "Module:languages" local load_module = "Module:load" local memoize_module = "Module:memoize" local pages_module = "Module:pages" local scripts_module = "Module:scripts" local script_utilities_module = "Module:script utilities" local string_encode_entities_module = "Module:string/encode entities" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local find = string.find local get_current_title = mw.title.getCurrentTitle local insert = table.insert local ipairs = ipairs local match = string.match local new_title = mw.title.new local pairs = pairs local remove = table.remove local sub = string.sub local toNFC = mw.ustring.toNFC local tostring = tostring local type = type local unstrip = mw.text.unstrip local NAMESPACE = get_current_title().nsText local function anchor_encode(...) anchor_encode = require(memoize_module)(mw.uri.anchorEncode, true) return anchor_encode(...) end local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function decode_entities(...) decode_entities = require(string_utilities_module).decode_entities return decode_entities(...) end local function decode_uri(...) decode_uri = require(string_utilities_module).decode_uri return decode_uri(...) end -- Can't yet replace, as the [[Module:string utilities]] version no longer has automatic double-encoding prevention, which requires changes here to account for. local function encode_entities(...) encode_entities = require(string_encode_entities_module) return encode_entities(...) end local function extend(...) extend = require(table_module).extend return extend(...) end local function find_best_script_without_lang(...) find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang return find_best_script_without_lang(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function format_genders(...) format_genders = require(gender_and_number_module).format_genders return format_genders(...) end local function format_decorations(...) format_decorations = require(decorations_module).format_decorations return format_decorations(...) end local function get_current_L2(...) get_current_L2 = require(pages_module).get_current_L2 return get_current_L2(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function language_anchor(...) language_anchor = require(anchors_module).language_anchor return language_anchor(...) end local function load_data(...) load_data = require(load_module).load_data return load_data(...) end local function request_script(...) request_script = require(script_utilities_module).request_script return request_script(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function tag_text(...) tag_text = require(script_utilities_module).tag_text return tag_text(...) end local function tag_translit(...) tag_translit = require(script_utilities_module).tag_translit return tag_translit(...) end local function trim(...) trim = require(string_utilities_module).trim return trim(...) end local function u(...) u = require(string_utilities_module).char return u(...) end local function ulower(...) ulower = require(string_utilities_module).lower return ulower(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end local m_headword_data local function get_headword_data() m_headword_data = load_data("Module:headword/data") return m_headword_data end local function track(page, code) local tracking_page = "links/" .. page debug_track(tracking_page) if code then debug_track(tracking_page .. "/" .. code) end end local function field_non_empty(list, field) if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `termobj.%s`=%s, should be %s\"table\""):format( field, mw.dumpObject(list), field == "q" or field == "qq" and "\"string\" or " or "")) end return list[1] end --[=[ Add any decorations (left or right regular or accent qualifiers, labels or references) to an item. `text` is the item's text (to which to add the decorations) and `itemobj` is the object specifying the item's decorations, which should optionally contain: * left regular qualifiers in `q` (an array of strings or a single string); an empty array will be ignored; * right regular qualifiers in `qq` (an array of strings or a single string); an empty array will be ignored; * left accent qualifiers in `a` (an array of strings); an empty array will be ignored; * right accent qualifiers in `aa` (an array of strings); an empty array will be ignored; * left labels in `l` (an array of strings); an empty array will be ignored; * right labels in `ll` (an array of strings); an empty array will be ignored; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; an empty array will be ignored. `lang` is a language object and is required if any accent qualifiers or labels are given. ]=] local function add_text_decorations(text, itemobj, lang) local q = itemobj.q if type(q) == "string" then q = { q } end local qq = itemobj.qq if type(qq) == "string" then qq = { qq } end if field_non_empty(q, "q") or field_non_empty(qq, "qq") or field_non_empty(itemobj.a, "a") or field_non_empty(itemobj.aa, "aa") or field_non_empty(itemobj.l, "l") or field_non_empty(itemobj.ll, "ll") or field_non_empty(itemobj.refs, "refs") then text = format_decorations { lang = lang, text = text, q = q, qq = qq, a = itemobj.a, aa = itemobj.aa, l = itemobj.l, ll = itemobj.ll, refs = itemobj.refs, } end return text end local function selective_trim(...) -- Unconditionally trimmed charset. local always_trim = "\194\128-\194\159" .. -- U+0080-009F (C1 control characters) "\194\173" .. -- U+00AD (soft hyphen) "\226\128\170-\226\128\174" .. -- U+202A-202E (directionality formatting characters) "\226\129\166-\226\129\169" -- U+2066-2069 (directionality formatting characters) -- Standard trimmed charset. local standard_trim = "%s" .. -- (default whitespace charset) "\226\128\139-\226\128\141" .. -- U+200B-200D (zero-width spaces) always_trim -- If there are non-whitespace characters, trim all characters in `standard_trim`. -- Otherwise, only trim the characters in `always_trim`. selective_trim = function(text) if text == "" then return text end local trimmed = trim(text, standard_trim) if trimmed ~= "" then return trimmed end return trim(text, always_trim) end return selective_trim(...) end local function escape(text, str) local rep repeat text, rep = text:gsub("\\\\(\\*" .. str .. ")", "\5%1") until rep == 0 return (text:gsub("\\" .. str, "\6")) end local function unescape(text, str) return (text :gsub("\5", "\\") :gsub("\6", str)) end -- Remove bold, italics, soft hyphens, strip markers and HTML tags. local function remove_formatting(str) str = str :gsub("('*)'''(.-'*)'''", "%1%2") :gsub("('*)''(.-'*)''", "%1%2") :gsub("­", "") return (unstrip(str) :gsub("<[^<>]+>", "")) end --[==[ Split `text` on double slashes (taking account of escaping backslashes). Return a list of split items. ]==] function export.split_on_slashes(text) if text:find("\\", nil, true) then track("escaped", "split_on_slashes") end text = split(escape(text, "//"), "//", true) or {} for i, v in ipairs(text) do text[i] = unescape(v, "//") if v == "" then text[i] = false end end return text end --[==[ If `text` is a wikilink (i.e. in the form `<nowiki>[[foo]]</nowiki>` or `<nowiki>[[foo|bar]]</nowiki>`), return the link target and display text. If the wikilink is single-part, the display text will be the same as the link target. If the text is not in wikilink form, return {nil} for both link target and display text. ]==] function export.get_wikilink_parts(text) -- TODO: replace `allow_bad_target` with `allow_unsupported`, with support for links to unsupported titles, including escape sequences. if ( -- Filters out anything but "[[...]]" with no intermediate "[[" or "]]". not match(text, "^()%[%[") or -- Faster than sub(text, 1, 2) ~= "[[". find(text, "[[", 3, true) or find(text, "]]", 3, true) ~= #text - 1 ) then return nil, nil end local pipe = find(text, "|", 3, true) local title, display if pipe then title, display = sub(text, 3, pipe - 1), sub(text, pipe + 1, -3) else title = sub(text, 3, -3) display = title end return title, display end -- Does the work of export.get_fragment, but can be called directly to avoid unnecessary checks for embedded links. local function get_fragment(text) text = escape(text, "#") -- Replace numeric character references with the corresponding character (&#39; → '), -- as they contain #, which causes the numeric character reference to be -- misparsed (wa'a → wa&#39;a → pagename wa&, fragment 39;a). text = decode_entities(text) local target, fragment = text:match("^(.-)#(.+)$") target = target or text target = unescape(target, "#") fragment = fragment and unescape(fragment, "#") or nil return target, fragment end --[==[ Given a link target possibly containing a fragment (i.e. a pound sign and following anchor), return two values, the actual target (minus the fragment) and the fragment, or {nil} if there is no fragment. ]==] function export.get_fragment(text) if text:find("\\", nil, true) then track("escaped", "get_fragment") end -- If there are no embedded links, process input. local open = find(text, "[[", nil, true) if not open then return get_fragment(text) end local close = find(text, "]]", open + 2, true) if not close then return get_fragment(text) -- If there is one, but it's redundant (i.e. encloses everything with no pipe), remove and process. elseif open == 1 and close == #text - 1 and not find(text, "|", 3, true) then return get_fragment(sub(text, 3, -3)) end -- Otherwise, return the input. return text, nil end --[==[ Given a link target as passed to `full_link()`, get the actual page that the target refers to. This removes bold, italics, strip markets and HTML; calls `makeEntryName()` for the language in question; converts targets beginning with `*` to the Reconstruction namespace; and converts appendix-constructed languages to the Appendix namespace. Returns up to three values: # the actual page to link to, or {nil} to not link to anything; # how the target should be displayed as, if the user didn't explicitly specify any display text; generally the same as the original target, but minus any anti-asterisk !!; # the value `true` if the target had a backslash-escaped * in it (FIXME: explain this more clearly). FIXME: This should not be calling deprecated `makeEntryName()` and should always return the same number of values. ]==] function export.get_link_page_with_auto_display(target, lang, sc, plain) local orig_target = target if not target then return nil elseif target:find("\\", nil, true) then track("escaped", "get_link_page") end target = remove_formatting(target) if target:sub(1, 1) == ":" then track("initial colon") -- FIXME, the auto_display (second return value) should probably remove the colon return target:sub(2), orig_target end local prefix = target:match("^(.-):") -- Convert any escaped colons target = target:gsub("\\:", ":") if prefix then -- If this is an a link to another namespace or an interwiki link, ensure there's an initial colon and then -- return what we have (so that it works as a conventional link, and doesn't do anything weird like add the term -- to a category.) prefix = ulower(trim(prefix)) if prefix ~= "" and ( load_data("Module:data/namespaces")[prefix] or load_data("Module:data/interwikis")[prefix] ) then return target, orig_target end end -- Check if the term is reconstructed and remove any asterisk. Also check for anti-asterisk (!!). -- Otherwise, handle the escapes. local reconstructed, escaped, anti_asterisk if not plain then target, reconstructed = target:gsub("^%*(.)", "%1") if reconstructed == 0 then target, anti_asterisk = target:gsub("^!!(.)", "%1") if anti_asterisk == 1 then -- Remove !! from original. FIXME! We do it this way because the call to remove_formatting() above -- may cause non-initial !! to be interpreted as anti-asterisks. We should surely move the -- remove_formatting() call later. orig_target = orig_target:gsub("^!!", "") end end end target, escaped = target:gsub("^(\\-)\\%*", "%1*") if not (sc and sc:getCode() ~= "None") then sc = lang:findBestScript(target) end -- Remove carets if they are used to capitalize parts of transliterations (unless they have been escaped). if (not sc:hasCapitalization()) and sc:isTransliterated() and target:match("%^") then target = escape(target, "^") :gsub("%^", "") target = unescape(target, "^") end -- Get the entry name for the language. target = lang:makeEntryName(target, sc, reconstructed == 1 or lang:hasType("appendix-constructed")) -- If the link contains unexpanded template parameters, then don't create a link. if target:match("{{{.-}}}") then -- FIXME: Should we return the original target as the default display value (second return value)? return nil end -- Link to appendix for reconstructed terms and terms in appendix-only languages. Plain links interpret * -- literally, however. if reconstructed == 1 then if lang:getFullCode() == "und" then -- Return the original target as default display value. If we don't do this, we wrongly get -- [Term?] displayed instead. return nil, orig_target end target = "Rekonstruksi:Bahasa " .. lang:getFullName() .. "/" .. target -- Reconstructed languages and substrates require an initial *. elseif anti_asterisk ~= 1 and (lang:hasType("reconstructed") or lang:getFamilyCode() == "qfa-sub") then error(("The specified language %s is unattested, while the term '%s' does not begin with '*' to indicate that it is reconstructed.") : format(lang:getCanonicalName(), orig_target)) elseif lang:hasType("appendix-constructed") then target = "Lampiran:Bahasa " .. lang:getFullName() .. "/" .. target else target = target end return target, orig_target, escaped > 0 end function export.get_link_page(target, lang, sc, plain) local target, auto_display, escaped = export.get_link_page_with_auto_display(target, lang, sc, plain) return target, escaped end -- Make a link from a given link's parts. local function make_link(link, lang, sc, id, isolated, cats, no_alt_ast, plain) -- Convert percent encoding to plaintext. link.target = link.target and decode_uri(link.target, "PATH") link.fragment = link.fragment and decode_uri(link.fragment, "PATH") -- Find fragments (if one isn't already set). -- Prevents {{l|en|word#Etymology 2|word}} from linking to [[word#Etymology 2#English]]. -- # can be escaped as \#. if link.target and link.fragment == nil then link.target, link.fragment = get_fragment(link.target) end -- Process the target local auto_display, escaped link.target, auto_display, escaped = export.get_link_page_with_auto_display(link.target, lang, sc, plain) -- Create a default display form. -- If the target is "" then it's a link like [[#English]], which refers to the current page. if auto_display == "" then auto_display = (m_headword_data or get_headword_data()).pagename end -- If the display is the target and the reconstruction * has been escaped, remove the escaping backslash. if escaped then auto_display = auto_display:gsub("\\([^\\]*%*)", "%1", 1) end -- Process the display form. if link.display then local orig_display = link.display link.display = lang:makeDisplayText(link.display, sc, true) if cats then auto_display = lang:makeDisplayText(auto_display, sc) -- If the alt text is the same as what would have been automatically generated, then the alt parameter is -- redundant (e.g. {{l|en|foo|foo}}, {{l|en|w:foo|foo}}, but not {{l|en|w:foo|w:foo}}). If they're -- different, but the alt text could have been entered as the term parameter without it affecting the target -- page, then the target parameter is redundant (e.g. {{l|ru|фу|фу́}}). If `no_alt_ast` is true, use pcall -- to catch the error which will be thrown if this is a reconstructed lang and the alt text doesn't have *. if link.display == auto_display then insert(cats, "Pautan dengan parameter alt lewah bahasa " .. lang:getFullName()) else local ok, check if no_alt_ast then ok, check = pcall(export.get_link_page, orig_display, lang, sc, plain) else ok = true check = export.get_link_page(orig_display, lang, sc, plain) end if ok and link.target == check then insert(cats, "Pautan dengan parameter sasaran lewah bahasa " .. lang:getFullName()) end end end else link.display = lang:makeDisplayText(auto_display, sc) end if not link.target then return link.display end -- If the target is the same as the current page, there is no sense id -- and either the language code is "und" or the current L2 is the current -- language then return a "self-link" like the software does. if link.target == get_current_title().prefixedText then local fragment, current_L2 = link.fragment, get_current_L2() if ( fragment and fragment == current_L2 or not (id or fragment) and (lang:getFullCode() == "und" or lang:getFullName() == current_L2) ) then return tostring(mw.html.create("strong") :addClass("selflink") :wikitext(link.display)) end end -- Add fragment. Do not add a section link to "Undetermined", as such sections do not exist and are invalid. -- TabbedLanguages handles links without a section by linking to the "last visited" section, but adding -- "Undetermined" would break that feature. For localized prefixes that make syntax error, please use the -- format: ["xyz"] = true. local prefix = link.target:match("^:*([^:]+):") prefix = prefix and ulower(prefix) if prefix ~= "category" and not (prefix and load_data("Module:data/interwikis")[prefix]) then if (link.fragment or link.target:sub(-1) == "#") and not plain then track("fragment", lang:getFullCode()) if cats then insert(cats, "Pautan dengan serpihan manual bahasa " .. lang:getFullName()) end end if not link.fragment then if id then link.fragment = lang:getFullCode() == "und" and anchor_encode(id) or language_anchor(lang, id) elseif lang:getFullCode() ~= "und" and not (link.target:match("^Lampiran:") or link.target:match("^Rekonstruksi:")) then link.fragment = anchor_encode("Bahasa " .. lang:getFullName()) end end end -- Put inward-facing square brackets around a link to isolated spacing character(s). if isolated and link.display[1] and not umatch(decode_entities(link.display), "%S") then link.display = "&#x5D;" .. link.display .. "&#x5B;" end link.target = link.target:gsub("^(:?)(.*)", function(m1, m2) return m1 .. encode_entities(m2, "#%&+/:<=>@[\\]_{|}") end) link.fragment = link.fragment and encode_entities(remove_formatting(link.fragment), "#%&+/:<=>@[\\]_{|}") return "[[" .. link.target:gsub("^[^:]", ":%0") .. (link.fragment and "#" .. link.fragment or "") .. "|" .. link.display .. "]]" end -- Split a link into its parts. local function parse_link(linktext) local link = { target = linktext } local target = link.target link.target, link.display = target:match("^(..-)|(.+)$") if not link.target then link.target = target link.display = target end -- There's no point in processing these, as they aren't real links. local target_lower = link.target:lower() for _, false_positive in ipairs({ "category", "cat", "file", "image" }) do if target_lower:match("^" .. false_positive .. ":") then return nil end end link.display = decode_entities(link.display) link.target, link.fragment = get_fragment(link.target) -- So that make_link does not look for a fragment again. if not link.fragment then link.fragment = false end return link end local function check_params_ignored_when_embedded(alt, lang, id, cats) if alt then track("alt-ignored") if cats then insert(cats, "Pautan dengan parameter alt dihirau bahasa " .. lang:getFullName()) end end if id then track("id-ignored") if cats then insert(cats, "Pautan dengan parameter id dihirau bahasa " .. lang:getFullName()) end end end -- Find embedded links and ensure they link to the correct section. local function process_embedded_links(text, alt, lang, sc, id, cats, no_alt_ast, plain) -- Process the non-linked text. text = lang:makeDisplayText(text, sc, true) -- If the text begins with * and another character, then act as if each link begins with *. However, don't do this if the * is contained within a link at the start. E.g. `|*[[foo]]` would set all_reconstructed to true, while `|[[*foo]]` would not. local all_reconstructed = false if not plain then -- anchor_encode removes links etc. if anchor_encode(text):sub(1, 1) == "*" then all_reconstructed = true end -- Otherwise, handle any escapes. text = text:gsub("^(\\-)\\%*", "%1*") end check_params_ignored_when_embedded(alt, lang, id, cats) local function process_link(space1, linktext, space2) local capture = "[[" .. linktext .. "]]" local link = parse_link(linktext) -- Return unprocessed false positives untouched (e.g. categories). if not link then return capture end if all_reconstructed then if link.target:find("^!!") then -- Check for anti-asterisk !! at the beginning of a target, indicating that a reconstructed term -- wants a part of the term to link to a non-reconstructed term, e.g. Old English -- {{ang-noun|m|head=*[[!!Crist|Cristes]] [[!!mæsseǣfen]]}}. link.target = link.target:sub(3) -- Also remove !! from the display, which may have been copied from the target (as in mæsseǣfen in -- the example above). link.display = link.display:gsub("^!!", "") elseif not link.target:match("^%*") then link.target = "*" .. link.target end end linktext = make_link(link, lang, sc, id, false, nil, no_alt_ast, plain) :gsub("^%[%[", "\3") :gsub("%]%]$", "\4") return space1 .. linktext .. space2 end -- Use chars 1 and 2 as temporary substitutions, so that we can use charsets. These are converted to chars 3 and 4 by process_link, which means we can convert any remaining chars 1 and 2 back to square brackets (i.e. those not part of a link). text = text :gsub("%[%[", "\1") :gsub("%]%]", "\2") -- If the script uses ^ to capitalize transliterations, make sure that any carets preceding links are on the inside, so that they get processed with the following text. if ( text:find("^", nil, true) and not sc:hasCapitalization() and sc:isTransliterated() ) then text = escape(text, "^") :gsub("%^\1", "\1%^") text = unescape(text, "^") end text = text:gsub("\1(%s*)([^\1\2]-)(%s*)\2", process_link) -- Remove the extra * at the beginning of a language link if it's immediately followed by a link whose display begins with * too. if all_reconstructed then text = text:gsub("^%*\3([^|\1-\4]+)|%*", "\3%1|*") end return (text :gsub("[\1\3]", "[[") :gsub("[\2\4]", "]]") ) end local function simple_link(term, fragment, alt, lang, sc, id, cats, no_alt_ast, suppress_redundant_wikilink_cat) local plain if lang == nil then lang, plain = get_lang("und"), true end -- Get the link target and display text. If the term is the empty string, treat the input as a link to the current page. if term == "" then term = get_current_title().prefixedText elseif term then local new_term, new_alt = export.get_wikilink_parts(term) if new_term then check_params_ignored_when_embedded(alt, lang, id, cats) -- [[|foo]] links are treated as plaintext "[[|foo]]". -- FIXME: Pipes should be handled via a proper escape sequence, as they can occur in unsupported titles. if new_term == "" then term, alt = nil, term else local title = new_title(new_term) if title then local ns = title.namespace -- File: and Category: links should be returned as-is. if ns == 6 or ns == 14 then return term end end term, alt = new_term, new_alt if cats then if not (suppress_redundant_wikilink_cat and suppress_redundant_wikilink_cat(term, alt)) then insert(cats, "Pautan bahasa " .. lang:getFullName() .. " dengan pautan wiki lewah") end end end end end if alt then alt = selective_trim(alt) if alt == "" then alt = nil end end -- If there's nothing to process, return nil. if not (term or alt) then return nil end -- If there is no script, get one. if not sc then sc = lang:findBestScript(alt or term) end -- Embedded wikilinks need to be processed individually. if term then local open = find(term, "[[", nil, true) if open and find(term, "]]", open + 2, true) then return process_embedded_links(term, alt, lang, sc, id, cats, no_alt_ast, plain) end term = selective_trim(term) end -- If not, make a link using the parameters. return make_link({ target = term, display = alt, fragment = fragment }, lang, sc, id, true, cats, no_alt_ast, plain) end --[==[ Create a basic link to the given term. It links to the language section (such as `==English==`), but it does not add language and script wrappers, so any code that uses this function should call `[[Module:script utilities#tag_text]]` to add such wrappers itself at some point. The first argument, `data`, may contain the following items, a subset of the items used in the `data` argument of `##full_link`. If any other items are included, they are ignored. { { term = "entry_to_link_to", alt = "link_text_or_displayed_text", lang = language_object, sc = script_object, fragment = "link_fragment", id = "sense_id", no_alt_ast = boolean, suppress_redundant_wikilink_cat = function(term, alt) -> boolean, cats = nil or {}, -- NOTE: If given, will be side-effected to return categories to add the page to. } } Specifically: * `term`: Term to turn into a link. This is generally the name of a page, possibly with extra diacritics added (e.g. length marks in Latin or Old English terms, accents in Russian terms, vowel diacritics in Arabic terms, etc.), which are stripped to determine the actual pagename. The term can contain wikilinks already embedded in it. These are processed individually just like a single link would be. The `alt` argument is ignored in this case. * `alt`: The alternative display for the link, if different from the linked page. If this is {nil}, the `term` argument is used instead (much like regular wikilinks). If `term` contains wikilinks in it, this argument is ignored and has no effect. (Links in which the alt is ignored are tracked with the tracking template {{whatlinkshere|tracking=links/alt-ignored}}.) * `lang` ('''required'''): The [[Module:languages#Language objects|language object]] for the term being linked. The link or links in `term` will normally have their fragment set to point to the canonical name (see {{tl|language data documentation}}) of `lang` (or, if it is an etymology-only language, to the canonical name of its L2 parent). (However, if `id` is defined, the fragment will point to a language-specific sense ID corresponding to this field, which in turn will be overridden by `fragment` if specified.) * `sc`: The [[Module:scripts#Script objects|script object]] for the term being linked. This rarely needs to be specified because it is autodetected based on `term` or `alt`, and the detection is usually correct. It is used to determine how to convert the term into a pagename, possibly by stripping certain diacritics from the term's text. * `fragment`: If specified, overrides the fragment in the generated link (i.e. the portion after `#`, which determines where on the page to go to when the link is clicked). If not specified, the fragment is generated from `id` (if given) or otherwise from `lang`. * `id`: Sense ID string. If this argument is defined, the link will point to a language-specific sense ID ({{ll|en|identifier|id=HTML}}) created by the template {{temp|senseid}}. The fragment for a sense ID consists of the language's canonical name, a hyphen (`-`), and the string that was supplied as the `id` argument. This is useful when a term has more than one sense in a language. If the `term` argument contains wikilinks, this argument is ignored. (Links in which the sense ID is ignored are tracked with the tracking template {{whatlinkshere|tracking=links/id-ignored}}.) * `no_alt_ast`: This is the same as `no_alt_ast` in `##full_link()`. See that function for more information. * `suppress_redundant_wikilink_cat`: This is the same as `suppress_redundant_wikilink_cat` in `##full_link()`. See that function for more information. * `cats`: This should be either {nil} or an empty list. In the latter case, tracking categories will be added to the list when appropriate. The caller can choose to add the page to those categories (as is done by ##full_link()`). The following special options are processed for each link (both simple terms and with embedded wikilinks): * The target page name will be processed by stripping certain diacritics (as mentioned above) and converting the resulting ''logical'' pagename to a ''physical'' pagename (which will be different from the logical pagename in the case of pages with unsupported characters in them and certain overly large pages, such as [[a]], that are split into parts). * If the term starts with `*`, then it is considered a reconstructed term, and a link to the `Reconstruction:` namespace will be created. If the text contains embedded wikilinks, then `*` is automatically applied to each one individually, while preserving the displayed form of each link as it was given. This allows linking to phrases containing multiple reconstructed terms, while only showing the `*` once at the beginning. * If the text starts with `:`, then the link is treated as "raw" and the above steps are skipped. This can be used in rare cases where the page name begins with `*` or if diacritics should not be stripped. For example: ** {{tl|l|en|*nix}} links to the nonexistent page [[Reconstruction:English/nix]] (`*` is interpreted as a reconstruction), but {{tl|l|en|:*nix}} links to [[*nix]]. ** {{tl|l|sl|Franche-Comté}} links to the nonexistent page [[Franche-Comte]] (`é` is converted to `e` by the diacritic-stripping process), but {{tl|l|sl|:Franche-Comté}} links to [[Franche-Comté]]. ]==] function export.language_link(data) if type(data) ~= "table" then error( "The first argument to the function language_link must be a table. See [[Module:links/documentation]] for more information.") elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then track("escaped", "language_link") end -- Categorize links to "und". local lang, cats = data.lang, data.cats if cats and lang:getCode() == "und" then insert(cats, "Pautan bahasa tidak ditentukan") end return simple_link( data.term, data.fragment, data.alt, lang, data.sc, data.id, cats, data.no_alt_ast, data.suppress_redundant_wikilink_cat ) end function export.plain_link(data) if type(data) ~= "table" then error( "The first argument to the function plain_link must be a table. See Module:links/documentation for more information.") elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then track("escaped", "plain_link") end return simple_link( data.term, data.fragment, data.alt, nil, data.sc, data.id, data.cats, data.no_alt_ast, data.suppress_redundant_wikilink_cat ) end --[==[Replace any links with links to the correct section, but don't link the whole text if no embedded links are found. Returns the display text form.]==] function export.embedded_language_links(data) if type(data) ~= "table" then error( "The first argument to the function embedded_language_links must be a table. See Module:links/documentation for more information.") elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then track("escaped", "embedded_language_links") end local term, lang, sc = data.term, data.lang, data.sc -- If we don't have a script, get one. if not sc then sc = lang:findBestScript(term) end -- Do we have embedded wikilinks? If so, they need to be processed individually. local open = find(term, "[[", nil, true) if open and find(term, "]]", open + 2, true) then return process_embedded_links(term, data.alt, lang, sc, data.id, data.cats, data.no_alt_ast) end -- If not, return the display text. term = selective_trim(term) -- FIXME: Double-escape any percent-signs, because we don't want to treat non-linked text as having percent-encoded -- characters. This is a hack: percent-decoding should come out of [[Module:languages]] and only dealt with in this -- module, as it's specific to links. term = term:gsub("%%", "%%25") return lang:makeDisplayText(term, sc, true) end function export.mark(text, item_type, face, lang) local tag = { "", "" } if item_type == "gloss" then tag = { '<span class="mention-gloss-double-quote">“</span><span class="mention-gloss">', '</span><span class="mention-gloss-double-quote">”</span>' } if type(text) == "string" and text:match("^''[^'].*''$") then -- Temporary tracking for mention glosses that are entirely italicized or bolded, which is probably -- wrong. (Note that this will also find bolded mention glosses since they use triple apostrophes.) track("italicized-mention-gloss", lang and lang:getFullCode() or nil) end elseif item_type == "tr" then if face == "term" then tag = { '<span lang="' .. lang:getFullCode() .. '" class="tr mention-tr Latn">', '</span>' } else tag = { '<span lang="' .. lang:getFullCode() .. '" class="tr Latn">', '</span>' } end elseif item_type == "ts" then -- \226\129\160 = word joiner (zero-width non-breaking space) U+2060 tag = { '<span class="ts mention-ts Latn">/\226\129\160', '\226\129\160/</span>' } elseif item_type == "pos" then tag = { '<span class="ann-pos">', '</span>' } elseif item_type == "non-gloss" then tag = { '<span class="ann-non-gloss">', '</span>' } elseif item_type == "annotations" then tag = { '<span class="mention-gloss-paren annotation-paren">(</span>', '<span class="mention-gloss-paren annotation-paren">)</span>' } elseif item_type == "infl" then tag = { '<span class="ann-infl">', '</span>' } end if type(text) == "string" then return tag[1] .. text .. tag[2] else return "" end end --[=[ Implementation of `format_transliteration` and `format_transcription`. The implementation is identical except that the field containing the transliteration or transcription may vary and is specified in `field`, and the way a given transliteration or transcription is tagged may vary and is controlled by `tag_fn`, which is passed three parameters: `text` (the transliteration or transcription), `lang` (the language passed in) and `face` (the face passed in). On input, `item` is the transliteration or transcription or list of such objects; `field` is either {"tr"} or {"ts"}; `tag_fn` is a function of three parameters to tag the item, as described above; `lang` is the language object of the term whose transliteration or transcription is specified; and `face` is a string indicating how to display the item (generally only the strings {"term"} and {"default"} are recognized). ]=] local function format_transliteration_or_transcription(item, field, tag_fn, lang, item_face) if type(item) == "string" then return tag_fn(item, lang, item_face) end local formatted_items = {} for _, itemobj in ipairs(item) do local tagged_item = tag_fn(itemobj[field], lang, item_face) tagged_item = add_text_decorations(tagged_item, item, lang) insert(formatted_items, tagged_item) end if formatted_items[2] then -- FIXME: This should be customizable. return concat(formatted_items, " <i>or</i> ") else return formatted_items[1] end end --[==[ Format a transliteration string or list of transliteration objects. `tr` contains the transliteration(s), which for forward compatibility reasons can only be either a single transliteration string or a list of transliteration objects (each of which has a `tr` field holding the transliteration and optional fields `q`, `qq`, `l`, `ll` and/or `refs`). This correctly handles multiple transliterations as well as decorations (qualifiers, labels or references) attached to transliterations. ]==] function export.format_transliteration(tr, lang, face) return format_transliteration_or_transcription(tr, "tr", tag_translit, lang, face) end local function tag_transcription(ts, _lang, _face) return export.mark(ts, "ts") end --[==[ Format a transcription string or list of transcription objects. `ts` contains the transcription(s), which for forward compatibility reasons can only be either a single transcription string or a list of transcription objects (each of which has a `ts` field holding the transcription and optional decoration fields `q`, `qq`, `l`, `ll` and/or `refs`). This correctly handles multiple transcriptions as well as decorations (qualifiers, labels or references) attached to transcriptions. ]==] function export.format_transcription(ts, lang, face) return format_transliteration_or_transcription(ts, "ts", tag_transcription, lang, face) end local pos_tags --[==[ Format the annotations that are displayed with a link created by `full_link()`. Annotations are the extra bits of information that are displayed following the linked term, and include things such as gender, transliteration, gloss, etc. The first argument is a table with some or all of the following keys (all are optional): * `interwiki`: An interwiki link. This is used for links in translation tables to the corresponding term in another Wiktionary. If specified, it should be a fully formatted link and is inserted as-is at the beginning of the output. See the `interwiki()` function in [[Module:translations]]. * `genders`: Table containing a list of gender specifications in the style of [[Module:gender and number]]. If specified, these are formatted using `format_genders()` in [[Module:gender and number]] and the result inserted at the beginning of the output, following any interwiki link and (in all cases) directly after a no-break space. * `tr`: Transliteration or transliterations. Currently, this is always a one-item list. It is a list because of potential support for per-alternant transliterations due to the multiple alternants (separated by `//`) that can be specified in `term` or `alt`. The item in the list can be either a string or a list of transliteration objects (see `format_transliteration()`). * `ts`: Transcription or transcriptions. Like `tr`, this is currently always a one-item list, with the item being either a single string or a list of transcription objects, as described in `format_transcription()`. * `gloss`: Gloss that translates the term in the link. * `pos`: Part of speech of the linked term. If the given argument matches one of the aliases in `pos_aliases` in [[Module:headword/data]], or consists of a part of speech or alias followed by `f` (for a non-lemma form), expand it appropriately. Otherwise, just show the given text as it is. * `infl`: A list of tags, each a string. Multiple tag sets may be encoded in this list by separating them with an element consisting of a semicolon. If there are multiple tag sets, they are formatted individually and separated by a semicolon + space. * `ng`: Arbitrary non-gloss descriptive text for the link. This should be used in preference to putting descriptive text in `gloss` or `pos`. * `lit`: Literal meaning of the term, if the usual meaning is figurative or idiomatic. * `postprocess_annotations`: A function to postprocess the annotations, after they have been formatted (see below). The `interwiki` and `genders` properties are formatted specially, and `postprocess_annotations` is a callback function rather than an item to display; all others are formatted (each in their own way), separated by commas and placed inside of parentheses (except that if both transliteration and transcription are present, they are separated by a space). The order of the annotations is `tr`+`ts`, `gloss`, `pos`, `infl`, `ng` and `lit`. The `postprocess_annotations` function, if supplied, is passed a single object, a table with two keys `data` (the `data` object passed into `format_link_annotations()`) and `annotations` (the formatted annotations, prior to being concatenated). It should side-effect the `annotations` list, e.g. by inserting more annotations. (It is used to handle nested inflections in [[Module:headword]]. FIXME: It should probably be generalized so that it can return the final formatted string, to allow for e.g. changing the way the annotations are formatted.) * The second argument is a string controlling the "face" that the terms are displayed in. Currently it only affects transliteration and transcription and only when the value {"term"} is passed in, in which case those annotations are displayed italicized. ]==] function export.format_link_annotations(data, face) local output = {} -- Interwiki link if data.interwiki then insert(output, data.interwiki) end -- Genders if type(data.genders) ~= "table" then data.genders = { data.genders } end if data.genders and data.genders[1] then local genders, gender_cats = format_genders(data.genders, data.lang) insert(output, "&nbsp;" .. genders) if gender_cats then local cats = data.cats if cats then extend(cats, gender_cats) end end end local annotations = {} -- Transliteration and transcription local tr = data.tr and data.tr[1] or nil local ts = data.ts and data.ts[1] or nil if tr or ts then local item_face if face == "term" then item_face = face else item_face = "default" end local formatted_tr = tr and export.format_transliteration(tr, data.lang, item_face) or nil local formatted_ts = ts and export.format_transcription(ts, data.lang, item_face) or nil if formatted_tr and formatted_ts then insert(annotations, formatted_tr .. " " .. formatted_ts) else insert(annotations, formatted_tr or formatted_ts) end end -- Gloss/translation if data.gloss then insert(annotations, export.mark(data.gloss, "gloss")) end -- Part of speech if data.pos then -- debug category for pos= containing transcriptions if data.pos:match("/[^><]-/") then data.pos = data.pos .. "[[Kategori:Pautan yang mungkin mengandungi transkripsi dalam pos]]" end -- Canonicalize part of speech aliases as well as non-lemma aliases like 'nf' or 'nounf' for "noun form". pos_tags = pos_tags or (m_headword_data or get_headword_data()).pos_aliases local pos = pos_tags[data.pos] if not pos and data.pos:find("f$") then local pos_form = data.pos:sub(1, -2) -- We only expand something ending in 'f' if the result is a recognized non-lemma POS. pos_form = "bentuk " .. (pos_tags[pos_form] or pos_form) if (m_headword_data or get_headword_data()).nonlemmas[pos_form] then pos = pos_form end end insert(annotations, export.mark(pos or data.pos, "pos")) end -- Inflection data if data.infl then local m_form_of = require(form_of_module) -- Split tag sets manually, since tagged_inflections creates a numbered list, and we do not want that. local infl_outputs = {} local tag_sets = m_form_of.split_tag_set(data.infl) for _, tag_set in ipairs(tag_sets) do table.insert(infl_outputs, m_form_of.tagged_inflections({ tags = tag_set, lang = data.lang, nocat = true, nolink = true, nowrap = true })) end insert(annotations, export.mark(table.concat(infl_outputs, "; "), "infl")) end -- Non-gloss text if data.ng then insert(annotations, export.mark(data.ng, "non-gloss")) end -- Literal/sum-of-parts meaning if data.lit then insert(annotations, "secara harfiah " .. export.mark(data.lit, "gloss")) end -- Provide a hook to insert additional annotations such as nested inflections. if data.postprocess_annotations then data.postprocess_annotations { data = data, annotations = annotations } end if annotations[1] then insert(output, " " .. export.mark(concat(annotations, ", "), "annotations")) end return concat(output) end -- Encode certain characters to avoid various delimiter-related issues at various stages. We need to encode < and > -- because they end up forming part of CSS class names inside of <span ...> and will interfere with finding the end -- of the HTML tag. I first tried converting them to URL encoding, i.e. %3C and %3E; they then appear in the URL as -- %253C and %253E, which get mapped back to %3C and %3E when passed to [[Module:accel]]. But mapping them to &lt; -- and &gt; somehow works magically without any further work; they appear in the URL as < and >, and get passed to -- [[Module:accel]] as < and >. I have no idea who along the chain of calls is doing the encoding and decoding. If -- someone knows, please modify this comment appropriately! local accel_char_map local function get_accel_char_map() accel_char_map = { ["%"] = ".", [" "] = "_", ["_"] = u(0xFFF0), ["<"] = "&lt;", [">"] = "&gt;", } return accel_char_map end local function encode_accel_param_chars(param) return (param:gsub("[%% <>_]", accel_char_map or get_accel_char_map())) end local function encode_accel_param(prefix, param) if not param then return "" end if type(param) == "table" then local filled_params = {} -- There may be gaps in the sequence, especially for translit params. local maxindex = 0 for k in pairs(param) do if type(k) == "number" and k > maxindex then maxindex = k end end for i = 1, maxindex do filled_params[i] = param[i] or "" end -- [[Module:accel]] splits these up again. param = concat(filled_params, "*~!") end -- This is decoded again by [[WT:ACCEL]]. return prefix .. encode_accel_param_chars(param) end local function insert_if_not_blank(list, item) if item == "" then return end insert(list, item) end local function get_css_classes(lang, tr, accel, nowrap) if not accel and not nowrap then return "" end local classes = {} if accel then insert(classes, "form-of lang-" .. lang:getFullCode()) local form = accel.form if form then insert(classes, encode_accel_param_chars(form) .. "-form-of") end insert_if_not_blank(classes, encode_accel_param("gender-", accel.gender)) insert_if_not_blank(classes, encode_accel_param("pos-", accel.pos)) insert_if_not_blank(classes, encode_accel_param("transliteration-", accel.translit or (tr ~= "-" and tr or nil))) insert_if_not_blank(classes, encode_accel_param("target-", accel.target)) insert_if_not_blank(classes, encode_accel_param("origin-", accel.lemma)) insert_if_not_blank(classes, encode_accel_param("origin_transliteration-", accel.lemma_translit)) if accel.no_store then insert(classes, "form-of-nostore") end end if nowrap then insert(classes, nowrap) end return concat(classes, " ") end --[==[ Creates a full link, with annotations (see `##format_link_annotations`), in the style of {{tl|l}} or {{tl|m}}. The first argument, `data`, must be a table. It contains the various elements that can be supplied as parameters to {{tl|l}} or {{tl|m}}: { { -- Basic link-related fields term = "entry_to_link_to", alt = "link_text_or_displayed_text", lang = language_object, sc = script_object, fragment = "link_fragment", id = "sense_id", accel = {accelerated_creation_tags}, -- Link annotation fields interwiki = "interwiki_link", genders = {"gender1", "gender2", ...} or {{spec = "gender1", q = {"left qualifier", ...}, qq = {"right qualifier"}, ...}, ...}, tr = "transliteration" or "-" or {{tr = "transliteration", q = {"left_qualifier", ...}, qq = {"right_qualifier", ...}, ..., genders = {gender_spec, ...}}, ...}, ts = "transcription" or {{ts = "transliteration", q = {"left_qualifier", ...}, qq = {"right_qualifier", ...}, ..., genders = {gender_spec, ...}}, ...}, gloss = "gloss", pos = "part_of_speech_tag", infl = {"infl1_tag1", "infl1_tag2", ..., ";", "infl2_tag1", "infl2_tag2", ...}, ng = "non-gloss text", lit = "literal_translation", postprocess_annotations = function({data = full_link_data, annotations = {"annotation1", "annotation2", ...}}) -> nil, -- Other transliteration-related fields respect_link_tr = boolean, never_call_transliteration_module = boolean, suppress_tr = boolean, -- Decoration fields q = { "left_qualifier1", "left_qualifier2", ...} or "left_qualifier", qq = { "right_qualifier1", "right_qualifier2", ...} or "right_qualifier", l = { "left_label1", "left_label2", ...}, ll = { "right_label1", "right_label2", ...}, a = { "left_accent_qualifier1", "left_accent_qualifier2", ...}, aa = { "right_accent_qualifier1", "right_accent_qualifier2", ...}, refs = { "formatted_ref1", "formatted_ref2", ...} or { {text = "text", name = "name", group = "group"}, ... }, pretext = "text_at_beginning", posttext = "text_at_end", show_decorations = boolean, -- Fields controlling tracking categories track_sc = boolean, no_nonstandard_sc_cat = boolean, suppress_redundant_wikilink_cat = function(term, alt) -> boolean, -- Miscellaneous fields no_alt_ast = boolean, no_generate_alternants = boolean, } } Any one of the items in the `data` table (except for `lang`) may be {nil}. If none of `term`, `alt` and `tr` is present, a term request will be shown. Thus, calling {full_link{ term = term, lang = lang, sc = sc }}, where `term` is the page to link to (which may have diacritics that will be stripped and/or embedded bracketed links) and `lang` is a [[Module:languages#Language objects|language object]] from [[Module:languages]], will give a plain link similar to the one produced by the template {{tl|l}}, and calling {full_link( { term = term, lang = lang, sc = sc }, "term" )} will give a link similar to the one produced by the template {{tl|m}}. The function will: * Try to determine the script, based on the characters found in the `term` or `alt` argument, if the script was not given. If a script is given and `track_sc` is {true}, it will check whether the input script is the same as the one which would have been automatically generated and add the category ` ``lang`` terms with redundant script codes` if yes, or ` ``lang`` terms with non-redundant manual script codes` if no. This should be used when the input script object is directly determined by a template's `sc` parameter. * Call `simple_link()` on the `term` or `alt` forms, to remove diacritics in the page name, process any embedded wikilinks and create links to Reconstruction or Appendix pages when necessary. (`simple_link()` is almost exactly the same as `##language_link()`; the latter is a simple wrapper around the former that adds a bit of extra tracking.) * Call `[[Module:script utilities#tag_text]]` to add the appropriate language and script tags to the term and italicize terms written in the Latin script if necessary. Accelerated creation tags, as used by [[WT:ACCEL]], are included. * Generate a transliteration, based on the `alt` or `term` arguments, if the script is not Latin, no transliteration was provided in `tr` and the combination of the term's language and script support automatic transliteration. The transliteration itself will be linked if both `.respect_link_tr` is specified and the language of the term has the `link_tr` property set for the script of the term; but not otherwise. * Add the annotations (transliteration, gender, gloss, etc.) after the link. * If `no_alt_ast` is specified, then the `alt` text does not need to contain an asterisk if the language is reconstructed. This should only be used by modules which really need to allow links to reconstructions that don't display asterisks (e.g. number boxes). * If `suppress_redundant_wikilink_cat` is specified, it should be a function that indicates whether to suppress the generation of the ` ``lang`` links with redundant wikilinks` tracking category. It is passed two arguments, the `term` and `alt` parameters. Normally, this tracking category is added whenever the `term` argument consists entirely of a one-part or two-part embedded link, which is considered "redundant" in that the link can be rewritten into separate `term` and `alt` arguments without any embedded links. For certain wrapping templates, however, otherwise "redundant" embedded links are necessary to prevent interpretation of certain characters as delimiters. For example, {{tl|col}} and related templates use `~` as a separator, as well as `,` when not followed by a space. In these templates, embedded links are required to correctly link to terms containing those delimiters, such as [[Micros~1]] and [[1,6-Cleves acid]], but will incorrectly trigger the addition of the tracking category unless the appropriate `suppress_redundant_wikilink_cat` function is given. * If `pretext` or `posttext` is specified, this is text to (respectively) prepend or append to the output, directly before processing decorations (qualifiers, labels and references). This can be used to add arbitrary extra text inside of the decorations. * If `show_decorations` is specified, then decorations specified in `data` (i.e. left and right qualifiers, accent qualifiers, labels and references) will be displayed, otherwise they will be ignored. (This is because a fair amount of code stores decorations in these fields and displays them itself, rather than expecting {full_link()} to display them.) * ]==] function export.full_link(data, face, allow_self_link, show_qualifiers) if type(data) ~= "table" then error("The first argument to the function full_link must be a table. " .. "See Module:links/documentation for more information.") elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then track("escaped", "full_link") end if show_qualifiers then -- FIXME: Eventually remove the error code. Added 2026-09-17, remove after 2026-10-17 or so. error("Can't pass fourth parameter `show_qualifiers` any more. Set `show_decorations = true` on data.") end if data.show_qualifiers then -- FIXME: Eventually remove the error code. Added 2026-09-18, remove after 2026-10-18 or so. error("Can't set field `show_qualifiers`; use `show_decorations`") end if data.no_generate_forms then -- FIXME: Eventually remove the error code. Added 2026-09-14, remove after 2026-10-14 or so. error("Can't set field `no_generate_forms`; use `no_generate_alternants`") end -- Prevent data from being destructively modified. data = shallow_copy(data) data.cats = {} -- Categorize links to "und". local lang, cats = data.lang, data.cats if cats and lang:getCode() == "und" then insert(cats, "Pautan bahasa tidak ditentukan") end local terms = { true } -- Generate multiple alternants if applicable. for _, param in ipairs { "term", "alt" } do if type(data[param]) == "string" and data[param]:find("//", nil, true) then data[param] = export.split_on_slashes(data[param]) elseif type(data[param]) == "string" and not (type(data.term) == "string" and data.term:find("//", nil, true)) then if not data.no_generate_alternants then data[param] = lang:generateAlternants(data[param]) else data[param] = { data[param] } end else data[param] = {} end end for _, param in ipairs { "sc", "tr", "ts" } do data[param] = { data[param] } end for _, param in ipairs { "term", "alt", "sc", "tr", "ts" } do for i in pairs(data[param]) do terms[i] = true end end -- Create the link local outparts = {} local id, no_alt_ast, suppress_redundant_wikilink_cat, accel, never_call_transliteration_module = data.id, data.no_alt_ast, data.suppress_redundant_wikilink_cat, data.accel, data.never_call_transliteration_module local link_tr = data.respect_link_tr and lang:link_tr(data.sc[1]) for i in ipairs(terms) do local link -- Is there any text to show? if (data.term[i] or data.alt[i]) then -- Try to detect the script if it was not provided local display_term = data.alt[i] or data.term[i] local best = lang:findBestScript(display_term) -- no_nonstandard_sc_cat is intended for use in [[Module:interproject]] if ( not data.no_nonstandard_sc_cat and best:getCode() == "None" and find_best_script_without_lang(display_term):getCode() ~= "None" ) then insert(cats, "Perkataan bahasa " .. lang:getFullName() .. " dalam bentuk tulisan tidak piawai") end if not data.sc[i] then data.sc[i] = best -- Track uses of sc parameter. elseif data.track_sc then if data.sc[i]:getCode() == best:getCode() then insert(cats, "Perkataan dengan kod tulisan lewah bahasa " .. lang:getFullName()) else insert(cats, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. lang:getFullName()) end end -- If using a discouraged character sequence, add to maintenance category if data.sc[i]:hasNormalizationFixes() == true then if (data.term[i] and data.sc[i]:fixDiscouragedSequences(toNFC(data.term[i])) ~= toNFC(data.term[i])) or (data.alt[i] and data.sc[i]:fixDiscouragedSequences(toNFC(data.alt[i])) ~= toNFC(data.alt[i])) then insert(cats, "Laman menggunakan jujukan aksara tidak digalakkan") end end link = simple_link( data.term[i], data.fragment, data.alt[i], lang, data.sc[i], id, cats, no_alt_ast, suppress_redundant_wikilink_cat ) end -- simple_link can return nil, so check if a link has been generated. if link then -- Add "nowrap" class to prefixes in order to prevent wrapping after the hyphen local nowrap local display_term = data.alt[i] or data.term[i] if display_term and (display_term:find("^%-") or display_term:find("^־")) then -- Hebrew maqqef -- FIXME, use hyphens from [[Module:affix]] nowrap = "nowrap" end link = tag_text(link, lang, data.sc[i], face, get_css_classes(lang, data.tr[i], accel, nowrap)) else --[[ No term to show. Is there at least a transliteration we can work from? ]] link = request_script(lang, data.sc[i]) -- No link to show, and no transliteration either. Show a term request (unless it's a substrate, as they rarely take terms). if (link == "" or (not data.tr[i]) or data.tr[i] == "-") and lang:getFamilyCode() ~= "qfa-sub" then -- If there are multiple terms, break the loop instead. if i > 1 then remove(outparts) break elseif NAMESPACE ~= "Templat" then insert(cats, "Permintaan perkataan bahasa " .. lang:getFullName()) end link = "<small>[Istilah?]</small>" end end insert(outparts, link) if i < #terms then insert(outparts, "<span class=\"Zsym mention\" style=\"font-size:100%;\">&nbsp;/ </span>") end end -- When suppress_tr is true, do not show or generate any transliteration if data.suppress_tr then data.tr[1] = nil else -- TODO: Currently only handles the first transliteration, pending consensus on how to handle multiple translits for multiple forms, as this is not always desirable (e.g. traditional/simplified Chinese). if data.tr[1] == "" or data.tr[1] == "-" then data.tr[1] = nil else local phonetic_extraction = load_data("Module:links/data").phonetic_extraction phonetic_extraction = phonetic_extraction[lang:getCode()] or phonetic_extraction[lang:getFullCode()] if phonetic_extraction then data.tr[1] = data.tr[1] or require(phonetic_extraction).getTranslit(export.remove_links(data.alt[1] or data.term[1])) elseif (data.term[1] or data.alt[1]) and data.sc[1]:isTransliterated() then -- Track whenever there is manual translit. The categories below like 'terms with redundant transliterations' -- aren't sufficient because they only work with reference to automatic translit and won't operate at all in -- languages without any automatic translit, like Persian and Hebrew. if data.tr[1] then local full_code = lang:getFullCode() track("manual-tr", full_code) end if not never_call_transliteration_module then -- Try to generate a transliteration. local text = data.alt[1] or data.term[1] if not link_tr then text = export.remove_links(text, true) end local automated_tr = lang:transliterate(text, data.sc[1]) if automated_tr then local manual_tr = data.tr[1] if manual_tr then if export.remove_links(manual_tr) == export.remove_links(automated_tr) then insert(cats, "Perkataan dengan transliterasi lewah bahasa " .. lang:getFullName()) else -- Prevents Arabic root categories from flooding the tracking categories. if NAMESPACE ~= "Kategori" then insert(cats, "Perkataan dengan transliterasi manual tidak lewah bahasa " .. lang:getFullName()) end end end if not manual_tr or lang:overrideManualTranslit(data.sc[1]) then data.tr[1] = automated_tr end end end end end end -- Link to the transliteration entry for languages that require this if data.tr[1] and link_tr and not data.tr[1]:match("%[%[(.-)%]%]") then data.tr[1] = simple_link( data.tr[1], nil, nil, lang, get_script("Latn"), nil, cats, no_alt_ast, suppress_redundant_wikilink_cat ) elseif data.tr[1] and not link_tr then -- Remove the pseudo-HTML tags added by remove_links. data.tr[1] = data.tr[1]:gsub("</?link>", "") end if data.tr[1] and not umatch(data.tr[1], "[^%s%p]") then data.tr[1] = nil end insert(outparts, export.format_link_annotations(data, face)) if data.pretext then insert(outparts, 1, data.pretext) end if data.posttext then insert(outparts, data.posttext) end local categories = cats[1] and format_categories(cats, lang, "-", nil, nil, data.sc) or "" local output = concat(outparts) if data.show_decorations then output = add_text_decorations(output, data, lang) end return output .. categories end --[==[ Strip links by replacing all wikilinks with their displayed text, and remove any categories. This function can be invoked either from a template or from another module. Specifically, this function deletes category links, the targets of piped links, and any double square brackets involved in links (other than file links, which are untouched). If `tag` is set, then any links removed will be given pseudo-HTML tags, which allow the substitution functions in [[Module:languages]] to properly subdivide the text in order to reduce the chance of substitution failures in modules which scrape pages like [[Module:zh-translit]]. (FIXME: This is quite hacky. We probably want this to be integrated into [[Module:languages]], but we can't do that until we know that nothing is pushing pipe linked transliterations through it for languages which don't have link_tr set.) * `<nowiki>[[page|displayed text]]</nowiki>` &rarr; `displayed text` * `<nowiki>[[page and displayed text]]</nowiki>` &rarr; `page and displayed text` * `<nowiki>[[Category:English lemmas|WORD]]</nowiki>` &rarr; ''(nothing)'' ]==] function export.remove_links(text, tag) if type(text) == "table" then text = text.args[1] end if not text or text == "" then return "" end text = text :gsub("%[%[", "\1") :gsub("%]%]", "\2") -- Parse internal links for the display text. text = text:gsub("(\1)([^\1\2]-)(\2)", function(c1, c2, c3) -- Don't remove files. for _, false_positive in ipairs({ "file", "image" }) do if c2:lower():match("^" .. false_positive .. ":") then return c1 .. c2 .. c3 end end -- Remove categories completely. for _, false_positive in ipairs({ "category", "cat" }) do if c2:lower():match("^" .. false_positive .. ":") then return "" end end -- In piped links, remove all text before the pipe, unless it's the final character (i.e. the pipe trick), in which case just remove the pipe. c2 = c2:match("^[^|]*|(.+)") or c2:match("([^|]+)|$") or c2 if tag then return "<link>" .. c2 .. "</link>" else return c2 end end) text = text :gsub("\1", "[[") :gsub("\2", "]]") return text end function export.section_link(link) if type(link) ~= "string" then error("The first argument to section_link was a " .. type(link) .. ", but it should be a string.") elseif link:find("\\", nil, true) then track("escaped", "section_link") end local target, section = get_fragment((link:gsub("_", " "))) if not section then error("No \"#\" delineating a section name") end return simple_link( target, section, target .. " §&nbsp;" .. section ) end return export q7sfw2xyccv8fplvxsqjqtahr7rn6v9 Modul:gender and number 828 9772 375355 185123 2026-09-22T03:13:44Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92717017|92717017]]) 375355 Scribunto text/plain local export = {} local decorations_module = "Module:decorations" local load_module = "Module:load" local parameters_module = "Module:parameters" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local insert = table.insert local function deep_copy(...) deep_copy = require(table_module).deepCopy return deep_copy(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function format_decorations(...) format_decorations = require(decorations_module).format_decorations return format_decorations(...) end local function load_data(...) load_data = require(load_module).load_data return load_data(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local gender_and_number_data local function get_gender_and_number_data() gender_and_number_data, get_gender_and_number_data = load_data("Module:gender and number/data"), nil return gender_and_number_data end --[==[ intro: This module creates standardised displays for gender and number. It converts a gender specification into Wiki/HTML format. A gender/number specification consists of one or more gender/number elements, separated by hyphens. Examples are: {"n"} (neuter gender), {"f-p"} (feminine plural), {"m-an-p"} (masculine animate plural), {"pf"} (perfective aspect). Each gender/number element has the following properties: # A code, as used in the spec, e.g. {"f"} for feminine, {"p"} for plural. # A type, e.g. `gender`, `number` or `animacy`. Each element in a given spec must be of a different type. # A display form, which in turn consists of a display code and a tooltip gloss. The display code may not be the same as the spec code, e.g. the spec code {"an"} has display code {"anim"} and tooltip gloss ''animate''. # A category into which lemmas of the right part of speech are placed if they have a gender/number spec containing the given element. For example, a noun with gender/number spec {"m-an-p"} is placed into the categories `<var>lang</var> masculine nouns`, `<var>lang</var> animate nouns` and `<var>lang</var> pluralia tantum`. ]==] --[==[ Version of format_genders() that can be invoked from a template. ]==] function export.show_list(frame) local params = { [1] = {list = true}, ["lang"] = {type = "language"}, } local iargs = process_params(frame.args, params) local text, cats = export.format_genders(iargs[1], iargs.lang) if not cats then return text end return text .. format_categories(cats, iargs.lang) end function export.format_list(...) -- FIXME: Added 2026-09-17. Remove after 2026-10-17. error("Use format_genders instead of format_list") end local function autoadd_abbr(display) if not display then error("Internal error: '.display' for gender/number code is missing") end if display:find("<abbr", nil, true) then return display end return ('%s'):format(display, display) end --[=[ Add decorations (qualifiers, labels and references) to a formatted gender/number spec. `spec` is the object describing the gender/number spec, which should optionally contain: * left qualifiers in `q`, an array of strings; * right qualifiers in `qq`, an array of strings; * left labels in `l`, an array of strings; * right labels in `ll`, an array of strings; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; `formatted` is the formatted version of the term itself, and `lang` is the optional language object passed into format_genders(). ]=] local function add_decorations(formatted, spec, lang) local function field_non_empty(field) local list = spec[field] if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `spec.%s`=%s, should be \"table\""):format( field, mw.dumpObject(list))) end return list[1] end if field_non_empty("q") or field_non_empty("qq") or field_non_empty("l") or field_non_empty("ll") or field_non_empty("refs") then formatted = format_decorations { lang = lang, text = formatted, q = spec.q, qq = spec.qq, l = spec.l, ll = spec.ll, refs = spec.refs, } end return formatted end --[==[ Format one or more gender/number specifications. Each spec is either a string, e.g. {"f-p"}, or a table of the form { {spec = "SPEC", q = {"LEFT_QUALIFIER", "LEFT_QUALIFIER", ...}, qq = {"RIGHT_QUALIFIER", "RIGHT_QUALIFIER", ...}, l = {"LEFT_LABEL", "LEFT_LABEL", ...}, ll = {"RIGHT_LABEL", "RIGHT_LABEL", ...}, refs = {"FORMATTED_REFERENCE_TEXT", {text = "FORMATTED_REFERENCE_TEXT", name = "REFNAME", group = "REFGROUP"}, ...}, }} In the table format, only `spec` is required, and in the `refs` object format, only `text` is required. If `lang` is not given, the funtion returns only one value, the formatted text. If `lang` is given, the function returns two values: # the formatted text; # a list of the categories to add, which will be {nil} if there are no categories to add. If `lang` (which should be a language object) and `pos_for_cat` (which should be a plural part of speech) are given, gender categories such as `German masculine nouns` or `Russian imperfective verbs` are added to the categories, and request categories such as `Requests for gender in <var>lang</var> entries` or `Requests for animacy in <var>lang</var> entries` may also be added. Otherwise, if only `lang` is given, only request categories may be returned. If both are omitted, only one return value is returned, as mentioned above. ]==] function export.format_genders(specs, lang, pos_for_cat) local formatted_specs, categories, seen_types = {} local all_is_nounclass = nil local full_langname = lang and lang:getFullName() or nil local function do_gender_spec(spec, parts) local types = {} local codes = (gender_and_number_data or get_gender_and_number_data()).codes for key, code in ipairs(parts) do -- Is this code valid? if not codes[code] then error('The tag "' .. code .. '" in the gender specification "' .. spec.spec .. '" is not valid. See [[Module:gender and number]] for a list of valid tags.') end -- Check for multiple genders/numbers/animacies in a single spec. local typ = codes[code].type if typ ~= "other" and types[typ] then error('The gender specification "' .. spec.spec .. '" contains multiple tags of type "' .. typ .. '".') end types[typ] = true parts[key] = autoadd_abbr(codes[code].display) -- Generate categories if called for. if lang and pos_for_cat then local cat = codes[code].cat if cat then if not categories then categories = {} end insert(categories, cat .. " bahasa " .. full_langname) end if not seen_types then seen_types = {} elseif seen_types[typ] and seen_types[typ] ~= code then cat = (gender_and_number_data or get_gender_and_number_data()).multicode_cats[typ] if cat then if not categories then categories = {} end insert(categories, cat .. " bahasa " .. full_langname) end end seen_types[typ] = code end if lang and codes[code].req then local type_for_req = typ if code == "?" then -- Keep in mind `pos_for_cat` may be nil here. type_for_req = pos_for_cat == "kata kerja" and "aspek" or "genus" end if not categories then categories = {} end insert(categories, "Permintaan untuk " .. type_for_req .. " dalam entri bahasa " .. full_langname) end end -- Add the processed codes together with non-breaking spaces if not parts[2] and parts[1] then return parts[1] else return concat(parts, "&nbsp;") end end for _, spec in ipairs(specs) do if type(spec) ~= "table" then spec = {spec = spec} end local spec_spec, is_nounclass = spec.spec -- If the specification starts with cX, then it is a noun class specification. if spec_spec:match("^%d") or spec_spec:match("^c[^-]") then is_nounclass = true local code = spec_spec:gsub("^c", "") local text if code == "?" then text = '<abbr class="noun-class" title="noun class missing">?</abbr>' if lang then if not categories then categories = {} end insert(categories, "Permintaan kelas kata nama dalam entri bahasa " .. full_langname) end else text = '<abbr class="noun-class" title="noun class ' .. code .. '">' .. code .. "</abbr>" if lang and pos_for_cat then if not categories then categories = {} end insert(categories, "POS" .. code .. " kelas bahasa " .. full_langname) end end local text_with_decorations = add_decorations(text, spec, lang) insert(formatted_specs, text_with_decorations) else -- Split the parts and iterate over each part, converting it into its display form local parts = split(spec.spec, "-", true, true) local combined_codes = (gender_and_number_data or get_gender_and_number_data()).combinations if lang then -- Check if the specification is valid --elseif langinfo.genders then -- local valid_genders = {} -- for _, g in ipairs(langinfo.genders) do valid_genders[g] = true end -- -- if not valid_genders[spec.spec] then -- local valid_string = {} -- for i, g in ipairs(langinfo.genders) do valid_string[i] = g end -- error('The gender specification "' .. spec.spec .. '" is not valid for ' .. langinfo.names[1] .. ". Valid are: " .. concat(valid_string, ", ")) -- end --end end local has_combined = false for _, code in ipairs(parts) do if combined_codes[code] then has_combined = true break end end if not has_combined then if formatted_specs[1] then insert(formatted_specs, "or") end insert(formatted_specs, add_decorations(do_gender_spec(spec, parts), spec, lang)) else -- This logic is to handle combined gender specs like 'mf' and 'mfbysense'. local all_parts = {{}} local extra_displays local this_formatted_specs = {} for _, code in ipairs(parts) do if combined_codes[code] then local new_all_parts = {} for _, one_parts in ipairs(all_parts) do for _, one_code in ipairs(combined_codes[code].codes) do local new_combined_parts = deep_copy(one_parts) insert(new_combined_parts, one_code) insert(new_all_parts, new_combined_parts) end end all_parts = new_all_parts if lang and pos_for_cat then local extra_cat = combined_codes[code].cat if extra_cat then if not categories then categories = {} end insert(categories, extra_cat .. " bahasa " .. full_langname) end end local extra_display = combined_codes[code].display if extra_display then if not extra_displays then extra_displays = {} end insert(extra_displays, autoadd_abbr(extra_display)) end else for _, one_parts in ipairs(all_parts) do insert(one_parts, code) end end end for _, this_parts in ipairs(all_parts) do if this_formatted_specs[1] then insert(this_formatted_specs, "or") end insert(this_formatted_specs, do_gender_spec(spec, this_parts)) end if extra_displays then for _, display in ipairs(extra_displays) do insert(this_formatted_specs, display) end end insert(formatted_specs, add_decorations( concat(this_formatted_specs, " "), spec, lang)) end is_nounclass = false end -- Ensure that the specifications are either all noun classes, or none are. if all_is_nounclass == nil then all_is_nounclass = is_nounclass elseif all_is_nounclass ~= is_nounclass then error("Noun classes and genders cannot be mixed. Please use either one or the other.") end end if categories and lang and pos_for_cat then for i, cat in ipairs(categories) do categories[i] = cat:gsub("POS", pos_for_cat) end end local formatted_retval if all_is_nounclass then -- Add the processed codes together with slashes formatted_retval = '<span class="gender">class ' .. concat(formatted_specs, "/") .. "</span>" else -- Add the processed codes together with spaces formatted_retval = '<span class="gender">' .. concat(formatted_specs, " ") .. "</span>" end if lang then return formatted_retval, categories else return formatted_retval end end return export 2wkm2desvqy89qy6hbwj4b0amm5ltiu Modul:links/templates 828 9792 375365 245095 2026-09-22T04:36:17Z Hakimi97 2668 Kemas kini 375365 Scribunto text/plain -- Prevent substitution. if mw.isSubsting() then return require("Module:unsubst") end local export = {} local links_module = "Module:links" local process_params = require("Module:parameters").process local remove = table.remove local upper = require("Module:string utilities").upper --[=[ Modules used: [[Module:links]] [[Module:languages]] [[Module:scripts]] [[Module:parameters]] [[Module:debug]] ]=] do local function get_args(frame) -- `compat` is a compatibility mode for {{term}}. -- If given a nonempty value, the function uses lang= to specify the -- language, and all the positional parameters shift one number lower. local iargs = frame.args iargs.compat = iargs.compat and iargs.compat ~= "" iargs.langname = iargs.langname and iargs.langname ~= "" iargs.notself = iargs.notself and iargs.notself ~= "" local alias_of_4 = {alias_of = 4} local boolean = {type = "boolean"} local params = { [1] = {required = true, type = "language", default = "und"}, [2] = true, [3] = true, [4] = true, g = {list = true, type = "genders", flatten = true}, gloss = alias_of_4, id = true, lit = true, ng = true, pos = true, sc = {type = "script"}, t = alias_of_4, tr = true, ts = true, q = {type = "qualifier"}, qq = {type = "qualifier"}, l = {type = "labels"}, ll = {type = "labels"}, ref = {type = "references"}, ["accel-form"] = true, ["accel-translit"] = true, ["accel-lemma"] = true, ["accel-lemma-translit"] = true, ["accel-gender"] = true, ["accel-nostore"] = boolean, } if iargs.compat then params.lang = {type = "language", default = "und"} remove(params, 1) alias_of_4.alias_of = 3 end if iargs.langname then params.w = boolean end return process_params(frame:getParent().args, params), iargs end -- Used in [[Template:l]] and [[Template:m]]. function export.l_term_t(frame) local args, iargs = get_args(frame) local compat = iargs.compat local lang = args[compat and "lang" or 1] -- Tracking for und. if not compat and lang:getCode() == "und" then require("Module:debug").track("link/und") end local term = args[(compat and 1 or 2)] local alt = args[(compat and 2 or 3)] term = term ~= "" and term or nil if not term and not alt and iargs.demo then term = iargs.demo end local langname = iargs.langname and ( args.w and lang:makeWikipediaLink() or lang:getCanonicalName() ) or nil if langname and term == "-" then return langname end -- Forward the information to full_link return (langname and langname .. " " or "") .. require(links_module).full_link( { lang = lang, sc = args.sc, track_sc = true, term = term, alt = alt, gloss = args[4], id = args.id, tr = args.tr, ts = args.ts, genders = args.g, pos = args.pos, ng = args.ng, lit = args.lit, q = args.q, qq = args.qq, l = args.l, ll = args.ll, refs = args.ref, show_decorations = true, accel = args["accel-form"] and { form = args["accel-form"], translit = args["accel-translit"], lemma = args["accel-lemma"], lemma_translit = args["accel-lemma-translit"], gender = args["accel-gender"], nostore = args["accel-nostore"], } or nil }, iargs.face, not iargs.notself ) end -- Used in [[Template:link-annotations]]. function export.l_annotations_t(frame) local args, iargs = get_args(frame) -- Forward the information to format_link_annotations return require(links_module).format_link_annotations( { lang = args[1], tr = { args.tr }, ts = { args.ts }, genders = args.g, pos = args.pos, ng = args.ng, lit = args.lit }, iargs.face ) end end -- Used in [[Template:ll]]. do local function get_args(frame) return process_params(frame:getParent().args, { [1] = {required = true, type = "language", default = "und"}, [2] = {allow_empty = true}, [3] = true, id = true, sc = {type = "script"}, }) end function export.ll(frame) local args = get_args(frame) local lang = args[1] local sc = args.sc local term = args[2] term = term ~= "" and term or nil return require(links_module).language_link{ lang = lang, sc = sc, term = term, alt = args[3], id = args.id } or "<small>[Istilah?]</small>" .. require("Module:utilities").format_categories( {"Permintaan istilah bahasa " .. lang:getFullName()}, lang, "-", nil, nil, sc ) end end function export.def_t(frame) local args = process_params(frame:getParent().args, { [1] = {required = true, default = ""}, }) local face = frame.args.face local ret = require("Module:script utilities").tag_definition(require(links_module).embedded_language_links{ term = args[1], lang = require("Module:languages").getByCode("ms"), sc = require("Module:scripts").getByCode("Latn") }, face) if face == "non-gloss" then return ret end return '<span class="mention-gloss-paren">(</span>' .. ret .. '<span class="mention-gloss-paren">)</span>' end function export.linkify_t(frame) local args = process_params(frame:getParent().args, { [1] = {required = true, default = ""}, }) args[1] = mw.text.trim(args[1]) if args[1] == "" or args[1]:find("[[", nil, true) then return args[1] end return "[[" .. args[1] .. "]]" end function export.cap_t(frame) local args = process_params(frame:getParent().args, { [1] = {required = true}, [2] = true, lang = {type = "language", default = "ms"}, }) local term = args[1] return require(links_module).full_link{ lang = args.lang, term = term, alt = term:gsub("^.[\128-\191]*", upper) .. (args[2] or "") } end function export.section_link_t(frame) local args = process_params(frame:getParent().args, { [1] = {}, }) return require(links_module).section_link(args[1]) end return export dhr763ljty2b15td9ft3hiqn36wxt5h Modul:languages/data/3/k 828 9814 375377 373727 2026-09-22T05:35:12Z Hakimi97 2668 Move "Dusun Tambunan" back to Module:languages/data/3/k 375377 Scribunto text/plain local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared local m = {} m["kaa"] = { "Karakalpak", 33541, "trk-kno", "Latn, Cyrl, Arab", dotted_dotless_i = true, strip_diacritics = { from = {"['’]"}, to = {"ʼ"} }, sort_key = { Latn = { from = { -- Sort the old orthography (using the apostrophe) after the new orthography (using the acute accent). "í", "iʼ", "i", -- Ensure "i" comes after "í", "iʼ", "ı". "sh", "ch", "á", "aʼ", "ǵ", "gʼ", "x", p[4], p[5], "ı", "q", "ń", "nʼ", "ó", "oʼ", "ú", "uʼ", "c" }, to = { p[4], p[5], "i" .. p[3], "z" .. p[1], "z" .. p[3], "a" .. p[1], "a" .. p[2], "g" .. p[1], "g" .. p[2], "h" .. p[1], "i", "i" .. p[1], "i" .. p[2], "k" .. p[1], "n" .. p[1], "n" .. p[2], "o" .. p[1], "o" .. p[2], "u" .. p[1], "u" .. p[2], "z" .. p[2] } }, Cyrl = { from = {"ә", "ғ", "ё", "қ", "ң", "ө", "ү", "ў", "ҳ"}, to = {"а" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1], "у" .. p[2], "х" .. p[1]} }, }, } m["kab"] = { "Kabyle", 35853, "ber", "Latn, Arab, Tfng", } m["kac"] = { "Jingpho", 33332, "sit-jnp", "Latn, Mymr", } m["kad"] = { "Kadara", 3914011, "nic-plc", "Latn", } m["kae"] = { "Ketangalan", 2779411, "map", } m["kaf"] = { "Katso", 246122, "tbq-kzh", } m["kag"] = { "Kajaman", 6348863, "poz", "Latn", } m["kah"] = { "Fer", 5443742, "csu-bgr", "Latn", } m["kai"] = { "Karekare", 3438770, "cdc-wst", "Latn", } m["kaj"] = { "Jju", 35401, "nic-plc", "Latn", } m["kak"] = { "Kayapa Kallahan", 3192220, "phi", "Latn", } m["kam"] = { "Kamba", 2574767, "bnt-kka", "Latn", } m["kao"] = { "Kassonke", 36905, "dmn-wmn", "Latn", } m["kap"] = { "Bezhta", 33054, "cau-ets", "Cyrl", translit = "cau-nec-translit", override_translit = true, display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]}, } m["kaq"] = { "Capanahua", 2937196, "sai-pan", "Latn", } m["kaw"] = { "Jawa Kuno", 49341, "poz", "Latn, Java, Kawi", translit = "jv-translit", --same as jv } m["kax"] = { "Kao", 3192799, "paa-gto", "Latn", } m["kay"] = { "Kamayurá", 3192336, "tup-gua", "Latn", } m["kba"] = { "Kalarko", 5517764, "aus-pam", "Latn", } m["kbb"] = { "Kaxuyana", 12953626, "sai-prk", "Latn", } m["kbc"] = { "Kadiwéu", 18168288, "sai-guc", "Latn", } m["kbd"] = { "Kabardia", 33522, "cau-cir", "Cyrl, Latn, Arab", translit = { Cyrl = "cau-cir-translit", Arab = "ar-translit", }, override_translit = true, display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, sort_key = { Cyrl = { from = { "кхъу", "къӏу", -- 4 chars "гъу", "джу", "дзу", "жъу", "къу", "кхъ", "къӏ", "кӏу", "кӏь", "лъу", "лӏу", "пӏу", "сӏу", "тӏу", "фӏу", "хъу", "цӏу", "чъу", "чӏу", "шъу", "шӏу", "щӏу", -- 3 chars "гу", "гъ", "гь", "дж", "дз", "ё", "жъ", "жь", "ку", "къ", "кь", "кӏ", "лъ", "ль", "лӏ", "пӏ", "сӏ", "тӏ", "фӏ", "ху", "хъ", "хь", "цу", "цӏ", "чу", "чъ", "чӏ", "шъ", "шӏ", "щӏ", "ӏу", "ӏь", -- 2 chars "э" -- 1 char }, to = { "к" .. p[5], "к" .. p[7], "г" .. p[3], "д" .. p[2], "д" .. p[4], "ж" .. p[2], "к" .. p[3], "к" .. p[4], "к" .. p[6], "к" .. p[10], "к" .. p[11], "л" .. p[2], "л" .. p[5], "п" .. p[2], "с" .. p[2], "т" .. p[2], "ф" .. p[2], "х" .. p[3], "ц" .. p[3], "ч" .. p[3], "ч" .. p[5], "ш" .. p[2], "ш" .. p[4], "щ" .. p[2], "г" .. p[1], "г" .. p[2], "г" .. p[4], "д" .. p[1], "д" .. p[3], "е" .. p[1], "ж" .. p[1], "ж" .. p[3], "к" .. p[1], "к" .. p[2], "к" .. p[8], "к" .. p[9], "л" .. p[1], "л" .. p[3], "л" .. p[4], "п" .. p[1], "с" .. p[1], "т" .. p[1], "ф" .. p[1], "х" .. p[1], "х" .. p[2], "х" .. p[4], "ц" .. p[1], "ц" .. p[2], "ч" .. p[1], "ч" .. p[2], "ч" .. p[4], "ш" .. p[1], "ш" .. p[3], "щ" .. p[1], "ӏ" .. p[1], "ӏ" .. p[2], "а" .. p[1] } }, }, } m["kbe"] = { "Kanju", 10543322, "aus-pam", "Latn", } m["kbh"] = { "Camsá", 2842667, "qfa-iso", "Latn", } m["kbi"] = { "Kaptiau", 6367294, "poz-oce", "Latn", } m["kbj"] = { "Kari", 6370438, "bnt-boa", "Latn", } m["kbk"] = { "Grass Koiari", 12952642, "ngf-koi", "Latn", } m["kbm"] = { "Iwal", 3156391, "poz-ocw", "Latn", } m["kbn"] = { "Kare (Africa)", 35554, "alv-mbm", "Latn", } m["kbo"] = { "Keliko", 11275553, "csu-mma", "Latn", } m["kbp"] = { "Kabiyé", 35475, "nic-gne", "Latn", } m["kbq"] = { "Kamano", 11732272, "ngf-kya", "Latn", } m["kbr"] = { "Kafa", 35481, "omv-gon", "Ethi, Latn", } m["kbs"] = { "Kande", 35556, "bnt-tso", "Latn", } m["kbt"] = { "Gabadi", 3291159, "poz-ocw", "Latn", } m["kbu"] = { "Kabutra", 10966761, "raj", } m["kbv"] = { "Kamberataro", 5261289, "paa-sng", "Latn", } m["kbw"] = { "Kaiep", 6347632, "poz-ocw", "Latn", } m["kbx"] = { "Ap Ma", 56298, "paa-eke", "Latn", } m["kbz"] = { "Duhwa", 56295, "cdc-wst", "Latn", } m["kcb"] = { "Kawacha", 11732302, "ngf-woj", "Latn", } m["kcc"] = { "Lubila", 3914381, "nic-uce", "Latn", } m["kcd"] = { "Ngkâlmpw Kanum", 12952566, "paa-ngk", "Latn", } m["kce"] = { "Kaivi", 6348685, "nic-kau", "Latn", } m["kcf"] = { "Ukaan", 36651, "nic-bco", "Latn", } m["kcg"] = { "Tyap", 3912765, "nic-plc", "Latn", } m["kch"] = { "Vono", 3913920, "nic-kau", "Latn", } m["kci"] = { "Kamantan", 3914019, "nic-plc", } m["kcj"] = { "Kobiana", 35609, "alv-nyn", "Latn", } m["kck"] = { "Kalanga", 33672, "bnt-sho", "Latn", } m["kcl"] = { "Kala", 6349982, "poz-ocw", "Latn", } m["kcm"] = { "Tar Gula", 277963, "csu-bba", } m["kcn"] = { "Nubi", 36388, "crp", "Latn, Arab", ancestors = "apd", strip_diacritics = {remove_diacritics = c.acute}, } m["kco"] = { "Kinalakna", 11732320, "ngf-dal", "Latn", } m["kcp"] = { "Kanga", 6362384, "qfa-kad", "Latn", } m["kcq"] = { "Kamo", 3914879, "alv-wjk", } m["kcr"] = { "Katla", 35688, "nic-ktl", "Latn", } m["kcs"] = { "Koenoem", 3438755, "cdc-wst", } m["kct"] = { "Kaian", 6347538, "paa-ott", "Latn", } m["kcu"] = { "Kikami", 3915212, "bnt-ruv", "Latn", } m["kcv"] = { "Kete", 3195598, "bnt-lub", "Latn", } m["kcw"] = { "Kabwari", 6344539, "bnt-glb", "Latn", } m["kcx"] = { "Kachama-Ganjule", 12634070, "omv-eom", } m["kcy"] = { "Korandje", 33427, "son", } m["kcz"] = { "Konongo", 11732345, "bnt-tkm", "Latn", } m["kda"] = { "Worimi", 3914062, "aus-pam", "Latn", } m["kdc"] = { "Kutu", 6448634, "bnt-ruv", } m["kdd"] = { "Yankunytjatjara", 34207, "aus-pam", "Latn", } m["kde"] = { "Makonde", 35172, "bnt-rvm", "Latn", } m["kdf"] = { "Mamusi", 6746036, "poz-ocw", "Latn", } m["kdg"] = { "Seba", 7442316, "bnt-sbi", "Latn", } m["kdh"] = { "Tem", 36531, "nic-gne", "Latn", } m["kdi"] = { "Kumam", 6443410, "sdv-los", } m["kdj"] = { "Karamojong", 56326, "sdv-ttu", "Latn", } m["kdk"] = { "Numee", 3346774, "poz-cln", "Latn", } m["kdl"] = { "Tsikimba", 3914404, "nic-kam", } m["kdm"] = { "Kagoma", 3914420, "nic-plc", } m["kdn"] = { "Kunda", 4121130, "bnt-sna", "Latn", } m["kdp"] = { "Kaningdon-Nindem", 3914956, "nic-nin", } m["kdq"] = { "Koch", 56431, "tbq-bdg", } m["kdr"] = { "Karaim", 33725, "trk-kcu", "Cyrl, Latn, Hebr", -- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["kdt"] = { "Kuy", 56310, "mkh-kat", "Thai, Khmr, Laoo", } m["kdu"] = { "Kadaru", 35441, "nub-hil", "Latn", } m["kdv"] = { "Kado", 7402721, "sit-luu", } m["kdw"] = { "Koneraw", 11732341, "ngf-mom", "Latn", } m["kdx"] = { "Kam", 36753, "alv-wjk", } m["kdy"] = { "Keder", 6383641, "paa-tor", "Latn", } m["kdz"] = { "Kwaja", 11128866, "nic-nka", "Latn", } m["kea"] = { "Kabuverdianu", 35963, "crp", "Latn", ancestors = "pt", } m["keb"] = { "Kélé", 35559, "bnt-kel", } m["kec"] = { "Keiga", 3409311, "qfa-kad", "Latn", } m["ked"] = { "Kerewe", 6393846, "bnt-haj", } m["kee"] = { "Keres Timur", 15649021, "nai-ker", "Latn", } m["kef"] = { "Kpessi", 35748, "alv-gbe", } m["keg"] = { "Tese", 16887296, "sdv", } m["keh"] = { "Keak", 6382110, "paa-nnd", "Latn", } m["kei"] = { "Kei", 2410352, "poz-cet", "Latn", } m["kej"] = { "Kadar", 6345179, "dra-mal", } m["kek"] = { "Q'eqchi", 35536, "myn", "Latn", } m["kel"] = { "Kela-Yela", 6385426, "bnt-mon", "Latn", } m["kem"] = { "Kemak", 35549, "poz-tim", "Latn", } m["ken"] = { "Kenyang", 35650, "nic-mam", "Latn", } m["keo"] = { "Kakwa", 3033547, "sdv-bri", "Latn", } m["kep"] = { "Kaikadi", 6347757, "dra-tam", } m["keq"] = { "Kamar", 14916877, "inc-hal", } m["ker"] = { "Kera", 56251, "cdc-est", "Latn", } m["kes"] = { "Kugbo", 3813394, "nic-cde", "Latn", } m["ket"] = { "Ket", 33485, "qfa-yke", "Cyrl", translit = "ket-utils", display_text = "ket-utils", strip_diacritics = "ket-utils", sort_key = { from = {"ӷ", "ё", "ӄ", "ӈ", "ө", "ә", "ʼ"}, to = {"г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "ъ" .. p[1], "ь" .. p[1]} }, } m["keu"] = { "Akebu", 35026, "alv-ktg", "Latn", } m["kev"] = { "Kanikkaran", 6363201, "dra-mal", "Taml, Mlym", -- Mlym translit in [[Module:scripts/data]] (NOTE: not present before, presumably an accidental omission) } m["kew"] = { "Kewa", 12952619, "ngf-ank", "Latn", } m["kex"] = { "Kukna", 5031131, "inc-bhi", wikipedia_article = "Dhodia–Kukna language", } m["key"] = { "Kupia", 6445354, "inc-eas", } m["kez"] = { "Kukele", 3915391, "nic-ucn", "Latn", } m["kfa"] = { "Kodava", 33531, "dra-kod", "Knda, Mlym", -- Knda translit in [[Module:scripts/data]] -- Mlym translit in [[Module:scripts/data]] } m["kfb"] = { "Kolami", 33479, "dra-knk", "Deva, Telu", translit = { Telu = "te-translit", }, } m["kfc"] = { "Konda-Dora", 35679, "dra-kki", "Orya, Telu", translit = { Orya = "gon-Orya-translit", Telu = "te-translit", }, } m["kfd"] = { "Korra Koraga", 12952655, "dra-kor", "Knda", -- Knda translit in [[Module:scripts/data]] } m["kfe"] = { "Kota (India)", 33483, "dra-tkt", "Taml", translit = "ta-translit", } m["kff"] = { "Koya", 33471, "dra-gon", "Telu, Orya, Deva, Latn", } m["kfg"] = { "Kudiya", 12952667, "dra-tlk", } m["kfh"] = { "Kurichiya", 12952676, "dra-mal", "Mlym", -- Mlym translit in [[Module:scripts/data]] } m["kfi"] = { "Kannada Kurumba", 56589, "dra-sdo", } m["kfj"] = { "Kemiehua", 27144776, "mkh-pal", } m["kfk"] = { "Kinnauri", 2383208, "sit-kin", "Takr, Deva, Latn", } m["kfl"] = { "Kung", 6444510, "nic-rnc", "Latn", } m["kfn"] = { "Kuk", 6442398, "nic-rnc", "Latn", } m["kfo"] = { "Koro (Afrika Barat)", 11160588, "dmn-mnk", "Latn, Nkoo", } m["kfp"] = { "Korwa", 6432786, "mun", } m["kfq"] = { "Korku", 33715, "mun", "Deva", } m["kfr"] = { "Kachchi", 56487, "inc-snd", "Gujr, Arab, Sind, Khoj", translit = { Gujr = "gu-translit", Sind = "Sind-translit", Arab = "sd-Arab-translit", }, strip_diacritics = { remove_diacritics = c.kashida .. c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.superalef, from = {u(0x0671)}, to = {u(0x0627)} }, } m["kfs"] = { "Bilaspuri", 12953397, "him", "Deva, Takr", translit = "hi-translit", } m["kft"] = { "Kanjari", 12953610, "inc-pan", ancestors = "pa", } m["kfu"] = { "Katkari", 6377671, "inc-sou", } m["kfv"] = { "Kurmukar", 6446193, "inc-eas", } m["kfw"] = { "Kharam Naga", 12952906, "tbq-kuk", } m["kfx"] = { "Kullu Pahari", 6443148, "him", "Deva", translit = "hi-translit", } m["kfy"] = { "Kumaoni", 33529, "inc-pac", "Deva, Shrd, Takr", -- Shrd translit in [[Module:scripts/data]] (NOTE: not present before, presumably an accidental omission) } m["kfz"] = { "Koromfé", 35701, "nic-gur", "Latn", } m["kga"] = { "Koyaga", 11155632, "dmn-mnk", } m["kgb"] = { "Kawe", 12952750, "poz-hce", "Latn", } m["kgd"] = { "Kataang", 12953622, "mkh", } m["kge"] = { "Komering", 49224, "poz-lgx", "Latn, Arab", } m["kgf"] = { "Kube", 11732359, "ngf-kto", "Latn", } m["kgg"] = { "Kusunda", 33630, "qfa-iso", -- central Nepal "Latn", } m["kgi"] = { "Selangor Sign Language", 33731, "sgn-asl", } m["kgj"] = { "Gamale Kham", 22236996, "sit-kha", "Deva", } m["kgk"] = { "Kaiwá", 3111883, "gn", "Latn", } m["kgl"] = { "Kunggari", 10550184, "aus-pam", "Latn", } m["kgn"] = { "Karingani", 6371041, "xme-ttc", "Arab, Latn", ancestors = "xme-ttc-nor", } m["kgo"] = { "Krongo", 6438927, "qfa-kad", "Latn", } m["kgp"] = { "Kaingang", 2665734, "sai-sje", "Latn", } m["kgq"] = { "Kamoro", 6359001, "ngf-ask", "Latn", } m["kgr"] = { "Abun", 56657, "qfa-iso", -- Papuan; isolate in Ethnologue, Glottolog and Palmer (2018); grouped with West Papuan by Ross (2005) "Latn", } m["kgs"] = { "Kumbainggar", 3915412, "aus-pam", "Latn", } m["kgt"] = { "Somyev", 3913354, "nic-mmb", "Latn", } m["kgu"] = { "Kobol", 11732325, "ngf-omo", "Latn", } m["kgv"] = { "Karas", 6368621, "qfa-dis", -- Divergent Papuan language; grouped with Mbaham-Iha by Glottolog to form a (mainland) West Bomberai -- family, but with Mbaham-Iha and Timor-Alor-Pantar by Wikipedia (following Usher and Schapper 2022) -- into a (Greater) West Bomberai family. "Latn", } m["kgw"] = { "Karon Dori", 56817, "paa-may", "Latn", } m["kgx"] = { "Kamaru", 12953604, "poz-wot", "Latn", } m["kgy"] = { "Kyerung", 12952691, "sit-kyk", } m["kha"] = { "Khasi", 33584, "aav-pkl", "Latn, as-Beng", } m["khb"] = { "Lü", 36948, "tai-swe", "Talu, Lana", translit = {Talu = "Talu-translit"}, strip_diacritics = {remove_diacritics = c.ZWNJ}, sort_key = { Talu = "Talu-sortkey", Lana = "Lana-sortkey", }, } m["khc"] = { "Tukang Besi Utara", 18611555, "poz", } m["khd"] = { "Bädi Kanum", 20888004, "paa-ngk", "Latn", } m["khe"] = { "Korowai", 6432598, "ngf-bda", "Latn", } m["khf"] = { "Khuen", 27144893, "mkh", } m["khh"] = { "Kehu", 10994953, } m["khj"] = { "Kuturmi", 3914490, "nic-plc", "Latn", } m["khl"] = { "Lusi", 3267788, "poz-ocw", "Latn", } m["kho"] = { "Khotan", 6583551, "xsc-sak", "Brah, Khar", -- Brah translit in [[Module:scripts/data]] } m["khp"] = { "Kapauri", 3502575, "qfa-dis", -- isolate per Glottolog, possibly Greater Kwerba per Wikipedia in Kapauri-Sause family "Latn", } m["khq"] = { "Koyra Chiini", 33600, "son", "Latn, Arab", } m["khr"] = { "Kharia", 3915562, "mun", "Deva, Orya, Latn", } m["khs"] = { "Kasua", 6374863, "ngf-bos", "Latn", } m["kht"] = { "Khamti", 3915502, "tai-swe", "Mymr", display_text = s["kht-displaytext"], strip_diacritics = s["kht-stripdiacritics"], } m["khu"] = { "Nkhumbi", 11019169, "bnt-swb", } m["khv"] = { "Khvarshi", 56425, "cau-wts", "Cyrl", translit = "khv-translit", display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]}, } m["khw"] = { "Khowar", 938216, "inc-chi", "Aran", strip_diacritics = { -- character "ۂ" code U+06C2 to "ه" and "هٔ" (U+0647 + U+0654) to "ه"; hamzatu l-waṣli to a regular alif from = {"هٔ", "ۂ", "ٱ"}, to = {"ہ", "ہ", "ا"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, } m["khx"] = { "Kanu", 12952571, "bnt-lgb", } m["khy"] = { "Ekele", 6385549, "bnt-ske", "Latn", } m["khz"] = { "Keapara", 12952603, "poz-ocw", "Latn", } m["kia"] = { "Kim", 35685, "alv-kim", } m["kib"] = { "Koalib", 35859, "alv-hei", } m["kic"] = { "Kickapoo", 20162127, "alg-sfk", "Latn", } m["kid"] = { "Koshin", 35632, "nic-beb", "Latn", } m["kie"] = { "Kibet", 56893, } m["kif"] = { "Kham Parbate Timur", 12953022, "sit-kha", "Deva", } m["kig"] = { "Kimaama", 11732321, "paa-kol", "Latn", } m["kih"] = { "Kilmeri", 6408020, "paa-bew", "Latn", } m["kii"] = { "Kitsai", 56627, "cdd", "Latn", } m["kij"] = { "Kilivila", 3196601, "poz-ocw", "Latn", } m["kil"] = { "Kariya", 3438708, "cdc-wst", } m["kim"] = { "Tofa", 36848, "trk-ssb", "Cyrl", } m["kio"] = { "Kiowa", 56631, "nai-kta", "Latn", } m["kip"] = { "Sheshi Kham", 12952622, "sit-kha", "Deva", } m["kiq"] = { "Kosadle", 6432994, "paa-kko", "Latn", } m["kis"] = { "Kis", 6416362, "poz-ocw", "Latn", } m["kit"] = { "Agob", 3332143, "paa-pah", "Latn", } m["kiv"] = { "Kimbu", 10997740, "bnt-tkm", } m["kiw"] = { "Kiwai Timur Laut", 11732324, "paa-kiw", "Latn", } m["kix"] = { "Naga Khiamniungan", 6401546, "sit-kch", "Latn", } m["kiy"] = { "Kirikiri", 6415159, "paa-wlp", "Latn", } m["kiz"] = { "Kisi", 3912772, "bnt-bki", } m["kja"] = { "Mlap", 6885683, "paa-nim", "Latn", } m["kjb"] = { "Q'anjob'al", 35551, "myn", "Latn", } m["kjc"] = { "Konjo Pesisir", 3198689, "poz", "Latn", } m["kjd"] = { "Kiwai Selatan", 11732322, "paa-kiw", "Latn", } m["kje"] = { "Kisar", 3197441, "poz", "Latn", } m["kjg"] = { "Khmu", 33335, "mkh", "Laoo", ancestors = "mkh-khm-pro", sort_key = "Laoo-sortkey", } m["kjh"] = { "Khakas", 33575, "trk-ssb", "Cyrl", translit = "kjh-translit", override_translit = true, } m["kji"] = { "Zabana", 379130, "poz-ocw", "Latn", } m["kjj"] = { "Khinalug", 35278, "cau-nec", "Cyrl, Latn", translit = "kjj-translit", override_translit = true, display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, } m["kjk"] = { "Highland Konjo", 3198688, "poz", } m["kjl"] = { "Kham", 22237017, "sit-kha", "Deva", } m["kjm"] = { "Kháng", 6403501, "mkh-pal", } m["kjn"] = { "Kunjen", 3200468, "aus-pmn", "Latn", } m["kjo"] = { "Harijan Kinnauri", 5657463, "him", "Takr, Deva", } m["kjp"] = { "Pwo Timur", 5330390, "kar", "Mymr, Leke, Thai", translit = "kjp-translit", override_translit = true, } m["kjq"] = { "Keres Barat", 12645568, "nai-ker", "Latn", } m["kjr"] = { "Kurudu", 12952678, "poz-hce", "Latn", } m["kjs"] = { "Kewa Timur", 20050949, "ngf-ank", "Latn", } m["kjt"] = { "Phrae Pwo", 7187991, "kar", "Thai", } m["kju"] = { "Kashaya", 3193689, "nai-pom", "Latn", } m["kjx"] = { "Ramopa", 56830, "paa-nbo", "Latn", } m["kjy"] = { "Erave", 12952416, "ngf-ank", "Latn", } m["kjz"] = { "Bumthangkha", 2786408, "sit-ebo", "Tibt", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["kka"] = { "Kakanda", 3915342, "alv-ngb", } m["kkb"] = { "Kwerisa", 56881, "paa-clp", "Latn", } m["kkc"] = { "Odoodee", 12952987, "ngf-est", "Latn", } m["kkd"] = { "Kinuku", 6414422, "nic-kau", } m["kke"] = { "Kakabe", 3913966, "dmn-mok", "Latn", } m["kkf"] = { "Monpa Kalaktang", 63257089, "sit-tsk", "Tibt, Latn, Deva", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["kkg"] = { "Kalinga Lembah Mabaka", 18753304, "phi", } m["kkh"] = { "Khün", 3545044, "tai-swe", "Lana, Thai", sort_key = { Lana = "Lana-sortkey", Thai = "Thai-sortkey" }, } m["kki"] = { "Kagulu", 12952537, "bnt-ruv", "Latn", } m["kkj"] = { "Kako", 35755, "bnt-kak", } m["kkk"] = { "Kokota", 3198399, "poz-ocw", "Latn", } m["kkl"] = { "Kosarek Yale", 6432995, "ngf-mek", "Latn", } m["kkm"] = { "Kiong", 6414512, "nic-ucr", "Latn", } m["kkn"] = { "Kon Keu", 6428686, "mkh-pal", } m["kko"] = { "Karko", 35529, "nub-hil", "Latn", } m["kkp"] = { "Koko-Bera", 6426699, "aus-pmn", "Latn", } m["kkq"] = { "Kaiku", 6347840, "bnt-kbi", "Latn", } m["kkr"] = { "Kir-Balar", 3440527, "cdc-wst", "Latn", } m["kks"] = { "Kirfi", 56242, "cdc-wst", "Latn", } m["kkt"] = { "Koi", 6426194, "sit-kiw", } m["kku"] = { "Tumi", 3913934, "nic-kau", } m["kkv"] = { "Kangean", 2071325, "poz-msa", "Latn", } m["kkw"] = { "Teke-Kukuya", 36560, "bnt-tek", } m["kkx"] = { "Kohin", 6425997, "poz-brw", } m["kky"] = { "Guugu Yimidhirr", 56543, "aus-pam", "Latn", } m["kkz"] = { "Kaska", 20823, "ath-nor", "Latn", } m["kla"] = { "Klamath-Modoc", 2669248, "nai-plp", "Latn", } m["klb"] = { "Kiliwa", 3182593, "nai-yuc", "Latn", } m["klc"] = { "Kolbila", 6427122, "alv-lek", } m["kld"] = { "Gamilaraay", 3111818, "aus-cww", "Latn", } m["kle"] = { "Kulung", 6443304, "sit-kic", } m["klf"] = { "Kendeje", 56895, } m["klg"] = { "Kalagan Tagakaulu", 18756514, "phi", "Latn", } m["klh"] = { "Weliki", 7981017, "ngf-uru", "Latn", } m["kli"] = { "Kalumpang", 13561407, "poz", } m["klj"] = { "Khalaj", 33455, "trk", "Arab, Latn", ancestors = "klj-arg", strip_diacritics = { remove_diacritics = c.kashida .. c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun, } } m["klk"] = { "Kono (Nigeria)", 6429589, "nic-kau", "Latn", } m["kll"] = { "Kalagan Kagan", 18748913, "phi", } m["klm"] = { "Kolom", 6844970, "ngf-rai", "Latn", } m["kln"] = { "Kalenjin", 637228, "sdv-nma", "Latn", } m["klo"] = { "Kapya", 6367410, "nic-ykb", } m["klp"] = { "Kamasa", 6356107, "ngf-woj", "Latn", } m["klq"] = { "Rumu", 7379420, "paa-tki", "Latn", } m["klr"] = { "Khaling", 56381, "sit-kiw", "Deva", } m["kls"] = { "Kalasha", 33416, "inc-chi", "Latn, Aran", } m["klt"] = { "Nukna", 7068874, "ngf-uru", "Latn", } m["klu"] = { "Klao", 3914866, "kro-wkr", } m["klv"] = { "Maskelynes", 3297282, "poz-vnc", "Latn", } m["klw"] = { "Lindu", 18390055, "poz-kal", "Latn", } m["klx"] = { "Koluwawa", 6427954, "poz-ocw", "Latn", } m["kly"] = { "Kalao", 6350643, "poz-wot", "Latn", } m["klz"] = { "Kabola", 11732258, "paa-alp", "Latn", } m["kma"] = { "Konni", 35680, "nic-buk", } m["kmb"] = { "Kimbundu", 35891, "bnt-kmb", "Latn", } m["kmc"] = { "Kam Selatan", 35379, "qfa-kms", "Latn", } m["kmd"] = { "Kalinga Madukayang", 18753305, "phi", } m["kme"] = { "Bakole", 35068, "bnt-kpw", "Latn", } m["kmf"] = { "Kare (New Guinea)", 11732286, "ngf-mab", "Latn", } m["kmg"] = { "Kâte", 3201059, "ngf-kma", "Latn", } m["kmh"] = { "Kalam", 12952550, "ngf-kak", "Latn", } m["kmi"] = { "Kami", 3915372, "alv-ngb", "Latn", } m["kmj"] = { "Kumarbhag Paharia", 3130374, "dra-mlo", "Beng, Deva", } m["kmk"] = { "Kalinga Limos", 18753303, "phi", "Latn", } m["kml"] = { "Kalinga Tanudan", 18753307, "phi", "Latn", } m["kmm"] = { "Kom (India)", 12952647, "tbq-kuk", } m["kmn"] = { "Awtuw", 3504217, "paa-sep", "Latn", } m["kmo"] = { "Kwoma", 11732376, "paa-sep", "Latn", } m["kmp"] = { "Gimme", 11152236, "alv-dur", } m["kmq"] = { "Kwama", 2591184, "ssa-kom", } m["kmr"] = { "Kurdi Utara", 36163, "ku", "Latn, Cyrl, Armn, Arab, Yezi", translit = { Cyrl = "kmr-translit", -- Armn translit in [[Module:scripts/data]] Arab = "ckb-translit", }, strip_diacritics = { Latn = { remove_diacritics = "'’", from = {"r̄", "R̄", "ẍ", "Ẍ"}, to = {"rr", "Rr", "x", "X"} }, }, wikimedia_codes = "ku", } m["kms"] = { "Kamasau", 6356117, "paa-mar", "Latn", } m["kmt"] = { "Kemtuik", 6387179, "paa-nim", "Latn", } m["kmu"] = { "Kanite", 12952567, "ngf-kya", "Latn", } m["kmv"] = { "Perancis Kreol Karipúna", 2523999, "crp", "Latn", ancestors = "fr", sort_key = s["roa-oil-sortkey"], } m["kmw"] = { "Kumu", 6428450, "bnt-kbi", "Latn", } m["kmx"] = { "Waboda", 7958705, "paa-kiw", "Latn", } m["kmy"] = { "Koma", 35634, "alv-dur", } m["kmz"] = { "Turki Khorasan", 35373, "trk-ogz", "Arab", ancestors = "trk-oat", } m["kna"] = { "Kanakuru", 56811, "cdc-wst", "Latn", } m["knb"] = { "Kalinga Lubuagan", 12953602, "phi", "Latn", } m["knd"] = { "Konda", 11732340, "ngf-sbh", "Latn", } m["kne"] = { "Kankanaey", 18753329, "phi", "Latn", strip_diacritics = { Latn = { remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer, } }, sort_key = { Latn = "tl-sortkey", }, standard_chars = { Latn = "AaBbKkDdEeGgHhIiLlMmNnOoPpRrSsTtUuWwYy" .. c.punc, }, } m["knf"] = { "Mankanya", 35789, "alv-pap", "Latn", } m["kni"] = { "Kanufi", 3913297, "nic-nin", "Latn", } m["knj"] = { "Akatek", 34923, "myn", "Latn", } m["knk"] = { "Kuranko", 3198896, "dmn-mok", "Latn", } m["knl"] = { "Keninjal", 6389309, "poz-mly", "Latn", } m["knm"] = { -- two unrelated lects have this name; this is the Katukinian one "Kanamari", 3438373, "sai-ktk", "Latn", } m["kno"] = { "Kono (Sierra Leone)", 35675, "dmn-vak", "Latn", } m["knp"] = { "Kwanja", 35641, "nic-mmb", "Latn", } m["knq"] = { "Kintaq", 6414335, "mkh-asl", "Latn", } m["knr"] = { "Kaningra", 6363253, "paa-sep", "Latn", } m["kns"] = { "Kensiu", 6391529, "mkh-asl", "Latn, Thai, Hani", } m["knt"] = { "Katukina", 3194265, "sai-pan", "Latn", } m["knu"] = { -- a dialect of 'kpe' "Kono (Guinea)", 3198703, "dmn-msw", "Latn, Kpel", ancestors = "kpe", } m["knv"] = { "Tabo", 7959888, "aav", } m["knx"] = { "Salako", 6388963, "poz-mly", "Latn", } m["kny"] = { "Kanyok", 11110766, "bnt-lub", "Latn", } m["knz"] = { "Kalamsé", 3914000, "nic-gnn", "Latn", } m["koa"] = { "Konomala", 3198732, "poz-ocw", "Latn", } m["koc"] = { "Kpati", 3913279, "nic-nge", "Latn", } m["kod"] = { "Kodi", 4577633, "poz-cet", "Latn", } m["koe"] = { "Kacipo-Balesi", 5364424, "sdv", "Latn", } m["kof"] = { "Kubi", 3438718, "cdc-wst", "Latn", } m["kog"] = { "Cogui", 3198286, "cba", "Latn", } m["koh"] = { "Koyo", 35649, "bnt-mbo", "Latn", } m["koi"] = { "Komi-Permyak", 56318, "kv", "Cyrl", translit = "kv-translit", strip_diacritics = {remove_diacritics = c.acute}, override_translit = true, } m["kok"] = { "Konkani", 34239, "inc-sou", "Deva, Knda, Mlym, Arab, Latn", translit = { Deva = "mr-translit", }, -- Knda translit in [[Module:scripts/data]] -- Mlym translit in [[Module:scripts/data]] strip_diacritics = { -- FIXME: Separate out the scripts from = {"च़", "ज़", "झ़", "ಚ಼", "ಜ಼", "ಝ಼"}, to = {"च", "ज", "झ", "ಚ", "ಜ", "ಝ"} } , } m["kol"] = { "Kol (New Guinea)", 4227542, } m["koo"] = { "Konzo", 2361829, "bnt-glb", "Latn", } m["kop"] = { "Waube", 11732373, "ngf-nur", "Latn", } m["koq"] = { "Kota (Gabon)", 35607, "bnt-kel", "Latn", } m["kos"] = { "Kosrae", 33464, "poz-mic", "Latn", } m["kot"] = { "Lagwan", 3502264, "cdc-cbm", "Latn", } m["kou"] = { "Koke", 797249, "alv-bua", } m["kov"] = { "Kudu-Camo", 3915850, "nic-jer", } m["kow"] = { "Kugama", 3913307, "alv-mye", } m["koy"] = { "Koyukon", 28304, "ath-nor", "Latn", } m["koz"] = { "Korak", 6431365, "ngf-kow", "Latn", } m["kpa"] = { "Kutto", 3437656, "cdc-wst", } m["kpb"] = { "Mullu Kurumba", 19573111, "dra-mal", } m["kpc"] = { "Curripaco", 2882543, "awd-nwk", "Latn", } m["kpd"] = { "Koba", 6424249, "poz", } m["kpe"] = { "Kpelle", 35673, "dmn-msw", "Latn, Kpel", } m["kpf"] = { "Komba", 6428239, "ngf-kab", "Latn", } m["kpg"] = { "Kapingamarangi", 35771, "poz-pnp", "Latn", } m["kph"] = { "Kplang", 35628, "alv-gng", } m["kpi"] = { "Kofei", 6425665, "paa-egb", "Latn", } m["kpj"] = { "Karajá", 10322066, "sai-mje", "Latn", } m["kpk"] = { "Kpan", 3915380, "nic-jkn", "Latn", } m["kpl"] = { "Kpala", 11154769, "nic-nkk", "Latn", } m["kpm"] = { "Koho", 3511919, "mkh-ban", "Latn", } m["kpn"] = { "Kepkiriwát", 3195366, "tup", "Latn", } m["kpo"] = { "Ikposo", 35029, "alv-ktg", "Latn", } m["kpq"] = { "Korupun-Sela", 6432769, "ngf-mek", "Latn", } m["kpr"] = { "Korafe-Yegha", 11732347, "ngf-gko", "Latn", } m["kps"] = { "Tehit", 7694851, "paa-wbh", "Latn", } m["kpt"] = { "Karata", 56636, "cau-and", "Cyrl", translit = "kpt-translit", override_translit = true, display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]}, } m["kpu"] = { "Kafoa", 6346151, "paa-alp", "Latn", } m["kpv"] = { "Komi-Zyrian", 34114, "kv", "Cyrl", translit = "kv-translit", override_translit = true, wikimedia_codes = "kv", } m["kpw"] = { "Kobon", 11732326, "ngf-kak", "Latn", } m["kpx"] = { "Koiari Gunung", 6925030, "ngf-koi", "Latn", } m["kpy"] = { "Koryak", 36199, "qfa-ckn", "Cyrl", strip_diacritics = { from = {"['’]"}, to = {"ʼ"} }, sort_key = { from = {"вʼ", "гʼ", "ё", "ӄ", "ӈ"}, to = {"в" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1]} }, translit = "kpy-translit", } m["kpz"] = { "Kupsabiny", 56445, "sdv-kln", "Latn", } m["kqa"] = { "Mum", 6935252, "ngf-nso", "Latn", } m["kqb"] = { "Kovai", 6434822, "ngf-ehu", "Latn", } m["kqc"] = { "Doromu-Koki", 5298175, "paa-man", "Latn", } m["kqd"] = { "Koy Sanjaq Surat", 33463, "sem-nna", } m["kqe"] = { "Kalagan", 18748906, "phi", "Latn", } m["kqf"] = { "Kakabai", 6349119, "poz-ocw", "Latn", } m["kqg"] = { "Khe", 3914015, "nic-gur", "Latn", } m["kqh"] = { "Kisankasa", 6416409, "sdv", } m["kqi"] = { "Koitabu", 6426363, "ngf-koi", "Latn", } m["kqj"] = { "Koromira", 6432520, "paa-sbo", "Latn", } m["kqk"] = { "Gbe Kotafon", 12952447, "alv-pph", } m["kql"] = { "Kyenele", 11732453, "paa-yua", "Latn", } m["kqm"] = { "Khisa", 3913955, "nic-gur", "Latn", } m["kqn"] = { "Kaonde", 33601, "bnt-lub", "Latn", } m["kqo"] = { "Krahn Timur", 3915374, "kro-wee", "Latn", } m["kqp"] = { "Kimré", 3441210, "cdc-est", "Latn", } m["kqq"] = { "Krenak", 6436747, "sai-cer", "Latn", } m["kqr"] = { "Kimaragang", 3196845, "poz-san", "Latn", } m["kqs"] = { "Kissi Utara", 19921576, "alv-kis", } m["kqt"] = { "Kadazan Sungai Klias", 12953594, "poz-san", } m["kqu"] = { "Seroa", 33127766, "khi-tuu", "Latn", } m["kqv"] = { "Okolod", 7082487, "poz-san", } m["kqw"] = { "Kandas", 3192590, "poz-ocw", "Latn", } m["kqx"] = { "Mser", 3502347, "cdc-cbm", } m["kqy"] = { "Koorete", 6430753, "omv-eom", "Ethi, Latn", } m["kqz"] = { "Korana", 2756709, "khi-khk", "Latn", } m["kra"] = { "Kumhali", 13580783, "inc-bih", "Deva", translit = "ne-translit", } m["krb"] = { "Karkin", 3193345, "nai-utn", "Latn", } m["krc"] = { "Karachay-Balkar", 33714, "trk-kcu", "Cyrl", translit = "krc-translit", sort_key = { from = {"гъ", "дж", "ё", "къ", "нг"}, to = {"г" .. p[1], "д" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1]} }, } m["krd"] = { "Kairui-Midiki", 12953277, "poz-tim", } m["kre"] = { "Panará", 3361895, "sai-cer", "Latn", } m["krf"] = { "Koro (Vanuatu)", 3198995, "poz-vnn", "Latn", } m["krh"] = { "Kurama", 35593, "nic-kau", } m["kri"] = { "Krio", 35744, "crp", "Latn", ancestors = "en", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ}, sort_key = { from = {"ɛ", "gb", "kp", "ɔ"}, to = {"e" .. p[1], "g" .. p[1], "k" .. p[1], "o" .. p[1]} }, } m["krj"] = { "Kinaray-a", 33720, "phi", "Latn", } m["krk"] = { "Kerek", 332792, "qfa-ckn", "Cyrl", } m["krl"] = { "Karelia", 33557, "urj-fin", "Latn", sort_key = { from = { "č", "š", "ž", "ü", "ä", "ö", -- 2 chars "z", "'" -- 1 char }, to = { "c" .. p[1], "s" .. p[1], "s" .. p[3], "y" .. p[1], "y" .. p[2], "y" .. p[3], "s" .. p[2], "y" .. p[4], } }, } m["krm"] = { "Krim", 35713, "alv", } m["krn"] = { "Sapo", 3915386, "kro-wee", } m["krp"] = { "Korop", 35626, "nic-ucr", "Latn", } m["krr"] = { "Kru'ng", 12953650, "mkh-ban", } m["krs"] = { "Kresh", 56674, "csu-bkr", } m["kru"] = { "Kurukh", 33492, "dra-kml", "Deva, Tols", translit = { Deva = "hi-translit", }, } m["krv"] = { "Kavet", 12953649, "sai-ktk", "Latn", } m["krw"] = { "Krahn Barat", 10975611, "kro-wee", "Latn", } m["krx"] = { "Karon", 35704, "alv-jol", } m["kry"] = { "Kryts", 35861, "cau-ssm", "Latn, Cyrl", display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = { Latn = s["cau-Latn-stripdiacritics"], Cyrl = s["cau-Cyrl-stripdiacritics"], }, } m["krz"] = { "Sota Kanum", 12952568, "paa-kan", "Latn", } m["ksa"] = { "Shuwa-Zamani", 3913929, "nic-kau", } m["ksb"] = { "Shambala", 3788739, "bnt-seu", "Latn", } m["ksc"] = { "Kalinga Selatan", 18753301, "phi", } m["ksd"] = { "Tolai", 35870, "poz-ocw", "Latn", } m["kse"] = { "Kuni", 6444619, "poz-ocw", "Latn", } m["ksf"] = { "Bafia", 34930, "bnt-baf", "Latn", } m["ksg"] = { "Kusaghe", 3200638, "poz-ocw", "Latn", } m["ksi"] = { "Krisa", 841704, "paa-sko", "Latn", } m["ksj"] = { "Uare", 6450052, "paa-kwa", "Latn", } m["ksk"] = { "Kansa", 3192772, "sio-dhe", "Latn", } m["ksl"] = { "Kumalu", 17584381, "poz-ocw", "Latn", } m["ksm"] = { "Kumba", 3913972, "alv-mye", } m["ksn"] = { "Kasiguranin", 6374525, "phi", } m["kso"] = { "Kofa", 56278, "cdc-cbm", } m["ksp"] = { "Kaba", 3915316, "csu-sar", } m["ksq"] = { "Kwaami", 3440525, "cdc-wst", } m["ksr"] = { "Borong", 4946263, "ngf-kbm", "Latn", } m["kss"] = { "Kissi Selatan", 11028974, "alv-kis", } m["kst"] = { "Winyé", 3913360, "nic-gnw", } m["ksu"] = { "Khamyang", 6583541, "tai-swe", } m["ksv"] = { "Kusu", 6448199, "bnt-tet", } m["ksw"] = { "Karen S'gaw", 56410, "kar", "Mymr", translit = "ksw-translit", } m["ksx"] = { "Kedang", 6382520, "poz", "Latn", } m["ksy"] = { "Kharia Thar", 6400661, "inc-eas", } m["ksz"] = { "Kodaku", 21179986, "mun", } m["kta"] = { "Katua", 6378404, "mkh-ban", } m["ktb"] = { "Kambaata", 35664, "cus-hec", "Latn", } m["ktc"] = { "Kholok", 3440464, "cdc-wst", } m["ktd"] = { "Kokata", 10547021, "aus-pam", "Latn", } m["ktf"] = { "Kwami", 12952687, "bnt-lgb", } m["ktg"] = { "Kalkatungu", 3914057, "aus-pam", "Latn", } m["kth"] = { "Karanga", 713643, } m["kti"] = { "Muyu Utara", 20857698, "ngf-lok", "Latn", } m["ktj"] = { "Plapo Krumen", 10975356, "kro-grb", } m["ktk"] = { "Kaniet", 3399050, "poz-aay", "Latn", } m["ktl"] = { "Koroshi", 3775265, "ira-nwi", ancestors = "bal", } m["ktm"] = { "Kurti", 3200615, "poz-aay", "Latn", } m["ktn"] = { "Karitiâna", 3112184, "tup", "Latn", } m["kto"] = { "Kuot", 56537, } m["ktp"] = { "Kaduo", 769809, "tbq-bka", } m["ktq"] = { "Katabaga", 3193895, } m["kts"] = { "Muyu Selatan", 42308820, "ngf-lok", "Latn", } m["ktt"] = { "Ketum", 12952616, "ngf-dum", "Latn", } m["ktu"] = { "Kituba", 35746, "crp", "Latn", ancestors = "kg", } m["ktv"] = { "Katu Timur", 22808951, "mkh-kat", "Latn", } m["ktw"] = { "Kato", 20831, "ath-pco", "Latn", } m["ktx"] = { "Kaxararí", 6380124, "sai-pan", "Latn", } m["kty"] = { "Kango", 6362818, "bnt-bta", "Latn", } m["ktz"] = { "Juǀ'hoan", 1192295, "khi-kxa", "Latn", } m["kub"] = { "Kutep", 35645, "nic-jkn", } m["kuc"] = { "Kwinsu", 6450460, "paa-tor", "Latn", } m["kud"] = { "Auhelawa", 5166, "poz-ocw", "Latn", } m["kue"] = { "Kuman", 137525, "ngf-sim", "Latn", } m["kuf"] = { "Katu Barat", 6378400, "mkh-kat", "Laoo, Tale, Latn", } m["kug"] = { "Kupa", 3915336, "alv-ngb", } m["kuh"] = { "Kushi", 3438747, "cdc-wst", } m["kui"] = { "Kuikúro", 3915522, "sai-kui", "Latn", } m["kuj"] = { "Kuria", 6445968, "bnt-lok", "Latn", } m["kuk"] = { "Kepo'", 6393217, "poz", } m["kul"] = { "Kulere", 3440506, "cdc-wst", } m["kum"] = { "Kumyk", 36209, "trk-kcu", "Cyrl", translit = "kum-translit", sort_key = { from = {"гъ", "гь", "ё", "къ", "нг", "оь", "уь"}, to = {"г" .. p[1], "г" .. p[2], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1]} }, } m["kun"] = { "Kunama", 36041, } m["kuo"] = { "Kumukio", 11732362, "ngf-dal", "Latn", } m["kup"] = { "Kunimaipa", 6444696, "paa-kun", "Latn", } m["kuq"] = { "Karipuna", 6371071, "tup-gua", "Latn", } m["kus"] = { "Kusaal", 35708, "nic-dag", "Latn", } m["kut"] = { "Kutenai", 33434, "qfa-iso", "Latn", } m["kuu"] = { "Upper Kuskokwim", 28062, "ath-nor", "Latn", } m["kuv"] = { "Kur", 12635082, "poz-cma", "Latn", } m["kuw"] = { "Kpagua", 11137573, "bad-cnt", } m["kux"] = { "Kukatja", 10549839, "aus-pam", "Latn", } m["kuy"] = { "Kuuku-Ya'u", 10550697, "aus-pmn", "Latn", } m["kuz"] = { "Kunza", 2669181, "qfa-iso", "Latn", } m["kva"] = { "Bagvalal", 56638, "cau-and", "Cyrl", translit = "cau-nec-translit", override_translit = true, display_text = {Cyrl = s["cau-Cyrl-displaytext"]}, strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]}, } m["kvb"] = { "Kubu", 6441341, "poz-mly", } m["kvc"] = { "Kove", 3199402, "poz-ocw", "Latn", } m["kvd"] = { "Kui (Indonesia)", 6442230, "paa-alp", "Latn", } m["kve"] = { "Kalabakan", 6350003, "poz-san", "Latn", } m["kvf"] = { "Kabalai", 3440427, "cdc-est", } m["kvg"] = { "Kuni-Boazi", 2907551, "paa-boa", "Latn", } m["kvh"] = { "Komodo", 3198565, "poz-cet", "Latn", } m["kvi"] = { "Kwang", 3440398, "cdc-est", "Latn", } m["kvj"] = { "Psikye", 56304, "cdc-cbm", } m["kvk"] = { "Bahasa Isyarat Korea", 3073428, "sgn-jsl", } m["kvl"] = { "Karen Brek", 12952577, "kar", } m["kvm"] = { "Kendem", 35751, "nic-mam", "Latn", } m["kvn"] = { "Kuna Sempadan", 31777873, "cba", } m["kvo"] = { "Dobel", 5286559, "poz", "Latn", } m["kvp"] = { "Kompane", 18343041, "poz", } m["kvq"] = { "Karen Geba", 12952581, "kar", "Latn, Mymr", } m["kvr"] = { "Kerinci", 3195442, "poz-mly", "Latn, Arab", -- Also Incung, which we don't have } m["kvt"] = { "Karen Lahta", 12952582, "kar", } m["kvu"] = { "Karen Yinbaw", 14426328, "kar", } m["kvv"] = { "Kola", 6426967, "poz", "Latn", } m["kvw"] = { "Wersing", 7983599, "paa-alp", "Latn", } m["kvx"] = { "Parkari Koli", 3244176, "inc-wes", } m["kvy"] = { "Karen Yintale", 14426329, "kar", } m["kvz"] = { "Tsakwambo", 7849438, "ngf-kts", "Latn", } m["kwa"] = { "Dâw", 3042278, "sai-nad", "Latn", } m["kwb"] = { "Baa", 34842, "alv-ada", } m["kwc"] = { "Likwala", 35597, "bnt-mbo", } m["kwd"] = { "Kwaio", 3200796, "poz-sls", "Latn", } m["kwe"] = { "Kwerba", 6450328, "paa-kwe", "Latn", } m["kwf"] = { "Kwara'ae", 3200829, "poz-sls", "Latn", } m["kwg"] = { "Sara Kaba Deme", 3915384, "csu-kab", } m["kwh"] = { "Kowiai", 6435028, "poz", "Latn", } m["kwi"] = { "Awa-Cuaiquer", 2603103, "sai-bar", "Latn", } m["kwj"] = { "Kwanga", 3438383, "paa-sep", "Latn", } m["kwk"] = { "Kwak'wala", 2640628, "wak", "Latn", } m["kwl"] = { "Kofyar", 3441382, "cdc-wst", "Latn", } m["kwm"] = { "Kwambi", 3487165, "bnt-ova", } m["kwn"] = { "Kwangali", 36334, "bnt-kav", "Latn", } m["kwo"] = { "Kwomtari", 3508116, "paa-kwo", "Latn", } m["kwp"] = { "Kodia", 3914867, "kro-ekr", } m["kwq"] = { "Kwak", 11014183, "nic-nka", ancestors = "yam", } m["kwr"] = { "Kwer", 12635137, "ngf-wok", "Latn", } m["kws"] = { "Kwese", 3200846, "bnt-pen", } m["kwt"] = { "Kwesten", 6450354, "paa-tor", "Latn", } m["kwu"] = { "Kwakum", 35624, "bnt-kak", } m["kwv"] = { "Sara Kaba Náà", 3915361, "csu-kab", "Latn", } m["kww"] = { "Kwinti", 721182, "crp", "Latn", ancestors = "en" } m["kwx"] = { "Khirwar", 12976968, "dra", } m["kwz"] = { "Kwadi", 2364661, "khi-kkw", "Latn", } m["kxa"] = { "Kairiru", 3398785, "poz-ocw", "Latn", } m["kxb"] = { "Krobu", 35586, "alv-ptn", "Latn", } m["kxc"] = { "Khonso", 56624, "cus-eas", "Ethi, Latn", } m["kxd"] = { "Melayu Brunei", 3182878, "poz-mly", "Latn, Arab", } m["kxe"] = { "Kakihum", 3914433, "nic-kam", ancestors = "tvd", } m["kxf"] = { "Karen Manumanaw", 12952592, "kar", "Mymr, Latn", } m["kxh"] = { "Karo", 3447116, "omv-aro", } m["kxi"] = { "Murut Keningau", 6389308, "poz-san", "Latn", } m["kxj"] = { "Kulfa", 713654, "csu-kab", } m["kxk"] = { "Karen Zayein", 14352960, "kar", } -- Nepali Kurux [kxl] treated as part of Kurux [kru], consistent with ISO merger in 2020 m["kxm"] = { "Khmer Utara", 3502234, "mkh-kmr", "Thai, Khmr", ancestors = "xhm", sort_key = { from = {"[%pๆ]", "[็-๎]", "([เแโใไ])([ก-ฮ])"}, to = {"", "", "%2%1"} }, } m["kxn"] = { "Kanowit", 6364300, "poz-bnn", "Latn", } m["kxo"] = { "Kanoé", 4356223, "qfa-iso", "Latn", } m["kxp"] = { "Wadiyara Koli", 12953645, "inc-wes", } m["kxq"] = { "Smärky Kanum", 12952569, "paa-kan", "Latn", } m["kxr"] = { "Koro (New Guinea)", 3198994, "poz-aay", "Latn", } m["kxs"] = { "Kangjia", 3182570, "xgn-shr", "Latn", } m["kxt"] = { "Koiwat", 6426388, "paa-nnd", "Latn", } m["kxu"] = { "Kui (India)", 33919, "dra-kki", "Orya", translit = "kxv-translit", strip_diacritics = { remove_diacritics = "୕", from = {"ଆଆ", "ଇଇ", "ଉଉ", "ଏଏ", "ଓଓ", "ିଇ", "ୁଉ", "େଏ", "ୋଓ"}, to = {"ଆ", "ଈ", "ଊ", "ଏ", "ଓ", "ୀ", "ୂ", "େ", "ୋ"}, }, } m["kxv"] = { "Kuvi", 3200721, "dra-kki", "Orya", translit = "kxv-translit", strip_diacritics = { remove_diacritics = "୕", from = {"ଆଆ", "ଇଇ", "ଉଉ", "ଏଏ", "ଓଓ", "([କ-ହ])ଆ", "ିଇ", "ୁଉ", "େଏ", "ୋଓ"}, to = {"ଆ", "ଈ", "ଊ", "ଏ", "ଓ", "%1ା", "ୀ", "ୂ", "େ", "ୋ"}, }, } m["kxw"] = { "Konai", 11732339, "ngf-est", "Latn", } m["kxx"] = { "Likuba", 35646, "bnt-bmo", } m["kxy"] = { "Kayong", 6380673, "mkh", } m["kxz"] = { "Kerewo", 6393847, "paa-kiw", "Latn", } m["kya"] = { "Kwaya", 6450276, "bnt-haj", "Latn", } m["kyb"] = { "Kalinga Butbut", 18753300, "phi", "Latn", } m["kyc"] = { "Kyaka", 12952690, "ngf-enc", "Latn", } m["kyd"] = { "Karey", 6370196, "poz", } m["kye"] = { "Krache", 35658, "alv-gng", } m["kyf"] = { "Kouya", 35595, "kro-bet", } m["kyg"] = { "Keyagana", 6398208, "ngf-kya", "Latn", } m["kyh"] = { "Karok", 1288440, "qfa-iso", -- or Hokan? "Latn", } m["kyi"] = { "Kiput", 3038653, "poz-swa", "Latn", } m["kyj"] = { "Karao", 3192950, "phi", "Latn", } m["kyk"] = { "Kamayo", 3192339, "phi", "Latn", } m["kyl"] = { "Kalapuya", 3192120, "nai-klp", } m["kym"] = { "Kpatili", 3913982, "znd", } m["kyn"] = { "Karolanos", 6373093, "phi", } m["kyo"] = { "Kelon", 6386414, "paa-alp", "Latn", } m["kyp"] = { "Kang", 25559558, "tai", } m["kyq"] = { "Kenga", 35707, "csu-bgr", } m["kyr"] = { "Kuruáya", 3200633, "tup", "Latn", } m["kys"] = { "Kayan Baram", 2883794, "poz", "Latn", } m["kyt"] = { "Kayagar", 6380394, "paa-kay", "Latn", } m["kyu"] = { "Kayah Barat", 12952596, "kar", "Kali, Mymr, Latn", translit = {Kali = "Kali-translit"}, } m["kyv"] = { "Kayort", 6380675, "inc-krd", "Deva", } m["kyw"] = { "Kudmali", 6446173, "inc-sad", "Deva, as-Beng, Orya, Chis", } m["kyx"] = { "Rapoisi", 7294279, "paa-nbo", "Latn", } m["kyy"] = { "Kambaira", 6356254, "ngf-kai", "Latn", } m["kyz"] = { "Kayabí", 6380372, "tup-gua", "Latn", } m["kza"] = { "Karaboro Barat", 36601, "alv-krb", } m["kzb"] = { "Kaibobo", 6347565, "poz-cma", "Latn", } m["kzc"] = { "Bondoukou Kulango", 11031321, "alv-kul", "Latn", } m["kzd"] = { "Kadai", 7679471, "poz-cma", "Latn", } --kze (Kosena) made an etym-only child of auy (Auyana) per [[Wiktionary:Language_treatment_requests#merge_Kosena_[kze]_into_Auyana_[auy]]] m["kzf"] = { "Da'a Kaili", 33103997, "poz-kal", "Latn", } m["kzg"] = { "Kikai", 3196527, "jpx-nry", "Jpan", translit = s["jpx-translit"], display_text = s["jpx-displaytext"], strip_diacritics = s["jpx-stripdiacritics"], sort_key = s["jpx-sortkey"], } m["kzh"] = { "Dongolawi", 5295991, "nub", "Latn", } m["kzi"] = { "Kelabit", 6385445, "poz-swa", "Latn", } m["kzj"] = { "Kadazan", 3307195, "poz-san", "Latn", } m["kzk"] = { "Kazukuru", 1089069, "poz-ocw", } m["kzl"] = { "Kayeli", 4207444, "poz-cma", "Latn", } m["kzm"] = { "Kais", 6348319, "ngf-sbh", "Latn", } m["kzn"] = { "Kokola", 11128329, "bnt-mak", "Latn", ancestors = "vmw", } m["kzo"] = { "Kaningi", 35683, "bnt-mbt", } m["kzp"] = { "Kaidipang", 6347611, "phi", "Latn", } m["kzq"] = { "Kaike", 10951226, "sit-tam", } m["kzr"] = { "Karang", 35681, "alv-mbm", "Latn", } m["kzs"] = { "Dusun Sugut", 12953510, "poz-san", "Latn", } m["kzt"] = { "Dusun Tambunan", 12953514, "poz-san", "Latn", } m["kzu"] = { "Kayupulau", 6380723, "poz-ocw", } m["kzv"] = { "Komyandaret", 6428671, "ngf-kts", "Latn", } m["kzw"] = { -- contrast xoo, sai-kat, sai-xoc, the last of which the ISO conflated into this code "Kariri", 12953620, "sai-mje", "Latn", } m["kzx"] = { "Kamarian", 6356040, "poz-cma", "Latn", } m["kzy"] = { "Kango-Sua", 11008360, "bnt-kbi", "Latn", ancestors = "bip", } m["kzz"] = { "Kalabra", 6350038, "paa-wbh", "Latn", } return require("Module:languages").finalizeData(m, "language") 3e9kmkh1hp86ecbfgihibcrgjsohcph Modul:languages/data/3/t 828 9821 375375 373732 2026-09-22T05:33:45Z Hakimi97 2668 Move "Temuan" back to Module:languages/data/3/t 375375 Scribunto text/plain local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared local m = {} m["taa"] = { "Lower Tanana", 28565, "ath-nor", "Latn", } m["tab"] = { "Tabasaran", 34079, "cau-esm", "Cyrl, Latn, Arab", translit = { Cyrl = "tab-translit", }, override_translit = true, display_text = { Cyrl = s["cau-Cyrl-displaytext"] }, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, sort_key = { Cyrl = "tab-sortkey", } } m["tac"] = { "Lowland Tarahumara", 15616384, "azc-trc", "Latn", } m["tad"] = { "Tause", 2356440, "paa-wlp", "Latn", } m["tae"] = { "Tariana", 732726, "awd-nwk", "Latn", } m["taf"] = { "Tapirapé", 7684673, "tup-gua", "Latn", } m["tag"] = { "Tagoi", 36537, "nic-ras", "Latn", } m["taj"] = { "Tamang Timur", 12953177, "sit-tam", "sit-tam-Tibt, Deva", -- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] -- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake. } m["tak"] = { "Tala", 3914494, "cdc-wst", "Latn", } m["tal"] = { "Tal", 3440387, "cdc-wst", "Latn", } m["tan"] = { "Tangale", 529921, "cdc-wst", "Latn", } m["tao"] = { "Yami", 715760, "phi", "Latn", } m["tap"] = { "Taabwa", 7673650, "bnt-sbi", "Latn", } m["tar"] = { "Tarahumara Tengah", 20090009, "azc-trc", "Latn", sort_key = {remove_diacritics = c.acute .. "ꞌ"}, } m["tas"] = { "Tây Bồi", 2233794, "crp", "Latn", ancestors = "fr", sort_key = s["roa-oil-sortkey"], } m["tau"] = { "Upper Tanana", 28281, "ath-nor", "Latn", } m["tav"] = { "Tatuyo", 2524007, "sai-tuc", "Latn", } m["taw"] = { "Tai", 7675861, "ngf-kak", "Latn", } m["tax"] = { "Tamki", 3449082, "cdc-est", "Latn", } m["tay"] = { "Atayal", 715766, "map-ata", "Latn", } m["taz"] = { "Tocho", 36680, "alv-tal", "Latn", } m["tba"] = { "Aikanã", 3409307, "qfa-iso", "Latn", } -- Tapeba [tbb] is spurious m["tbc"] = { "Takia", 3514336, "poz-oce", "Latn", } m["tbd"] = { "Kaki Ae", 6349417, "qfa-iso", -- isolate in Glottolog and Pawley and Hammarström (2018); tentatively in Eleman family by Ross (2005) -- (and Usher?), but they don't address counterarguments of Clifton 1997 "Latn", } m["tbe"] = { "Tanimbili", 3515188, "poz-tem", "Latn", } m["tbf"] = { "Mandara", 3285424, "poz-ocw", "Latn", } m["tbg"] = { "Tairora Utara", 20210398, "ngf-tai", "Latn", } m["tbh"] = { "Thurawal", 3537135, "aus-yuk", "Latn", } m["tbi"] = { "Gaam", 35338, "sdv-eje", "Latn", } m["tbj"] = { "Tiang", 3528020, "poz-ocw", "Latn", } m["tbk"] = { "Calamian Tagbanwa", 3915487, "phi-kal", "Tagb, Latn", } m["tbl"] = { "Tboli", 7690594, "phi", "Latn", } m["tbm"] = { "Tagbu", 7675188, "nic-ser", } m["tbn"] = { "Barro Negro Tunebo", 12953943, "cba", } m["tbo"] = { "Tawala", 7689206, "poz-ocw", "Latn", } m["tbp"] = { "Taworta", 7689337, "paa-elp", "Latn", } m["tbr"] = { "Tumtum", 3407029, "qfa-kad", } m["tbs"] = { "Tanguat", 7683166, "paa-ata", "Latn", } m["tbt"] = { "Kitembo", 13123561, "bnt-shh", "Latn", } m["tbu"] = { "Tubar", 56730, "azc-trc", "Latn", } m["tbv"] = { -- considered a dialect of Kulungtfu-Yuanggeng-Tobo [kgf] by Glottolog "Tobo", 7811712, "ngf-kto", "Latn", } m["tbw"] = { "Tagbanwa", 3915475, "phi", "Latn", } m["tbx"] = { "Kapin", 6366665, "poz-ocw", "Latn", } m["tby"] = { "Tabaru", 11732670, "paa-gto", "Latn", } m["tbz"] = { "Ditammari", 35186, "nic-eov", "Latn", } m["tca"] = { "Ticuna", 1815205, "sai-tyu", "Latn", } m["tcb"] = { "Tanacross", 28268, "ath-nor", "Latn", } m["tcc"] = { "Datooga", 35327, "sdv-nis", "Latn", } m["tcd"] = { "Tafi", 36545, "alv-ktg", } m["tce"] = { "Tutchone Selatan", 31091048, "ath-nor", "Latn", } m["tcf"] = { "Tlapanec Malinaltepec", 25559732, "omq", "Latn", } m["tcg"] = { "Tamagario", 7680531, "paa-kay", "Latn", } m["tch"] = { "Inggeris Kreol Turks dan Caicos", 7855478, "crp", "Latn", ancestors = "en", } m["tci"] = { "Wára", 20825638, "paa-wko", "Latn", } m["tck"] = { "Tchitchege", 36595, "bnt-tek", } m["tcl"] = { "Taman (Myanmar)", 15616518, "sit-jnp", "Latn", } m["tcm"] = { "Tanahmerah", 3514927, "qfa-dis", -- Papuan; isolate per Glottolog and Palmer (2018), considered an independent branch of TNG by Usher -- (2020); seems based only on some pronoun correspondences "Latn", } m["tco"] = { "Taungyo", 12953186, "tbq-brm", ancestors = "obr", } m["tcp"] = { "Chin Tawr", 7689338, "tbq-kuk", } m["tcq"] = { "Kaiy", 6348709, "paa-clp", "Latn", } m["tcs"] = { "Kreol Selat Torres", 36648, "crp", "Latn", ancestors = "en", } m["tct"] = { "T'en", 3442330, "qfa-kms", } m["tcu"] = { "Tarahumara Tenggara", 36807, "azc-trc", "Latn", } m["tcw"] = { "Tecpatlán Totonac", 7692795, "nai-ttn", "Latn", } m["tcx"] = { "Toda", 34042, "dra-tkt", "Taml", --translit = {Taml = "Taml-translit"}, } m["tcy"] = { "Tulu", 34251, "dra-tlk", "Tutg, Mlym, Knda", -- Mlym is nearer than Knda but both lack ɛ/ɛː. translit = { Tutg = "tcy-Tutg-translit", }, -- Knda translit in [[Module:scripts/data]] -- Mlym translit in [[Module:scripts/data]] } m["tcz"] = { "Chin Thado", 6583558, "tbq-kuk", } m["tda"] = { "Tagdal", 36570, "son", } m["tdb"] = { "Panchpargania", 21946879, "inc-sad", "Deva, as-Beng, Orya, Chis", } m["tdc"] = { "Emberá-Tadó", 3052041, "sai-chc", "Latn", } m["tdd"] = { "Tai Nüa", 36556, "tai-swe", "Tale", translit = "Tale-translit", strip_diacritics = {remove_diacritics = c.ZWNJ .. c.ZWJ}, } m["tde"] = { "Tiranige Diga Dogon", 5313387, "nic-dgw", } m["tdf"] = { "Talieng", 37525108, "mkh-ban", } m["tdg"] = { "Tamang Barat", 12953178, "sit-tam", "sit-tam-Tibt, Deva", -- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] -- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake. } m["tdh"] = { "Thulung", 56553, "sit-kiw", } m["tdi"] = { "Tomadino", 7818197, "poz-btk", "Latn", } m["tdj"] = { "Tajio", 7676870, "poz", "Latn", } m["tdk"] = { "Tambas", 3440392, "cdc-wst", } m["tdl"] = { "Sur", 3914453, "nic-tar", } m["tdm"] = { "Taruma", 5559094, } m["tdn"] = { "Tondano", 3531514, "phi", "Latn", } m["tdo"] = { "Teme", 3913994, "alv-mye", } m["tdq"] = { "Tita", 3914899, "nic-bco", } m["tdr"] = { "Todrah", 7812881, "mkh", } m["tds"] = { "Doutai", 5302331, "paa-clp", "Latn", } m["tdt"] = { "Tetun Dili", 12643484, "poz-tim", "Latn", } m["tdv"] = { "Toro", 3438367, "nic-alu", "Latn", } m["tdy"] = { "Tadyawan", 7674700, "phi", "Latn", } m["tea"] = { "Temiar", 3914693, "mkh-asl", "Latn", } m["teb"] = { "Tetete", 7706087, "sai-tuc", "Latn", } m["tec"] = { "Terik", 3518379, "sdv-nma", } m["ted"] = { "Tepo Krumen", 11152243, "kro-grb", } m["tee"] = { "Tepehua Huehuetla", 56455, "nai-ttn", "Latn", } m["tef"] = { "Teressa", 3518362, "aav-nic", } m["teg"] = { "Teke-Tege", 36478, "bnt-tek", } m["teh"] = { "Tehuelche", 33930, "sai-cho", "Latn", } m["tei"] = { "Torricelli", 3450788, "paa-kom", "Latn", } m["tek"] = { "Ibali Teke", 2802914, "bnt-tek", } m["tem"] = { "Temne", 36613, "alv-mel", "Latn", } m["ten"] = { "Tama (Colombia)", 3832969, "sai-tuc", "Latn", } m["teo"] = { "Ateso", 29474, "sdv-ttu", "Latn", } m["tep"] = { "Tepecano", 3915525, "azc-pim", "Latn", } m["teq"] = { "Temein", 7698064, "sdv", } m["ter"] = { "Tereno", 3314742, "awd", "Latn", } m["tes"] = { "Tengger", 12473479, "poz", "Latn, Java", } m["tet"] = { "Tetum", 34125, "poz-tim", "Latn", } m["teu"] = { "Soo", 3437607, "ssa-klk", } m["tev"] = { "Teor", 12953198, "poz-cma", "Latn", } m["tew"] = { "Tewa", 56492, "nai-kta", "Latn", } m["tex"] = { "Tennet", 56346, "sdv", } m["tey"] = { "Tulishi", 12911106, "qfa-kad", "Latn", } m["tez"] = { "Tetserret", 7706841, "ber", "Latn", } m["tfi"] = { "Tofin Gbe", 3530330, "alv-pph", } m["tfn"] = { "Dena'ina", 27785, "ath-nor", "Latn", } m["tfo"] = { "Tefaro", 7694618, "paa-egb", "Latn", } m["tfr"] = { "Teribe", 36533, "cba", "Latn", } m["tft"] = { "Ternate", 3518492, "paa-tti", "Latn, Arab", } m["tga"] = { "Sagalla", 12953082, "bnt-cht", } m["tgb"] = { "Tobilung", 12953913, "poz-san", "Latn", } m["tgc"] = { "Tigak", 3528276, "poz-ocw", "Latn", } m["tgd"] = { "Ciwogai", 3438799, "cdc-wst", "Latn", } m["tge"] = { "Tamang Gorkha Timur", 12953175, "sit-tam", "sit-tam-Tibt, Deva", -- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] -- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake. } m["tgf"] = { "Chali", 3695197, "sit-ebo", "Tibt, Latn", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["tgh"] = { "Inggeris Kreol Tobago", 7811541, "crp", ancestors = "en", } m["tgi"] = { "Lawunuia", 3219937, "poz-ocw", } m["tgn"] = { "Tandaganon", 63311769, "phi", "Latn", } m["tgo"] = { "Sudest", 7675351, "poz-ocw", } m["tgp"] = { "Tangoa", 2410276, "poz-vnn", "Latn", } m["tgq"] = { "Tring", 7842360, "poz-swa", } m["tgr"] = { "Tareng", 25559541, "mkh", } m["tgs"] = { "Nume", 3346290, "poz-vnn", "Latn", } m["tgt"] = { "Tagbanwa Tengah", 3915515, "phi", "Tagb", } m["tgu"] = { "Tanggu", 7682930, "paa-ata", "Latn", } m["tgv"] = { "Tingui-Boto", 7808195, "sai-mje", "Latn", } m["tgw"] = { "Tagwana Senoufo", 36514, "alv-tdj", } m["tgx"] = { "Tagish", 28064, "ath-nor", "Latn", } m["tgy"] = { "Togoyo", 36825, "nic-ser", } m["thc"] = { "Tai Hang Tong", 7675753, "tai-nor", } m["thd"] = { "Kuuk Thaayorre", 6448718, "aus-pmn", "Latn", } m["the"] = { "Chitwania Tharu", 22083804, "inc-tha", "Deva", } m["thf"] = { "Thangmi", 7710314, "sit-new", "Deva", } m["thh"] = { "Tarahumara Utara", 15616395, "azc-trc", "Latn", } m["thi"] = { "Tai Long", 25559562, "tai-swe", } m["thk"] = { "Tharaka", 15407179, "bnt-kka", } m["thl"] = { "Dangaura Tharu", 22083815, "inc-tha", "Deva", } m["thm"] = { "Thavung", 34780, "mkh-vie", "Thai", --Laoo is feasible but no evidence yet. sort_key = "Thai-sortkey", } m["thn"] = { "Thachanadan", 7708880, "dra-mal", } m["thp"] = { "Thompson", 1755054, "sal", "Latn, Dupl", } m["thq"] = { "Kochila Tharu", 22083826, "inc-tha", } m["thr"] = { "Rana Tharu", 12953920, "inc-tha", "Deva", } m["ths"] = { "Thakali", 7709348, "sit-tam", } m["tht"] = { "Tahltan", 30125, "ath-nor", "Latn", } m["thu"] = { "Thuri", 7799291, "sdv-lon", } m["thy"] = { "Tha", 3915849, "alv-bwj", } m["tic"] = { "Tira", 36677, "alv-hei", } m["tif"] = { "Tifal", 11732691, "ngf-mok", "Latn", } m["tig"] = { "Tigre", 34129, "sem-eth", "Ethi", translit = "Ethi-translit", } m["tih"] = { "Murut Timugon", 7807680, "poz-san", "Latn", } m["tii"] = { "Tiene", 36469, "bnt-tek", } m["tij"] = { "Tilung", 7803037, "sit-kiw", } m["tik"] = { "Tikar", 36483, "nic-bdn", "Latn", } m["til"] = { "Tillamook", 2109432, "sal", "Latn", } m["tim"] = { "Timbe", 7804599, "ngf-kab", "Latn", } m["tin"] = { "Tindi", 36860, "cau-and", "Cyrl", display_text = s["cau-Cyrl-displaytext"], strip_diacritics = s["cau-Cyrl-stripdiacritics"], } m["tio"] = { "Teop", 3518239, "poz-ocw", "Latn", } m["tip"] = { "Trimuris", 7842270, "paa-kwe", "Latn", } m["tiq"] = { "Tiéfo", 3914874, "alv-sav", } m["tis"] = { "Masadiit Itneg", 18748769, "phi", } m["tit"] = { "Tinigua", 3029805, "sai-tin", "Latn", } m["tiu"] = { "Adasen", 11214797, "phi", "Latn", } m["tiv"] = { "Tiv", 34131, "nic-tvc", "Latn", } m["tiw"] = { "Tiwi", 1656014, "qfa-iso", "Latn", } m["tix"] = { "Tiwa Selatan", 7570552, "nai-kta", "Latn", } m["tiy"] = { "Tiruray", 7809425, "phi", "Latn", } m["tiz"] = { "Tai Hongjin", 3915716, "tai-swe", } m["tja"] = { "Tajuasohn", 3915326, "kro-wkr", } m["tjg"] = { "Tunjung", 3542117, "poz", "Latn", } m["tji"] = { "Tujia Utara", 12953229, "sit-tja", "Latn", } m["tjl"] = { "Tai Laing", 7675773, "tai-swe", "Mymr", } m["tjm"] = { "Timucua", 638300, "qfa-iso", "Latn", } m["tjn"] = { "Tonjon", 3913372, "dmn-jje", } m["tjs"] = { "Tujia Selatan", 12633994, "sit-tja", "Latn", } m["tju"] = { "Tjurruru", 3913834, "aus-nga", "Latn", } m["tjw"] = { "Chaap Wuurong", 5285187, "aus-pam", "Latn", } m["tka"] = { "Truká", 7847648, } m["tkb"] = { "Buksa", 20983638, "inc-eas", "Deva", } m["tkd"] = { "Tukudede", 36863, "poz-tim", "Latn", } m["tke"] = { "Takwane", 11030092, "bnt-mak", "Latn", ancestors = "vmw", } m["tkf"] = { "Tukumanféd", 42330115, "tup-gua", "Latn", } m["tkl"] = { "Tokelau", 34097, "poz-pnp", "Latn", } m["tkm"] = { "Takelma", 56710, } m["tkn"] = { "Toku-No-Shima", 3530484, "jpx-nry", "Jpan", translit = s["jpx-translit"], display_text = s["jpx-displaytext"], strip_diacritics = s["jpx-stripdiacritics"], sort_key = s["jpx-sortkey"], } m["tkp"] = { "Tikopia", 36682, "poz-pnp", "Latn", } m["tkq"] = { "Tee", 3075144, "nic-ogo", "Latn", } m["tkr"] = { "Tsakhur", 36853, "cau-wsm", "Cyrl, Latn, Arab", translit = "tkr-translit", override_translit = true, display_text = { Cyrl = s["cau-Cyrl-displaytext"] }, strip_diacritics = { Cyrl = s["cau-Cyrl-stripdiacritics"], Latn = s["cau-Latn-stripdiacritics"], }, } m["tks"] = { "Ramandi", 25261947, "xme-ttc", "Arab", ancestors = "xme-ttc-sou", } m["tkt"] = { "Kathoriya Tharu", 22083822, "inc-tha", } m["tku"] = { "Upper Necaxa Totonac", 56343, "nai-ttn", "Latn", } m["tkv"] = { "Mur Pano", 16939373, "poz-ocw", "Latn", } m["tkw"] = { "Teanu", 3516731, "poz-tem", "Latn", } m["tkx"] = { "Tangko", 7682993, "ngf-tna", "Latn", } m["tkz"] = { "Takua", 7678544, "mkh", } m["tla"] = { "Tepehuan Barat Daya", 3518245, "azc-pim", "Latn", } m["tlb"] = { "Tobelo", 1142333, "paa-gto", "Latn", } m["tlc"] = { "Misantla Totonac", 56460, "nai-ttn", "Latn", } m["tld"] = { "Talaud", 7678964, "phi", "Latn", } m["tlf"] = { "Telefol", 7696150, "ngf-mok", "Latn", } m["tlg"] = { "Tofanma", 4461493, "paa-nto", "Latn", } m["tlh"] = { "Klingon", 10134, "art", "Latn", type = "appendix-constructed", } m["tli"] = { "Tlingit", 27792, "xnd", "Latn, Cyrl", } m["tlj"] = { "Talinga-Bwisi", 7679530, "bnt-haj", } m["tlk"] = { "Taloki", 3514563, "poz-btk", } m["tll"] = { "Tetela", 2613465, "bnt-tet", } m["tlm"] = { "Tolomako", 3130514, "poz-vnn", "Latn", } m["tln"] = { "Talondo'", 7680293, "poz-ssw", } m["tlo"] = { "Talodi", 36525, "alv-tal", } m["tlp"] = { "Filomena Mata-Coahuitlán Totonac", 5449202, "nai-ttn", "Latn", } m["tlq"] = { "Tai Loi", 7675784, "mkh-pal", } m["tlr"] = { "Talise", 3514510, "poz-sls", "Latn", } m["tls"] = { "Tambotalo", 7681065, "poz-vnn", "Latn", } m["tlt"] = { "Teluti", 12953194, "poz-cma", } m["tlu"] = { "Tulehu", 7852006, "poz-cma", } m["tlv"] = { "Taliabu", 3514498, "poz-cma", "Latn", } m["tlx"] = { "Khehek", 3196124, "poz-aay", } m["tly"] = { "Talysh", 34318, "xme-ttc", "Latn, Cyrl, Arab", } m["tma"] = { "Tama (Chad)", 57001, "sdv-tmn", } m["tmb"] = { "Avava", 2157461, "poz-vnc", "Latn", } m["tmc"] = { "Tumak", 3121045, "cdc-est", } m["tmd"] = { "Haruai", 12632146, "paa-pia", "Latn", } m["tme"] = { "Tremembé", 5246937, } m["tmf"] = { "Toba-Maskoy", 3033544, "sai-mas", "Latn", } m["tmg"] = { "Ternateño", 7232597, } m["tmh"] = { "Tuareg", 34065, "ber", "Latn, Tfng, Arab", strip_diacritics = { Latn = {remove_diacritics = c.grave .. c.acute .. c.circ}, }, } m["tmi"] = { "Tutuba", 7857052, "poz-vnn", "Latn", } m["tmj"] = { "Samarokena", 7408865, "paa-saa", "Latn", } m["tml"] = { "Tamnim Citak", 12643315, "ngf-asm", "Latn", } m["tmm"] = { "Tai Thanh", 7675842, "tai-swe", } m["tmn"] = { "Taman (Indonesia)", 7680671, "poz", "Latn", } m["tmo"] = { "Temoq", 7698205, "mkh-asl", } m["tmq"] = { "Tumleo", 7852641, "poz-ocw", } m["tms"] = { "Tima", 36684, "nic-ktl", } m["tmt"] = { "Tasmate", 7687571, "poz-vnn", "Latn", } m["tmu"] = { "Iau", 56867, "paa-lpl", "Latn", } m["tmv"] = { "Motembo", 11013108, "bnt-bun", } m["tmw"] = { "Temuan", 3025610, "poz-mly", "Latn", } m["tmy"] = { "Tami", 3514812, "poz-oce", } m["tmz"] = { "Tamanaku", 3441435, "sai-ven", "Latn", } m["tna"] = { "Tacana", 3182551, "sai-tac", "Latn", } m["tnb"] = { "Tunebo Barat", 3181238, "cba", } m["tnc"] = { "Tanimuca-Retuarã", 36535, "sai-tuc", "Latn", } m["tnd"] = { "Angosturas Tunebo", 25559604, "cba", } m["tne"] = { "Tinoc Kallahan", 3192219, } m["tng"] = { "Tobanga", 3440501, "cdc-est", } m["tnh"] = { "Maiani", 6735243, "ngf-kau", "Latn", } m["tni"] = { "Tandia", 7682454, "poz-hce", "Latn", } m["tnk"] = { "Kwamera", 3200806, "poz-vns", "Latn", } m["tnl"] = { "Lenakel", 3229429, "poz-vns", "Latn", } m["tnm"] = { "Tabla", 7673105, "paa-sen", "Latn", } m["tnn"] = { "Tanna Utara", 957945, "poz-vns", "Latn", } m["tno"] = { "Toromono", 510544, "sai-tac", "Latn", } m["tnp"] = { "Whitesands", 3063761, "poz-vns", "Latn", } m["tnq"] = { "Taíno", 5232952, "awd-taa", "Latn", } m["tnr"] = { "Bedik", 35096, "alv-ten", "Latn", } m["tns"] = { "Tenis", 7699870, "poz-stm", "Latn", } m["tnt"] = { "Tontemboan", 3531666, "phi", "Latn", } m["tnu"] = { "Tay Khang", 6362363, "tai", } m["tnv"] = { "Tanchangya", 7682361, "inc-bas", "Cakm", ancestors = "inc-obn", } m["tnw"] = { "Tonsawang", 3531660, "phi", "Latn", } m["tnx"] = { "Tanema", 2106984, "poz-tem", "Latn", } m["tny"] = { "Tongwe", 7821200, "bnt", } m["tnz"] = { "Ten'edn", 3073453, "mkh-asl", "Latn", } m["tob"] = { "Toba", 3113756, "sai-guc", "Latn", } m["toc"] = { "Coyutla Totonac", 15615591, "nai-ttn", "Latn", } m["tod"] = { "Toma", 11055484, "dmn-msw", "Latn, Loma" } m["tof"] = { "Gizrra", 5565941, "paa-etf", "Latn", } m["tog"] = { "Tonga (Malawi)", 3847648, "bnt-nys", "Latn", } m["toh"] = { "Tonga (Mozambique)", 7820988, "bnt-bso", } m["toi"] = { "Tonga (Zambia)", 34101, "bnt-bot", "Latn", } m["toj"] = { "Tojolabal", 36762, "myn", "Latn", } m["tok"] = { "Toki Pona", 36846, "art", "Latn", type = "appendix-constructed", } m["tol"] = { "Tolowa", 20827, "ath-pco", "Latn", } m["tom"] = { "Tombulu", 3531199, "phi", "Latn", } m["too"] = { "Xicotepec de Juárez Totonac", 8044353, "nai-ttn", "Latn", } m["top"] = { "Papantla Totonac", 56329, "nai-ttn", "Latn", } m["toq"] = { "Toposa", 3033588, "sdv-ttu", } m["tor"] = { "Togbo-Vara Banda", 11002922, "bad-cnt", } m["tos"] = { "Highland Totonac", 13154149, "nai-ttn", "Latn", } m["tou"] = { "Tho", 22694631, "mkh-vie", "Latn", } m["tov"] = { "Upper Taromi", 12953183, "xme-ttc", ancestors = "xme-ttc-cen", } m["tow"] = { "Jemez", 3912876, "nai-kta", "Latn", } m["tox"] = { "Tobian", 34022, "poz-mic", } m["toy"] = { "Topoiyo", 7824977, "poz-kal", } m["toz"] = { "To", 7811216, "alv-mbm", } m["tpa"] = { "Taupota", 7688832, "poz-ocw", } m["tpc"] = { "Azoyú Me'phaa", 25559730, "omq", "Latn", } m["tpe"] = { "Tippera", 16115423, "tbq-bdg", } m["tpf"] = { "Tarpia", 12953185, "poz-ocw", "Latn", } m["tpg"] = { "Kula", 6442714, "paa-alp", "Latn", } m["tpi"] = { "Tok Pisin", 34159, "crp", "Latn", ancestors = "en", } m["tpj"] = { "Tapieté", 3121063, "gn", "Latn", } m["tpk"] = { "Tupinikin", 33924, "tup-gua", } m["tpl"] = { "Tlacoapa Me'phaa", 16115511, "omq", } m["tpm"] = { "Tampulma", 36590, "nic-gnw", } m["tpn"] = { "Tupinambá", 31528147, "tup-gua", "Latn", } m["tpo"] = { "Tai Pao", 7675795, "tai-nor", } m["tpp"] = { "Pisaflores Tepehua", 56349, "nai-ttn", } m["tpq"] = { "Tukpa", 12953230, "sit-las", } m["tpr"] = { "Tuparí", 3542217, "tup", "Latn", } m["tpt"] = { "Tlachichilco Tepehua", 56330, "nai-ttn", } m["tpu"] = { "Tampuan", 3514882, "mkh-ban", "Khmr", } m["tpv"] = { "Tanapag", 3397371, "poz-mic", } m["tpw"] = { "Tupi Kuno", 56944, "tup-gua", "Latn", } m["tpx"] = { "Acatepec Me'phaa", 31157882, "omq", "Latn", } m["tpy"] = { "Trumai", 12294279, "qfa-iso", } m["tpz"] = { "Tinputz", 3529205, "poz-ocw", } m["tqb"] = { "Tembé", 10322157, "tup-gua", "Latn", } m["tql"] = { "Lehali", 3229119, "poz-vnn", "Latn", } m["tqm"] = { "Turumsa", 7856508, "paa-dtu", "Latn", } m["tqn"] = { "Tenino", 15699255, "nai-shp", "Latn", ancestors = "nai-spt", } m["tqo"] = { "Toaripi", 7811403, "paa-eel", "Latn", } m["tqp"] = { "Tomoip", 3531388, "poz-ocw", } m["tqq"] = { "Tunni", 3514343, "cus-som", } m["tqr"] = { "Torona", 36679, "alv-tal", } m["tqt"] = { "Totonac Barat", 7116691, "nai-ttn", "Latn", } m["tqu"] = { "Touo", 56750, } m["tqw"] = { "Tonkawa", 2454881, "qfa-iso", "Latn", } m["tra"] = { "Tirahi", 3812406, "inc-koh", } m["trb"] = { "Terebu", 7701797, "poz-ocw", } m["trc"] = { "Copala Triqui", 12953935, "omq-tri", "Latn", } m["trd"] = { "Turi", 7854914, "mun", } m["tre"] = { "Tarangan Timur", 18609750, "poz", } m["trf"] = { "Inggeris Kreol Trinidad", 7842493, "crp", "Latn", ancestors = "en", } m["trg"] = { "Lishán Didán", 56473, "sem-nna", "Hebr", } m["trh"] = { "Turaka", 12953237, "ngf-dag", "Latn", } m["tri"] = { "Trió", 56885, "sai-tar", "Latn", } m["trj"] = { "Toram", 3441225, "cdc-est", } m["trl"] = { "Traveller Scottish", 3915671, "qfa-mix", "Latn", ancestors = "rom, sco", } m["trm"] = { "Tregami", 34081, "nur-sou", } m["trn"] = { "Trinitario", 3539279, "awd", } m["tro"] = { "Tarao", 3515603, "tbq-kuk", "Latn", } m["trp"] = { "Kokborok", 35947, "tbq-bdg", "Beng, Latn" -- WP lists 2 more } m["trq"] = { "San Martín Itunyoso Triqui", 12953934, "omq-tri", "Latn", } m["trr"] = { "Taushiro", 1957508, nil, "Latn", } m["trs"] = { "Chicahuaxtla Triqui", 3539587, "omq-tri", "Latn", } m["trt"] = { "Tunggare", 615071, "paa-egb", "Latn", } m["tru"] = { "Turoyo", 34040, "sem-cna", "Syrc, Latn", translit = { Syrc = "tru-translit", }, strip_diacritics = { Syrc = "Syrc-stripdiacritics", }, } m["trv"] = { "Taroko", 716686, "map-ata", "Latn", } m["trw"] = { "Torwali", 2665246, "inc-koh", "Aran", } m["trx"] = { "Tringgus", 7842365, "day", } m["try"] = { "Turung", 7856514, "tai-swe", "as-Beng", } m["trz"] = { "Torá", 7827518, "sai-cpc", } m["tsa"] = { "Tsaangi", 36675, "bnt-nze", } m["tsb"] = { "Tsamai", 2371358, "cus-eas", } m["tsc"] = { "Tswa", 2085051, "bnt-tsr", } m["tsd"] = { "Tsakonian", 220607, "grk", "Grek", ancestors = "grc-dor", translit = "el-translit", -- Grek display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["tse"] = { "Bahasa Isyarat Tunisia", 7853191, "sgn", } m["tsg"] = { "Suluk", 34142, "phi", "Latn, Arab", } m["tsh"] = { "Tsuvan", 3502326, "cdc-cbm", "Latn", } m["tsi"] = { "Tsimshian", 20085721, "nai-tsi", "Latn", } m["tsj"] = { "Tshangla", 36840, "sit-tsk", "Tibt, Latn, Deva", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["tsl"] = { "Ts'ün-Lao", 3446675, "tai", } m["tsm"] = { "Bahasa Isyarat Turki", 36885, "sgn", } m["tsp"] = { "Toussiana Utara", 11155635, "alv-sav", } m["tsq"] = { "Bahasa Isyarat Thai", 7709156, "sgn-asl", "Sgnw", } m["tsr"] = { "Akei", 2828964, "poz-vnn", "Latn", } m["tss"] = { "Bahasa Isyarat Taiwan", 34019, "sgn-jsl", } m["tsu"] = { "Tsou", 716681, "map", "Latn", } m["tsv"] = { "Tsogo", 36674, "bnt-tso", } m["tsw"] = { "Tsishingini", 13123571, "nic-kam", } m["tsx"] = { "Mubami", 6930815, "paa-wig", "Latn", } m["tsy"] = { "Bahasa Isyarat Tebul", 7692090, "sgn", } m["tta"] = { "Tutelo", 2311602, "sio-ohv", "Latn", } m["ttb"] = { "Gaa", 3438361, "nic-dak", "Latn", } m["ttc"] = { "Tektiteko", 36686, "myn", "Latn", } m["ttd"] = { "Tauade", 7688634, "qfa-dis", -- Papuan; isolate per Glottolog; Glottolog says "A Goilalan family uniting Kunimaipan, Tauade and Fuyug -- is often posited based on the lexicostatistical figures reported in Tom E. Dutton 1975: 631-632" but -- goes on to say the data "is clearly insufficient, as the lexical links so far proposed are few and -- show irregular one-consonant correspondences". "Latn", } m["tte"] = { "Bwanabwana", 5003667, "poz-ocw", "Latn", } m["ttf"] = { "Tuotomb", 7853459, "nic-mbw", "Latn", } m["ttg"] = { "Tutong", 3507990, "poz-swa", "Latn", } m["tth"] = { "Upper Ta'oih", 3512660, "mkh-kat", } m["tti"] = { "Tobati", 7811556, "poz-ocw", "Latn", } m["ttj"] = { "Tooro", 7824218, "bnt-nyg", "Latn", } m["ttk"] = { "Totoro", 3532756, "sai-bar", "Latn", } m["ttl"] = { "Totela", 10962316, "bnt-bot", "Latn", } m["ttm"] = { "Tutchone Utara", 20822, "ath-nor", "Latn", } m["ttn"] = { "Towei", 7829606, "paa-wpw", "Latn", } m["tto"] = { "Lower Ta'oih", 25559539, "mkh-kat", } m["ttp"] = { "Tombelala", 6799663, "poz-kal", } m["ttr"] = { "Tera", 56267, "cdc-cbm", } m["tts"] = { "Isan", 33417, "tai-swe", "Thai", -- also Tai Noi/Lao Buhan script sort_key = "Thai-sortkey", } m["ttt"] = { "Tat", 56489, "ira-swi", "Cyrl, Latn, Armn, Arab", -- Armn translit in [[Module:scripts/data]] (NOTE: formerly not present, probably an accidental omission) ancestors = "fa", } m["ttu"] = { "Torau", 3532208, "poz-ocw", } m["ttv"] = { "Titan", 3445811, "poz-aay", "Latn", } m["ttw"] = { "Long Wat", 7856961, "poz-swa", } m["tty"] = { "Sikaritai", 7513600, "paa-clp", "Latn", } m["ttz"] = { "Tsum", 12953223, "sit-kyk", } m["tua"] = { "Wiarumus", 7998045, "paa-mmu", "Latn", } m["tub"] = { "Tübatulabal", 56704, "azc", "Latn", } m["tuc"] = { "Mutu", 3331003, "poz-ocw", "Latn", } m["tud"] = { "Tuxá", 7857217, } m["tue"] = { "Tuyuca", 2520538, "sai-tuc", "Latn", } m["tuf"] = { "Tunebo Tengah", 12953942, "cba", "Latn", } m["tug"] = { "Tunia", 863721, "alv-bua", } m["tuh"] = { "Taulil", 3516141, } m["tui"] = { "Tupuri", 36646, "alv-mbm", "Latn", } m["tuj"] = { "Tugutil", 12953228, "paa-gto", "Latn", } m["tul"] = { "Tula", 3914907, "alv-wjk", } m["tum"] = { "Tumbuka", 34138, "bnt-nys", "Latn", } m["tun"] = { "Tunica", 56619, "qfa-iso", "Latn", } m["tuo"] = { "Tucano", 3541834, "sai-tuc", "Latn", } m["tuq"] = { "Tedaga", 36639, "ssa-sah", "Latn", } m["tus"] = { "Tuscarora", 36944, "iro-nor", "Latn", } m["tuu"] = { "Tututni", 20627, "ath-pco", "Latn", } m["tuv"] = { "Turkana", 36958, "sdv-ttu", "Latn", } m["tux"] = { "Tuxináwa", 7857204, "sai-pan", "Latn", } m["tuy"] = { "Tugen", 3541935, "sdv-nma", } m["tuz"] = { "Turka", 36643, "nic-gur", "Latn", } m["tva"] = { "Vaghua", 3553248, "poz-ocw", "Latn", } m["tvd"] = { "Tsuvadi", 3914936, "nic-kam", } m["tve"] = { "Te'un", 7690709, "poz-cet", "Latn", } m["tvk"] = { "Ambrym Tenggara", 252411, "poz-vnc", "Latn", } m["tvl"] = { "Tuvalu", 34055, "poz-pnp", "Latn", } m["tvm"] = { "Tela-Masbuar", 7695666, "poz-tim", } m["tvn"] = { "Tavoyan", 7689158, "tbq-brm", "Mymr", ancestors = "obr", } m["tvo"] = { "Tidore", 3528199, "paa-tti", "Latn, Arab", } m["tvs"] = { "Taveta", 15632387, "bnt-par", "Latn", } m["tvt"] = { "Tutsa Naga", 7856987, "sit-tno", } m["tvu"] = { "Tunen", 36632, "nic-mbw", } m["tvw"] = { "Sedoa", 7445362, "poz-kal", } m["tvx"] = { "Taivoan", 1975271, "map", "Latn", } m["tvy"] = { "Timor Pidgin", 4904029, "crp", ancestors = "pt", } m["twa"] = { "Twana", 7857412, "sal", "Latn", } m["twb"] = { "Tawbuid Barat", 12953912, "phi", } m["twc"] = { "Teshenawa", 3436597, "cdc-wst", "Latn", } m["twe"] = { "Teiwa", 3519302, "paa-alp", "Latn", } m["twf"] = { "Taos", 7684320, "nai-kta", "Latn", } m["twg"] = { "Tereweng", 12953200, "paa-alp", "Latn", } m["twh"] = { "Tai Dón", 7675751, "tai-swe", "Tavt", --translit = "Tavt-translit", sort_key = { from = {"[꪿ꫀ꫁ꫂ]", "([ꪵꪶꪹꪻꪼ])([ꪀ-ꪯ])"}, to = {"", "%2%1"} }, } m["twm"] = { "Tawang Monpa", 36586, "sit-ebo", "Tibt", override_translit = true, -- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["twn"] = { "Twendi", 7857682, "nic-mmb", } m["two"] = { "Tswapong", 3446241, "bnt-sts", } m["twp"] = { "Ere", 3056045, "poz-aay", "Latn", } m["twq"] = { "Tasawaq", 36564, "son", } m["twr"] = { "Tarahumara Barat Daya", 12953909, "azc-trc", "Latn", } m["twt"] = { "Turiwára", 3542307, "tup-gua", "Latn", } m["twu"] = { "Termanu", 7702572, "poz-tim", "Latn", } m["tww"] = { "Tuwari", 7857159, "paa-wal", "Latn", } m["twy"] = { "Tawoyan", 3513542, "poz-bre", "Latn", } m["txa"] = { "Tombonuo", 7818692, "poz-san", "Latn", } m["txb"] = { "Tocharia B", 3199353, "ine-toc", "Latn", standard_chars = "AaÄäĀāCcEeIiKkLlMmṂṃNnṄṅÑñOoPpRrSsŚśṢṣTtUuWwYy" .. c.punc, } m["txc"] = { "Tsetsaut", 20829, "ath-nor", "Latn", } m["txe"] = { "Totoli", 7828387, "poz-tot", "Latn", } m["txg"] = { "Tangut", 2727930, "ero", "Tang", -- Tang translit in [[Module:scripts/data]] } m["txh"] = { "Thracian", 36793, "ine", "Latn, Polyt", -- Polyt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]] } m["txi"] = { "Ikpeng", 9344891, "sai-pek", "Latn", } m["txj"] = { "Tarjumo", 24906088, "ssa-sah", "Latn, Arab", } m["txm"] = { "Tomini", 7818911, "poz", "Latn", } m["txn"] = { "Tarangan Barat", 3515594, "poz", "Latn", } m["txo"] = { "Toto", 36709, "sit-dhi", "Beng, Toto" } m["txq"] = { "Tii", 7801784, "poz-tim", } m["txr"] = { "Tartessian", 36795, "qfa-unc", -- extinct, no consensus on classification } m["txs"] = { "Tonsea", 3531659, "phi", "Latn", } m["txt"] = { "Citak", 3447279, "ngf-asm", "Latn", } m["txu"] = { "Kayapó", 3101212, "sai-nje", "Latn", } m["txx"] = { "Tatana", 18643518, "poz-san", "Latn", } m["tya"] = { "Tauya", 7688978, "ngf-rai", "Latn", } m["tye"] = { "Kyenga", 3913304, "dmn-bbu", "Latn", } m["tyh"] = { "O'du", 3347428, "mkh", } m["tyi"] = { "Teke-Tsaayi", 33123613, "bnt-nze", } m["tyj"] = { "Tai Do", 7675746, "tai-nor", -- Chamberlain (1991), but Pittayaporn (2009) suggests tai-swe "Latn, Tayo", -- Vietnam } m["tyl"] = { "Thu Lao", 12953921, "tai-cen", } m["tyn"] = { "Kombai", 6428241, "ngf-nde", "Latn", } m["typ"] = { "Kuku-Thaypan", 3915693, "aus-pmn", "Latn", } m["tyr"] = { "Tai Daeng", 3915207, "tai-swe", "Tavt", } m["tys"] = { "Sapa", 3446668, "tai-sap", "Latn", } m["tyt"] = { "Tày Tac", 7862029, "tai-swe", } m["tyu"] = { "Kua", 3832933, "khi-kal", "Latn", } m["tyv"] = { "Tuva", 34119, "trk-ssb", "Cyrl", translit = "tyv-translit", override_translit = true, sort_key = "tyv-sortkey", } m["tyx"] = { "Teke-Tyee", 36634, "bnt-nze", "Latn", } m["tyz"] = { "Tày", -- This does not mean its family "Tai" languages. 2511476, "tai-tay", "Latn, Hani", sort_key = { Hani = "Hani-sortkey" }, } m["tza"] = { "Bahasa Isyarat Tanzania", 7684177, "sgn", } m["tzh"] = { "Tzeltal", 36808, "myn", "Latn", } m["tzj"] = { "Tz'utujil", 36941, "myn", "Latn", } m["tzl"] = { "Talossan", 1063911, "art", "Latn", type = "appendix-constructed", sort_key = "tzl-sortkey", } m["tzm"] = { "Tamazight Atlas Tengah", 49741, "ber", "Tfng, Arab, Latn", translit = { Tfng = "Tfng-translit", }, } m["tzn"] = { "Tugun", 12953225, "poz-tim", "Latn", } m["tzo"] = { "Tzotzil", 36809, "myn", "Latn", } m["tzx"] = { "Tabriak", 56872, "paa-lse", "Latn", } return require("Module:languages").finalizeData(m, "language") seo5fecw6ixaf75o4763th5in98arnd Modul:translations 828 9941 375372 344653 2026-09-22T05:12:48Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708794|92708794]]) 375372 Scribunto text/plain local export = {} local anchors_module = "Module:anchors" local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local pages_module = "Module:pages" local parameters_module = "Module:parameters" local string_utilities_module = "Module:string utilities" local templatestyles_module = "Module:TemplateStyles" local utilities_module = "Module:utilities" local wikimedia_languages_module = "Module:wikimedia languages" local concat = table.concat local html_create = mw.html.create local insert = table.insert local load_data = mw.loadData local new_title = mw.title.new local require = require --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==] local function decode_uri(...) decode_uri = require(string_utilities_module).decode_uri return decode_uri(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function get_wikimedia_lang(...) get_wikimedia_lang = require(wikimedia_languages_module).getByCode return get_wikimedia_lang(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function normalize_anchor(...) normalize_anchor = require(anchors_module).normalize_anchor return normalize_anchor(...) end local function plain_link(...) plain_link = require(links_module).plain_link return plain_link(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function remove_links(...) remove_links = require(links_module).remove_links return remove_links(...) end local function split_on_slashes(...) split_on_slashes = require(links_module).split_on_slashes return split_on_slashes(...) end local function templatestyles(...) templatestyles = require(templatestyles_module) return templatestyles(...) end local function track(...) track = require(debug_track_module) return track(...) end --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local en local function get_en() en, get_en = require(languages_module).getByCode("ms"), nil return en end local headword_data local function get_headword_data() headword_data, get_headword_data = load_data("Module:headword/data"), nil return headword_data end local parameters_data local function get_parameters_data() parameters_data, get_parameters_data = load_data("Module:parameters/data"), nil return parameters_data end local translations_data local function get_translations_data() translations_data, get_translations_data = load_data("Module:translations/data"), nil return translations_data end local function is_translation_subpage(pagename) if (headword_data or get_headword_data()).page.namespace ~= "" then return false elseif not pagename then pagename = (headword_data or get_headword_data()).encoded_pagename end return pagename:match("./translations$") and true or false end local function canonical_pagename() local pagename = (headword_data or get_headword_data()).encoded_pagename return is_translation_subpage(pagename) and pagename:sub(1, -14) or pagename end local function interwiki(terminfo, term, lang, langcode) -- No interwiki link if term is empty/missing if not term or #term < 1 then terminfo.interwiki = false return end -- Percent-decode the term. term = decode_uri(terminfo.term, "PATH") -- Don't show an interwiki link if it's an invalid title. if not new_title(term) then terminfo.interwiki = false return end local interwiki_langcode = (translations_data or get_translations_data()).interwiki_langs[langcode] local wmlangs = interwiki_langcode and {get_wikimedia_lang(interwiki_langcode)} or lang:getWikimediaLanguages() -- Don't show the interwiki link if the language is not recognised by Wikimedia. if #wmlangs == 0 then terminfo.interwiki = false return end local sc = terminfo.sc local target_page = get_link_page(term, lang, sc) local split = split_on_slashes(target_page) if not split[1] then terminfo.interwiki = false return end target_page = split[1] local wmlangcode = wmlangs[1]:getCode() local interwiki_link = language_link{ lang = lang, sc = sc, term = wmlangcode .. ":" .. target_page, alt = "(" .. wmlangcode .. ")", tr = "-" } terminfo.interwiki = tostring(html_create("span") :addClass("tpos") :wikitext("&nbsp;" .. interwiki_link) ) end function export.show_terminfo(terminfo, check) local lang = terminfo.lang local langcode, langname = lang:getCode(), lang:getCanonicalName() -- Translations must be for mainspace languages. if not lang:hasType("regular") then error("Translations must be for attested and approved main-namespace languages.") else local disallowed = (translations_data or get_translations_data()).disallowed local err_msg = disallowed[langcode] if err_msg then error("Translations not allowed in " .. langname .. " (" .. langcode .. "). " .. langname .. " translations should " .. err_msg) end local fullcode = lang:getFullCode() if fullcode ~= langcode then err_msg = disallowed[fullcode] if err_msg then langname = lang:getFullName() error("Translations not allowed in " .. langname .. " (" .. fullcode .. "). " .. langname .. " translations should " .. err_msg) end end end if langcode == "ms" then if terminfo.interwiki then error("Interwiki translations not allowed for English; they should always link to a different Wiktionary") end local current_L2 = require(pages_module).get_current_L2() if current_L2 ~= "Rentas bahasa" and mw.title.getCurrentTitle().nsText ~= "Wikikamus" then if current_L2 then error("English translations only allowed in Translingual section, not in " .. current_L2) else error("English translations only allowed in Translingual section, not outside of any L2") end end end local term = terminfo.term -- Check if there is a term. Don't show the interwiki link if there is nothing to link to. if not term then -- Track entries that don't provide a term. -- FIXME: This should be a category. track("translations/no term") track("translations/no term/" .. langcode) end if terminfo.interwiki then interwiki(terminfo, term, lang, langcode) end langcode = lang:getFullCode() if (translations_data or get_translations_data()).need_super[langcode] then local tr = terminfo.tr if tr ~= nil then terminfo.tr = tr:gsub("%d[%d%*%-]*%f[^%d%*]", "<sup>%0</sup>") end end terminfo.show_decorations = true local link = full_link(terminfo, "translation") local canonical_name = lang:getCanonicalName() local full_name = lang:getFullName() local categories = {"Perkataan dengan terjemahan bahasa " .. canonical_name} if canonical_name ~= full_name then insert(categories, "Perkataan dengan terjemahan bahasa " .. full_name) end if check then link = tostring(html_create("span") :addClass("ttbc") :tag("sup") :addClass("ttbc") :wikitext("(sila [[WT:Terjemahan#Terjemahan untuk disemak|sahkan]])") :done() :wikitext(" " .. link) ) insert(categories, "Permintaan pengesahan terjemahan bahasa " .. langname ) end return link .. format_categories(categories, en or get_en(), nil, canonical_pagename()) end -- Implements {{t}}, {{t+}}, {{t-check}} and {{t+check}}. function export.show(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation"]) local check = frame.args.check return export.show_terminfo({ lang = args[1], sc = args.sc, track_sc = true, term = args[2], alt = args.alt, id = args.id, genders = args[3], tr = args.tr, ts = args.ts, lit = args.lit, q = args.q, qq = args.qq, l = args.l, ll = args.ll, refs = args.ref, interwiki = frame.args.interwiki, }, check and check ~= "") end local function add_id(div, id) return id and div:attr("id", normalize_anchor("Translations-" .. id)) or div end -- Implements {{ter-atas}} and part of {{ter-atas-juga}}. local function top(args, title, id, navhead) local column_width = (args["column-width"] == "wide" or args["column-width"] == "narrow") and "-" .. args["column-width"] or "" local div = html_create("div") :addClass("NavFrame") :node(navhead) :tag("div") :addClass("NavContent") :tag("table") :addClass("translations") :attr("role", "presentation") :attr("data-gloss", title or "") :tag("tr") :tag("td") :addClass("translations-cell") :addClass("multicolumn-list" .. column_width) :attr("colspan", "3") :allDone() div = add_id(div, id) local categories = {} if not title then insert(categories, "Pengepala jadual terjemahan kekurangan padanan kata") end local pagename = canonical_pagename() if is_translation_subpage() then insert(categories, "Sublaman terjemahan") end return (tostring(div):gsub("</td></tr></table></div></div>$", "")) .. (#categories > 0 and format_categories(categories, en or get_en(), nil, pagename) or "") .. -- Category to trigger [[MediaWiki:Gadget-TranslationAdder.js]]; we want this even on -- user pages and such. format_categories("Perkataan dengan kotak terjemahan", nil, nil, nil, true) .. templatestyles("Modul:translations/styles.css") end -- Entry point for {{ter-atas}}. function export.top(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas"]) local title = args[1] local id = args.id or title title = title and remove_links(title) return top(args, title, id, html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(title or "Terjemahan") ) end -- Entry point for {{checktrans-top}}. function export.check_top(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["checktrans-top"]) local text = "\n:''The translations below need to be checked and inserted above into the appropriate translation tables. See instructions at " .. frame:expandTemplate{ title = "section link", args = {"Wikikamus:Susun atur entri#Terjemahan"} } .. ".''\n" local header = html_create("div") :addClass("checktrans") :wikitext(text) local subtitle = args[1] local title = "Terjemahan untuk disemak" if subtitle then title = title .. "&zwnj;: \"" .. subtitle .. "\"" end -- No ID, since these should always accompany proper translation tables, and can't be trusted anyway (i.e. there's no use-case for links). return tostring(header) .. "\n" .. top(args, title, nil, html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(title or "Terjemahan") ) end -- Implements {{ter-bawah}}. function export.bottom(frame) -- Check nothing is being passed as a parameter. process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-bawah"]) return "</table></div></div>" end -- Implements {{ter-lihat}} and part of {{ter-atas-juga}}. local function see(args, see_text) local navhead = html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(args[1] .. " ") :tag("span") :css("font-weight", "normal") :wikitext("— ") :tag("i") :wikitext(see_text) :allDone() local terms, id = args[2], args.id if #terms == 0 then terms[1] = args[1] end for i = 1, #terms do local term_id = id[i] or id.default local data = { term = terms[i], id = term_id and "Translations-" .. term_id or "Translations", } terms[i] = plain_link(data) end return navhead:wikitext(concat(terms, ",&lrm; ")) end -- Entry point for {{ter-lihat}}. function export.see(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-lihat"]) local div = html_create("div") :addClass("pseudo") :addClass("NavFrame") :node(see(args, "see ")) return tostring(add_id(div, args.id.default or args[1])) end -- Entry point for {{ter-atas-juga}}. function export.top_also(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas-juga"]) local navhead = see(args, "lihat juga ") local title = args[1] local id = args.id.default or title title = remove_links(title) return top(args, title, id, navhead) end -- Implements {{translation subpage}}. function export.subpage(frame) process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation subpage"]) if not is_translation_subpage() then error("This template should only be used on translation subpages, which have titles that end with '/translations'.") end -- "Translation subpages" category is handled by {{trans-top}}. return ("''This page contains translations for ''%s''. See the main entry for more information.''"):format(full_link{ lang = en or get_en(), term = canonical_pagename(), }) end -- Implements {{t-needed}}. function export.needed(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["t-needed"]) local lang, category = args[1], "" local span = html_create("span") :addClass("trreq") :attr("data-lang", lang:getCode()) :tag("i") :wikitext("sila tambah terjemahan ini jika boleh") :done() if not args.nocat then local type, sort = args[2], args.sort if type == "quote" then category = "Permintaan terjemahan petikan bahasa " .. lang:getCanonicalName() elseif type == "usex" then category = "Permintaan terjemahan contoh penggunaan bahasa " .. lang:getCanonicalName() else category = "Permintaan terjemahan ke dalam bahasa " .. lang:getCanonicalName() lang = en or get_en() end category = format_categories(category, lang, sort, not sort and canonical_pagename() or nil) end return tostring(span) .. category end -- Implements {{no equivalent translation}}. function export.no_equivalent(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no equivalent translation"]) local text = "tiada padanan kata dalam bahasa " .. args[1]:getCanonicalName() if not args.noend then text = text .. ", tetapi lihat" end return tostring(html_create("i"):wikitext(text)) end -- Implements {{no attested translation}}. function export.no_attested(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no attested translation"]) local langname = args[1]:getCanonicalName() local text = "tiada perkataan [[WT:ATTEST|disahkan]] dalam bahasa " .. langname local category = "" if not args.noend then text = text .. ", but see" local sort = args.sort category = format_categories("Terjemahan bahasa " .. langname .. " yang tidak disahkan", en or get_en(), sort, not sort and canonical_pagename() or nil) end return tostring(html_create("i"):wikitext(text)) .. category end -- Implements {{not used}}. function export.not_used(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["not used"]) return tostring(html_create("i"):wikitext((args[2] or "tidak digunakan") .. " dalam " .. args[1]:getCanonicalName())) end return export hj4av85y0jqga4vvfg9ql3zt8i96adp 375388 375372 2026-09-22T07:30:23Z Hakimi97 2668 Betulkan nama English ke Malay 375388 Scribunto text/plain local export = {} local anchors_module = "Module:anchors" local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local pages_module = "Module:pages" local parameters_module = "Module:parameters" local string_utilities_module = "Module:string utilities" local templatestyles_module = "Module:TemplateStyles" local utilities_module = "Module:utilities" local wikimedia_languages_module = "Module:wikimedia languages" local concat = table.concat local html_create = mw.html.create local insert = table.insert local load_data = mw.loadData local new_title = mw.title.new local require = require --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==] local function decode_uri(...) decode_uri = require(string_utilities_module).decode_uri return decode_uri(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function get_wikimedia_lang(...) get_wikimedia_lang = require(wikimedia_languages_module).getByCode return get_wikimedia_lang(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function normalize_anchor(...) normalize_anchor = require(anchors_module).normalize_anchor return normalize_anchor(...) end local function plain_link(...) plain_link = require(links_module).plain_link return plain_link(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function remove_links(...) remove_links = require(links_module).remove_links return remove_links(...) end local function split_on_slashes(...) split_on_slashes = require(links_module).split_on_slashes return split_on_slashes(...) end local function templatestyles(...) templatestyles = require(templatestyles_module) return templatestyles(...) end local function track(...) track = require(debug_track_module) return track(...) end --[==[ Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==] local en local function get_en() en, get_en = require(languages_module).getByCode("ms"), nil return en end local headword_data local function get_headword_data() headword_data, get_headword_data = load_data("Module:headword/data"), nil return headword_data end local parameters_data local function get_parameters_data() parameters_data, get_parameters_data = load_data("Module:parameters/data"), nil return parameters_data end local translations_data local function get_translations_data() translations_data, get_translations_data = load_data("Module:translations/data"), nil return translations_data end local function is_translation_subpage(pagename) if (headword_data or get_headword_data()).page.namespace ~= "" then return false elseif not pagename then pagename = (headword_data or get_headword_data()).encoded_pagename end return pagename:match("./translations$") and true or false end local function canonical_pagename() local pagename = (headword_data or get_headword_data()).encoded_pagename return is_translation_subpage(pagename) and pagename:sub(1, -14) or pagename end local function interwiki(terminfo, term, lang, langcode) -- No interwiki link if term is empty/missing if not term or #term < 1 then terminfo.interwiki = false return end -- Percent-decode the term. term = decode_uri(terminfo.term, "PATH") -- Don't show an interwiki link if it's an invalid title. if not new_title(term) then terminfo.interwiki = false return end local interwiki_langcode = (translations_data or get_translations_data()).interwiki_langs[langcode] local wmlangs = interwiki_langcode and {get_wikimedia_lang(interwiki_langcode)} or lang:getWikimediaLanguages() -- Don't show the interwiki link if the language is not recognised by Wikimedia. if #wmlangs == 0 then terminfo.interwiki = false return end local sc = terminfo.sc local target_page = get_link_page(term, lang, sc) local split = split_on_slashes(target_page) if not split[1] then terminfo.interwiki = false return end target_page = split[1] local wmlangcode = wmlangs[1]:getCode() local interwiki_link = language_link{ lang = lang, sc = sc, term = wmlangcode .. ":" .. target_page, alt = "(" .. wmlangcode .. ")", tr = "-" } terminfo.interwiki = tostring(html_create("span") :addClass("tpos") :wikitext("&nbsp;" .. interwiki_link) ) end function export.show_terminfo(terminfo, check) local lang = terminfo.lang local langcode, langname = lang:getCode(), lang:getCanonicalName() -- Translations must be for mainspace languages. if not lang:hasType("regular") then error("Translations must be for attested and approved main-namespace languages.") else local disallowed = (translations_data or get_translations_data()).disallowed local err_msg = disallowed[langcode] if err_msg then error("Translations not allowed in " .. langname .. " (" .. langcode .. "). " .. langname .. " translations should " .. err_msg) end local fullcode = lang:getFullCode() if fullcode ~= langcode then err_msg = disallowed[fullcode] if err_msg then langname = lang:getFullName() error("Translations not allowed in " .. langname .. " (" .. fullcode .. "). " .. langname .. " translations should " .. err_msg) end end end if langcode == "ms" then if terminfo.interwiki then error("Interwiki translations not allowed for Malay; they should always link to a different Wiktionary") end local current_L2 = require(pages_module).get_current_L2() if current_L2 ~= "Rentas bahasa" and mw.title.getCurrentTitle().nsText ~= "Wikikamus" then if current_L2 then error("Malay translations only allowed in Translingual section, not in " .. current_L2) else error("Malay translations only allowed in Translingual section, not outside of any L2") end end end local term = terminfo.term -- Check if there is a term. Don't show the interwiki link if there is nothing to link to. if not term then -- Track entries that don't provide a term. -- FIXME: This should be a category. track("translations/no term") track("translations/no term/" .. langcode) end if terminfo.interwiki then interwiki(terminfo, term, lang, langcode) end langcode = lang:getFullCode() if (translations_data or get_translations_data()).need_super[langcode] then local tr = terminfo.tr if tr ~= nil then terminfo.tr = tr:gsub("%d[%d%*%-]*%f[^%d%*]", "<sup>%0</sup>") end end terminfo.show_decorations = true local link = full_link(terminfo, "translation") local canonical_name = lang:getCanonicalName() local full_name = lang:getFullName() local categories = {"Perkataan dengan terjemahan bahasa " .. canonical_name} if canonical_name ~= full_name then insert(categories, "Perkataan dengan terjemahan bahasa " .. full_name) end if check then link = tostring(html_create("span") :addClass("ttbc") :tag("sup") :addClass("ttbc") :wikitext("(sila [[WT:Terjemahan#Terjemahan untuk disemak|sahkan]])") :done() :wikitext(" " .. link) ) insert(categories, "Permintaan pengesahan terjemahan bahasa " .. langname ) end return link .. format_categories(categories, en or get_en(), nil, canonical_pagename()) end -- Implements {{t}}, {{t+}}, {{t-check}} and {{t+check}}. function export.show(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation"]) local check = frame.args.check return export.show_terminfo({ lang = args[1], sc = args.sc, track_sc = true, term = args[2], alt = args.alt, id = args.id, genders = args[3], tr = args.tr, ts = args.ts, lit = args.lit, q = args.q, qq = args.qq, l = args.l, ll = args.ll, refs = args.ref, interwiki = frame.args.interwiki, }, check and check ~= "") end local function add_id(div, id) return id and div:attr("id", normalize_anchor("Translations-" .. id)) or div end -- Implements {{ter-atas}} and part of {{ter-atas-juga}}. local function top(args, title, id, navhead) local column_width = (args["column-width"] == "wide" or args["column-width"] == "narrow") and "-" .. args["column-width"] or "" local div = html_create("div") :addClass("NavFrame") :node(navhead) :tag("div") :addClass("NavContent") :tag("table") :addClass("translations") :attr("role", "presentation") :attr("data-gloss", title or "") :tag("tr") :tag("td") :addClass("translations-cell") :addClass("multicolumn-list" .. column_width) :attr("colspan", "3") :allDone() div = add_id(div, id) local categories = {} if not title then insert(categories, "Pengepala jadual terjemahan kekurangan padanan kata") end local pagename = canonical_pagename() if is_translation_subpage() then insert(categories, "Sublaman terjemahan") end return (tostring(div):gsub("</td></tr></table></div></div>$", "")) .. (#categories > 0 and format_categories(categories, en or get_en(), nil, pagename) or "") .. -- Category to trigger [[MediaWiki:Gadget-TranslationAdder.js]]; we want this even on -- user pages and such. format_categories("Perkataan dengan kotak terjemahan", nil, nil, nil, true) .. templatestyles("Modul:translations/styles.css") end -- Entry point for {{ter-atas}}. function export.top(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas"]) local title = args[1] local id = args.id or title title = title and remove_links(title) return top(args, title, id, html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(title or "Terjemahan") ) end -- Entry point for {{checktrans-top}}. function export.check_top(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["checktrans-top"]) local text = "\n:''The translations below need to be checked and inserted above into the appropriate translation tables. See instructions at " .. frame:expandTemplate{ title = "section link", args = {"Wikikamus:Susun atur entri#Terjemahan"} } .. ".''\n" local header = html_create("div") :addClass("checktrans") :wikitext(text) local subtitle = args[1] local title = "Terjemahan untuk disemak" if subtitle then title = title .. "&zwnj;: \"" .. subtitle .. "\"" end -- No ID, since these should always accompany proper translation tables, and can't be trusted anyway (i.e. there's no use-case for links). return tostring(header) .. "\n" .. top(args, title, nil, html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(title or "Terjemahan") ) end -- Implements {{ter-bawah}}. function export.bottom(frame) -- Check nothing is being passed as a parameter. process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-bawah"]) return "</table></div></div>" end -- Implements {{ter-lihat}} and part of {{ter-atas-juga}}. local function see(args, see_text) local navhead = html_create("div") :addClass("NavHead") :css("text-align", "left") :wikitext(args[1] .. " ") :tag("span") :css("font-weight", "normal") :wikitext("— ") :tag("i") :wikitext(see_text) :allDone() local terms, id = args[2], args.id if #terms == 0 then terms[1] = args[1] end for i = 1, #terms do local term_id = id[i] or id.default local data = { term = terms[i], id = term_id and "Translations-" .. term_id or "Translations", } terms[i] = plain_link(data) end return navhead:wikitext(concat(terms, ",&lrm; ")) end -- Entry point for {{ter-lihat}}. function export.see(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-lihat"]) local div = html_create("div") :addClass("pseudo") :addClass("NavFrame") :node(see(args, "see ")) return tostring(add_id(div, args.id.default or args[1])) end -- Entry point for {{ter-atas-juga}}. function export.top_also(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas-juga"]) local navhead = see(args, "lihat juga ") local title = args[1] local id = args.id.default or title title = remove_links(title) return top(args, title, id, navhead) end -- Implements {{translation subpage}}. function export.subpage(frame) process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation subpage"]) if not is_translation_subpage() then error("This template should only be used on translation subpages, which have titles that end with '/translations'.") end -- "Translation subpages" category is handled by {{trans-top}}. return ("''This page contains translations for ''%s''. See the main entry for more information.''"):format(full_link{ lang = en or get_en(), term = canonical_pagename(), }) end -- Implements {{t-needed}}. function export.needed(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["t-needed"]) local lang, category = args[1], "" local span = html_create("span") :addClass("trreq") :attr("data-lang", lang:getCode()) :tag("i") :wikitext("sila tambah terjemahan ini jika boleh") :done() if not args.nocat then local type, sort = args[2], args.sort if type == "quote" then category = "Permintaan terjemahan petikan bahasa " .. lang:getCanonicalName() elseif type == "usex" then category = "Permintaan terjemahan contoh penggunaan bahasa " .. lang:getCanonicalName() else category = "Permintaan terjemahan ke dalam bahasa " .. lang:getCanonicalName() lang = en or get_en() end category = format_categories(category, lang, sort, not sort and canonical_pagename() or nil) end return tostring(span) .. category end -- Implements {{no equivalent translation}}. function export.no_equivalent(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no equivalent translation"]) local text = "tiada padanan kata dalam bahasa " .. args[1]:getCanonicalName() if not args.noend then text = text .. ", tetapi lihat" end return tostring(html_create("i"):wikitext(text)) end -- Implements {{no attested translation}}. function export.no_attested(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no attested translation"]) local langname = args[1]:getCanonicalName() local text = "tiada perkataan [[WT:ATTEST|disahkan]] dalam bahasa " .. langname local category = "" if not args.noend then text = text .. ", but see" local sort = args.sort category = format_categories("Terjemahan bahasa " .. langname .. " yang tidak disahkan", en or get_en(), sort, not sort and canonical_pagename() or nil) end return tostring(html_create("i"):wikitext(text)) .. category end -- Implements {{not used}}. function export.not_used(frame) local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["not used"]) return tostring(html_create("i"):wikitext((args[2] or "tidak digunakan") .. " dalam " .. args[1]:getCanonicalName())) end return export hg552qndw0fl79qt8gujt8rt8an4ytr Modul:IPA 828 9946 375359 227150 2026-09-22T03:14:30Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92723663|92723663]]) 375359 Scribunto text/plain local export = {} local force_cat = false -- for testing local decorations_module = "Module:decorations" local pages_module = "Module:pages" local qualifier_module = "Module:qualifier" local string_utilities_module = "Module:string utilities" local syllables_module = "Module:syllables" local utilities_module = "Module:utilities" local m_data = mw.loadData("Module:IPA/data") local m_str_utils = require(string_utilities_module) local m_syllables -- [[Module:syllables]]; loaded below if needed local m_symbols = mw.loadData("Module:IPA/data/symbols") local concat = table.concat local decode_entities = m_str_utils.decode_entities local find = string.find local gcodepoint = m_str_utils.gcodepoint local gmatch = m_str_utils.gmatch local gsub = string.gsub local insert = table.insert local is_preview = require(pages_module).is_preview local len = m_str_utils.len local listToText = mw.text.listToText local match = string.match local pattern_escape = m_str_utils.pattern_escape local sub = string.sub local u = m_str_utils.char local ugsub = m_str_utils.gsub local umatch = m_str_utils.match local usub = m_str_utils.sub local function with_codepoints(s) if find(s, "%S%s") then local parts = {} for ch in gmatch(s, "%S") do parts[#parts + 1] = with_codepoints(ch) end return concat(parts, ", ") end local cps = {} for cp in gcodepoint(s) do cps[#cps + 1] = ("U+%04X"):format(cp) end return s .. " [" .. concat(cps, " ") .. "]" end local namespace = mw.title.getCurrentTitle().nsText local function is_content_page(lang, namespace) return namespace == "" or namespace == "Rekonstruksi" or lang and lang:hasType("appendix-constructed") and namespace == "Lampiran" end -- Etymology-only languages are not L2 entry languages; IPA should use the parent full language. local function assert_not_etymology_only_lang(lang) if lang and lang.hasType and lang:hasType("language", "etymology-only") then local parent_code = lang.getParentCode and lang:getParentCode() or nil error(("Cannot use IPA with the etymology-only language %q; use the parent full language %q instead."):format(lang:getCode(), parent_code)) end end local function track(page) require("Module:debug/track")("IPA/" .. page) return true end local function process_maybe_split_categories(split_output, categories, prontext, lang, errtext) if split_output ~= "raw" then if categories[1] then categories = require(utilities_module).format_categories(categories, lang, nil, nil, force_cat) else categories = "" end end if split_output then -- for use of IPA in links, etc. if errtext then return prontext, categories, errtext else return prontext, categories end else return prontext .. (errtext or "") .. categories end end --[==[ Format a line of one or more IPA pronunciations as {{tl|IPA}} would do it, i.e. with a preceding {"IPA:"} followed by the word {"key"} linking to an Appendix page describing the language's phonology, and with an added category ` ``lang`` terms with IPA pronunciation`. Other than the extra preceding text and category, this is identical to {format_IPA_multiple()}, and the considerations described there in the documentation apply here as well. There is a single parameter `data`, an object with the following fields: * `lang`: Object representing the language of the pronunciations, which is used when adding cleanup categories for pronunciations with invalid phonemes; for determining how many syllables the pronunciations have in them, in order to add a category such as [[:Category:Italian 2-syllable words]] (for certain languages only); for adding a category ` ``lang`` terms with IPA pronunciation`; and for determining the proper sort keys for categories. Unlike for {format_IPA_multiple()}, `lang` may not be {nil}. * `items`: List of pronunciations, in exactly the same format as for {format_IPA_multiple()}. * `err`: If not {nil}, a string containing an error message to use in place of the link to the language's phonology. * `separator`: The default separator to use when separating formatted items. Defaults to {", "}. Does not apply to the first item, where the default separator is always the empty string. Overridden by the per-item `separator` field in `items`. * `sort_key`: Explicit sort key used for categories. * `no_count`: Suppress adding a {#-syllable words} category such as [[:Category:Italian 2-syllable words]]. Note that only certain languages add such categories to begin with, because it depends on knowing how to count syllables in a given language, which depends on the phonology of the language. Also, this does not suppress the addition of cleanup or other categories. If you need them suppressed, use `split_output` to return the categories separately and ignore them. * `split_output`: If not given, the return value is a concatenation of the formatted pronunciation and formatted categories. Otherwise, two values are returned: the formatted pronunciation and the categories. If `split_output` is the value {"raw"}, the categories are returned in list form, where the list elements are a combination of category strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]]. If `split_output` is any other value besides {nil}, the categories are returned as a pre-formatted concatenated string. * `include_langname`: If specified, prefix the result with the language name, followed by a colon. * `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display at the beginning, before the formatted pronunciations and preceding {"IPA:"}. * `qq`: {nil} or a list of right qualifiers to display after all formatted pronunciations. * `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display at the beginning, before the formatted pronunciations and preceding {"IPA:"}. * `aa`: {nil} or a list of right accent qualifiers to display after all formatted pronunciations. ]==] function export.format_IPA_full(data) if type(data) ~= "table" or data.getCode then error("Must now supply a table of arguments to format_IPA_full(); first argument should be that table, not a language object") end local lang = data.lang local items = data.items local err = data.err local separator = data.separator local sort_key = data.sort_key local no_count = data.no_count local split_output = data.split_output local q = data.q local qq = data.qq local a = data.a local aa = data.aa local include_langname = data.include_langname if data.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("overall `.qualifiers` is no longer supported; change the code to use `.q` or `.qq`") end local hasKey = m_data.langs_with_infopages if not lang or not lang.getCode then error("Must specify language to format_IPA_full()") end assert_not_etymology_only_lang(lang) local langname = lang:getCanonicalName() local prefix_text if err then prefix_text = '<span class="error">' .. err .. '</span>' else if hasKey[lang:getCode()] then prefix_text = "Lampiran:Sebutan bahasa " .. langname else prefix_text = "wikipedia:Fonologi bahasa " .. langname end prefix_text = "[[" .. prefix_text .. "|kekunci]]" end local prefix = "[[Wikikamus:Abjad Fonetik Antarabangsa|AFA]]<sup>(" .. prefix_text .. ")</sup>:&#32;" local IPAs, categories = export.format_IPA_multiple(lang, items, separator, no_count, "raw") if is_content_page(lang, namespace) then insert(categories, { cat = "Perkataan dengan sebutan AFA bahasa " .. langname, sort_key = sort_key }) end local prontext = prefix .. IPAs if q and q[1] or qq and qq[1] or a and a[1] or aa and aa[1] then prontext = require(decorations_module).format_decorations { lang = lang, text = prontext, q = q, qq = qq, a = a, aa = aa, } end if include_langname then prontext = langname .. ": " .. prontext end return process_maybe_split_categories(split_output, categories, prontext, lang) end local function split_phonemic_phonetic(pron) local reconstructed, phonemic, phonetic = match(pron, "^(%*?)(/.-/)%s+(%[.-%])$") if reconstructed then return reconstructed .. phonemic, reconstructed .. phonetic else return pron, nil end end local function determine_repr(pron) local reconstructed -- Temporarily remove any initial asterisk before representation marks, -- which avoids having to account for it in the data, but set the -- `reconstructed` flag. if sub(pron, 1, 1) == "*" then reconstructed = true pron = sub(pron, 2) end -- Some representation types have aliases for convenience (e.g. "// //" is -- an alias for "⫽ ⫽"). and these need to be substituted in before checking -- for other data. local opening, n = match(pron, "^.[\128-\191]*") local subs_data = m_data.representation_subs[opening] if subs_data then pron, n = ugsub(pron, subs_data[1], subs_data[2]) -- If the substitution was made, `opening` needs to be changed to the -- new opening character. if n ~= 0 then opening = subs_data[3] end end -- Get the type data based on the opening character (if any), and set the -- representation type if the closing character matches. local type_data, repr, closing = m_data.representation_types[opening] if type_data then closing = type_data[2] if type_data and match(pron, pattern_escape(closing) .. "$", #opening + 1) then repr = type_data[1] end end -- Default to the empty string. if not repr then opening, closing = "", "" end -- Reattach the asterisk if reconstructed. if reconstructed then pron = "*" .. pron end return pron, repr, opening, closing, reconstructed end local function hasInvalidSeparators(transcription) -- Escape certain characters as well as pauses, which have the format "(...)" (with any number of dots), to avoid false-positives. transcription = transcription:gsub(".[\128-\191]*", m_symbols.separator_escapes) :gsub("%(%.+%)", "\3") :gsub("[()]+", "") return ( transcription:find("..", nil, true) or transcription:match("%.%f[%z \1\2\3,:;]") or transcription:match("\1%f[%z \2\3,:;]") or transcription:match("\2%f[%z \1\3,:;]") or transcription:match("\3[:;]") or transcription:match("%f[^%z \1\2\3,]%.") ) and true or false end --[==[ Format a line of one or more bare IPA pronunciations (i.e. without any preceding {"IPA:"} and without adding to a category ` ``lang`` terms with IPA pronunciation`). Individual pronunciations are formatted using {format_IPA()} and are combined with separators, decorations, pre-text, post-text, etc. to form a line of pronunciations. Parameters accepted are: * `lang` is an object representing the language of the pronunciations, which is used when adding cleanup categories for pronunciations with invalid phonemes; for determining how many syllables the pronunciations have in them, in order to add a category such as [[:Category:Italian 2-syllable words]] (for certain languages only); and for computing the proper sort keys for categories. `lang` may be {nil}. * `items` is a list of pronunciations, each of which is an object with the following properties: ** `pron`: the pronunciation, in the same format as is accepted by {format_IPA()}, i.e. it should be either phonemic (surrounded by {/.../}), phonetic (surrounded by {[...]}), orthographic (surrounded by {⟨...⟩}) or a rhyme (beginning with a hyphen); ** `pretext`: text to display directly before the formatted pronunciation, inside of any qualifiers or accent qualifiers; ** `posttext`: text to display directly after the formatted pronunciation, inside of any qualifiers or accent qualifiers; ** `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display before the formatted pronunciation; ** `qq`: {nil} or a list of right qualifiers to display after the formatted pronunciation; ** `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display before the formatted pronunciation; ** `aa`: {nil} or a list of right accent qualifiers to after before the formatted pronunciation; ** `refs`: {nil} or a list of references or reference specs to add after the pronunciation and any posttext and qualifiers; the value of a list item is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}}) and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the actual reference, located in the {{cd|<nowiki><references /></nowiki>}} section; ** `gloss`: {nil} or a gloss (definition) for this item, if different definitions have different pronunciations; ** `pos`: {nil} or a part of speech for this item, if different parts of speech have different pronunciations; ** `separator`: the separator text to insert directly before the formatted pronunciation and all decorations and pre-text; defaults to the outer `separator` parameter. * `separator`: The default separator to use when separating formatted items. Defaults to {", "}. Does not apply to the first item, where the default separator is always the empty string. Overridden by the per-item `separator` field in `items`. * `no_count`: Suppress adding a {#-syllable words} category such as [[:Category:Italian 2-syllable words]]. Note that only certain languages add such categories to begin with, because it depends on knowing how to count syllables in a given language, which depends on the phonology of the language. Also, this does not suppress the addition of cleanup categories. If you need them suppressed, use `split_output` to return the categories separately and ignore them. * `split_output`: If not given, the return value is a concatenation of the formatted pronunciation and formatted categories. Otherwise, two values are returned: the formatted pronunciation and the categories. If `split_output` is the value {"raw"}, the categories are returned in list form, where the list elements are a combination of category strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]]. If `split_output` is any other value besides {nil}, the categories are returned as a pre-formatted concatenated string. ]==] function export.format_IPA_multiple(lang, items, separator, no_count, split_output) local categories = {} separator = separator or ", " if not lang then track("format-multiple-nolang") else assert_not_etymology_only_lang(lang) end -- Format if not items[1] then if namespace == "Templat" then insert(items, {pron = "/aɪ piː ˈeɪ/"}) else insert(categories, "Templat sebutan tanpa sebutan") end end local bits = {} for i, item in ipairs(items) do local bit -- If the pronunciation is entirely empty, allow this and don't do anything, so that e.g. the pretext and/or -- posttext can be specified to force something like ''unknown'' to appear in place of the pronunciation -- (as happens e.g. when ? is used as a respelling in [[Module:ca-IPA]]; see [[guèiser]] for an example). if item.pron == "" then bit = "" else local item_categories, errtext bit, item_categories, errtext = export.format_IPA(lang, item.pron, "raw") bit = bit .. errtext for _, cat in ipairs(item_categories) do insert(categories, cat) end end if item.pretext then bit = item.pretext .. bit end if item.posttext then bit = bit .. item.posttext end if item.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("`.qualifiers` is no longer supported; change the code to use `.q` or `.qq`") end local has_decorations = item.q and item.q[1] or item.qq and item.qq[1] or item.a and item.a[1] or item.aa and item.aa[1] or item.refs and item.refs[1] local has_gloss_or_pos = item.gloss or item.pos if has_decorations or has_gloss_or_pos then -- FIXME: Currently we tack the gloss and POS (in that order) onto the end of the regular left qualifiers. -- Should we do something different? local q = item.q if has_gloss_or_pos then q = mw.clone(item.q) or {} if item.gloss then local m_qualifier = require(qualifier_module) insert(q, m_qualifier.wrap_qualifier_css("“", "quote") .. item.gloss .. m_qualifier.wrap_qualifier_css("”", "quote")) end if item.pos then -- FIXME: Consider expanding aliases as found in [[Module:headword/data]] or similar. insert(q, item.pos) end end bit = require(decorations_module).format_decorations { lang = lang, text = bit, q = q, qq = item.qq, a = item.a, aa = item.aa, refs = item.refs, } end bit = (item.separator or (i == 1 and "" or separator)) .. bit insert(bits, bit) --[=[ [[Special:WhatLinksHere/Wiktionary:Tracking/IPA/syntax-error]] The length or gemination symbol should not appear after a syllable break or stress symbol. ]=] -- The nature of the following pattern match is such that we don't have to split a combined '/.../ [...]' spec -- into its parts in order to process. if match(item.pron, "[.\203][\136\140]?\203[\144\145]") then -- [.ˈˌ][ːˑ] track("syntax-error") end if lang then -- Add syllable count if the language's diphthongs are listed in [[Module:syllables]]. -- Don't do this if the term has spaces, a liaison mark (‿) or isn't in mainspace. if not no_count and namespace == "" then m_syllables = m_syllables or require(syllables_module) local langcode = lang:getCode() if m_data.langs_to_generate_syllable_count_categories[langcode] then local raw_phonemic, phonetic, use_it = split_phonemic_phonetic(item.pron) local phonemic, repr = determine_repr(raw_phonemic) if not phonetic then -- not a '/.../ [...]' combined pronunciation if m_data.langs_to_use_phonetic_or_phonemic_notation[langcode] then use_it = phonemic elseif m_data.langs_to_use_phonetic_notation[langcode] then use_it = repr == "phonetic" and phonemic or nil else use_it = repr == "phonemic" and phonemic or nil end elseif repr == "phonetic" then use_it = phonetic elseif repr == "phonemic" then use_it = phonemic end -- Note: two uses of find with plain patterns is much faster than umatch with [ ‿]. if use_it and not (find(use_it, " ") or find(use_it, "‿")) then local syllable_count = m_syllables.getVowels(use_it, lang) if syllable_count then insert(categories, "Perkataan " .. syllable_count .. " suku kata bahasa " .. lang:getCanonicalName()) end end end end end end return process_maybe_split_categories(split_output, categories, concat(bits), lang) end --[=[ Format a single IPA pronunciation, which cannot be a combined spec (such as {/.../ [...]}). This has been extracted from {format_IPA()} to allow the latter to handle such combined specs. This works like {format_IPA()} but requires that pre-created {err} (for error messages) and {categories} lists be passed in, and adds any generated error messages and categories to those lists. A single value is returned, the pronunciation, which is usually the same as passed in, but may have HTML added surrounding invalid characters so they appear in red. ]=] local function format_one_IPA(lang, raw_pron, err, categories) -- Disallow wikilinks. if match(raw_pron, "%[%[.-%]%]") then error("IPA input must not contain wikilinks.") end raw_pron = decode_entities(raw_pron) -- Detect the type of transcription. local pron, repr, opening, closing, reconstructed = determine_repr(raw_pron) -- Strip any reconstruction asterisk and representation marks. pron = sub(pron, #opening + 1 + (reconstructed and 1 or 0), -#closing - 1) if not repr then insert(categories, "Sebutan AFA dengan tanda perwakilan tidak sah") -- insert(err, "tanda perwakilan tidak sah") -- Removed because it's annoying when previewing pronunciation pages. end if repr ~= "orthographic" and lang and lang:getCode() == "en" and hasInvalidSeparators(pron) then insert(categories, "Sebutan AFA bahasa Inggeris dengan pemisah tidak sah") end if pron == "" then insert(categories, "Sebutan AFA dengan tiada sebutan") end -- Check for obsolete and nonstandard symbols for _, symbol in ipairs(m_data.nonstandard) do local result for nonstandard in gmatch(pron, symbol) do if not result then result = {} end insert(result, nonstandard) insert(categories, {cat = "Sebutan AFA dengan aksara usang atau tidak standard", sort_key = nonstandard} ) end if result then insert(err, "aksara usang atau tidak standard (" .. concat(result) .. ")") break end end --[[ Check for invalid symbols after removing the following: 1. wikilinks (handled above) 2. paired HTML tags 3. bolding 4. italics 5. asterisk at beginning of transcription 6. comma followed by spacing characters 7. superscripts enclosed in superscript parentheses ]] local found_HTML local result = gsub(pron, "<(%a+)[^>]*>([^<]+)</%1>", function(tagName, content) found_HTML = true return content end) result = gsub(result, "'''([^']*)'''", "%1") result = gsub(result, "''([^']*)''", "%1") result = gsub(result, "^%*", "") result = ugsub(result, ",%s+", "") -- VS15 local vs15_class = "[" .. m_symbols.add_vs15 .. "]" if umatch(pron, vs15_class) then local vs15 = u(0xFE0E) if find(result, vs15) then result = gsub(result, vs15, "") pron = gsub(pron, vs15, "") end pron = ugsub(pron, vs15_class, "%0" .. vs15) end if result ~= "" then local content_page = is_content_page(lang, namespace) if lang then -- Get the per_lang_valid data, and convert any per-language valid sequences to spaces. local per_lang_valid = m_symbols.per_lang_valid[lang:getCode()] if per_lang_valid then if type(per_lang_valid) == "table" then for _, pattern in pairs(per_lang_valid) do result = ugsub(result, pattern, " ") end else -- Should be a string. result = ugsub(result, per_lang_valid, " ") end end end local suggestions = {} -- Check for any invalid sequences, excluding anything in the per-language lookup table. for k, v in pairs(m_symbols.invalid) do if find(result, k, nil, true) then insert(suggestions, with_codepoints(k) .. " dengan " .. with_codepoints(v)) end end if suggestions[1] then local replacements = "menggantikan " .. listToText(suggestions) if content_page then error("Invalid IPA: " .. replacements) end insert(err, replacements) end -- Convert any valid character sequences to spaces for _, pattern in pairs(m_symbols.valid) do result = ugsub(result, pattern, " ") end if not match(result, "^ *$") then local category = "Sebutan AFA dengan aksara AFA tidak sah" if not content_page then category = category .. "/non_mainspace" end insert(categories, category) insert(err, "aksara AFA tidak sah: " .. with_codepoints(result)) end end if found_HTML then insert(categories, "Sebutan AFA dengan tag HTML berpasangan") end if (repr == "phonemic" or repr == "rhyme") and lang and m_data.phonemes[lang:getCode()] then local valid_phonemes = m_data.phonemes[lang:getCode()] local rest = pron local phonemes = {} while #rest > 0 do local longestmatch, longestmatch_len = "", 0 local rest_init = sub(rest, 1, 1) if rest_init == "(" or rest_init == ")" then longestmatch = rest_init longestmatch_len = 1 else for _, phoneme in ipairs(valid_phonemes) do local phoneme_len = len(phoneme) if phoneme_len > longestmatch_len and usub(rest, 1, phoneme_len) == phoneme then longestmatch = phoneme longestmatch_len = len(longestmatch) end end end if longestmatch_len > 0 then insert(phonemes, longestmatch) rest = usub(rest, longestmatch_len + 1) else local phoneme = usub(rest, 1, 1) insert(phonemes, "<span style=\"color: var(--wikt-palette-red,red)\">" .. phoneme .. "</span>") rest = usub(rest, 2) insert(categories, "Sebutan AFA dengan fonem tidak sah/" .. lang:getCode()) track("fonem tidak sah/" .. phoneme) end end pron = concat(phonemes) end return (reconstructed and "*" or "") .. opening .. pron .. closing end --[==[ Format an IPA pronunciation. This wraps the pronunciation in appropriate CSS classes and adds cleanup categories and error messages as needed. The pronunciation `pron` should be either phonemic (surrounded by {/.../}), phonetic (surrounded by {[...]}), orthographic (surrounded by {⟨...⟩}), a rhyme (beginning with a hyphen) or a combined phonemic/phonetic spec (of the form {/.../ [...]}). `lang` indicates the language of the pronunciation and can be {nil}. If not {nil}, and the specified language has data in [[Module:IPA/data]] indicating the allowed phonemes, then the page will be added to a cleanup category and an error message displayed next to the outputted pronunciation. Note that {lang} also determines sort key processing in the added cleanup categories. If `split_output` is not given, the return value is a concatenation of the formatted pronunciation, error messages and formatted cleanup categories. Otherwise, three values are returned: the formatted pronunciation, the cleanup categories and the concatenated error messages. If `split_output` is the value {"raw"}, the cleanup categories are returned in list form, where the list elements are a combination of category strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]]. If `split_output` is any other value besides {nil}, the cleanup categories are returned as a pre-formatted concatenated string. ]==] function export.format_IPA(lang, pron, split_output) local err = {} local categories = {} -- `pron` shouldn't contain ref tags. if match(pron, "\127'\"`UNIQ%-%-ref%-[%dA-F]+%-QINU`\"'\127") then error("<ref> tags found inside pronunciation parameter.") end if not lang then track("format-nolang") else assert_not_etymology_only_lang(lang) end local phonemic, phonetic = split_phonemic_phonetic(pron) pron = format_one_IPA(lang, phonemic, err, categories) if phonetic then track("phonemic-phonetic") -- There's no benefit to supporting the "/.../ [...]" format within one parameter. phonetic = format_one_IPA(lang, phonetic, err, categories) pron = pron .. " " .. phonetic end if err[1] and is_preview() then err = '<span class="error" style="font-size: small;>&#32;' .. concat(err, ", ") .. "</span>" else err = "" end return process_maybe_split_categories(split_output, categories, '<span class="IPA nowrap">' .. pron .. "</span>", lang, err) end --[==[ Format a line of one or more enPR pronunciations as {{tl|enPR}} would do it, i.e. with a preceding {"enPR:"} (linked to [[Appendix:English pronunciation]]) followed by one or more formatted, comma-separated enPR pronunciations. The pronunciations are formatted by wrapping them in the `AHD` and `enPR` CSS classes and adding any decorations (qualifiers, accent qualifiers and references). In addition, the overall result is wrapped in any overall decorations. There is a single parameter `data`, an object with the following fields: * `items` is a list of enPR pronunciations, each of which is an object with the following properties: ** `pron`: the enPR pronunciation; ** `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display before the formatted pronunciation; ** `qq`: {nil} or a list of right qualifiers to display after the formatted pronunciation; ** `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display before the formatted pronunciation; ** `aa`: {nil} or a list of right accent qualifiers to after before the formatted pronunciation. * `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display at the beginning, before the formatted pronunciations and preceding {"enPR:"}. * `qq`: {nil} or a list of right qualifiers to display after all formatted pronunciations. * `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display at the beginning, before the formatted pronunciations and preceding {"enPR:"}. * `aa`: {nil} or a list of right accent qualifiers to display after all formatted pronunciations. ]==] function export.format_enPR_full(data) local prefix = "[[Appendix:English pronunciation|enPR]]: " local lang = require("Module:languages").getByCode("en") local parts = {} for _, item in ipairs(data.items) do local part = '<span class="AHD enPR">' .. item.pron .. "</span>" if item.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("`.qualifiers` is no longer supported; change the code to use `.q` or `.qq`") end if item.q and item.q[1] or item.qq and item.qq[1] or item.a and item.a[1] or item.aa and item.aa[1] then part = require(decorations_module).format_decorations { lang = lang, text = part, q = item.q, qq = item.qq, a = item.a, aa = item.aa, } end insert(parts, part) end local prontext = prefix .. concat(parts, ", ") if data.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("overall `.qualifiers` is no longer supported; change the code to use `.q` or `.qq`") end if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then prontext = require(decorations_module).format_decorations { lang = lang, text = prontext, q = data.q, qq = data.qq, a = data.a, aa = data.aa, } end return prontext end return export 1qhkiy8wyqpetrqk0p5yanxnrkmiezk Modul:IPA/data 828 10199 375367 218692 2026-09-22T04:48:13Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91333763|91333763]]) 375367 Scribunto text/plain local list_to_set = require("Module:table").listToSet local data = {} --[=[ A list of representation types (e.g. /foo/ for phonemic and [bar] for phonetic), given as a table. The key is the opening character, the first value the representation type, and the second value the closing symbol.]=] data.representation_types = { ["/"] = {"phonemic", "/"}, ["["] = {"phonetic", "]"}, ["⫽"] = {"morphophonemic", "⫽"}, ["⟨"] = {"orthographic", "⟩"}, ["-"] = {"rhyme", ""}, } --[=[ A list of convenience inputs for certain representation types. The key is the opening character, and the table is a three-item array consisting of (1) an mw.ustring.gsub pattern which is anchored to the start and end of the string, with a single capture group that excludes the characters to be substituted, (2) a corresponding replacement pattern to be used with the pattern, and (3) the replacement opening character.]=] data.representation_subs = { ["<"] = {"^<(.*)>$", "⟨%1⟩", "⟨"}, ["/"] = {"^//(.*)//$", "⫽%1⫽", "⫽"}, } --[=[ This should list the language codes of all languages that have a pronunciation page in the appendix of the form ''Appendix:LANG pronunciation'', e.g. [[Appendix:Russian pronunciation]]. For these languages, the text "key" next to the generated pronunciation links to such pages; for other languages, it links to the "LANG phonology" page in Wikipedia (which may or may not exist). [[Module:IPA]] is responsible for this linking; see format_IPA_full().]=] data.langs_with_infopages = list_to_set{ "acw", "ady", "ang", "arc", "ba", "bg", "bo", "ca", "cho", "cmn", "cs", "cv", "cy", "da", "de", "dsb", "dz", "egl", "egy", "el", "en", "enm", "eo", "es", "fa", "fi", "fo", "fr", "fy", "ga", "gd", "ghc", "gmh", "gmw-msc", "got", "he", "hi", "hrx", "hu", "hy", "id", "ii", "is", "it", "iu", "ja", "jbo", "ka", "kls", "ko", "kw", "la", "lb", "liv", "lt", "lv", "mdf", "mfe", "mic", "mk", "mns-nor", "ms", "mt", "mul", "my", "nan", "nci", "nl", "nn", "no", "nov", "nv", "pjt", "pl", "ps", "pt", "ro", "ru", "scn", "sco", "sga", "sh", "sl", "sq", "sv", "sw", "syc", "szl", "tg", "th", "tl", "tpw", "tr", "tyv", "ug", "uk", "vi", "vo", "wlm", "yi", "yrl", "yue", "zlw-mas" } --[=[ This should list the diphthongs of a language (in the form of Lua patterns), provided they do *NOT* contain semivowel symbols such as /j w ɰ ɥ/ or vowels with nonsyllabic diacritics such as /i̯ u̯/. For example, list /au/ or /aʊ/, but do not list /aw/ or /au̯/. The data in this table is used to count the number of syllables in a word. [[Module:syllables]] automatically knows how to correctly handle semivowel symbols and nonsyllabic diacritics. Any language listed here will automatically have categories of the form "LANG #-syllable words" generated. In addition, any language listed below under `langs_to_generate_syllable_count_categories` will also have such categories generated. NOTE: There are some additional languages that have these categories. For example: * Thai words have these categories added by [[Module:th-pron]].]=] data.diphthongs = { ["cs"] = { -- [[w:Czech phonology#Diphthongs]] "[aeo]u", }, ["de"] = { "a[ɪʊ]", "ɔ[ʏɪ]", }, ["en"] = { -- from [[Appendix:English pronunciation]] mostly, but /ʌɪ/ is from the OED "[aɑæeɛoɔʌ][ɪi]", "[ɑɒæo]e", "[əɐ]ʉ", "[aɒəoɔæ]ʊ", "æo", "[ɛeɪiɔʊʉ]ə", -- /iə/ is a diphthong in NZE, but a disyllabic sequence in GA. -- /ɪə/ is both a disyllabic sequence and a diphthong in old-fashioned RP. "[aʌ][ʊɪ]ə", -- May be a disyllabic sequence in some or all dialects? }, ["grc"] = { "[aeyo]i", "[ae]u", "[ɛɔa]ː[iu]", }, ["hrx"] = { "aɪ̯", "aʊ̯", "oɪ̯", "eʊ̯", }, ["is"] = { -- [[w:Icelandic phonology#Vowels]] "[aeɔœʏ]i", -- diphthongs as the module generates them "[ao]u", -- diphthongs as the module generates them "ø[iɪy]", -- additional forms that may occur; Wikipedia is oddly specific about the second element: ei and ai, but øɪ. }, ["it"] = { "[aeɛoɔu]i", "[aeɛioɔ]u", }, ["lb"] = { "[iu]ə", "[ɜoæɑ]ɪ", "[əæɑ]ʊ", }, ["lt"] = { "ɐɪ", "ɒʊ", "ɛɪ", "ɛʊ", "ʊɪ", "ɔɪ", "ɔʊ", -- Simple diphthongs (unstressed forms) "iɛ", "uɔ", -- Complex diphthongs "ɑˑɪ", "ɑˑʊ", "æˑɪ", "æˑʊ", "oˑɪ", -- Falling tone (acute) "ɐɪˑ", "ɒʊˑ", "ɛɪˑ", "ɛʊˑ", "ʊɪˑ", -- Rising tone (tilde) - lengthened second element -- Note: Mixed diphthongs (e.g., ɐlˑ, æˑn, ʊl, etc.) are omitted since they are inherently monosyllabic }, } --[=[ This should list any languages for which categories of the form "LANG #-syllable words", e.g. [[:Category:Russian 3-syllable words]], should be generated. Do not list languages here if they have an entry above under `data.diphthongs`; such languages are automatically added to this list.]=] local langs_to_generate_syllable_count_categories = list_to_set{ "ar", -- Arabic has diphthongs, but they are transcribed -- with semivowel symbols. "ary", -- Moroccan Arabic has diphthongs, but they are transcribed -- with semivowel symbols. "bg", -- Bulgarian has diphthongs with /j/ and marginally with /w/, -- but these are semivowels. "ca", -- Catalan has diphthongs, but they are generally transcribed using -- /w/ and /j/, so do not need to be listed (see [[w:Catalan language#Diphthongs and triphthongs]]. "eo", "es", -- Spanish has diphthongs, but they are transcribed with i̯ etc. "eu", -- Basque has dipthongs, but they are transcribed with i̯ and u̯. "fi", -- Finnish has diphthongs, but they are now automatically transcribed with -- the nonsyllabic diacritic "fr", -- French has diphthongs, but they are transcribed -- with semivowel symbols: [[w:French phonology#Glides and diphthongs]]. "hnn", "id", -- Indonesian has diphthongs, but they are transcribed with i̯ or /j/ etc. "ka", "kne", "kmr", "ku", "la", -- All diphthongs transcribed with e̯ or /j/ etc. "mk", "ms", -- Malay has diphthongs, but they are transcribed with i̯ or /j/ etc. "mt", -- Maltese has diphthongs, but they are transcribed -- with semivowel symbols. "pl", -- No diphthongs, properly speaking; sequences of a vowel and /w/ or /j/ though. "pt", -- Portuguese has diphthongs, but they are transcribed with i̯ or /j/ etc. "rsk", -- No diphthongs but there are sequences of vowel and /j/ or /w/. "ru", -- No diphthongs, properly speaking; sequences of a vowel and /j/ though. "sk", -- Slovak has rising diphthongs, /i̯e, i̯a, i̯u, u̯o/, which are probably always spelled with the nonsyllabic diacritic, so do not need to be listed. "sl", -- No diphthongs, properly speaking; sequences of a vowel, /j/ and /w/ though "sq", -- [[w:Albanian language#Vowels]] doesn't mention anything about diphthongs. "szy", -- All diphthongs are transcribed with /j/ or /w/ "tl", -- Tagalog has diphthongs, but they are transcribed with i̯ or /j/ etc "tsg", "ug", -- No diphthongs. } -- Also add languages listed under `data.diphthongs`. for langcode, _ in pairs(data.diphthongs) do langs_to_generate_syllable_count_categories[langcode] = true end data.langs_to_generate_syllable_count_categories = langs_to_generate_syllable_count_categories -- Languages to use the phonetic not phonemic notation to compute syllable counts. data.langs_to_use_phonetic_notation = list_to_set{ "bg", "es", "id", "la", "lt", "mk", "ms", "rsk", "ru", } -- Languages to use the phonetic or phonemic notation to compute syllable counts, whichever is available. data.langs_to_use_phonetic_or_phonemic_notation = list_to_set{ -- [[Module:is-IPA]] generates [...] but many manual pronuns use /.../. "is", } -- Non-standard or obsolete IPA symbols. data.nonstandard = { --[[ The following symbols consist of more than one character, so we can't put them in the line below. ]] "ɑ̢", "ɔ̗", "ɔ̖", "[?ƍσƺƪƞƛłščžǰǧǯẋⱻʚω∅ØȣᴀᴇⱻQKPT]" } -- See valid IPA characters at [[Module:IPA/data/symbols]]. data.phonemes = {} data.phonemes["dz"] = { "m", "n", "ŋ", "p", "t", "ʈ", "k", "pʰ", "tʰ", "ʈʰ", "kʰ", "t͡s", "t͡ɕ", "t͡sʰ", "t͡ɕʰ", "w", "s", "z", "ɬ", "l", "r", "ɕ", "ʑ", "j", "h", "ɑ", "e", "i", "o", "u", "ɑː", "eː", "ɛː", "iː", "oː", "øː", "uː", "yː", "ɑ˥", "e˥", "i˥", "o˥", "u˥", "ɑː˥", "eː˥", "ɛː˥", "iː˥", "oː˥", "øː˥", "uː˥", "yː˥", "m˥", "n˥", "ŋ˥", "p˥", "k˥", "k̚˥", "w˥", "l˥", "r˥", "ɕ˥", "j˥", ")˥", "ɑ˩", "e˩", "i˩", "o˩", "u˩", "ɑː˩", "eː˩", "ɛː˩", "iː˩", "oː˩", "øː˩", "uː˩", "yː˩", "m˩", "n˩", "ŋ˩", "p˩", "k˩", "k̚˩", "w˩", "l˩", "r˩", "ɕ˩", "j˩", ")˩", ".", ",", "-", } data.phonemes["eo"] = { "a", "b", "d", "d͡ʒ", "d͡z", "e", "f", "h", "i", "j", "k", "l", "m", "n", "o", "p", "r", "s", "t", "t͡s", "t͡ʃ", "u", "u̯", "v", "w", "x", "z", "ɡ", "ʃ", "ʒ", "ˈ", ".", " ", "-", "u̯", "i̯" } data.phonemes["hy"] = { "ɑ", "b", "ɡ", "d", "e", "z", "ə", "tʰ", "ʒ", "i", "l", "χ", "t͡s", "k", "h", "d͡z", "ʁ", "t͡ʃ", "m", "j", "n", "ʃ", "ɔ", "t͡ʃʰ", "p", "d͡ʒ", "r", "s", "v", "t", "ɾ", "t͡sʰ", "v", "pʰ", "kʰ", "o", "f", "ŋɡ", "ŋk", "ŋχ", "u", "œ", "ʏ", "ˈ", "ˌ", ".", " ", "ː", } data.phonemes["nl"] = { "m", "n", "ŋ", "p", "b", "t", "d", "k", "ɡ", "f", "v", "s", "z", "ʃ", "ʒ", "x", "ɣ", "ɦ", "ʋ", "l", "j", "r", "ɪ", "ʏ", "ɛ", "ə", "ɔ", "ɑ", "i", "iː", "y", "yː", "u", "uː", "eː", "øː", "oː", "ɛː", "œː", "ɔː", "aː", "ɛi̯", "œy̯", "ɔi̯", "ɑu̯", "ɑi̯", "iu̯", "yu̯", "ui̯", "eːu̯", "oːi̯", "aːi̯", "ˈ", "ˌ", ".", " ", "-", } data.phonemes["mt"] = { "m", "n", "p", "t", "k", "ʔ", "b", "d", "ɡ", "t͡s", "t͡ʃ", "d͡z", "d͡ʒ", "f", "s", "ʃ", "ħ", "v", "z", "ʒ", "ɣ", "l", "j", "w", "r", "ɪ", "ɛ", "ɔ", "a", "u", "ɛˤ", "ɔˤ", "aˤ", "əˤ", "ɛˤː", "ɔˤː", "aˤː", "əˤː", "ɪˤː", "iː", "ɪː", "ɛː", "ɔː", "aː", "uː", "ˈ", "ˌ", ".", " ", "‿", "-" } return data 9j3yfr5dzmr70htnh6pfy97jzb11aph Modul:columns 828 10232 375356 229188 2026-09-22T03:13:54Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708759|92708759]]) 375356 Scribunto text/plain local export = {} local collation_module = "Module:collation" local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local headword_data_module = "Module:headword/data" local JSON_module = "Module:JSON" local languages_module = "Module:languages" local links_module = "Module:links" local pages_module = "Module:pages" local parameter_utilities_module = "Module:parameter utilities" local parameters_module = "Module:parameters" local parse_utilities_module = "Module:parse utilities" local qualifier_module = "Module:qualifier" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local utilities_module = "Module:utilities" local yesno_module = "Module:yesno" local m_str_utils = require(string_utilities_module) local concat = table.concat local html = mw.html.create local is_substing = mw.isSubsting local insert = table.insert local rmatch = m_str_utils.match local remove = table.remove local sub = string.sub local trim = m_str_utils.trim local u = m_str_utils.char local dump = mw.dumpObject local function track(page) require(debug_track_module)("columns/" .. page) return true end local function deepEquals(...) deepEquals = require(table_module).deepEquals return deepEquals(...) end local function term_already_linked(term) return term == "?" or -- signals an unknown term -- optimization to avoid unnecessarily loading [[Module:parse utilities]] (term:find("[<{]") and require(parse_utilities_module).term_already_linked(term)) end local function convert_delimiter_to_separator(item, itemind, args) if itemind == 1 then item.separator = nil elseif item.delimiter == " " then item.separator = args.space_delim elseif item.delimiter == "~" then item.separator = args.tilde_delim else item.separator = args.comma_delim end end local function get_horizontal_separator(args_horiz, embedded_comma) return args_horiz == "bullet" and " · " or embedded_comma and "; " or ", " end -- Suppress false positives in categories like [[Category:English links with redundant wikilinks]] so people won't -- be tempted to "correct" them; terms like embedded ~ like [[Micros~1]] or embedded comma not followed by a space -- such as [[1,6-Cleves acid]] need to have a link around them to avoid the tilde or comma being interpreted as a -- delimiter. local function suppress_redundant_wikilink_cat(term, _alt) return term:find("~") or term:find(",%S") end local function full_link_and_track_self_links(item, face, nochecktr) if item.term then local pagename = mw.loadData(headword_data_module).pagename local term_is_pagename = item.term == pagename local term_contains_pagename = item.term:find("%[%[" .. m_str_utils.pattern_escape(pagename) .. "[|%]]") if term_is_pagename or term_contains_pagename then local current_L2 = require(pages_module).get_current_L2() if current_L2 then local current_L2_lang = require(languages_module).getByCanonicalName(current_L2) if current_L2_lang and current_L2_lang:getCode() == item.lang:getCode() then if term_is_pagename then track("term-is-pagename") else track("term-contains-pagename") end end end end end item.suppress_redundant_wikilink_cat = suppress_redundant_wikilink_cat item.never_call_transliteration_module = nochecktr return require(links_module).full_link(item, face) end local function format_subitem(subitem, lang, face, compute_embedded_comma, nochecktr) local embedded_comma = false local text if subitem.term and term_already_linked(subitem.term) then text = subitem.term if compute_embedded_comma then embedded_comma = not not require(utilities_module).get_plaintext(text):find(",") end else text = full_link_and_track_self_links(subitem, face, nochecktr) if compute_embedded_comma then -- We don't check decoration text for commas as it's inside parens or displayed elsewhere. local subitem_plaintext = subitem.alt or subitem.term if subitem_plaintext then embedded_comma = not not subitem_plaintext:find(",") end end end -- We could use the `show_decorations` field to full_link() but not when term_already_linked(). if subitem.q and subitem.q[1] or subitem.qq and subitem.qq[1] or subitem.l and subitem.l[1] or subitem.ll and subitem.ll[1] or subitem.refs and subitem.refs[1] then text = require(decorations_module).format_decorations { lang = subitem.lang or lang, text = text, q = subitem.q, qq = subitem.qq, l = subitem.l, ll = subitem.ll, refs = subitem.refs, } end return text, embedded_comma end function export.format_item(item, args, face) local compute_embedded_comma = args.horiz == "comma" local embedded_comma = false local nochecktr = args.noautotr if type(item) == "table" then if item.terms then local parts = {} local is_first = true for _, subitem in ipairs(item.terms) do if subitem == false then -- omitted subitem; do nothing else local separator = subitem.separator or not is_first and (args.subitem_separator or ", ") if separator then if compute_embedded_comma then embedded_comma = embedded_comma or not not separator:find(",") end insert(parts, separator) end local formatted, this_embedded_comma = format_subitem(subitem, args.lang, face, compute_embedded_comma, nochecktr) embedded_comma = embedded_comma or this_embedded_comma insert(parts, formatted) is_first = false end end return concat(parts), embedded_comma else return format_subitem(item, args.lang, face, compute_embedded_comma, nochecktr) end else if compute_embedded_comma then embedded_comma = not not require(utilities_module).get_plaintext(item):find(",") end if args.lang and not term_already_linked(item) then return full_link_and_track_self_links({lang = args.lang, term = item, sc = args.sc}, face, nochecktr), embedded_comma else return item, embedded_comma end end end function export.construct_old_style_header(header, horiz) local old_style_header local function ib_colon() return tostring(html("span"):addClass("ib-colon"):addClass("ib-content"):wikitext(":")) end if horiz then old_style_header = require(qualifier_module).format_qualifiers { qualifiers = header, open = false, close = false, } .. ib_colon() .. " " else old_style_header = require(qualifier_module).format_qualifiers { qualifiers = header } .. ib_colon() old_style_header = tostring(html("div"):wikitext(old_style_header)) end return old_style_header end -- Construct the sort base of a single term. As a hack, sort appendices after mainspace items. local function term_sortbase(val) if not val then -- This should not normally happen. return u(0x10FFFF) elseif val:find("^%[*Appendix:") then return u(0x10FFFE) .. val else return val end end -- Construct the sort base of a single item, using the display form preferentially, otherwise the term itself. -- As a hack, sort appendices after mainspace items. local function item_sortbase(item) return term_sortbase(item.alt or item.term) end local function make_sortbase(item) if item == false then return "*" -- doesn't matter, will be omitted in create_list() elseif type(item) == "table" then if item.terms then -- Optimize for the common case of only a single term if item.terms[2] then local parts = {} -- multiple terms local first = true for _, subitem in ipairs(item.terms) do if subitem ~= false then if not first then insert(parts, ", ") end insert(parts, item_sortbase(subitem)) first = false end end if parts[1] then return concat(parts) end else local subitem = item.terms[1] if subitem ~= false then return item_sortbase(subitem) end end return "*" -- doesn't matter, entire group will be omitted in create_list() else return item_sortbase(item) end else return item end end local function make_node_sortbase(node) return make_sortbase(node.item) end -- Sort a sublist of `list` in place, keeping the first `keepfirst` and last `keeplast` items fixed. -- `lang` is the language of the items and `make_sortbase` creates the appropriate sort base. local function sort_sublist(list, lang, make_sortbase_fn, keepfirst, keeplast) if keepfirst == 0 and keeplast == 0 then require(collation_module).sort(list, lang, make_sortbase_fn) else local sublist = {} for i = keepfirst + 1, #list - keeplast do sublist[i - keepfirst] = list[i] end require(collation_module).sort(sublist, lang, make_sortbase_fn) for i = keepfirst + 1, #list - keeplast do list[i] = sublist[i - keepfirst] end end end --[=[ Unused but could be useful in the future -- URL-encode only the characters that serve as template delimiters (left and right brace, vertical bar, equal sign -- and percent sign since it's the escape character). local function bot_url_encode(txt) return (txt:gsub("[%%|{}=&]", {["%"] = "%25", ["|"] = "%7C", ["{"] = "%7B", ["}"] = "%7D", ["="] = "%3D", ["&"] = "%26"})) end ]=] -- Reverse the action of bot_url_encode(). local function bot_url_decode(txt) return (txt:gsub("%%7([BCD])", {B = "{", C = "|", D = "}"}):gsub("%%3D", "="):gsub("%%26", "&"):gsub("%%25", "%%")) end --[==[ Bot-callable function to generate a number of sortkeys simultaneously. {{para|1}} contains the langcode, and remaining numeric parameters contain "bot-URL-encoded" strings whose sort keys will be computed and returned as a JSON array. Here, "bot-URL-encoded" means that the six characters `{ | } = & %` should be converted to their URL-encoded representation (respectively `%7B %7C %7D %3D %26 %25`), and will be decoded appropriately before computing the sortkey. ]==] function export.make_sortkey(frame) local iparams = { [1] = {type = "language"}, [2] = {list = true}, } local iargs = require(parameters_module).process(frame.args, iparams) local make_sortkey = require(collation_module).make_lang_sortkey_function(iargs[1], term_sortbase) local retval = {} for _, arg in ipairs(iargs[2]) do arg = bot_url_decode(arg) insert(retval, make_sortkey(arg)) end return require(JSON_module).toJSON(retval) end local large_text_scripts = { ["Arab"] = true, ["Beng"] = true, ["Deva"] = true, ["Gujr"] = true, ["Guru"] = true, ["Hebr"] = true, ["Khmr"] = true, ["Knda"] = true, ["Laoo"] = true, ["Mlym"] = true, ["Mong"] = true, ["Mymr"] = true, ["Orya"] = true, ["Sinh"] = true, ["Syrc"] = true, ["Taml"] = true, ["Telu"] = true, ["Tfng"] = true, ["Thai"] = true, ["Tibt"] = true, } --[==[ Format a list of items using HTML. `args` is an object specifying the items to add and related properties, with the following fields: * `content`: A list of the items to format. See below for the format of the items. * `lang`: The language object of the items to format, if the items in `content` are strings. * `sc`: The script object of the items to format, if the items in `content` are strings. * `raw`: If true, return the list raw, without any collapsing or columns. * `class`: The CSS class of the surrounding {<div>}. * `column_count`: Number of columns to format the list into. * `alphabetize`: If true, sort the items in the table. * `collapse`: If true, make the table partially collapsed by default, with a "Show more" button at the bottom. * `toggle_category`: Value of `data-toggle-category` property grouping collapsible elements. * `header`: If specified, Wikicode to prepend to the output. * `title_new_style`: If true, the header is treated as a title and displayed in a new style. This is ignored if `horiz` is non-nil. * `subitem_separator`: Separator used between subitems when multiple subitems occur on a line, if not specified in the subitem itself (using the `separator` field). Defaults to {", "}. * `keepfirst`: If > 0, keep this many rows unsorted at the beginning of the top level. * `keeplast`: If > 0, keep this many rows unsorted at the end of the top level. * `horiz`: If non-nil, format the items horizontally. If the value is "bullet", put a center dot/bullet (·) between items. If the value is "comma", put a comma between items (but if there is an embedded comma in any item, put a semicolon between all items). Each item in `content` is in one of the following formats: * A string. This is for compatibility and should not be used by new callers. * An object describing an item to format, in the format expected by full_link() in [[Module:links]], including decorations (left or right qualifiers, left or right labels, or references). * An object describing a list of subitems to format, displayed side-by-side, separated by a comma or other separator. This format is identified by the presence of a key `terms` specifying the list of subitems. Each subitem is in the same format as for a single top-level item, except that it should also have a `separator` field specifying the separator to display before each item (which will typically be a blank string before the first item). ]==] function export.create_list(args) if type(args) ~= "table" then error("expected table, got " .. type(args)) end local column_count = args.column_count or 1 local toggle_category = args.toggle_category or "kata terbitan" local keepfirst = args.keepfirst or 0 local keeplast = args.keeplast or 0 if keepfirst > 0 then track("keepfirst") end if keeplast > 0 then track("keeplast") end -- maybe construct old-style header local old_style_header = nil if args.header and (args.horiz or not args.title_new_style) then old_style_header = export.construct_old_style_header(args.header, args.horiz) end if args.horiz then old_style_header = "* " .. (old_style_header or "") end local list local any_extra_indented_item = false for _, item in ipairs(args.content) do if item == false then -- do nothing elseif type(item) == "table" and item.extra_indent and item.extra_indent > 0 then any_extra_indented_item = true break end end -- If any extra indented item, convert the items to a nested structure, which is necessary both for sorting and -- for converting to HTML. if any_extra_indented_item then local function make_node(item) return { item = item } end local root_node = make_node(nil) local node_stack = {root_node} local last_indent = 0 local function append_subnode(node, subnode) if not node.subnodes then node.subnodes = {} end insert(node.subnodes, subnode) end for i, item in ipairs(args.content) do if item == false then -- do nothing else local this_indent if type(item) ~= "table" then this_indent = 1 else this_indent = (item.extra_indent or 0) + 1 end local node = make_node(item) if this_indent == last_indent then append_subnode(node_stack[#node_stack], node) elseif this_indent > last_indent + 1 then error(("Element #%s (%s) has indent %s, which is more than one greater than the previous item with indent %s"):format( i, make_sortbase(item), this_indent, last_indent)) elseif this_indent > last_indent then -- Start a new sublist attached to the last item of the sublist one level up; but we need special -- handling for the root node (last_indent == 0). if last_indent > 0 then local subnodes = node_stack[#node_stack].subnodes if not subnodes then error(("Internal error: Not first item and no subnodes at preceding level %s: %s"):format( #node_stack, dump(node_stack))) end insert(node_stack, subnodes[#subnodes]) end append_subnode(node_stack[#node_stack], node) last_indent = this_indent else while last_indent > this_indent do local finished_node = table.remove(node_stack) if args.alphabetize then require(collation_module).sort(finished_node.subnodes, args.lang, make_node_sortbase) end last_indent = last_indent - 1 end append_subnode(node_stack[#node_stack], node) end end end if args.alphabetize then while node_stack[1] do local finished_node = table.remove(node_stack) if node_stack[1] then -- We're sorting something other than the root node. require(collation_module).sort(finished_node.subnodes, args.lang, make_node_sortbase) else -- We're sorting the root node; honor `keepfirst` and `keeplast`. sort_sublist(finished_node.subnodes, args.lang, make_node_sortbase, keepfirst, keeplast) end end end local function format_node(node, depth) local sublist local embedded_comma = false if node.subnodes then if args.horiz then sublist = {} else sublist = html("ul") end local prevnode = nil for _, subnode in ipairs(node.subnodes) do local thisnode, this_embedded_comma = format_node(subnode, depth + 1) embedded_comma = embedded_comma or this_embedded_comma if not prevnode or not args.alphabetize or not deepEquals(prevnode, thisnode) then if args.horiz then table.insert(sublist, thisnode) else sublist = sublist:node(thisnode) end prevnode = thisnode end end if args.horiz then sublist = table.concat(sublist, get_horizontal_separator(args.horiz, embedded_comma)) end end if not node.item then -- At the root. return sublist, embedded_comma end local formatted, listitem -- Ignore embedded commas in subitems inside of parens or square brackets. formatted, embedded_comma = export.format_item(node.item, args) if args.horiz then listitem = formatted if sublist then -- Use parens for the first, third, fifth, etc. sublists and square brackets for the remainder. if depth % 2 == 1 then listitem = ("%s (%s)"):format(listitem, sublist) else listitem = ("%s [%s]"):format(listitem, sublist) end end else listitem = html("li"):wikitext(formatted) if sublist then listitem = listitem:node(sublist) end end return listitem, embedded_comma end list = format_node(root_node, 0) else if args.alphabetize then sort_sublist(args.content, args.lang, make_sortbase, keepfirst, keeplast) end if args.horiz then list = {} else list = html("ul") end local previtem = nil local embedded_comma = false for _, item in ipairs(args.content) do if item == false then -- omitted item; do nothing else local thisitem, this_embedded_comma = export.format_item(item, args) embedded_comma = embedded_comma or this_embedded_comma if not previtem or not args.alphabetize or previtem ~= thisitem then if args.horiz then table.insert(list, thisitem) else list = list:node(html("li"):wikitext(thisitem)) end previtem = thisitem end end end if args.horiz then list = table.concat(list, get_horizontal_separator(args.horiz, embedded_comma)) end end local output if args.horiz then output = list else output = html("div"):addClass("term-list"):node(list) if args.class then output:addClass(args.class) end if not args.raw then output:addClass("ul-column-count") :attr("data-column-count", column_count) if args.collapse then output = html("div") :node(output) :addClass("list-switcher") :attr("data-toggle-category", toggle_category) -- identify commonly used scripts that use large text and -- provide a special CSS class to make the template bigger local sc = args.sc if sc == nil then local scripts = args.lang:getScripts() if #scripts > 0 then sc = scripts[1] end end if sc ~= nil then local scriptcode = sc:getParentCode() if scriptcode == "top" then scriptcode = sc:getCode() end if large_text_scripts[scriptcode] then output:addClass("list-switcher-large-text") end end end end if args.collapse or args.title_new_style then -- wrap in wrapper to prevent interference from floating elements local list_switcher_wrapper = html("div") :addClass("list-switcher-wrapper") if args.title_new_style then list_switcher_wrapper :node( html("div") :addClass("list-switcher-header") :wikitext(args.header) ) end list_switcher_wrapper:node(output) output = list_switcher_wrapper end output = tostring(output) end return (old_style_header or "") .. output end -- This function is for compatibility with earlier version of [[Module:columns]] -- (now found in [[Module:columns/old]]). function export.create_table(...) -- Earlier arguments to create_table: -- n_columns, content, alphabetize, bg, collapse, class, title, column_width, line_start, lang local args = {} args.column_count, args.content, args.alphabetize, args.collapse, args.class, args.header, args.column_width, args.line_start, args.lang = ... return export.create_list(args) end function export.display_from(frame_args, parent_args, frame) local boolean = {type = "boolean"} local iparams = { ["class"] = true, -- Default for auto-collapse. Overridable by template |collapse= param. ["collapse"] = boolean, -- If specified, this specifies the number of columns, and no columns parameter is available on the template. -- Otherwise, the columns parameter is named |n=. ["columns"] = {type = "number"}, -- If specified, this specifies the default language code, which can be overridden using |lang= in the template. -- Otherwise, the language-code parameter is required and normally found in |1=, but for compatibility can be -- specified as |lang= (which leads to deprecation handling). ["lang"] = {type = "language"}, -- Default for auto-sort. Overridable by template |sort= param. ["sort"] = boolean, ["toggle_category"] = true, -- Minimum number of rows required to format into a multicolumn list. If below this, the list is displayed "raw" -- (no columns, no collapsbility). ["minrows"] = {type = "number", default = 5}, -- Disables automatic transliteration; entries without a manual transliteration will have none at all. -- Used on large pages, especially Chinese ones, because zh-translit works by fetching and parsing -- the target of the page, which is a performance killer on large pages with potentially thousands -- of link targets. -- Note: noautotr also disables redundant transliteration checks. ["noautotr"] = boolean, } local iargs = require(parameters_module).process(frame_args, iparams) local langcode_in_lang = iargs.lang or parent_args.lang local lang_param = langcode_in_lang and "lang" or 1 local deprecated = not iargs.lang and langcode_in_lang local ret = export.handle_display_from_or_topic_list(iargs, parent_args, nil) return deprecated and frame:expandTemplate{title = "check deprecated lang param usage", -- FIXME: Accessing undefined global var args = {ret, lang = args[lang_param]}} or ret end --[==[ Implement `display_from()` [the internal entry point for {{tl|col}} and variants, which enter originally through `display()`] as well as regular (column-oriented) topic lists, invoked through [[Module:topic list]]. `iargs` are the invocation args of {{tl|col}}, and `raw_item_args` are the arguments specifying the values of each row as well as other properties, corresponding to the user-specified template arguments of {{tl|col}}. Note that `show()` in [[Module:topic list]] is normally invoked directly by a topic list template, whose invocation arguments are passed in using `raw_item_args` and are similar to the template arguments of {{tl|col}}. `iargs` for topic-list invocations is hard-coded, and template arguments to a topic-list template are processed in [[Module:topic list]] itself. Note that the handling of topic lists is currently implemented almost entirely through callbacks in `topic_list_data` (which is nil if we're processing {{tl|col}} rather than a topic list) in an attempt to reduce the coupling and keep the topic-list-specific code in [[Module:topic list]], but IMO the coupling is still too tight. Probably the control structure should be reversed and the following function split up into subfunctions, which are invoked as needed by {{tl|col}} and/or [[Module:topic list]]. ]==] function export.handle_display_from_or_topic_list(iargs, raw_item_args, topic_list_data) local boolean = {type = "boolean"} local langcode_in_lang = iargs.lang or raw_item_args.lang local lang_param = langcode_in_lang and "lang" or 1 local first_content_param = langcode_in_lang and 1 or 2 local params = { [lang_param] = {required = not iargs.lang, type = "language", template_default = not iargs.lang and "und" or nil}, ["n"] = not iargs.columns and {type = "number"} or nil, [first_content_param] = {list = true, allow_holes = true}, ["title"] = {}, ["collapse"] = boolean, ["sort"] = boolean, ["sc"] = {type = "script"}, -- used when calling from [[Module:saurus]] so the page displaying the synonyms/antonyms doesn't occur in the -- list ["omit"] = {list = true}, ["keepfirst"] = {type = "number", default = 0}, ["keeplast"] = {type = "number", default = 0}, ["horiz"] = {}, ["notr"] = boolean, ["noautotr"] = boolean, ["allow_space_delim"] = boolean, ["tilde_delim"] = {}, ["space_delim"] = {}, ["comma_delim"] = {}, } if topic_list_data then topic_list_data.add_topic_list_params(params) end local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {default = true, require_index = true}, {group = "link"}, -- sc has separate_no_index = true; that's the only one -- It makes no sense to have overall l=, ll=, q= or qq= params for columnar display. {group = {"ref", "l", "q"}, require_index = true}, } m_param_utils.augment_params_with_modifiers(params, param_mods) local processed_args = require(parameters_module).process(raw_item_args, params) local horiz = processed_args.horiz if horiz and horiz ~= "comma" and horiz ~= "bullet" then horiz = require(yesno_module)(horiz) if horiz == nil then error(("Unrecognized value |horiz=%s; should be 'comma', 'bullet' or a recognized Boolean value such " .. "as 'yes' or '1' (same as 'bullet') or 'no' or '0'"):format(processed_args.horiz)) end if horiz == true then horiz = "bullet" end processed_args.horiz = horiz end -- If default argument values specified, set them after parsing the caller-specified arguments in `raw_item_args`. if topic_list_data then topic_list_data.set_default_arguments(processed_args) end -- Now set defaults for the various delimiters, depending in some cases on whether horiz was set. -- We can't set these defaults (even regardless of their dependency on horiz=) in `local params` above -- because we want any defaults specified in `default_props` to override these. if not processed_args.tilde_delim then local tilde_with_abbr = '<abbr title="near equivalent">~</abbr>' processed_args.tilde_delim = processed_args.horiz and tilde_with_abbr or " " .. tilde_with_abbr .. " " end if not processed_args.space_delim then processed_args.space_delim = "&nbsp;" end if not processed_args.comma_delim then processed_args.comma_delim = processed_args.horiz and "/" or ", " end -- Check for extra term indent. Do this before calling parse_list_with_inline_modifiers_and_separate_params() -- because sometimes space is a delimiter and the space in the indent will confuse things and get interpreted as a -- delimiter. local extra_indent_by_termno = {} local termargs = processed_args[first_content_param] for i = 1, termargs.maxindex do local term = termargs[i] if term then local extra_indent, actual_term = rmatch(term, "^(%*+)%s+(.-)$") if extra_indent then termargs[i] = actual_term extra_indent_by_termno[i] = #extra_indent end end end local groups, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { param_mods = param_mods, processed_args = processed_args, termarg = first_content_param, parse_lang_prefix = true, allow_multiple_lang_prefixes = true, disallow_custom_separators = true, track_module = "columns", lang = iargs.lang or lang_param, sc = "sc.default", splitchar = processed_args.allow_space_delim and "[,~ ]" or "[,~]", no_show_decorations = true, -- since we handle them ourselves in format_subitem() } local lang = iargs.lang or args[lang_param] local langcode = lang:getCode() local sc = args.sc.default local sort = iargs.sort if args.sort ~= nil then if not args.sort then track("nosort") end sort = args.sort else -- HACK! For Japanese-script languages (Japanese, Okinawan, Miyako, etc.), sorting doesn't yet work properly, so -- disable it. for _, langsc in ipairs(lang:getScriptCodes()) do if langsc == "Jpan" then sort = false break end end end local collapse = iargs.collapse if args.collapse ~= nil then if not args.collapse then track("nocollapse") end collapse = args.collapse end local title = args.title local formatted_cats if topic_list_data then title, formatted_cats = topic_list_data.get_title_and_formatted_cats(args, lang, sc, topic_list_data) end local number_of_groups = 0 for i, group in ipairs(groups) do local number_of_items = 0 group.extra_indent = extra_indent_by_termno[group.orig_index] for j, item in ipairs(group.terms) do convert_delimiter_to_separator(item, j, args) if args.notr then item.tr = "-" elseif args.noautotr then item.tr = item.tr or "-" end -- If a separate language code was given for the term, display the language name as a right qualifier. -- (Briefly we made them labels but this leads to non-obvious behavior e.g. "French" becoming "France" under -- some circumstances.) Otherwise it may not be obvious that the term is in a separate language (e.g. if the -- main language is 'zh' and the term language is a Chinese lect such as Min Nan). But don't do this for -- Translingual terms, which are often added to the list of English and other-language terms. if item.termlangs then local qqs = {} for _, termlang in ipairs(item.termlangs) do local termlangcode = termlang:getCode() if termlangcode ~= langcode and termlangcode ~= "mul" then insert(qqs, termlang:getCanonicalName()) end end if item.qq then for _, qq in ipairs(item.qq) do insert(qqs, qq) end end item.qq = qqs end local omitted = false for _, omitted_item in ipairs(args.omit) do if omitted_item == item.term then omitted = true break end end if omitted then -- signal create_list() to omit this item group.terms[j] = false else number_of_items = number_of_items + 1 end end if number_of_items == 0 then -- omit the whole group groups[i] = false else number_of_groups = number_of_groups + 1 end end local column_count = iargs.columns or args.n -- FIXME: This needs a total rewrite. if column_count == nil then column_count = number_of_groups <= 3 and 1 or number_of_groups <= 9 and 2 or number_of_groups <= 27 and 3 or number_of_groups <= 81 and 4 or 5 end local raw = number_of_groups < iargs.minrows local horiz_edit_button if topic_list_data and args.horiz then -- append edit button to title horiz_edit_button = topic_list_data.make_horiz_edit_button(topic_list_data.topic_list_template) end return export.create_list { column_count = column_count, raw = raw, content = groups, alphabetize = sort, header = title, title_new_style = (title ~= nil and title ~= ''), collapse = collapse, toggle_category = iargs.toggle_category, -- columns-bg (in [[MediaWiki:Gadget-Site.css]]) provides the background color class = (iargs.class and iargs.class .. " columns-bg" or "columns-bg"), lang = lang, sc = sc, subitem_separator = ", ", keepfirst = args.keepfirst, keeplast = args.keeplast, horiz = args.horiz, noautotr = args.noautotr, } .. (horiz_edit_button or "") .. (formatted_cats or "") end function export.display(frame) if not is_substing() then return export.display_from(frame.args, frame:getParent().args, frame, false) end -- If substed, unsubst template with newlines between each term, redundant wikilinks removed, and remove duplicates + sort terms if sort is enabled. local m_table = require("Module:table") local m_template_parser = require("Module:template parser") local parent = frame:getParent() local elems = m_table.shallowCopy(parent.args) local code = remove(elems, 1) code = code and trim(code) local lang = require("Module:languages").getByCode(code, 1) local i = 1 while true do local elem = elems[i] while elem do elem = trim(elem, "%s") if elem ~= "" then break end remove(elems, i) elem = elems[i] end if not elem then break elseif not ( -- Strip redundant wikilinks. not elem:match("^()%[%[") or elem:find("[[", 3, true) or elem:find("]]", 3, true) ~= #elem - 1 or elem:find("|", 3, true) ) then elem = sub(elem, 3, -3) elem = trim(elem, "%s") end elems[i] = elem .. "\n" i = i + 1 end -- If sort is enabled, remove duplicates then sort elements. if require("Module:yesno")(frame.args.sort) then elems = m_table.removeDuplicates(elems) require("Module:collation").sort(elems, lang) end -- Readd the langcode. insert(elems, 1, code .. "\n") -- TODO: Place non-numbered parameters after 1 and before 2. local template = m_template_parser.getTemplateInvocationName(mw.title.new(parent:getTitle())) return "{{" .. concat(m_template_parser.buildTemplate(template, elems), "|") .. "}}" end return export ayrhhchd4gj67psdazy9iw6oj4i10ai Modul:ca-headword 828 10361 375387 223898 2026-09-22T07:28:03Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714447|92714447]]) 375387 Scribunto text/plain local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local require_when_needed = require("Module:utilities/require when needed") local m_table = require("Module:table") local com = require("Module:ca-common") local ca_IPA_module = "Module:ca-IPA" local ca_verb_module = "Module:ca-verb" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local headword_utilities_module = "Module:headword utilities" local inflection_utilities_module = "Module:inflection utilities" local parse_utilities_module = "Module:parse utilities" local romut_module = "Module:romance utilities" local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local m_string_utilities = require_when_needed("Module:string utilities") local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local lang = require("Module:languages").getByCode("ca") local langname = lang:getCanonicalName() local list_to_text = mw.text.listToText local insert = table.insert local concat = table.concat local rfind = m_string_utilities.find local rmatch = m_string_utilities.match local rsplit = m_string_utilities.split local usub = m_string_utilities.sub local rsub = com.rsub local function track(page) require("Module:debug/track")("ca-headword/" .. page) return true end local list_param = {list = true, disallow_holes = true} local boolean_param = {type = "boolean"} ----------------------------------------------------------------------------------------- -- Main entry point -- ----------------------------------------------------------------------------------------- -- The main entry point. -- This is the only function that can be invoked from a template. function export.show(frame) local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.") local params = { ["head"] = list_param, ["id"] = true, ["splithyph"] = boolean_param, ["nolinkhead"] = boolean_param, ["json"] = boolean_param, ["pagename"] = true, -- for testing } if pos_functions[poscat] then for key, val in pairs(pos_functions[poscat].params) do params[key] = val end end local args = require("Module:parameters").process(frame:getParent().args, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local user_specified_heads = args.head local heads = user_specified_heads if args.nolinkhead then if #heads == 0 then heads = {pagename} end else local romut = require(romut_module) local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph) if #heads == 0 then heads = {auto_linked_head} else for i, head in ipairs(heads) do if head:find("^~") then head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2)) heads[i] = head end if head == auto_linked_head then track("redundant-head") end end end end local data = { lang = lang, pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat, categories = {}, heads = heads, user_specified_heads = user_specified_heads, no_redundant_head_cat = #user_specified_heads == 0, genders = {}, inflections = {}, pagename = pagename, id = args.id, force_cat_output = force_cat, checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true, } if pagename:find("^%-") and poscat ~= "bentuk akhiran" then data.is_suffix = true data.pos_category = "suffixes" data.checkredlinks = true local singular_poscat = require(en_utilities_module).singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end if pos_functions[poscat] then pos_functions[poscat].func(args, data) end if args.json then return require("Module:JSON").toJSON(data) end local post_note = data.post_note and "; " .. data.post_note or "" return require("Module:headword").full_headword(data) .. post_note end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local function replace_hash_with_lemma(term, lemma) -- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture replace -- expression. lemma = lemma:gsub("%%", "%%%%") -- Assign to a variable to discard second return value. term = term:gsub("#", lemma) return term end -- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come -- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given; -- `accel` is the accelerator form, or nil. local function parse_and_insert_inflection(data, args, field, label, accel) m_headword_utilities.parse_and_insert_inflection { headdata = data, forms = args[field], paramname = field, splitchar = ",", label = label, accel = accel and {form = accel} or nil, } end -- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result. -- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list -- into which the plurals are inserted (which inherit their decorations from `plobj`). local function insert_defpls(defpls, plobj, dest) if not defpls then -- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing. return end if #defpls == 1 then plobj.term = defpls[1] insert(dest, plobj) else for _, defpl in ipairs(defpls) do local newplobj = m_table.shallowCopy(plobj) newplobj.term = defpl insert(dest, newplobj) end end end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function do_adjective(args, data, is_superlative) local feminines = {} local masculine_plurals = {} local feminine_plurals = {} -- Use "participle" not "past participle" for categories such as 'invariable paticiples' local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) if args.sp then local romut = require(romut_module) if not romut.allowed_special_indicators[args.sp] then local indicators = {} for indic, _ in pairs(romut.allowed_special_indicators) do insert(indicators, "'" .. indic .. "'") end table.sort(indicators) error("Special inflection indicator beginning can only be " .. list_to_text(indicators) .. ": " .. args.sp) end end local lemma = data.pagename local function fetch_inflections(field) local retval = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", } if not retval[1] then return {{term = "+"}} end return retval end local function insert_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if args.f[1] == "ind" or args.f[1] == "inv" then -- invariable adjective insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) if args.sp or args.f[2] or args.pl[1] or args.mpl[1] or args.fpl[1] then error("Can't specify inflections with an invariable " .. category_pos) end elseif args.fonly then -- feminine-only if args.f[1] then error("Can't specify explicit feminines with feminine-only " .. category_pos) end if args.pl[1] then error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=") end if args.mpl[1] then error("Can't specify explicit masculine plurals with feminine-only " .. category_pos) end local argsfpl = fetch_inflections("fpl") for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- Generate default feminine plural. local defpls = com.make_plural(lemma, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, fpl, feminine_plurals) else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end insert(data.inflections, {label = "feminine-only"}) insert_inflection(feminine_plurals, "feminine plural", "f|p") else -- Gather feminines. for _, f in ipairs(fetch_inflections("f")) do if f.term == "mf" then f.term = lemma elseif f.term == "+" then -- Generate default feminine. f.term = com.make_feminine(lemma, args.sp) else f.term = replace_hash_with_lemma(f.term, lemma) end insert(feminines, f) end local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and not m_headword_utilities.termobj_has_decorations(feminines[1]) if fem_like_lemma then insert(data.categories, langname .. " epicene " .. category_plpos) end local mpl_field = "mpl" local fpl_field = "fpl" if args.pl[1] then if args.mpl[1] or args.fpl[1] then error("Can't specify both pl= and mpl=/fpl=") end mpl_field = "pl" fpl_field = "pl" end local argsmpl = fetch_inflections(mpl_field) local argsfpl = fetch_inflections(fpl_field) for _, mpl in ipairs(argsmpl) do if mpl.term == "+" then -- Generate default masculine plural. local defpls -- First, some special hacks based on the feminine singular. if not fem_like_lemma and not args.sp and not lemma:find(" ") then for _, f in ipairs(feminines) do if f.term:find("ssa$") then -- If the feminine ends in -ssa, assume that the -ss- is also in the -- masculine plural form defpls = {rsub(f.term, "a$", "os")} break elseif f.term == lemma .. "na" then defpls = {lemma .. "ns"} break elseif lemma:find("ig$") and f.term:find("ja$") then -- Adjectives in -ig have two masculine plural forms, one derived from -- the m.sg. and the other derived from the f.sg. defpls = {lemma .. "s", rsub(f.term, "ja$", "jos")} break end end end defpls = defpls or com.make_plural(lemma, "m", args.sp) if not defpls then error("Unable to generate default plural of '" .. lemma .. "'") end insert_defpls(defpls, mpl, masculine_plurals) else mpl.term = replace_hash_with_lemma(mpl.term, lemma) insert(masculine_plurals, mpl) end end for _, fpl in ipairs(argsfpl) do if fpl.term == "+" then -- First, some special hacks based on the feminine singular. if fem_like_lemma and not args.sp and not lemma:find(" ") and lemma:find("[çx]$") then -- Adjectives ending in -ç or -x behave as mf-type in the singular, but -- regular type in the plural. local defpls = com.make_plural(lemma .. "a", "f") if not defpls then error("Unable to generate default plural of '" .. lemma .. "a'") end insert_defpls(defpls, fpl, feminine_plurals) else for _, f in ipairs(feminines) do -- Generate default feminine plural; f is a table. local defpls = com.make_plural(f.term, "f", args.sp) if not defpls then error("Unable to generate default plural of '" .. f.term .. "'") end for _, defpl in ipairs(defpls) do local fplobj = m_table.shallowCopy(fpl) fplobj.term = defpl m_headword_utilities.combine_termobj_decorations(fplobj, f) insert(feminine_plurals, fplobj) end end end else fpl.term = replace_hash_with_lemma(fpl.term, lemma) insert(feminine_plurals, fpl) end end local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and m_table.deepEquals(masculine_plurals, feminine_plurals) local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and not m_headword_utilities.termobj_has_decorations(masculine_plurals[1]) if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then -- actually invariable insert(data.inflections, {label = glossary_link("invariable")}) insert(data.categories, langname .. " indeclinable " .. category_plpos) else -- Make sure there are feminines given and not same as lemma. if not fem_like_lemma then insert_inflection(feminines, "feminine", "f|s") elseif args.gneut then data.genders = {"gneut"} else data.genders = {"mf"} end if fem_pl_like_masc_pl then if args.gneut then insert_inflection(masculine_plurals, "plural", "p") else insert_inflection(masculine_plurals, "masculine and feminine plural", "p") end else insert_inflection(masculine_plurals, "masculine plural", "m|p") insert_inflection(feminine_plurals, "feminine plural", "f|p") end end end parse_and_insert_inflection(data, args, "comp", "comparative") parse_and_insert_inflection(data, args, "sup", "superlative") parse_and_insert_inflection(data, args, "dim", "diminutive") parse_and_insert_inflection(data, args, "aug", "augmentative") if args.irreg and is_superlative then insert(data.categories, langname .. " irregular superlative " .. category_plpos) end end local function get_adjective_params(adjtype) local params = { ["sp"] = true, -- special indicator: "first", "first-last", etc. ["f"] = list_param, --feminine form(s) [1] = {alias_of = "f", list = false}, ["pl"] = list_param, --plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["fpl"] = list_param, --feminine plural override(s) } if adjtype == "base" then params["comp"] = list_param --comparative(s) params["sup"] = list_param --superlative(s) params["dim"] = list_param --diminutive(s) params["aug"] = list_param --augmentative(s) params["fonly"] = boolean_param -- feminine only params["hascomp"] = {} -- has comparative end if adjtype == "sup" then params["irreg"] = boolean_param end return params end -- Display additional inflection information for an adjective pos_functions["adjectives"] = { params = get_adjective_params("base"), func = do_adjective, } pos_functions["past participles"] = { params = get_adjective_params("part"), func = do_adjective, redlink_pos = "participles", } pos_functions["determiners"] = { params = get_adjective_params("det"), func = do_adjective, } pos_functions["pronouns"] = { params = get_adjective_params("pron"), func = do_adjective, } ----------------------------------------------------------------------------------------- -- Nouns -- ----------------------------------------------------------------------------------------- local allowed_genders = m_table.listToSet( {"m", "f", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"} ) local function validate_genders(genders) for _, g in ipairs(genders) do if type(g) == "table" then g = g.spec end if not allowed_genders[g] then error("Unrecognized gender: " .. g) end end end local function do_noun(args, data, is_proper) local is_plurale_tantum = false local has_singular = false local category_plpos = data.checkredlinks if category_plpos == true then category_plpos = data.pos_category end local category_pos = m_en_utilities.singularize(category_plpos) validate_genders(args[1]) data.genders = args[1] local saw_m = false local saw_f = false local saw_gneut = false local gender_for_irreg_ending, gender_for_default_plural -- Check for specific genders and pluralia tantum. for _, g in ipairs(args[1]) do if type(g) == "table" then g = g.spec end if g:find("-p$") then is_plurale_tantum = true else has_singular = true if g == "m" or g == "mf" or g == "mfbysense" then saw_m = true end if g == "f" or g == "mf" or g == "mfbysense" then saw_f = true end if g == "gneut" then saw_gneut = true end end end if saw_m and saw_f then gender_for_irreg_ending = "mf" elseif saw_f then gender_for_irreg_ending = "f" else gender_for_irreg_ending = "m" end gender_for_default_plural = saw_gneut and "gneut" or gender_for_irreg_ending == "mf" and "m" or gender_for_irreg_ending local lemma = data.pagename -- Plural local plurals = {} local function insert_noun_inflection(terms, label, accel) m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, } end if is_plurale_tantum and not has_singular then if args[2][1] then error("Can't specify plurals of plurale tantum " .. category_pos) end insert(data.inflections, {label = glossary_link("plural only")}) else plurals = m_headword_utilities.parse_term_list_with_modifiers { paramname = {2, "pl"}, forms = args[2], splitchar = ",", } -- Check for special plural signals local mode = nil local pl1 = plurals[1] if pl1 and #pl1.term == 1 then mode = pl1.term if mode == "?" or mode == "!" or mode == "-" or mode == "~" then pl1.term = nil if next(pl1) then error(("Can't specify inline modifiers with plural code '%s'"):format(mode)) end table.remove(plurals, 1) -- Remove the mode parameter elseif mode ~= "+" and mode ~= "#" then error(("Unexpected plural code '%s'"):format(mode)) end end if is_plurale_tantum then -- both singular and plural insert(data.inflections, {label = "sometimes " .. glossary_link("plural only") .. ", in variation"}) end if mode == "?" then -- Plural is unknown insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals") elseif mode == "!" then -- Plural is not attested insert(data.inflections, {label = "plural not attested"}) insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals") if plurals[1] then error("Can't specify any plurals along with unattested plural code '!'") end elseif mode == "-" then -- Uncountable noun; may occasionally have a plural insert(data.categories, langname .. " uncountable " .. category_plpos) -- If plural forms were given explicitly, then show "usually" if plurals[1] then insert(data.inflections, {label = "usually " .. glossary_link("uncountable")}) insert(data.categories, langname .. " countable " .. category_plpos) else insert(data.inflections, {label = glossary_link("uncountable")}) end else -- Countable or mixed countable/uncountable if not plurals[1] and not is_proper then plurals[1] = {term = "+"} end if mode == "~" then -- Mixed countable/uncountable noun, always has a plural insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")}) insert(data.categories, langname .. " uncountable " .. category_plpos) insert(data.categories, langname .. " countable " .. category_plpos) elseif plurals[1] then -- Countable nouns insert(data.categories, langname .. " countable " .. category_plpos) else -- Uncountable nouns insert(data.categories, langname .. " uncountable " .. category_plpos) end end -- Gather plurals, handling requests for default plurals. local has_default_or_hash = false for _, pl in ipairs(plurals) do if pl.term:find("^%+") or pl.term:find("#") then has_default_or_hash = true break end end if has_default_or_hash then local newpls = {} for _, pl in ipairs(plurals) do if pl.term == "+" then local default_pls = com.make_plural(lemma, gender_for_default_plural) insert_defpls(default_pls, pl, newpls) elseif pl.term:find("^%+") then pl.term = require(romut_module).get_special_indicator(pl.term) local default_pls = com.make_plural(lemma, gender_for_default_plural, pl.term) insert_defpls(default_pls, pl, newpls) else pl.term = replace_hash_with_lemma(pl.term, lemma) insert(newpls, pl) end end plurals = newpls end local pl1 = plurals[1] if pl1 and not plurals[2] and pl1.term == lemma then insert(data.inflections, {label = glossary_link("invariable"), q = pl1.q, qq = pl1.qq, l = pl1.l, ll = pl1.ll, refs = pl1.refs }) insert(data.categories, langname .. " indeclinable " .. category_plpos) else insert_noun_inflection(plurals, "plural", "p") end if plurals[2] then insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals") end end -- Gather masculines/feminines. For each one, generate the corresponding plural. `field` is the name of the field -- containing the masculine or feminine forms (normally "m" or "f"); `inflect` is a function of one or two arguments -- to generate the default masculine or feminine from the lemma (the arguments are the lemma and optionally a -- "special" flag to indicate how to handle multiword lemmas, and the function is normally make_feminine or -- make_masculine from [[Module:ca-common]]); and `default_plurals` is a list into which the corresponding default -- plurals of the gathered or generated masculine or feminine forms are stored. local function handle_mf(field, inflect, default_plurals) local function call_inflect(special) if inflect then -- Generate default feminine. return inflect(lemma, special) else -- FIXME error("Can't generate default masculine currently") end end local mfs = m_headword_utilities.parse_term_list_with_modifiers { paramname = field, forms = args[field], splitchar = ",", frob = function(term) if term == "+" then -- Generate default masculine/feminine. term = call_inflect() else term = replace_hash_with_lemma(term, lemma) end local special = require(romut_module).get_special_indicator(term) if special then term = call_inflect(special) end return term end } for _, mf in ipairs(mfs) do local mfpls = com.make_plural(mf.term, gender, special) if mfpls then for _, mfpl in ipairs(mfpls) do local plobj = m_table.shallowCopy(mf) plobj.term = mfpl -- Add an accelerator for each masculine/feminine plural whose lemma -- is the corresponding singular, so that the accelerated entry -- that is generated has a definition that looks like -- # {{plural of|ca|MFSING}} plobj.accel = {form = "p", lemma = mf.term} table.insert(default_plurals, plobj) end end end return mfs end local feminine_plurals = {} local feminines = handle_mf("f", com.make_feminine, feminine_plurals) local masculine_plurals = {} local masculines = handle_mf("m", com.make_masculine, masculine_plurals) local function handle_mf_plural(mfplfield, default_plurals, singulars) local mfpl = m_headword_utilities.parse_term_list_with_modifiers { paramname = mfplfield, forms = args[mfplfield], splitchar = ",", } local new_mfpls = {} local saw_plus for i, mfpl in ipairs(mfpl) do local accel if #mfpl == #singulars then -- If same number of overriding masculine/feminine plurals as singulars, assume each plural goes with -- the corresponding singular and use each corresponding singular as the lemma in the accelerator. The -- generated entry will have -- # {{plural of|ca|SINGULAR}} -- as the definition. accel = {form = "p", lemma = singulars[i].term} else accel = nil end if mfpl.term == "+" then -- We should never see + twice. If we do, it will lead to problems since we overwrite the values of -- default_plurals the first time around. if saw_plus then error(("Saw + twice when handling %s="):format(mfplfield)) end saw_plus = true if not default_plurals[1] then -- FIXME: Can this happen? Not in corresponding Spanish code and the old Portuguese code tried to -- handle this condition by generating the default plural from the lemma. error("Internal error: Something wrong, no generated default m/f plurals at this stage") end for _, defpl in ipairs(default_plurals) do -- defpl is already a table and has an accel field m_headword_utilities.combine_termobj_decorations(defpl, mfpl) insert(new_mfpls, defpl) end elseif mfpl.term:find("^%+") then mfpl.term = require(romut_module).get_special_indicator(mfpl.term) for _, mf in ipairs(singulars) do local default_mfpls = com.make_plural(mf.term, gender, mfpl.term) for _, defp in ipairs(default_mfpls) do local mfplobj = m_table.shallowCopy(mfpl) mfplobj.term = defp mfplobj.accel = accel m_headword_utilities.combine_termobj_decorations(mfplobj, mf) insert(new_mfpls, mfplobj) end end else mfpl.accel = accel mfpl.term = replace_hash_with_lemma(mfpl.term, lemma) insert(new_mfpls, mfpl) end end return new_mfpls end if args.fpl[1] then -- Override any existing feminine plurals. feminine_plurals = handle_mf_plural("fpl", feminine_plurals, feminines) end if args.mpl[1] then -- Override any existing masculine plurals. masculine_plurals = handle_mf_plural("mpl", masculine_plurals, masculines) end local function parse_and_insert_noun_inflection(field, label, accel) parse_and_insert_inflection(data, args, field, label, accel) end insert_noun_inflection(feminines, "feminine", "f") insert_noun_inflection(feminine_plurals, "feminine plural") insert_noun_inflection(masculines, "masculine") insert_noun_inflection(masculine_plurals, "masculine plural") parse_and_insert_noun_inflection("dim", "diminutive") parse_and_insert_noun_inflection("aug", "augmentative") parse_and_insert_noun_inflection("pej", "pejorative") parse_and_insert_noun_inflection("dem", "demonym") parse_and_insert_noun_inflection("fdem", "female demonym") -- Is this a noun with an unexpected ending (for its gender)? -- Only check if the term is one word (there are no spaces in the term). local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word if (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf") and irreg_gender_lemma:find("a$") then insert(data.categories, langname .. " masculine " .. category_plpos .. " ending in -a") elseif (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf") and not ( irreg_gender_lemma:find("a$") or irreg_gender_lemma:find("ió$") or irreg_gender_lemma:find("tat$") or irreg_gender_lemma:find("tud$") or irreg_gender_lemma:find("[dt]riu$")) then insert(data.categories, langname .. " feminine " .. category_plpos .. " with no feminine ending") end end local function get_noun_params(is_proper) return { [1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders", flatten = true}, -- gender(s) [2] = {list = "pl", disallow_holes = true}, --plural override(s) ["f"] = list_param, --feminine form(s) ["m"] = list_param, --masculine form(s) ["fpl"] = list_param, --feminine plural override(s) ["mpl"] = list_param, --masculine plural override(s) ["dim"] = list_param, --diminutive(s) ["aug"] = list_param, --diminutive(s) ["pej"] = list_param, --pejorative(s) ["dem"] = list_param, --demonym(s) ["fdem"] = list_param, --female demonym(s) } end pos_functions["Kata nama"] = { params = get_noun_params(), func = do_noun, } pos_functions["Kata nama khas"] = { params = get_noun_params("is proper"), func = function(args, data) do_noun(args, data, "is proper") end, } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["Kata kerja"] = { params = { [1] = true, ["pres"] = list_param, --present ["pres_qual"] = {list = "pres\1_qual", allow_holes = true}, ["pres3s"] = list_param, --third-singular present ["pres3s_qual"] = {list = "pres3s\1_qual", allow_holes = true}, ["pret"] = list_param, --preterite ["pret_qual"] = {list = "pret\1_qual", allow_holes = true}, ["part"] = list_param, --participle ["part_qual"] = {list = "part\1_qual", allow_holes = true}, ["short_part"] = list_param, --short participle ["short_part_qual"] = {list = "short_part\1_qual", allow_holes = true}, ["noautolinktext"] = boolean_param, ["noautolinkverb"] = boolean_param, ["attn"] = boolean_param, ["pres_1_sg"] = true, -- accept any ignore old-style param ["past_part"] = true, -- accept any ignore old-style param ["root"] = true, -- FIXME: Implement root-stressed vowel quality }, func = function(args, data, tracking_categories, frame) local preses, preses_3s, prets, parts, short_parts if args.attn then insert(tracking_categories, "Requests for attention concerning " .. langname) return end local ca_verb = require(ca_verb_module) local alternant_multiword_spec = ca_verb.do_generate_forms(args, "ca-verb", data.heads[1]) local specforms = alternant_multiword_spec.forms local function slot_exists(slot) return specforms[slot] and #specforms[slot] > 0 end local function do_finite(slot_tense, label_tense) -- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs); -- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]). local has_1s = slot_exists(slot_tense .. "_1s") local has_3s = slot_exists(slot_tense .. "_3s") local has_3p = slot_exists(slot_tense .. "_3p") if has_1s or (not has_3s and not has_3p) then return { slot = slot_tense .. "_1s", label = ("first-person singular %s"):format(label_tense), }, true elseif has_3s then return { slot = slot_tense .. "_3s", label = ("third-person singular %s"):format(label_tense), }, false else return { slot = slot_tense .. "_3p", label = ("third-person plural %s"):format(label_tense), }, false end end local did_pres_1s preses, did_pres_1s = do_finite("pres", "present") preses_3s = { slot = "pres_3s", label = "third-person singular present", } prets = do_finite("pret", "preterite") parts = { slot = "pp_ms", label = "past participle", } short_parts = { slot = "short_pp_ms", label = "short past participle", } if args.pres[1] or args.pres3s[1] or args.pret[1] or args.part[1] or args.short_part[1] then track("verb-old-multiarg") end local function strip_brackets(qualifiers) if not qualifiers then return nil end local stripped_qualifiers = {} for _, qualifier in ipairs(qualifiers) do local stripped_qualifier = qualifier:match("^%[(.*)%]$") if not stripped_qualifier then error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier) end insert(stripped_qualifiers, stripped_qualifier) end return stripped_qualifiers end local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty) local forms local to_insert if #args == 0 then forms = specforms[slot_desc.slot] if not forms or #forms == 0 then if skip_if_empty then return end forms = {{form = "-"}} end elseif #args == 1 and args[1] == "-" then forms = {{form = "-"}} else forms = {} for i, arg in ipairs(args) do local qual = qualifiers[i] if qual then -- FIXME: It's annoying we have to add brackets and strip them out later. The inflection -- code adds all footnotes with brackets around them; we should change this. qual = {"[" .. qual .. "]"} end local form = arg if not args.noautolinkverb then -- [[Module:inflection utilities]] already loaded by [[Module:ca-verb]] form = require(inflection_utilities_module).add_links(form) end insert(forms, {form = form, footnotes = qual}) end end if forms[1].form == "-" then to_insert = {label = "no " .. slot_desc.label} else local into_table = {label = slot_desc.label} for _, form in ipairs(forms) do local qualifiers = strip_brackets(form.footnotes) -- Strip redundant brackets surrounding entire form. These may get generated e.g. -- if we use the angle bracket notation with a single word. local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form -- Don't include accelerators if brackets remain in form, as the result will be wrong. -- FIXME: For now, don't include accelerators. We should use the new {{ca-verb form of}}. -- local this_accel = not stripped_form:find("%[%[") and accel or nil local this_accel = nil insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel}) end to_insert = into_table end insert(data.inflections, to_insert) end local skip_pres_if_empty if alternant_multiword_spec.no_pres1_and_sub then insert(data.inflections, {label = "no first-person singular present"}) insert(data.inflections, {label = "no present subjunctive"}) end if alternant_multiword_spec.no_pres_stressed then insert(data.inflections, {label = "no stressed present indicative or subjunctive"}) skip_pres_if_empty = true end if alternant_multiword_spec.only3s then insert(data.inflections, {label = glossary_link("impersonal")}) elseif alternant_multiword_spec.only3sp then insert(data.inflections, {label = "third-person only"}) elseif alternant_multiword_spec.only3p then insert(data.inflections, {label = "third-person plural only"}) end local has_vowel_alt if alternant_multiword_spec.vowel_alt then for _, vowel_alt in ipairs(alternant_multiword_spec.vowel_alt) do if vowel_alt ~= "+" and vowel_alt ~= "í" and vowel_alt ~= "ú" then has_vowel_alt = true break end end end do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty) -- We want to include both the pres_1s and pres_3s if there is a vowel alternation in the present singular. But we -- don't want to redundantly include the pres_3s if we already included it. if did_pres_1s and has_vowel_alt then do_verb_form(args.pres3s, args.pres3s_qual, preses_3s, skip_pres_if_empty) end do_verb_form(args.pret, args.pret_qual, prets) do_verb_form(args.part, args.part_qual, parts) do_verb_form(args.short_part, args.short_part_qual, short_parts, "skip if empty") -- Add categories. for _, cat in ipairs(alternant_multiword_spec.categories) do insert(data.categories, cat) end -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:ca-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Catalan equivalent). if #data.user_specified_heads == 0 or ( #data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end if args.root then local m_ca_IPA = require(ca_IPA_module) local parsed_respellings = {} local function set_parsed_respelling(dialect, parsed) -- Validate the individual root vowel specs. for _, termobj in ipairs(parsed.terms) do if not rfind(termobj.words[1].term, "^" .. m_ca_IPA.mid_vowel_hint_c .. "$") then error(("Root vowel spec '%s' should be one of the vowels %s"):format( termobj.words[1].term, m_ca_IPA.mid_vowel_hints)) end end if not dialect then for _, dial in ipairs(m_ca_IPA.dialects) do -- Need to clone as we destructively modify each one later with the pronun. parsed_respellings[dial] = m_table.deepCopy(parsed) end elseif m_ca_IPA.dialect_groups[dialect] then for _, dial in ipairs(m_ca_IPA.dialect_groups[dialect]) do -- Need to clone as we destructively modify each one later with the pronun. parsed_respellings[dial] = m_table.deepCopy(parsed) end else parsed_respellings[dialect] = parsed end end local function check_dialect_or_dialect_group(dialect) if not m_table.contains(m_ca_IPA.dialects, dialect) and not m_ca_IPA.dialect_groups[dialect] then local dialect_list = {} for _, dial in ipairs(m_ca_IPA.dialects) do insert(dialect_list, "'" .. dial .. "'") end dialect_list = list_to_text(dialect_list, nil, " or ") local dialect_group_list = {} for dialect_group, _ in pairs(m_ca_IPA.dialect_groups) do insert(dialect_group_list, "'" .. dialect_group .. "'") end dialect_group_list = list_to_text(dialect_group_list, nil, " or ") error(("Unrecognized dialect '%s': Should be a dialect %s or a dialect group %s"):format( dialect, dialect_list, dialect_group_list)) end end -- Parse the root vowel specs. if args.root:find("[<%[]") then local put = require(parse_utilities_module) -- Parse balanced segment runs involving either [...] (substitution notation) or <...> (inline -- modifiers). We do this because we don't want commas or semicolons inside of square or angle brackets -- to count as respelling delimiters. However, we need to rejoin square-bracketed segments with nearby -- ones after splitting alternating runs on comma and semicolon. local segments = put.parse_multi_delimiter_balanced_segment_run(args.root, {{"<", ">"}, {"[", "]"}}) local semicolon_separated_groups = put.split_alternating_runs(segments, "%s*;%s*") for _, group in ipairs(semicolon_separated_groups) do local first_element = group[1] local dialect if first_element:find("^[a-z]+:") then -- a dialect-specific spec local rest dialect, rest = first_element:match("^([a-z]+):(.*)$") check_dialect_or_dialect_group(dialect) group[1] = rest end local comma_separated_groups = put.split_alternating_runs_on_comma(group) -- Process each value. local outer_container = m_ca_IPA.parse_comma_separated_groups(comma_separated_groups, true, args.root, "root") set_parsed_respelling(dialect, outer_container) end else for _, dialect_spec in ipairs(rsplit(args.root, "%s*;%s*")) do local dialect if dialect_spec:find("^[a-z]+:") then -- a dialect-specific spec local rest dialect, rest = dialect_spec:match("^([a-z]+):(.*)$") check_dialect_or_dialect_group(dialect) dialect_spec = rest end local termobjs = {} for _, word in ipairs(rsplit(dialect_spec, ",")) do insert(termobjs, {words = {{term = word}}}) end set_parsed_respelling(dialect, { terms = termobjs, }) end end -- Convert each canonicalized respelling to phonemic/phonetic IPA. m_ca_IPA.generate_phonemic_phonetic(parsed_respellings) -- Group the results. local grouped_pronuns = m_ca_IPA.group_pronuns_by_dialect(parsed_respellings) -- Format for display. for _, grouped_pronun_spec in pairs(grouped_pronuns) do local pronunciations = {} local function ins(text) insert(pronunciations, text) end -- Loop through each pronunciation. For each one, format the phonetic version "raw". for j, pronun in ipairs(grouped_pronun_spec.pronuns) do -- Add dialect tags to left accent qualifiers if first one local as = pronun.a if j == 1 then if as then as = m_table.deepCopy(as) else as = {} end for _, dialect in ipairs(grouped_pronun_spec.dialects) do insert(as, m_ca_IPA.dialects_to_names[dialect]) end else ins(", ") end local slash_pron = "/" .. pronun.phonetic:gsub("ˈ", "") .. "/" if as or pronun.q or pronun.qq or pronun.aa then ins(require(decorations_module).format_decorations { lang = lang, text = slash_pron, q = pronun.q, a = as, qq = pronun.qq, aa = pronun.aa }) else ins(slash_pron) end if pronun.refs then -- FIXME: Copied from [[Module:IPA]]. Should be in a module. local refs = {} if #pronun.refs > 0 then for _, refspec in ipairs(pronun.refs) do if type(refspec) ~= "table" then refspec = {text = refspec} end local refargs if refspec.name or refspec.group then refargs = {name = refspec.name, group = refspec.group} end insert(refs, mw.getCurrentFrame():extensionTag("ref", refspec.text, refargs)) end ins(concat(refs)) end end end grouped_pronun_spec.formatted = concat(pronunciations) end -- Concatenate formatted results. local formatted = {} for _, grouped_pronun_spec in ipairs(grouped_pronuns) do insert(formatted, grouped_pronun_spec.formatted) end data.post_note = "''root stress'': " .. concat(formatted, "; ") end end } ----------------------------------------------------------------------------------------- -- Numerals -- ----------------------------------------------------------------------------------------- -- Display additional inflection information for a numeral pos_functions["Kata bilangan"] = { params = { [1] = true, [2] = true, }, func = function(args, data) if args[1] then insert(data.genders, "m") parse_and_insert_inflection(data, args, 1, "feminine") parse_and_insert_inflection(data, args, 2, "noun form") else insert(data.genders, "m") insert(data.genders, "f") end end } ----------------------------------------------------------------------------------------- -- Phrases -- ----------------------------------------------------------------------------------------- pos_functions["frasa"] = { params = { ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, ["m"] = list_param, ["f"] = list_param, }, func = function(args, data) validate_genders(args.g) data.genders = args.g parse_and_insert_inflection(data, args, "m", "masculine") parse_and_insert_inflection(data, args, "f", "feminine") end, } ----------------------------------------------------------------------------------------- -- Suffix forms -- ----------------------------------------------------------------------------------------- pos_functions["bentuk akhiran"] = { params = { [1] = {required = true, list = true, disallow_holes = true}, ["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true}, }, func = function(args, data) validate_genders(args.g) data.genders = args.g local suffix_type = {} for _, typ in ipairs(args[1]) do insert(suffix_type, typ .. "-forming suffix") end insert(data.inflections, {label = "non-lemma form of " .. m_table.serialCommaJoin(suffix_type, {conj = "or"})}) end, } return export oqhw4n4h2xa5p5vi46bv3kr1mv8txpb Modul:affix 828 10384 375358 373580 2026-09-22T03:14:21Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708528|92708528]]) 375358 Scribunto text/plain local export = {} local debug_force_cat = false -- if set to true, always display categories even on userspace pages local m_links = require("Module:links") local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local etymology_module = "Module:etymology" local scripts_module = "Module:scripts" local utilities_module = "Module:utilities" -- Export this so the category code in [[Module:category tree/etymology]] can access it. export.affix_lang_data_module_prefix = "Module:affix/lang-data/" local ulen = m_str_utils.len local rfind = m_str_utils.find local rmatch = m_str_utils.match local pluralize = require(en_utilities_module).pluralize local singularize = require(en_utilities_module).singularize local u = m_str_utils.char local ucfirst = m_str_utils.ucfirst local unpack = unpack or table.unpack -- Lua 5.2 compatibility function export.affix_variants(canonical, variants) local mappings = {} for _, variant in ipairs(variants) do mappings[variant] = canonical end return mappings end function export.id_mapping(default, ids) local mapping = { default = default } if ids then for id, target in pairs(ids) do mapping[id] = target end end return mapping end function export.id_mapping_with_affix_variants(base, id_variants) local mappings = {} for id, variants in pairs(id_variants) do for _, variant in ipairs(variants) do mappings[variant] = export.id_mapping(base, {[id] = base}) end end return mappings end function export.merge_tables(...) local result = {} for i = 1, select('#', ...) do local t = select(i, ...) if t then for k, v in pairs(t) do result[k] = v end end end return result end -- Export this so the category code in [[Module:category tree/etymology]] can access it. export.langs_with_lang_specific_data = { ["az"] = true, ["fi"] = true, ["fr"] = true, ["izh"] = true, ["la"] = true, ["sah"] = true, ["tr"] = true, ["trk-pro"] = true, } local default_pos = "perkataan" local function pluralize_pos(pos) return pluralize(singularize(pos)) end --[==[ intro: ===About different types of hyphens ("template", "display" and "lookup"):=== * The "template hyphen" is the per-script hyphen character that is used in template calls to indicate that a term is an affix. This is always a single Unicode char, but there may be multiple possible hyphens for a given script. Normally this is just the regular hyphen character "-", but for some non-Latin-script languages (currently only right-to-left languages), it is different. * The "display hyphen" is the string (which might be an empty string) that is added onto a term as displayed and linked, to indicate that a term is an affix. Currently this is always either the same as the template hyphen or an empty string, but the code below is written generally enough to handle arbitrary display hyphens. Specifically: *# For East Asian languages, the display hyphen is always blank. *# For Arabic-script languages, either tatweel (ـ) or ZWNJ (zero-width non-joiner) are allowed as template hyphens, where ZWNJ is supported primarily for Farsi, because some suffixes have non-joining behavior. The display hyphen corresponding to tatweel is also tatweel, but the display hyphen corresponding to ZWNJ is blank (tatweel is also the default display hyphen, for calls to {{tl|prefix}}/{{tl|suffix}}/etc. that don't include an explicit hyphen). * The "lookup hyphen" is the hyphen that is used when looking up language-specific affix mappings. (These mappings are discussed in more detail below when discussing link affixes.) It depends only on the script of the affix in question. Most scripts (including East Asian scripts) use a regular hyphen "-" as the lookup hyphen, but Hebrew and Arabic have their own lookup hyphens (respectively maqqef and tatweel). Note that for Arabic in particular, there are three possible template hyphens that are recognized (tatweel, ZWNJ and regular hyphen), but mappings must use tatweel. ===About different types of affixes ("template", "display", "link", "lookup" and "category"):=== * A "template affix" is an affix in its source form as it appears in a template call. Generally, a template affix has an attached template hyphen (see above) to indicate that it is an affix and indicate what type of affix it is (prefix, suffix, interfix or circumfix), but some of the older-style templates such as {{tl|suffix}}, {{tl|prefix}}, {{tl|confix}}, etc. have "positional" affixes where the presence of the affix in a certain position (e.g. the second or third parameter) indicates that it is a certain type of affix, whether or not it has an attached template hyphen. * A "display affix" is the corresponding affix as it is actually displayed to the user. The display affix may differ from the template affix for various reasons: *# The display affix may be specified explicitly using the {{para|alt<var>N</var>}} parameter, the `<alt:...>` inline modifier or a piped link of the form e.g. `<nowiki>[[-kas|-käs]]</nowiki>` (here indicating that the affix should display as `-käs` but be linked as `-kas`). Here, the template affix is arguably the entire piped link, while the display affix is `-käs`. *# Even in the absence of {{para|alt<var>N</var>}} parameters, `<alt:...>` inline modifiers and piped links, certain languages have differences between the "template hyphen" specified in the template (which always needs to be specified somehow or other in templates like {{tl|affix}}, to indicate that the term is an affix and what type of affix it is) and the display hyphen (see above), with corresponding differences between template and display affixes. * A (regular) "link affix" is the affix that is linked to when the affix is shown to the user. The link affix is usually the same as the display affix, but will differ in one of three circumstances: *# The display and link affixes are explicitly made different using {{para|alt<var>N</var>}} parameters, `<alt:...>` inline modifiers or piped links, as described above under "display affix". *# For certain languages, certain affixes are mapped to canonical form using language-specific mappings. For example, in Finnish, the adjective-forming suffix {{m|fi|-kas}} appears as {{m|fi|-käs}} after front vowels, but logically both forms are the same suffix and should be linked and categorized the same. Similarly, in Latin, the negative and intensive prefixes spelled {{m|la|in-}} (etymologically two distinct prefixes) appear variously as {{m|la|il-}}, {{m|la|im-}} or {{m|la|ir-}} before certain consonants. Mappings are supplied in [[Module:affix/lang-data/LANGCODE]] to convert Finnish {{m|fi|-käs}} to {{m|fi|-kas}} for linking and categorization purposes. Note that the affixes in the mappings use "lookup hyphens" to indicate the different types of affixes, which is usually the same as the template hyphen but differs for Arabic scripts, because there are multiple possible template hyphens recognized but only one lookup hyphen (tatweel). The form of the affix as used to look up in the mapping tables is called the "lookup affix"; see below. * A "stripped link affix" is a link affix that has been passed through the language's `stripDiacritics()` function, which may strip certain diacritics: e.g. macrons in Latin and Old English (indicating length); acute and grave accents in Russian and various other Slavic languages (indicating stress); vowel diacritics in most Arabic-script languages; and also tatweel in some Arabic-script languages (currently, for example, Persian, Arabic and Urdu strip tatweel, but Ottoman Turkish does not). Stripped link affixes are currently what are used in category names. * A "lookup affix" is the form of the affix as it is looked up in the language-specific lookup mappings described above under link affixes. There are actually two lookup stages: *# First, the affix is looked up in a modified display form (specifically, the same as the display affix but using lookup hyphens). Note that this lookup does not occur if an explicit display form is given using {{para|alt<var>N</var>}} or an `<alt:...>` inline modifier, or if the template affix contains a piped or embedded link. *# If no entry is found, the affix is then looked up in a modified link form (specifically, the modified display form passed through the language's `stripDiacritics()` function, which strips out certain diacritics, but with the lookup hyphen re-added if it was stripped out, as in the case of tatweel in many Arabic-script languages). The reason for this double lookup procedure is to allow for mappings that are sensitive to the extra diacritics, but also allow for mappings that are not sensitive in this fashion (e.g. Russian {{m|ru|-ливый}} occurs both stressed and unstressed, but is the same prefix either way). * A "category affix" is the affix as it appears in categories such as [[:Category:Finnish terms suffixed with -kas| Category:Finnish terms suffixed with ''-kas'']]. The category affix is currently always the same as the stripped link affix. This means that for Arabic-script languages, it may or may not have a tatweel, even if the correponding display affix and regular link affix have a tatweel. As mentioned above, stripDiacritics() strips tatweel for Arabic, Persian and Urdu, but not for Ottoman Turkish. Hence affix categories for Arabic, Persian and Urdu will be missing the tatweel, but affix categories for Ottoman Turkish will have it. An additional complication is that if the template affix contains a ZWNJ, the display (and hence the link and category affixes) will have no hyphen attached in any case. ]==] ----------------------------------------------------------------------------------------- -- Template and display hyphens -- ----------------------------------------------------------------------------------------- --[=[ Per-script template hyphens. The template hyphen is what appears in the {{affix}}/{{prefix}}/{{suffix}}/etc. template (in the wikicode). See above. They key below is a script code, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong' and 'xwo-Mong' will match 'Mong'. The value below is a string consisting of one or more hyphen characters. If there is more than one character, the default hyphen must come last and a non-default function must be specified for the script in display_hyphens[] so the correct display hyphen will be specified when no template hyphen is given (in {{suffix}}/{{prefix}}/etc.). Script detection is normally done when linking, but we need to do it earlier. However, under most circumstances we don't need to do script detection. Specifically, we only need to do script detection for a given language if (a) the language has multiple scripts; and (b) at least one of those scripts is listed below or in display_hyphens. ]=] local ZWNJ = u(0x200C) -- zero-width non-joiner local template_hyphens = { -- This covers all Arabic scripts. See above. ["Arab"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen ["Aran"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen ["Hebr"] = "־", -- Hebrew-specific hyphen termed "maqqef" ["Mong"] = "᠊", -- FIXME! What about the following right-to-left scripts? -- Adlm (Adlam) -- Armi (Imperial Aramaic) -- Avst (Avestan) -- Cprt (Cypriot) -- Khar (Kharoshthi) -- Mand (Mandaic/Mandaean) -- Mani (Manichaean) -- Mend (Mende/Mende Kikakui) -- Narb (Old North Arabian) -- Nbat (Nabataean/Nabatean) -- Nkoo (N'Ko) -- Orkh (Orkhon runes) -- Phli (Inscriptional Pahlavi) -- Phlp (Psalter Pahlavi) -- Phlv (Book Pahlavi) -- Phnx (Phoenician) -- Prti (Inscriptional Parthian) -- Rohg (Hanifi Rohingya) -- Samr (Samaritan) -- Sarb (Old South Arabian) -- Sogd (Sogdian) -- Sogo (Old Sogdian) -- Syrc (Syriac) -- Thaa (Thaana) } -- Hyphens used when looking up an affix in a lang-specific affix mapping. Defaults to regular hyphen (-). The keys -- are script codes, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong' and 'xwo-Mong' -- will match 'Mong'. The value should be a single character. local lookup_hyphens = { ["Hebr"] = "־", -- This covers all Arabic scripts. See above. ["Arab"] = "ـ", ["Aran"] = "ـ", } -- Default display-hyphen function. local function default_display_hyphen(script, hyph) if not hyph then return template_hyphens[script] or "-" end return hyph end local function arab_get_display_hyphen(_script, hyph) if not hyph then return "ـ" -- tatweel elseif hyph == ZWNJ then return "" else return hyph end end local function no_display_hyphen(_script, _hyph) return "" end -- Per-script function to return the correct display hyphen given the script and template hyphen. The function should -- also handle the case where the passed-in template hyphen is nil, corresponding to the situation in -- {{prefix}}/{{suffix}}/etc. where no template hyphen is specified. The key is the script code after removing a hyphen -- and anything preceding, so 'mnc-Mong', 'xwo-Mong' etc. will match 'Mong'. local display_hyphens = { -- This covers all Arabic scripts. See above. ["Arab"] = arab_get_display_hyphen, ["Aran"] = arab_get_display_hyphen, ["Bopo"] = no_display_hyphen, ["Hani"] = no_display_hyphen, ["Hans"] = no_display_hyphen, ["Hant"] = no_display_hyphen, -- The following is a mixture of several scripts. Hopefully the specs here are correct! ["Jpan"] = no_display_hyphen, ["Jurc"] = no_display_hyphen, ["Kitl"] = no_display_hyphen, ["Kits"] = no_display_hyphen, ["Laoo"] = no_display_hyphen, ["Nshu"] = no_display_hyphen, ["Shui"] = no_display_hyphen, ["Tang"] = no_display_hyphen, ["Thaa"] = no_display_hyphen, ["Thai"] = no_display_hyphen, ["Tibt"] = no_display_hyphen, } ----------------------------------------------------------------------------------------- -- Basic Utility functions -- ----------------------------------------------------------------------------------------- local function glossary_link(entry, text) text = text or entry return "[[Lampiran:Glosari#" .. entry .. "|" .. text .. "]]" end local function track(page) if type(page) == "table" then for i, pg in ipairs(page) do page[i] = "affix/" .. pg end else page = "affix/" .. page end require("Module:debug/track")(page) end local function ine(val) return val ~= "" and val or nil end ----------------------------------------------------------------------------------------- -- Compound types -- ----------------------------------------------------------------------------------------- local function make_compound_type(typ, alttext) return { text = "kata majmuk " .. glossary_link(typ, alttext), cat = "Kata majmuk " .. typ, } end -- Make a compound type entry with a simple rather than glossary link. -- These should be replaced with a glossary link when the entry in the glossary -- is created. local function make_non_glossary_compound_type(typ, alttext) local link = alttext and "[[" .. typ .. "|" .. alttext .. "]]" or "[[" .. typ .. "]]" return { text = "kata majmuk " .. link, cat = "Kata majmuk " .. typ, } end local function make_raw_compound_type(typ, alttext) return { text = glossary_link(typ, alttext), cat = pluralize(typ), } end local function make_borrowing_type(typ, alttext) return { text = glossary_link(typ, alttext), borrowing_type = pluralize(typ), } end export.etymology_types = { ["adapted borrowing"] = make_borrowing_type("pinjaman tersuai"), ["adap"] = "pinjaman tersuai", ["abor"] = "pinjaman tersuai", ["alliterative"] = make_non_glossary_compound_type("alliterative"), ["allit"] = "aliteratif", ["antonymous"] = make_non_glossary_compound_type("antonim"), ["ant"] = "antonim", ["bahuvrihi"] = make_compound_type("bahuvrihi", "bahuvrīhi"), ["bahu"] = "bahuvrihi", ["bv"] = "bahuvrihi", ["coordinative"] = make_compound_type("koordinatif"), ["coord"] = "coordinative", ["descriptive"] = make_compound_type("deskriptif"), ["desc"] = "deskriptif", ["determinative"] = make_compound_type("determinative"), ["det"] = "determinative", ["dvandva"] = make_compound_type("dvandva"), ["dva"] = "dvandva", ["dvigu"] = make_compound_type("dvigu"), ["dvi"] = "dvigu", ["endocentric"] = make_compound_type("endosentrik"), ["endo"] = "endocentric", ["exocentric"] = make_compound_type("eksosentrik"), ["exo"] = "exocentric", ["izafet I"] = make_compound_type("izafet I"), ["iz1"] = "izafet I", ["izafet II"] = make_compound_type("izafet II"), ["iz2"] = "izafet II", ["izafet III"] = make_compound_type("izafet III"), ["iz3"] = "izafet III", ["karmadharaya"] = make_compound_type("karmadharaya", "karmadhāraya"), ["karma"] = "karmadharaya", ["kd"] = "karmadharaya", ["kenning"] = make_raw_compound_type("kenning"), ["ken"] = "kenning", ["rhyming"] = make_non_glossary_compound_type("berima"), ["rhy"] = "berima", ["synonymous"] = make_non_glossary_compound_type("sinonim"), ["syn"] = "sinonim", ["tatpurusa"] = make_compound_type("tatpurusa", "tatpuruṣa"), ["tat"] = "tatpurusa", ["tp"] = "tatpurusa", } local function process_etymology_type(typ, nocap, notext, has_parts) local text_sections = {} local categories = {} local borrowing_type if typ then local typdata = export.etymology_types[typ] if type(typdata) == "string" then typdata = export.etymology_types[typdata] end if not typdata then error("Internal error: Unrecognized type '" .. typ .. "'") end local text = typdata.text if not nocap then text = ucfirst(text) end local cat = typdata.cat borrowing_type = typdata.borrowing_type local oftext = typdata.oftext or " bagi" if not notext then table.insert(text_sections, text) if has_parts then table.insert(text_sections, oftext) table.insert(text_sections, " ") end end if cat then table.insert(categories, cat) end end return text_sections, categories, borrowing_type end ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- -- Iterate an array up to the greatest integer index found. local function ipairs_with_gaps(t) local indices = m_table.numKeys(t) local max_index = #indices > 0 and math.max(unpack(indices)) or 0 local i = 0 return function() if i < max_index then i = i + 1 return i, t[i] end end end export.ipairs_with_gaps = ipairs_with_gaps --[==[ Join formatted parts (in `parts_formatted`) together with any overall {{para|lit}} spec (in `lit`) plus categories, which are formatted by prepending the language name as found in `lang`. The value of an entry in `categories` can be either a string (which is formatted using `sort_key`) or a table of the form `{ {cat=<var>category</var>, sort_key=<var>sort_key</var>, sort_base=<var>sort_base</var>}`, specifying the sort key and sort base to use when formatting the category. If `nocat` is given, no categories are added; otherwise, `force_cat` causes categories to be added even on userspace pages. ]==] function export.join_formatted_parts(data) local cattext local lang = data.data.lang local force_cat = data.data.force_cat or debug_force_cat if data.data.nocat then cattext = "" else for i, cat in ipairs(data.categories) do if type(cat) == "table" then data.categories[i] = require(utilities_module).format_categories(cat.cat .. " bahasa " .. lang:getFullName(), lang, cat.sort_key, cat.sort_base, force_cat) else data.categories[i] = require(utilities_module).format_categories(cat .. " bahasa " .. lang:getFullName(), lang, data.data.sort_key, nil, force_cat) end end cattext = table.concat(data.categories) end local result = table.concat(data.parts_formatted, not data.separator_already_added and " +&lrm; " or nil) .. (data.data.lit and ", secara harfiah " .. m_links.mark(data.data.lit, "gloss") or "") local q = data.data.q local qq = data.data.qq local l = data.data.l local ll = data.data.ll if q and q[1] or qq and qq[1] or l and l[1] or ll and ll[1] then result = require(decorations_module).format_decorations { lang = lang, text = result, q = q, qq = qq, l = l, ll = ll, } end return result .. cattext end -- Remove links and call lang:stripDiacritics(term). local function strip_diacritics_no_links(lang, term) return lang:stripDiacritics(m_links.remove_links(term)) end --[=[ Convert a raw part as passed into an entry point into a part ready for linking. `lang` and `sc` are the overall language and script objects. This uses the overall language and script objects as defaults for the part and parses off any fragment from the term. We need to do the latter so that fragments don't end up in categories and so that we correctly do affix mapping even in the presence of fragments. ]=] local function canonicalize_part(part, lang, sc) if not part then return end -- Save the original (user-specified, part-specific) value of `lang`. If such a value is specified, we don't insert -- a '*fixed with' category, and we format the part using format_derived() in [[Module:etymology]] rather than -- full_link() in [[Module:links]]. part.part_lang = part.lang part.lang = part.lang or lang part.sc = part.sc or sc local term = part.term if not term then return elseif not part.fragment then part.term, part.fragment = m_links.get_fragment(term) else part.term = m_links.get_fragment(term) end end --[==[ Construct a single linked part based on the information in `part`, for use by `show_affix()` and other entry points. This should be called after `canonicalize_part()` is called on the part. This is a thin wrapper around `full_link()` in [[Module:links]] unless `part.part_lang` is specified (indicating that a part-specific language was given), in which case `format_derived()` in [[Module:etymology]] is called to display a term in a language other than the language of the overall term (specified in `data.lang`). `data` contains the entire object passed into the entry point and is used to access information for constructing the categories added by `format_derived()`. ]==] function export.link_term(part, data, include_separator) local result if part.part_lang then result = require(etymology_module).format_derived { lang = data.lang, terms = {part}, sources = {part.lang}, sort_key = data.sort_key, nocat = data.nocat, template_name = "affix", decorations_on_outside = true, borrowing_type = data.borrowing_type, force_cat = data.force_cat or debug_force_cat, } else result = m_links.full_link(part, "term") end if include_separator and part.separator then return part.separator .. result else return result end end local function canonicalize_script_code(scode) -- Convert 'mnc-Mong', 'xwo-Mong' etc. to 'Mong'. return (scode:gsub("^.*%-", "")) end ----------------------------------------------------------------------------------------- -- Affix-handling functions -- ----------------------------------------------------------------------------------------- -- Figure out the appropriate script for the given affix and language (unless the script is explicitly passed in), and -- return the values of template_hyphens[], display_hyphens[] and lookup_hyphens[] for that script, substituting -- default values as appropriate. Four values are returned: -- DETECTED_SCRIPT, TEMPLATE_HYPHEN, DISPLAY_HYPHEN, LOOKUP_HYPHEN local function detect_script_and_hyphens(text, lang, sc) local scode -- 1. If the script is explicitly passed in, use it. if sc then scode = sc:getCode() else local possible_script_codes = lang:getScriptCodes() -- YUCK! `possible_script_codes` comes from loadData() so #possible_scripts doesn't work (always returns 0). local num_possible_script_codes = m_table.length(possible_script_codes) if num_possible_script_codes == 0 then -- This shouldn't happen; if the language has no script codes, -- the list {"None"} should be returned. error("Something is majorly wrong! Language " .. lang:getCanonicalName() .. " has no script codes.") end if num_possible_script_codes == 1 then -- 2. If the language has only one possible script, use it. scode = possible_script_codes[1] else -- 3. Check if any of the possible scripts for the language have non-default values for template_hyphens[] -- or display_hyphens[]. If so, we need to do script detection on the text. If not, just use "Latn", -- which may not be technically correct but produces the right results because Latn has all default -- values for template_hyphens[] and display_hyphens[]. local may_have_nondefault_hyphen = false for _, script_code in ipairs(possible_script_codes) do script_code = canonicalize_script_code(script_code) if template_hyphens[script_code] or display_hyphens[script_code] then may_have_nondefault_hyphen = true break end end if not may_have_nondefault_hyphen then scode = "Latn" else scode = lang:findBestScript(text):getCode() end end end scode = canonicalize_script_code(scode) local template_hyphen = template_hyphens[scode] or "-" local lookup_hyphen = lookup_hyphens[scode] or "-" local display_hyphen = display_hyphens[scode] or default_display_hyphen return scode, template_hyphen, display_hyphen, lookup_hyphen end --[=[ Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string, specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen, or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix). ]=] local function reconstruct_term_per_hyphens(term, affix_type, scode, thyph_re, new_hyphen) local function get_hyphen(hyph) if type(new_hyphen) == "string" then return new_hyphen end return new_hyphen(scode, hyph) end if affix_type == "non-affix" then return term elseif affix_type == "apitan" then local before, before_hyphen, after_hyphen, after = rmatch(term, "^(.*)" .. thyph_re .. " " .. thyph_re .. "(.*)$") if not before or ulen(term) <= 3 then -- Unlike with other types of affixes, don't try to add hyphens in the middle of the term to convert it to -- a circumfix. Also, if the term is just hyphen + space + hyphen, return it. return term end return before .. get_hyphen(before_hyphen) .. " " .. get_hyphen(after_hyphen) .. after elseif affix_type == "sisipan" or affix_type == "jalinan" then local before_hyphen, middle, after_hyphen = rmatch(term, "^" .. thyph_re .. "(.*)" .. thyph_re .. "$") if before_hyphen and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return get_hyphen(before_hyphen) .. (middle or term) .. get_hyphen(after_hyphen) elseif affix_type == "awalan" then local middle, after_hyphen = rmatch(term, "^(.*)" .. thyph_re .. "$") if middle and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return (middle or term) .. get_hyphen(after_hyphen) elseif affix_type == "akhiran" then local before_hyphen, middle = rmatch(term, "^" .. thyph_re .. "(.*)$") if before_hyphen and ulen(term) <= 1 then -- If the term is just a hyphen, return it. return term end return get_hyphen(before_hyphen) .. (middle or term) else error(("Internal error: Unrecognized affix type '%s'"):format(affix_type)) end end --[=[ Look up a mapping from a given affix variant to the canonical form used in categories and links. The lookup tables are language-specific according to `lang`, and may be ID-specific according to `affix_id`. The affixes as they appear in the lookup tables (both the variant and the canonical form) are in "lookup affix" format (approximately speaking, they use a regular hyphen for most scripts, but a tatweel for Arabic-script entries and a maqqef for Hebrew-script entries), but the passed-in `affix` param is in "template affix" format (which differs from the lookup affix for Arabic-script entries, because more types of hyphens are allowed in template affixes; see the comments at the top of the file). The remaining parameters to this function are used to convert from template affixes to lookup affixes; see the reconstruct_term_per_hyphens() function above. If the affix contains brackets, no lookup is done. Otherwise, a two-stage process is used, first looking up the affix directly and then stripping diacritics and looking it up again. The reason for this is documented above in the comments at the top of the file (specifically, the comments describing lookup affixes). The value of a mapping can either be a string (do the mapping regardless of affix ID) or a table indexed by affix ID (where the special value `false` indicates no affix ID). The values of entries in this table can also be strings, or tables with keys `affix` and `id` (again, use `false` to indicate no ID). This allows an affix mapping to map from one ID to another (for example, this is used in English to map the [[an-]] prefix with no ID to the [[a-]] prefix with the ID 'not'). The Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string, specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen, or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix). ]=] local function lookup_affix_mapping(affix, affix_type, lang, scode, thyph_re, lookup_hyph, affix_id) local function do_lookup(afx) -- Ensure that the affix uses lookup hyphens regardless of whether it used a different type of hyphens before -- or no hyphens. local lookup_affix = reconstruct_term_per_hyphens(afx, affix_type, scode, thyph_re, lookup_hyph) local function do_lookup_for_langcode(langcode) if export.langs_with_lang_specific_data[langcode] then local langdata = mw.loadData(export.affix_lang_data_module_prefix .. langcode) if langdata.affix_mappings then local mapping = langdata.affix_mappings[lookup_affix] if mapping then if type(mapping) == "table" then mapping = mapping[affix_id] or mapping.default or mapping[affix_id or false] if mapping then return mapping end else return mapping end end end end end -- If `lang` is an etymology-only language, look for a mapping both for it and its full parent. local langcode = lang:getCode() local mapping = do_lookup_for_langcode(langcode) if mapping then return mapping end local full_langcode = lang:getFullCode() if full_langcode ~= langcode then mapping = do_lookup_for_langcode(full_langcode) if mapping then return mapping end end return nil end if affix:find("%[%[") then return nil end return do_lookup(affix) or do_lookup(lang:stripDiacritics(affix)) or nil end --[==[ For a given template term in a given language (see the definition of "template affix" near the top of the file), possibly in an explicitly specified script `sc` (but usually nil), return the term's affix type ({"prefix"}, {"interfix"}, {"suffix"}, {"circumfix"} or {"non-affix"}) along with the corresponding link and display affixes (see definitions near the top of the file); also the corresponding lookup affix (if `return_lookup_affix` is specified). The term passed in should already have any fragment (after the # sign) parsed off of it. Four values are returned: `affix_type`, `link_term`, `display_term` and `lookup_term`. The affix type can be passed in instead of autodetected; in this case, the template term need not have any attached hyphens, and the appropriate hyphens will be added in the appropriate places. If `do_affix_mapping` is specified, look up the affix in the lang-specific affix mappings, as described in the comment at the top of the file; otherwise, the link and display terms will always be the same. (They will be the same in any case if the template term has a bracketed link in it or is not an affix.) If `return_lookup_affix` is given, the fourth return value contains the term with appropriate lookup hyphens in the appropriate places; otherwise, it is the same as the display term. (This functionality is used in [[Module:category tree/affixes and compounds]] to convert link affixes into lookup affixes so that they can be looked up in the affix mapping tables.) Exported because used by [[Module:headword utilities]] to determine the affix type of a given pagename. ]==] function export.parse_term_for_affixes(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) if not term then return "non-affix", nil, nil, nil end if term == "^" then -- Indicates a null term to emulate the behavior of {{suffix|foo||bar}}. term = "" return "non-affix", term, term, term end if term:find("^%^") then -- HACK! ^ at the beginning of Korean languages has a special meaning, triggering capitalization of the -- transliteration. Don't interpret it as "force non-affix" for those languages. local langcode = lang:getCode() if langcode ~= "ko" and langcode ~= "okm" and langcode ~= "jje" then -- Formerly we allowed ^ to force non-affix type; this is now handled using an inline modifier -- <naf>, <root>, etc. Throw an error for the moment when the old way is encountered. error("Use of ^ to force non-affix status is no longer supported; use an inline modifier <naf> or <root> " .. "after the component") end end -- Remove an asterisk if the morpheme is reconstructed and add it back at the end. local reconstructed = "" if term:find("^%*") then reconstructed = "*" term = term:gsub("^%*", "") end local scode, thyph, dhyph, lhyph = detect_script_and_hyphens(term, lang, sc) thyph = "([" .. thyph .. "])" if not affix_type then if rfind(term, thyph .. " " .. thyph) then affix_type = "apitan" else local has_beginning_hyphen = rfind(term, "^" .. thyph) local has_ending_hyphen = rfind(term, thyph .. "$") if has_beginning_hyphen and has_ending_hyphen then affix_type = "jalinan" elseif has_ending_hyphen then affix_type = "awalan" elseif has_beginning_hyphen then affix_type = "akhiran" else affix_type = "non-affix" end end end local link_term, display_term, lookup_term if affix_type == "non-affix" then link_term = term display_term = term lookup_term = term else display_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, dhyph) if do_affix_mapping then link_term = lookup_affix_mapping(term, affix_type, lang, scode, thyph, lhyph, affix_id) -- The return value of lookup_affix_mapping() may be an affix mapping with lookup hyphens if a mapping -- was found, otherwise nil if a mapping was not found. We need to convert to display hyphens in -- either case, but in the latter case we can reuse the display term, which has already been converted. if link_term then link_term = reconstruct_term_per_hyphens(link_term, affix_type, scode, thyph, dhyph) else link_term = display_term end else link_term = display_term end if return_lookup_affix then lookup_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, lhyph) else lookup_term = display_term end end link_term = reconstructed .. link_term display_term = reconstructed .. display_term lookup_term = reconstructed .. lookup_term return affix_type, link_term, display_term, lookup_term end --[==[ Add a hyphen to a term in the appropriate place, based on the specified affix type, stripping off any existing hyphens in that place. For example, if `affix_type` == {"prefix"}, we'll add a hyphen onto the end if it's not already there (or is of the wrong type). Three values are returned: the link term, display term and lookup term. This function is a thin wrapper around `parse_term_for_affixes`; see the comments above that function for more information. Note that this function is exposed externally because it is called by [[Module:category tree/affixes and compounds]]; see the comment in `parse_term_for_affixes` for more information. ]==] function export.make_affix(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) if not (affix_type == "awalan" or affix_type == "akhiran" or affix_type == "apitan" or affix_type == "sisipan" or affix_type == "jalinan" or affix_type == "non-affix") then error("Internal error: Invalid affix type " .. (affix_type or "(nil)")) end local _, link_term, display_term, lookup_term = export.parse_term_for_affixes(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id) return link_term, display_term, lookup_term end ----------------------------------------------------------------------------------------- -- Main entry points -- ----------------------------------------------------------------------------------------- --[==[ Core categorization logic for affixes. This is shared between show_affix(), show_compound_like() and get_affix_categories_only(). Returns the categories array and other metadata needed for formatting. ]==] local function generate_affix_categories(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) local text_sections, categories, borrowing_type = process_etymology_type(data.type, data.surface_analysis or data.nocap, data.notext, #data.parts > 0) data.borrowing_type = borrowing_type -- Process each part local whole_words = 0 local is_affix_or_compound = false -- Canonicalize and generate links for all the parts first; then do categorization in a separate step, because when -- processing the first part for categorization, we may access the second part and need it already canonicalized. for i, part in ipairs_with_gaps(data.parts) do part = part or {} data.parts[i] = part canonicalize_part(part, data.lang, data.sc) -- Determine affix type and get link and display terms (see text at top of file). Store them in the part -- (in fields that won't clash with fields used by full_link() in [[Module:links]] or link_term()), so they -- can be used in the loop below when categorizing. part.affix_type, part.affix_link_term, part.affix_display_term = export.parse_term_for_affixes(part.term, part.lang, part.sc, part.type, not part.alt, nil, part.id) -- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with inline -- modifiers. The intention in either case is not to link the term. part.term = ine(part.affix_link_term) -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. part.alt = part.alt or (part.affix_display_term ~= part.affix_link_term and part.affix_display_term) or nil end if not data.noaffixcat then -- Now do categorization. for i, part in ipairs_with_gaps(data.parts) do local affix_type = part.affix_type if affix_type ~= "non-affix" then is_affix_or_compound = true -- Make a sort key. For the first part, use the second part as the sort key; the intention is that if the -- term has a prefix, sorting by the prefix won't be very useful so we sort by what follows, which is -- presumably the root. local part_sort_base = nil local part_sort = part.sort or data.sort_key if i == 1 and data.parts[2] and data.parts[2].term then local part2 = data.parts[2] -- If the second-part link term is empty, the user requested an unlinked term; avoid a wikitext error -- by using the alt value if available. part_sort_base = ine(part2.affix_link_term) or ine(part2.alt) if part_sort_base then part_sort_base = strip_diacritics_no_links(part2.lang, part_sort_base) end end if part.pos and rfind(part.pos, "patronym") then table.insert(categories, {cat = "Patronim", sort_key = part_sort, sort_base = part_sort_base}) end if data.pos ~= "perkataan" and part.pos and rfind(part.pos, "diminutive") then table.insert(categories, {cat = ucfirst(data.pos) .. " diminutif", sort_key = part_sort, sort_base = part_sort_base}) end -- Don't add a '*fixed with' category if the link term is empty or is in a different language. if ine(part.affix_link_term) and not part.part_lang then table.insert(categories, {cat = ucfirst(data.pos) .. " dengan " .. affix_type .. " " .. strip_diacritics_no_links(part.lang, part.affix_link_term) .. (part.id and " (" .. part.id .. ")" or ""), sort_key = part_sort, sort_base = part_sort_base}) end else whole_words = whole_words + 1 if whole_words == 2 then is_affix_or_compound = true table.insert(categories, ucfirst(data.pos) .. " majmuk") end end end -- Make sure there was either an affix or a compound (two or more non-affix terms). if not is_affix_or_compound and not data.allow_no_affixes_or_compounds then error("The parameters did not include any affixes, and the term is not a compound. Please provide at least one affix.") end end return text_sections, categories, borrowing_type end --[==[ Implementation of {{tl|affix}} and {{tl|surface analysis}}. `data` contains all the information describing the affixes to be displayed, and contains the following: * `.lang` ('''required'''): Overall language object. Different from term-specific language objects (see `.parts` below). * `.sc`: Overall script object (usually omitted). Different from term-specific script objects. * `.parts` ('''required'''): List of objects describing the affixes to show. The general format of each object is as would be passed to `full_link()`, except that the `.lang` field should be missing unless the term is of a language different from the overall `.lang` value (in such a case, the language name is shown along with the term and an additional "derived from" category is added). '''WARNING''': The data in `.parts` will be destructively modified. * `.pos`: Overall part of speech (used in categories, defaults to {"terms"}). Different from term-specific part of speech. * `.sort_key`: Overall sort key. Normally omitted except e.g. in Japanese. * `.type`: Type of compound, if the parts in `.parts` describe a compound. Strictly optional, and if supplied, the compound type is displayed before the parts (normally capitalized, unless `.nocap` is given). * `.nocap`: Don't capitalize the first letter of text displayed before the parts (relevant only if `.type` or `.surface_analysis` is given). * `.notext`: Don't display any text before the parts (relevant only if `.type` or `.surface_analysis` is given). * `.nocat`: Disable all categorization. * `.noaffixcat`: Disable affix (and compound) categorization. Relevant for e.g. blends, which may otherwise be incorrectly categorized as compound terms. * `.lit`: Overall literal definition. Different from term-specific literal definitions. * `.force_cat`: Always display categories, even on userspace pages. * `.surface_analysis`: Implement {{surface analysis}}; adds `By surface analysis, ` before the parts. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_affix(data) local text_sections, categories, _ = generate_affix_categories(data) -- Process each part for display local parts_formatted = {} for i, part in ipairs_with_gaps(data.parts) do -- Make a link for the part table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if data.surface_analysis then local text = "mengikut " .. glossary_link("analisis permukaan") .. ", " if not data.nocap then text = ucfirst(text) end table.insert(text_sections, 1, text) end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Get only the categories that would be generated by show_affix(), without any text output or formatting. This is used by Module:etymon to get affix categorization. Returns an array of category objects, where each entry is either a string (simple category name) or a table with keys `cat`, `sort_key`, and `sort_base` for more complex categorization. `data` should have the same structure as passed to show_affix(): * `.lang` (required): Overall language object * `.parts` (required): Array of affix part objects with `.term`, `.lang`, `.id`, etc. * `.pos`: Part of speech (defaults to "terms") * `.sort_key`: Overall sort key for categories '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.get_affix_categories_only(data) local _, categories, _ = generate_affix_categories(data) return categories end function export.show_surface_analysis(data) data.surface_analysis = true data.allow_no_affixes_or_compounds = true return export.show_affix(data) end --[==[ Implementation of {{tl|compound}}. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_compound(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) local text_sections, categories, borrowing_type = process_etymology_type(data.type, data.nocap, data.notext, #data.parts > 0) data.borrowing_type = borrowing_type local parts_formatted = {} table.insert(categories, ucfirst(data.pos) .. " majmuk") -- Make links out of all the parts local whole_words = 0 for i, part in ipairs(data.parts) do canonicalize_part(part, data.lang, data.sc) -- Determine affix type and get link and display terms (see text at top of file). local affix_type, link_term, display_term = export.parse_term_for_affixes(part.term, part.lang, part.sc, part.type, not part.alt, nil, part.id) -- If the term is an interfix or the type was explicitly given, recognize it as such (which means e.g. that we -- will display the term without hyphens for East Asian languages). Otherwise, ignore the fact that it looks -- like an affix and display as specified in the template (but pay attention to the detected affix type for -- certain tracking purposes). if affix_type == "jalinan" or (part.type and part.type ~= "non-affix") then -- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with -- inline modifiers. The intention in either case is not to link the term. Don't add a '*fixed with' -- category in this case, or if the term is in a different language. -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. if link_term and link_term ~= "" and not part.part_lang then table.insert(categories, {cat = ucfirst(data.pos) .. " dengan " .. affix_type .. " " .. strip_diacritics_no_links(part.lang, link_term), sort_key = part.sort or data.sort_key}) end part.term = link_term ~= "" and link_term or nil part.alt = part.alt or (display_term ~= link_term and display_term) or nil else if affix_type ~= "non-affix" then local langcode = data.lang:getCode() -- If `data.lang` is an etymology-only language, track both using its code and its full parent's code. track { affix_type, affix_type .. "/lang/" .. langcode } local full_langcode = data.lang:getFullCode() if langcode ~= full_langcode then track(affix_type .. "/lang/" .. full_langcode) end else whole_words = whole_words + 1 end end table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if whole_words == 1 then track("one whole word") elseif whole_words == 0 then track("looks like confix") end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Implementation of {{tl|blend}}, {{tl|univerbation}} and similar "compound-like" templates. '''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`. ]==] function export.show_compound_like(data) data.allow_no_affixes_or_compounds = true local text_sections, categories, _ = generate_affix_categories(data) if data.cat then table.insert(categories, data.cat) end -- Process each part for display local parts_formatted = {} for i, part in ipairs_with_gaps(data.parts) do -- Make a link for the part table.insert(parts_formatted, export.link_term(part, data, "include_separator")) end if #data.parts > 0 and data.oftext then table.insert(text_sections, 1, " " .. data.oftext .. " ") end if data.text then table.insert(text_sections, 1, data.text) end table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories, separator_already_added = true }) return table.concat(text_sections) end --[==[ Make `part` (a structure holding information on an affix part) into an affix of type `affix_type`, and apply any relevant affix mappings. For example, if the desired affix type is "suffix", this will (in general) add a hyphen onto the beginning of the term, alt, tr and ts components of the part if not already present. The hyphen that's added is the "display hyphen" (see above) and may be script-specific. (In the case of East Asian scripts, the display hyphen is an empty string whereas the template hyphen is the regular hyphen, meaning that any regular hyphen at the beginning of the part will be effectively removed.) `lang` and `sc` hold overall language and script objects. Note that this also applies any language-specific affix mappings, so that e.g. if the language is Finnish and the user specified [[-käs]] in the affix and didn't specify an `.alt` value, `part.term` will contain [[-kas]] and `part.alt` will contain [[-käs]]. This function is used by the "legacy" templates ({{tl|prefix}}, {{tl|suffix}}, {{tl|confix}}, etc.) where the nature of the affix is specified by the template itself rather than auto-determined from the affix, as is the case with {{tl|affix}}. '''WARNING''': This destructively modifies `part`. ]==] local function make_part_into_affix(part, lang, sc, affix_type) canonicalize_part(part, lang, sc) local link_term, display_term = export.make_affix(part.term, part.lang, part.sc, affix_type, not part.alt, nil, part.id) part.term = link_term -- When we don't specify `do_affix_mapping` to make_affix(), link and display terms (first and second retvals of -- make_affix()) are the same. -- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being -- redundant alt text. part.alt = part.alt and export.make_affix(part.alt, part.lang, part.sc, affix_type) or (display_term ~= link_term and display_term) or nil local Latn = require(scripts_module).getByCode("Latn") part.tr = export.make_affix(part.tr, part.lang, Latn, affix_type) part.ts = export.make_affix(part.ts, part.lang, Latn, affix_type) end local function track_wrong_affix_type(template, part, expected_affix_type) if part and not part.type then local affix_type = export.parse_term_for_affixes(part.term, part.lang, part.sc) if affix_type ~= expected_affix_type then local part_name = expected_affix_type or "base" local langcode = part.lang:getCode() local full_langcode = part.lang:getFullCode() require("Module:debug/track") { template, template .. "/" .. part_name, template .. "/" .. part_name .. "/" .. (affix_type or "none"), template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. langcode } -- If `part.lang` is an etymology-only language, track both using its code and its full parent's code. if full_langcode ~= langcode then require("Module:debug/track")( template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. full_langcode ) end end end end local function insert_affix_category(categories, pos, affix_type, part, sort_key, sort_base) -- Don't add a '*fixed with' category if the link term is empty or is in a different language. if part.term and not part.part_lang then local cat = ucfirst(pos) .. " dengan " .. affix_type .. " " .. strip_diacritics_no_links(part.lang, part.term) .. (part.id and " (" .. part.id .. ")" or "") if sort_key or sort_base then table.insert(categories, {cat = cat, sort_key = sort_key, sort_base = sort_base}) else table.insert(categories, cat) end end end --[==[ Implementation of {{tl|circumfix}}. '''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`. ]==] function export.show_circumfix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.prefix, data.lang, data.sc, "awalan") make_part_into_affix(data.suffix, data.lang, data.sc, "akhiran") track_wrong_affix_type("apitan", data.prefix, "awalan") track_wrong_affix_type("apitan", data.base, nil) track_wrong_affix_type("apitan", data.suffix, "akhiran") -- Create circumfix term. local circumfix = nil if data.prefix.term and data.suffix.term then circumfix = data.prefix.term .. " " .. data.suffix.term data.prefix.alt = data.prefix.alt or data.prefix.term data.suffix.alt = data.suffix.alt or data.suffix.term data.prefix.term = circumfix data.suffix.term = circumfix end -- Make links out of all the parts. local parts_formatted = {} local categories = {} local sort_base if data.base.term then sort_base = strip_diacritics_no_links(data.base.lang, data.base.term) end table.insert(parts_formatted, export.link_term(data.prefix, data)) table.insert(parts_formatted, export.link_term(data.base, data)) table.insert(parts_formatted, export.link_term(data.suffix, data)) -- Insert the categories, but don't add a '*fixed with' category if the link term is in a different language. if not data.prefix.part_lang then table.insert(categories, {cat=ucfirst(data.pos) .. " dengan apitan " .. strip_diacritics_no_links(data.prefix.lang, circumfix), sort_key=data.sort_key, sort_base=sort_base}) end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|confix}}. '''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`. ]==] function export.show_confix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.prefix, data.lang, data.sc, "awalan") make_part_into_affix(data.suffix, data.lang, data.sc, "akhiran") track_wrong_affix_type("confix", data.prefix, "awalan") track_wrong_affix_type("confix", data.base, nil) track_wrong_affix_type("confix", data.suffix, "akhiran") -- Make links out of all the parts. local parts_formatted = {} local prefix_sort_base if data.base and data.base.term then prefix_sort_base = strip_diacritics_no_links(data.base.lang, data.base.term) elseif data.suffix.term then prefix_sort_base = strip_diacritics_no_links(data.suffix.lang, data.suffix.term) end -- Insert the categories and parts. local categories = {} table.insert(parts_formatted, export.link_term(data.prefix, data)) insert_affix_category(categories, data.pos, "awalan", data.prefix, data.sort_key, prefix_sort_base) if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) end table.insert(parts_formatted, export.link_term(data.suffix, data)) -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "akhiran", data.suffix) return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|infix}}. '''WARNING''': This destructively modifies both `data` and `.base` and `.infix`. ]==] function export.show_infix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. make_part_into_affix(data.infix, data.lang, data.sc, "sisipan") track_wrong_affix_type("sisipan", data.base, nil) track_wrong_affix_type("sisipan", data.infix, "sisipan") -- Make links out of all the parts. local parts_formatted = {} local categories = {} table.insert(parts_formatted, export.link_term(data.base, data)) table.insert(parts_formatted, export.link_term(data.infix, data)) -- Insert the categories. -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "sisipan", data.infix) return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|prefix}}. '''WARNING''': This destructively modifies both `data` and the structures within `.prefixes`, as well as `.base`. ]==] function export.show_prefix(data) data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. for i, prefix in ipairs(data.prefixes) do make_part_into_affix(prefix, data.lang, data.sc, "awalan") end for i, prefix in ipairs(data.prefixes) do track_wrong_affix_type("awalan", prefix, "awalan") end track_wrong_affix_type("awalan", data.base, nil) -- Make links out of all the parts. local parts_formatted = {} local first_sort_base = nil local categories = {} if data.prefixes[2] then first_sort_base = ine(data.prefixes[2].term) or ine(data.prefixes[2].alt) if first_sort_base then first_sort_base = strip_diacritics_no_links(data.prefixes[2].lang, first_sort_base) end elseif data.base then first_sort_base = ine(data.base.term) or ine(data.base.alt) if first_sort_base then first_sort_base = strip_diacritics_no_links(data.base.lang, first_sort_base) end end for i, prefix in ipairs(data.prefixes) do table.insert(parts_formatted, export.link_term(prefix, data)) insert_affix_category(categories, data.pos, "awalan", prefix, data.sort_key, i == 1 and first_sort_base or nil) end if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) else table.insert(parts_formatted, "") end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end --[==[ Implementation of {{tl|suffix}}. '''WARNING''': This destructively modifies both `data` and the structures within `.suffixes`, as well as `.base`. ]==] function export.show_suffix(data) local categories = {} data.pos = data.pos or default_pos data.pos = pluralize_pos(data.pos) canonicalize_part(data.base, data.lang, data.sc) -- Hyphenate the affixes and apply any affix mappings. for i, suffix in ipairs(data.suffixes) do make_part_into_affix(suffix, data.lang, data.sc, "akhiran") end track_wrong_affix_type("akhiran", data.base, nil) for i, suffix in ipairs(data.suffixes) do track_wrong_affix_type("akhiran", suffix, "akhiran") end -- Make links out of all the parts. local parts_formatted = {} if data.base then table.insert(parts_formatted, export.link_term(data.base, data)) else table.insert(parts_formatted, "") end for i, suffix in ipairs(data.suffixes) do table.insert(parts_formatted, export.link_term(suffix, data)) end -- Insert the categories. for i, suffix in ipairs(data.suffixes) do -- FIXME, should we be specifying a sort base here? insert_affix_category(categories, data.pos, "akhiran", suffix) if suffix.pos and rfind(suffix.pos, "patronym") then table.insert(categories, "Patronim") end end return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories } end return export 00ctt3fkh5mg0bo9nx60necwwv3han4 Modul:homophones 828 10482 375351 227296 2026-09-22T03:12:54Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708691|92708691]]) 375351 Scribunto text/plain local export = {} local decorations_module = "Module:decorations" local links_module = "Module:links" local parameter_utilities_module = "Module:parameter utilities" local function track(page) return require("Module:debug/track")("homophones/" .. page) end --[==[ Meant to be called from a module. `data` is a table containing the following fields: * `lang`: language object for the homophones; * `homophones`: a list of homophones, each described by an object which can contain all the fields in the object passed to {full_link()} in [[Module:links]] except for `lang` and `sc` (which are copied from the outer level), including decoration fields; ** `term`: the homophone itself; ** `separator`: {nil} or the string used to separate this homophone from the preceding one when displayed; defaults to the top-level `separator`; ** `alt`: display text for the homophone, as in {{tl|l}}; ** `gloss`: gloss for the homophone, as in {{tl|l}}; ** `tr`: transliteration for the homophone, as in {{tl|l}}; ** `ts`: transcription for the homophone, as in {{tl|l}}; ** `genders`: list of genders for the homophone, as in {{tl|l}}; ** `pos`: part of speech of the homophone, as in {{tl|l}}; ** `ng`: non-gloss text for the homophone, as in {{tl|l}}; ** `lit`: literal meaning of the homophone, as in {{tl|l}}; ** `id`: sense ID for the homophone, as in {{tl|l}}; ** `lang`: optional lang code, overriding the lang code in the top-level `lang` field; ** `sc`: optional script code, overriding the script code in the top-level `sc` field; ** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the homophone in question; ** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the homophone in question; ** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]) and displayed directly before the homophone in question; ** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the homophone in question; ** `refs`: {nil} or a list of references or reference specs to add after the pronunciation and any posttext and qualifiers; the value of a list item is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the reference, as in `<nowiki><ref name="foo">...</ref></nowiki>` or `<nowiki><ref name="foo" /></nowiki>`) and/or `group` (the group of the reference, as in `<nowiki><ref name="foo" group="bar">...</ref></nowiki>` or `<nowiki><ref name="foo" group="bar"/></nowiki>`); this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the actual reference, located in the `<nowiki><references /></nowiki>` section; * `separator`: {nil} or a string, specifying the separator displayed before all homophones but the first; by default, {", "}; overridable at the individual homophone level; * `q`: {nil} or a list of left regular qualifier strings, displayed before the initial caption; * `qq`: {nil} or a list of right regular qualifier strings, displayed after all homophones; * `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the initial caption; * `aa`: {nil} or a list of right accent qualifier strings, displayed after all homophones; * `sc`: {nil} or script object for the homophones; * `sort`: {nil} or sort key; * `caption`: {nil} or string specifying the caption to use, in place of {"Homophone"} (if there is a single homophone), or {"Homophones"} (otherwise); a colon and space is automatically added after the caption; * `nocaption`: If true, suppress the caption display. * `nocat`: If true, suppress categorization. If both regular and accent qualifiers on the same side and at the same level are specified, the accent qualifiers precede the regular qualifiers on both left and right. '''WARNING''': Destructively modifies the objects inside the `homophones` field. ]==] function export.format_homophones(data) local hmptexts = {} local hmpcats = {} local m_links = require(links_module) local overall_sep = data.separator or ", " for i, hmp in ipairs(data.homophones) do hmp.lang = hmp.lang or data.lang hmp.sc = hmp.sc or data.sc hmp.show_decorations = true -- make full_link() display decorations (`.q`, `.qq`, `.a`, `.aa` and `.refs`) if hmp.qualifiers then error("`.qualifiers` is no longer supported; change the code to use `.qq` or `.q`") end local text = m_links.full_link(hmp) table.insert(hmptexts, hmp.separator or i > 1 and overall_sep or "") table.insert(hmptexts, text) end table.insert(hmpcats, "Perkataan dengan homofon bahasa " .. data.lang:getFullName()) local text = table.concat(hmptexts) local caption = data.nocaption and "" or ( data.caption or "[[Lampiran:Glosari#homofon|Homofon" .. (#data.homophones > 1 and "" or "") .. "]]" ) .. ": " text = caption .. text if data.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("overall `.qualifiers` is no longer supported; change the code to use `.qq` or `.q`") end if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then text = require(decorations_module).format_decorations { lang = data.lang, text = text, q = data.q, qq = data.qq, a = data.a, aa = data.aa, } end text = "<span class=\"homophones\">" .. text .. "</span>" if not data.nocat then local categories = require("Module:utilities").format_categories(hmpcats, data.lang, data.sort) text = text .. categories end return text end --[==[ Entry point for {{tl|homophones}} template (also written {{tl|homophone}} and {{tl|hmp}}). ]==] function export.show(frame) local parent_args = frame:getParent().args local compat = parent_args.lang local offset = compat and 0 or 1 local lang_arg = compat and "lang" or 1 local params = { [lang_arg] = {required = true, type = "language", default = "ms"}, [1 + offset] = {list = true, required = true, allow_holes = true, default = "perkataan"}, ["caption"] = {}, ["nocaption"] = {type = "boolean"}, ["nocat"] = {type = "boolean"}, ["sort"] = {}, } local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "a", "q"}}, } local homophones, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 1 + offset, parse_lang_prefix = true, track_module = "homophones", lang = lang_arg, sc = "sc.default", } local data = { lang = args[lang_arg], homophones = homophones, caption = args.caption, nocaption = args.nocaption, nocat = args.nocat, sc = args.sc.default, sort = args.sort, q = args.q.default, qq = args.qq.default, a = args.a.default, aa = args.aa.default, } return export.format_homophones(data) end return export 1wlerteggfy10doldx9e5ckzpge6vq4 Modul:ar-headword 828 11149 375386 266015 2026-09-22T07:16:37Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714469|92714469]]) 375386 Scribunto text/plain -- Author: primarily Benwing2; some work by Fenakhay, Erutuon; early version by Rua local export = {} local pos_functions = {} local force_cat = false -- for testing; if true, categories appear in non-mainspace pages local ar_translit = require("Module:ar-translit") local ar_verb_module = "Module:ar-verb" local ar_utilities_module = "Module:ar-utilities" local ar = require(ar_utilities_module) local en_utilities_module = "Module:en-utilities" local headword_module = "Module:headword" local headword_utilities_module = "Module:headword utilities" local links_module = "Module:links" local inflection_utilities_module = "Module:inflection utilities" local parse_utilities_module = "Module:parse utilities" local require_when_needed = require("Module:utilities/require when needed") local remove_links = require_when_needed(links_module, "remove_links") local m_table = require("Module:table") local m_str_utils = require("Module:string utilities") local m_en_utilities = require_when_needed(en_utilities_module) local m_headword_utilities = require_when_needed(headword_utilities_module) local glossary_link = require_when_needed(headword_utilities_module, "glossary_link") local boolean_param = {type = "boolean"} local list_to_set = m_table.listToSet local rfind = m_str_utils.find local rmatch = m_str_utils.match local rsubn = m_str_utils.gsub local u = m_str_utils.char local rsplit = m_str_utils.split local insert = table.insert local concat = table.concat local unpack = unpack or table.unpack -- Lua 5.2 compatibility local langcode = "ar" local lang = require("Module:languages").getByCode(langcode) local langname = lang:getCanonicalName() local TEMPCOMMA = u(0xFFF0) local TEMPARCOMMA = u(0xFFF1) local misc_pos_with_gender = list_to_set { "suffixes", "adjective forms", "noun forms", "proper noun forms", "pronoun forms", "determiner forms", } ----------------------------------------------------------------------------------------- -- Utility functions -- ----------------------------------------------------------------------------------------- local dump = mw.dumpObject -- version of mw.ustring.gsub() that discards all but the first return value local function rsub(term, foo, bar) local retval = rsubn(term, foo, bar) return retval end local function ine(val) if val == "" then return nil else return val end end -- Replace comma with a temporary char in comma + whitespace. local function escape_comma_whitespace(run) local escaped = false if run:find("\\,") then run = run:gsub("\\,", "\\" .. TEMPCOMMA) escaped = true end if run:find("\\،") then run = run:gsub("\\،", "\\" .. TEMPARCOMMA) escaped = true end if run:find(",%s") then run = run:gsub(",(%s)", TEMPCOMMA .. "%1") escaped = true end if run:find("،%s") then run = run:gsub("،(%s)", TEMPARCOMMA .. "%1") escaped = true end return run, escaped end -- Undo replacement of comma with a temporary char in comma + whitespace. local function unescape_comma_whitespace(run) return (run:gsub(TEMPCOMMA, ","):gsub(TEMPARCOMMA, "،")) end -- Split an argument on comma or Arabic comma, but not either type of comma followed by whitespace. local function split_on_comma(val) if rfind(val, "[,،]%s") or val:find("\\") then return export.split_escaping(val, "[,،]", false, escape_comma_whitespace, unescape_comma_whitespace) else return rsplit(val, "[,،]") end end local function replace_tr_ending(tr, from, to) if not tr then return nil end local pref = tr:match("^(.*)" .. from .. "$") if not pref then error(("Translit '%s' does not end in -%s, as expected"):format(tr, from)) end return pref .. to end ----------------------------------------------------------------------------------------- -- Tracking functions -- ----------------------------------------------------------------------------------------- local trackfn = require("Module:debug/track") local function track(page) trackfn(langcode .. "-headword/" .. page) return true end --[==[ Examples of what you can find by looking at what links to the given pages: [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized]] all unvocalized pages [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/pl]] all unvocalized pages where the plural is unvocalized, whether specified using pl=, pl2=, etc. [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head]] all unvocalized pages where the head is unvocalized [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/nouns]] all nouns excluding proper nouns, collective nouns, singulative nouns where the head is unvocalized [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/proper]] nouns all proper nouns where the head is unvocalized [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/not]] proper nouns all words that are not proper nouns where the head is unvocalized [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/adjectives]] all adjectives where any parameter is unvocalized; currently only works for heads, so equivalent to .../unvocalized/head/adjectives [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-empty-head]] all pages with an empty head [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-manual-translit]] all unvocalized pages with manual translit [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-manual-translit/head/nouns]] all nouns where the head is unvocalized but has manual translit [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-no-translit]] all unvocalized pages without manual translit [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab]] all pages with any parameter containing i3rab of either -un, -u, -a or -i [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-un]] all pages with any parameter containing an -un i3rab ending [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-un/pl]] all pages where a form specified using pl=, pl2=, etc. contains an -un i3rab ending [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-u/head]] all pages with a head containing an -u i3rab ending [[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab/head/proper]] nouns (all proper nouns with a head containing i3rab of either -un, -u, -a or -i) In general, the format is one of the following: Wiktionary:Tracking/ar-headword/FIRSTLEVEL Wiktionary:Tracking/ar-headword/FIRSTLEVEL/ARGNAME Wiktionary:Tracking/ar-headword/FIRSTLEVEL/POS Wiktionary:Tracking/ar-headword/FIRSTLEVEL/ARGNAME/POS FIRSTLEVEL can be one of "unvocalized", "unvocalized-empty-head" or its opposite "unvocalized-specified", "unvocalized-manual-translit" or its opposite "unvocalized-no-translit", "i3rab", "i3rab-un", "i3rab-u", "i3rab-a", or "i3rab-i". ARGNAME is either "head" or an argument such as "pl", "f", "cons", etc. This automatically includes arguments specified as head2=, pl3=, etc. POS is a part of speech, lowercase and singular, e.g. "kata nama", "Kata sifat", "Kata nama khas", "collective nouns", etc. or "not proper noun", which includes all parts of speech but proper nouns. ]==] local function track_form(argname, form, translit, pos) form = ar.reorder_shadda(remove_links(form)) function dotrack(page) track(page) track(page .. "/" .. argname) if pos then track(page .. "/" .. pos) track(page .. "/" .. argname .. "/" .. pos) if pos ~= "Kata nama khas" then track(page .. "/not proper noun") track(page .. "/" .. argname .. "/not proper noun") end end end function track_i3rab(arabic, tr) if rfind(form, arabic .. "$") then dotrack("i3rab") dotrack("i3rab-" .. tr) end end track_i3rab(ar.UN, "un") track_i3rab(ar.U, "u") track_i3rab(ar.A, "a") track_i3rab(ar.I, "i") if form == "" or not (lang:transliterate(form)) then dotrack("unvocalized") if form == "" then dotrack("unvocalized-empty-head") else dotrack("unvocalized-specified") end if translit then dotrack("unvocalized-manual-translit") else dotrack("unvocalized-no-translit") end end end ----------------------------------------------------------------------------------------- -- Inflection-parsing functions -- ----------------------------------------------------------------------------------------- -- Construct the default construct state or informal form of a term in lemma format. Usually this is the same as the -- lemma but is different for final-weak nouns and adjectives ending in -n in their lemma. NOTE: Input must be -- shadda-reordered for this to work properly. local function default_construct_state_or_informal(term, tr) local pref = term:match("^(.*)" .. ar.HAMZA .. ar.IN .."$") -- Hamza on the line with -in changes to hamza-on-yā with -ī. if pref then return pref .. ar.HAMZA_ON_YA .. ar.II, replace_tr_ending(tr, "in", "ī") end -- Otherwise just change -in to -ī. pref = term:match("^(.*)" .. ar.IN .. "$") if pref then return pref .. ar.II, replace_tr_ending(tr, "in", "ī") end -- Change -an with alif maqṣūra to -ā with alif maqṣūra. pref = term:match("^(.*)" .. ar.AN .. ar.AMAQ .. "$") if pref then return pref .. ar.AAMAQ, replace_tr_ending(tr, "an", "ā") end -- Change -an with tall alif (e.g. عَصًا) to -ā with tall alif. pref = term:match("^(.*)" .. ar.AN .. ar.ALIF .. "$") if pref then return pref .. ar.AA, replace_tr_ending(tr, "an", "ā") end return term, tr end local function generate_construct_state_or_informal_default(data, args) local heads = data.heads local consobjs = {} local different_cons = false for _, headobj in ipairs(data.heads) do local consterm, constr = default_construct_state_or_informal(headobj.term, headobj.tr) different_cons = different_cons or consterm ~= headobj.term or constr ~= headobj.tr local consobj = m_table.shallowCopy(headobj) consobj.term = consterm consobj.tr = constr insert(consobjs, consobj) end if different_cons then return consobjs else return {} end end local noun_field_cons = { field = "cons", label = "<<construct state>>", generate_default = generate_construct_state_or_informal_default, default_when_not_explicit = function(args, data) return true end, } local noun_field_inf = {field = "inf", label = "informal"} local noun_field_obl = {field = "obl", label = "<<oblique>>"} local noun_field_def = {field = "def", label = "<<definite>> state"} local noun_inflections = { noun_field_cons, noun_field_inf, noun_field_obl, noun_field_def, } local adj_field_inf = { field = "inf", label = "informal", generate_default = generate_construct_state_or_informal_default, default_when_not_explicit = function(args, data) return true end, } local adj_field_obl = noun_field_obl local adj_field_def = noun_field_def local adjective_inflections = { adj_field_inf, adj_field_obl, adj_field_def, } local function has_construct_state(data) return data.pos_category ~= "Kata sifat" end local function parse_nominal_inflection(paramname, val, parse_err) return m_headword_utilities.parse_term_with_modifiers { val = val, paramname = paramname, splitchar = ",", include_mods = {"tr", "g"}, } end local function make_nominal_inflection_param_mod_spec(paramname) return {convert = function(val, parse_err) return parse_nominal_inflection(paramname, val, parse_err) end} end -- Parse an inflection. The raw arguments come from `args[field]`, which is parsed for inline modifiers. Multiple -- comma-separated values are allowed. local function parse_inflection(data, args, field, is_head) local argfield = field local argpref = field if type(argfield) == "table" then argpref = argfield[2] argfield = argfield[1] end local include_mods if is_head then include_mods = {"tr"} else include_mods = {"tr", "g"} for _, spec in ipairs(has_construct_state(data) and noun_inflections or adjective_inflections) do insert(include_mods, {spec.field, make_nominal_inflection_param_mod_spec(argpref .. "." .. spec.field)}) end end if is_head then local retval if args[argfield] then retval = m_headword_utilities.parse_term_with_modifiers { val = args[argfield], paramname = field, splitchar = ",", is_head = is_head, include_mods = include_mods, } end return retval or {} else return m_headword_utilities.parse_term_list_with_modifiers { forms = args[argfield], paramname = field, splitchar = ",", is_head = is_head, include_mods = include_mods, } end end local function insert_inflection(data, terms, label, accel, defgender, track_field, no_label, usually_no_label) local track_pos = m_en_utilities.singularize(data.pos_category) for _, termobj in ipairs(terms) do -- If the user supplied a construct state or informal form for the term with a value of "+", substitute the -- default value for the term. If the user supplied a value of "--", they want no value displayed. Otherwise, -- if the user didn't supply any value, we check to see if the default construct state or informal form is -- different from the lemma and display it if so; this applies particularly to terms in '-in' and '-an', where -- the default construct state or informal form is almost always correct. local field = has_construct_state(data) and "cons" or "inf" if not termobj[field] then local defcons, defconstr = default_construct_state_or_informal(termobj.term, termobj.tr) if termobj.term ~= defcons or termobj.tr ~= defconstr then -- We don't want to copy decorations from the term object because we're a subinflection of the term -- object. termobj[field] = {{term = defcons, tr = defconstr}} end elseif termobj[field][1].term == "--" then if termobj[field][2] then error("Can't specify more than one value for <" .. field .. ":...> if first value is '--', meaning \"don't insert anything\"") end termobj[field] = nil else for i, consobj in ipairs(termobj[field]) do if consobj.term == "+" then if consobj.tr then error("Can't specify translit for default value '+'") end consobj.term, consobj.tr = default_construct_state_or_informal(termobj.term, termobj.tr) elseif consobj.term == "~" then if consobj.tr then error("Can't specify translit for term-requesting value '~'") end consobj.term, consobj.tr = termobj.term, termobj.tr end end end if defgender and not termobj.genders then termobj.genders = {{spec = defgender}} end local function insert_nested_inflection(field, label) if termobj[field] then m_headword_utilities.insert_inflection { headdata = data, inflobj = termobj, terms = termobj[field], label = label } end end for _, spec in ipairs(has_construct_state(data) and noun_inflections or adjective_inflections) do insert_nested_inflection(spec.field, spec.label) end track_form(track_field, termobj.term, termobj.tr, track_pos) end m_headword_utilities.insert_inflection { headdata = data, terms = terms, label = label, accel = accel and {form = accel} or nil, no_label = no_label, usually_no_label = usually_no_label, } end ----------------------------------------------------------------------------------------- -- Main entry point -- ----------------------------------------------------------------------------------------- function export.show(frame) local iparams = { [1] = true, } local iargs = require("Module:parameters").process(frame.args, iparams) local parargs = frame:getParent().args local poscat = iargs[1] local pos_in_1 = not poscat if pos_in_1 then poscat = ine(parargs[1]) or mw.title.getCurrentTitle().fullText == "Template:" .. langcode .. "-head" and "interjection" or error("Part of speech must be specified in 1=") poscat = require(headword_module).canonicalize_pos(poscat) end local indexing_poscat = pos_in_1 and (misc_pos_with_gender[poscat] and "head_with_gender" or "head") or poscat local params = { ["suffix"] = boolean_param, ["nosuffix"] = boolean_param, ["id"] = true, ["json"] = boolean_param, ["pagename"] = {}, -- for testing } if pos_in_1 then params[1] = {required = true} -- required but ignored as already processed above end local head_is_head = pos_functions[indexing_poscat] and pos_functions[indexing_poscat].head_is_not_1 local headfield = head_is_head and "head" or pos_in_1 and 2 or 1 params[headfield] = head_is_head and true or {default = "+"} params.head2 = {replaced_by = false, instead = "use multiple comma-separated values in |" .. headfield .. "="} local tr_replaced_by = {replaced_by = false, instead = "use <tr:...> inline modifier on |" .. headfield .. "="} params.tr = tr_replaced_by params.tr2 = tr_replaced_by if pos_functions[indexing_poscat] then for key, val in pairs(pos_functions[indexing_poscat].params()) do params[key] = val end end local parargs = frame:getParent().args local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local data = { lang = lang, pos_category = poscat, orig_pos_category = poscat, categories = {}, heads = {}, genders = {}, inflections = {enable_auto_translit = true}, pagename = pagename, id = args.id, sort_key = args.sort, force_cat_output = force_cat, -- We expect a head always so the redundant head cat will be inaccurate. no_redundant_head_cat = true, } data.heads = parse_inflection(data, args, headfield, "is_head") for _, headobj in ipairs(data.heads) do if headobj.term == "+" then headobj.term = pagename end end data.is_suffix = false if args.suffix or ( not args.nosuffix and pagename:find("^%-") and poscat ~= "suffixes" and poscat ~= "suffix forms" ) then data.is_suffix = true data.pos_category = "suffixes" local singular_poscat = m_en_utilities.singularize(poscat) insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes") insert(data.inflections, {label = singular_poscat .. "-forming suffix"}) end if pos_functions[indexing_poscat] then pos_functions[indexing_poscat].func(data, args) end -- Do this after calling pos_functions[poscat].func() as it may modify data.heads (as verbs do). local irreg_translit = false for _, head in ipairs(data.heads) do if ar_translit.irregular_translit(head.term, head.tr) then irreg_translit = true break end end if irreg_translit then insert(data.categories, langname .. " terms with irregular pronunciations") end if args.json then return require("Module:JSON").toJSON(data) end return require(headword_module).full_headword(data) end ----------------------------------------------------------------------------------------- -- Gender handling -- ----------------------------------------------------------------------------------------- local valid_bare_genders = {false, "m", "f", "mf", "mfbysense", "mfequiv"} local valid_bare_numbers = {false, "d", "p"} local valid_bare_animacies = {false, "pr", "np"} local valid_genders = {} for _, gender in ipairs(valid_bare_genders) do for _, number in ipairs(valid_bare_numbers) do for _, animacy in ipairs(valid_bare_animacies) do local parts = {} local function ins_part(part) if part then insert(parts, part) end end ins_part(gender) ins_part(number) ins_part(animacy) local full_gender = concat(parts, "-") valid_genders[full_gender == "" and "?" or full_gender] = true end end end local function is_masc_sg(g) return g == "m" or g == "m-pr" or g == "m-np" end local function is_fem_sg(g) return g == "f" or g == "f-pr" or g == "f-np" end local function is_masc_fem_sg(g) g = g:gsub("%-pr", ""):gsub("%-np", "") return g == "mf" or g == "mfequiv" or g == "mfbysense" end local function add_gender_params(params, default) params[2] = {type = "genders", default = default or "?", template_default = "m"} params["g2"] = {replaced_by = false, instead = "use comma-separated values in |g="} end -- Handle gender in params 2=, inserting into `data.genders`. Also, if a lemma, insert categories into `data.categories` -- if the gender is unexpected for the form of the noun. (Note: If there are multiple genders, -- [[Module:gender and number]] will automatically insert 'Arabic POS with multiple genders'.) local function handle_gender(data, args, nonlemma, field) if not args[field or 2] then return end for _, gspec in ipairs(args[field or 2]) do if not valid_genders[gspec.spec] then error("Unrecognized gender: " .. gspec.spec) end end data.genders = args[field or 2] if nonlemma then return end for _, gspec in ipairs(data.genders) do local g = gspec.spec if is_masc_sg(g) or is_fem_sg(g) or is_masc_fem_sg(g) then local head = data.heads[1] if head then head = rsub(ar.reorder_shadda(remove_links(head.term)), ar.UNUOPT .. "$", "") local ends_with_tam = rfind(head, "^[^ ]*" .. ar.TAM .. "$") or rfind(head, "^[^ ]*" .. ar.TAM .. " ") if (is_masc_sg(g) or is_masc_fem_sg(g)) and ends_with_tam then insert(data.categories, langname .. " masculine terms with feminine ending") elseif (is_fem_sg(g) or is_masc_fem_sg(g)) and not ends_with_tam and not rfind(head, "[" .. ar.ALIF .. ar.AMAQ .. "]$") and not rfind(head, ar.ALIF .. ar.HAMZA .. "$") then insert(data.categories, langname .. " feminine terms lacking feminine ending") end end end end end ----------------------------------------------------------------------------------------- -- Inflection handlers -- ----------------------------------------------------------------------------------------- -- Add list parameters to `params` (a structure as passed to [[Module:parameters]]) for a parameter named `argpref`. -- If `argpref` is "*", add the nominal inflection parameters for construct state, definite state, etc. Related -- transliteration and gender parameters are no longer supported in favor of inline modifiers, and error messages are -- output if these parameters are used. local function add_infl_params(params, argpref) params[argpref] = {list = true, disallow_holes = true} params[argpref .. "tr"] = {replaced_by = false, instead = "use <tr:...> inline modifier on |" .. argpref .. "="} params[argpref .. "g"] = {replaced_by = false, instead = "use <g:...> inline modifier on |" .. argpref .. "="} end --[=[ Fetch a list of inflections from the arguments in `args` based on argument `field` (e.g. "pl"). Label with `label` (e.g. "plural"), which will appear in the headword. Insert into `data.inflections`, where `data` is the structure passed to [[Module:headword]]. If `generate_default` is specified, it should be a function of two arguments (`data`, `args`), which should generate the default value if no values are specified or if "+" is explicitly given. If `generate_default` isn't specified and the user gave no values, no inflection will be inserted. ]=] local function handle_infl(data, args, spec) local newinfls = parse_inflection(data, args, spec.field, false) if not newinfls[1] and spec.default_when_not_explicit and spec.default_when_not_explicit(data, args) then newinfls = {{term = "+"}} end if spec.handle then spec.handle(data, args, newinfls) end local default_specs = spec.allowed_defspecs if not default_specs then default_specs = spec.generate_default and {["+"] = true} or {} end local saw_defspec = false for _, newinfl in ipairs(newinfls) do if default_specs[newinfl.term] or newinfl.term == "~" then saw_defspec = true break end end if saw_defspec then local newnewinfls = {} for _, newinfl in ipairs(newinfls) do if default_specs[newinfl.term] then if newinfl.tr then error("Can't specify translit for default value '" .. newinfl.term .. "'") end local definfls = spec.generate_default(data, args, newinfl.term) for _, definfl in ipairs(definfls) do m_headword_utilities.combine_termobj_decorations(definfl, newinfl) insert(newnewinfls, definfl) end elseif newinfl.term == "~" then if newinfl.tr then error("Can't specify translit for head-requesting value '~'") end for _, headobj in ipairs(data.heads) do headobj = m_table.shallowCopy(headobj) m_headword_utilities.combine_termobj_decorations(headobj, newinfl) insert(newnewinfls, headobj) end else insert(newnewinfls, newinfl) end end newinfls = newnewinfls end if newinfls[1] then if newinfls[1].term == "--" then if newinfls[2] then error("Can't specify more than one term if first term is '--', meaning \"don't insert anything\"") end else insert_inflection(data, newinfls, spec.label, nil, spec.defgender, spec.field, spec.no_label, spec.usually_no_label) end end end local function add_infl_list_params(params, infl_list) for _, infl in ipairs(infl_list) do add_infl_params(params, infl.field) end end local function handle_infl_list_args(data, args, infl_list) for _, infl in ipairs(infl_list) do handle_infl(data, args, infl) end end ----------------------------------------------------------------------------------------- -- Default ending generators -- ----------------------------------------------------------------------------------------- local function make_conditional_default(specs) return function(data, args) local heads = data.heads if not heads[1] then heads = {{term = data.pagename}} end local newobjs = {} for _, headobj in ipairs(heads) do local term = ar.reorder_shadda(headobj.term) local tr = headobj.tr local matched = false for _, spec in ipairs(specs) do local from, fromtr, to, totr = unpack(spec) if from:find("^%^") then pref = rmatch(term, from .. "$") else pref = rmatch(term, "^(.*)" .. from .. "$") end if pref then term = pref .. to tr = replace_tr_ending(tr, fromtr, totr) matched = true headobj = m_table.shallowCopy(headobj) headobj.term = ar.undo_reorder_shadda(term) headobj.tr = tr insert(newobjs, headobj) break end end if not matched then error(("Internal error: No matching spec: head=%s"):format(dump(headobj))) end end return newobjs end end local default_feminine = make_conditional_default { {ar.AN .. ar.AMAQ, "an", ar.AAH, "āh"}, {ar.AN .. ar.ALIF, "an", ar.AAH, "āh"}, -- e.g. مُحْيًا {ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAH, "iya"}, {ar.IN, "in", ar.IYAH, "iya"}, {"", "", ar.AH, "a"}, } local default_masculine = make_conditional_default { -- tall alif substitutes for alif maqṣūra after a yāʔ {ar.Y .. ar.AAH, "āh", ar.AN .. ar.ALIF, "an"}, {ar.AAH, "āh", ar.AN .. ar.AMAQ, "an"}, -- handle the common case of final-weak feminine active participle with preceding hamza; -- the hamza-on-yāʔ always converts back to hamza on the line when preceded by ā (alif) but -- may not otherwise, so we just leave it alone in that case {ar.ALIF .. ar.HAMZA_ON_YA .. ar.IYAH, "iya", ar.HAMZA .. ar.IN, "in"}, {ar.IYAH, "iya", ar.IN, "in"}, {ar.AH, "a", "", ""}, {"", "", "", ""}, } local default_masculine_plural = make_conditional_default { {ar.AN .. ar.AMAQ, "an", ar.AWN, "awn"}, {ar.AN .. ar.ALIF, "an", ar.AWN, "awn"}, -- e.g. مُحْيًا {ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_WAW .. ar.UUN, "ūn"}, {ar.IN, "in", ar.UUN, "ūn"}, {"", "", ar.UUN, "ūn"}, } local default_feminine_plural = make_conditional_default { -- صَلَاة pl. صَلَوَات and أَدَاة pl. أَدَوَات and similar; but نَوَاة and وَفَاة with a و in them become نَوَيَات and وَفَيَات; -- and longer terms like مُبَارَاة and كُمَّثْرَاة invariably form their plural in -يَات. {"^([^و]" .. ar.A .. "[^و])" .. ar.AAH, "āh", ar.A .. ar.W .. ar.AAT, "awāt"}, {ar.AAH, "āh", ar.AYAAT, "ayāt"}, {ar.AN .. ar.AMAQ, "an", ar.AYAAT, "ayāt"}, {ar.AN .. ar.ALIF, "an", ar.AYAAT, "ayāt"}, -- e.g. مُحْيًا {ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAAT, "iyāt"}, {ar.IN, "in", ar.IYAAT, "iyāt"}, {ar.AH, "a", ar.AAT, "āt"}, {"", "", ar.AAT, "āt"}, } local default_masculine_dual = make_conditional_default { {ar.AN .. ar.AMAQ, "an", ar.AYAAN, "ayān"}, {ar.AN .. ar.ALIF, "an", ar.AYAAN, "ayān"}, -- e.g. مُحْيًا {ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAAN, "iyān"}, {ar.IN, "in", ar.IYAAN, "iyān"}, {"", "", ar.AAN, "ān"}, } local default_feminine_dual = make_conditional_default { {ar.AN .. ar.AMAQ, "an", ar.AATAAN, "ātān"}, {ar.AN .. ar.ALIF, "an", ar.AATAAN, "ātān"}, -- e.g. مُحْيًا {ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IY .. ar.ATAAN, "iyatān"}, {ar.IN, "in", ar.IY .. ar.ATAAN, "iyatān"}, {"", "", ar.ATAAN, "atān"}, } -- Return whether `term` is a nisba noun or adjective, ending in -iyy or -iyyah. `nisba_val` is the value of -- args.nisba; if non-nil, it overrides any auto-determination based on the shape of the term. local function term_is_nisba(term, nisba_val) if nisba_val ~= nil then return nisba_val end term = ar.reorder_shadda(term) -- necessary to avoid issues with e.g. أُورُوبِّيّ. local pref = rmatch(term, "^(.*)" .. ar.IYY .. ar.UN .. "?$") if not pref then pref = rmatch(term, "^(.*)" .. ar.IYYAH .. ar.UN .. "?$") end -- Avoid false positives for words like قَوِيّ "strong" and صَبِيّ "boy". There may be other false positives -- but this should catch most of them and will avoid very many false negatives. return pref and not rfind(pref, "^[^ا]" .. ar.A .. ".$") end ----------------------------------------------------------------------------------------- -- Adjectives -- ----------------------------------------------------------------------------------------- local function is_defaulting_adjective(data, args) return data.orig_pos_category == "defaulting adjectives" end local adj_field_elative = {field = "el", label = "<<elative>>"} local adj_inflections = { adj_field_inf, adj_field_obl, adj_field_def, {field = "f", label = "feminin", generate_default = default_feminine, default_when_not_explicit = is_defaulting_adjective}, {field = "d", label = "maskulin duaan", generate_default = default_masculine_dual}, {field = "fd", label = "feminin duaan", generate_default = default_feminine_dual}, {field = "cpl", label = "am jamak"}, {field = "pl", label = "maskulin jamak", generate_default = default_masculine_plural, default_when_not_explicit = is_defaulting_adjective}, {field = "fpl", label = "feminin jamak", generate_default = default_feminine_plural, default_when_not_explicit = is_defaulting_adjective}, } local function get_adj_params() local params = {} add_infl_list_params(params, adj_inflections) add_infl_params(params, "el") params.nisba = boolean_param return params end local function handle_adj_args(data, args) handle_infl_list_args(data, args, adj_inflections) handle_infl(data, args, adj_field_elative) for _, headobj in ipairs(data.heads) do if term_is_nisba(headobj.term, args.nisba) then insert(data.categories, "Kata sifat relatif (nisba) bahasa " .. langname) break end end end pos_functions["Kata sifat"] = { params = get_adj_params, func = handle_adj_args, } pos_functions["defaulting adjectives"] = { params = get_adj_params, func = function(data, args) data.pos_category = "Kata sifat" handle_adj_args(data, args) end, } ----------------------------------------------------------------------------------------- -- Nouns, etc. -- ----------------------------------------------------------------------------------------- local function get_masc_or_feminine_gender(data, default_type) local saw_m, saw_f, saw_mf for _, gender in ipairs(data.genders) do if is_masc_sg(gender.spec) then saw_m = true elseif is_fem_sg(gender.spec) then saw_f = true elseif is_masc_fem_sg(gender.spec) then saw_mf = true end end if saw_mf or saw_m and saw_f then error("Can't generate default for " .. default_type .. " when gender is both masculine and feminine") elseif saw_m then return "m" elseif saw_f then return "f" else error("Can't generate default for " .. default_type .. " when gender is not specified as " .. "maskulin or feminin tunggal") end end local function is_defaulting_noun(data, args) return data.orig_pos_category == "defaulting nouns" end local noun_field_dual = { field = "d", label = "dual", generate_default = function(data, args) local gender = get_masc_or_feminine_gender(data, "noun dual") if gender == "m" then return default_masculine_dual(data, args) else return default_feminine_dual(data, args) end end, } local noun_field_plural = { field = "pl", label = "jamak", generate_default = function(data, args, defspec) local gender = get_masc_or_feminine_gender(data, "noun plural") if gender == "m" then if defspec == "+f" then return default_feminine_plural(data, args) else return default_masculine_plural(data, args) end elseif defspec == "+f" then error("Can't specify '+f' with feminine gender; just use '+'") else return default_feminine_plural(data, args) end end, -- Handle the case where pl=-, indicating an uncountable noun. handle = function(data, args, terms) if terms[1] and terms[1] == "-" then insert(data.categories, langname .. " uncountable nouns") if args.pauc and args.pauc[1] then error("Can't specify paucals when pl=-") end end end, allowed_defspecs = {["+"] = true, ["+f"] = true}, default_when_not_explicit = is_defaulting_noun, no_label = "<<uncountable>>", usually_no_label = "usually <<uncountable>>", } local noun_field_paucal = { field = "pauc", label = "<<paucal>>", generate_default = default_feminine_plural, } local noun_field_feminine = { field = "f", label = "feminin", generate_default = default_feminine, default_when_not_explicit = function(data, args) if data.orig_pos_category ~= "defaulting nouns" then return nil end local gender = get_masc_or_feminine_gender(data, "defaulting-if-masculine noun feminine") return gender == "m" end, } local noun_field_masculine = { field = "m", label = "maskulin", generate_default = default_masculine, default_when_not_explicit = function(data, args) if data.orig_pos_category ~= "defaulting nouns" then return nil end local gender = get_masc_or_feminine_gender(data, "defaulting-if-feminine noun masculine") return gender == "f" end, } local noun_basic_inflections = { noun_field_cons, noun_field_inf, noun_field_obl, noun_field_def, } local noun_shared_inflections = { noun_field_dual, noun_field_plural, } local noun_extra_inflections = { noun_field_paucal, noun_field_feminine, noun_field_masculine, } local function get_noun_params() local params = {} add_gender_params(params) add_infl_list_params(params, noun_basic_inflections) add_infl_list_params(params, noun_shared_inflections) add_infl_list_params(params, noun_extra_inflections) params.nisba = boolean_param return params end local function handle_noun_args(data, args) handle_gender(data, args) handle_infl_list_args(data, args, noun_basic_inflections) handle_infl_list_args(data, args, noun_shared_inflections) handle_infl_list_args(data, args, noun_extra_inflections) for _, headobj in ipairs(data.heads) do if term_is_nisba(headobj.term, args.nisba) then insert(data.categories, langname .. " relative nouns (nisba)") break end end end pos_functions["Kata nama"] = { params = get_noun_params, func = handle_noun_args, } pos_functions["defaulting nouns"] = { params = get_noun_params, func = function(data, args) data.pos_category = "Kata nama" handle_noun_args(data, args) end, } local noun_field_singulative = {field = "sing", label = "<<singulative>>", defgender = "f", generate_default = default_feminine} local noun_field_collective = {field = "coll", label = "<<collective>>", defgender = "m", generate_default = default_masculine} local function handle_sing_coll_noun_infls(data, args, otherinfl, otherlabel, othergender) -- Handle sing= (corresponding singulative noun) or coll= (corresponding collective noun) and their gender handle_infl(data, args, otherinfl, otherlabel, nil, othergender) handle_infl_list_args(data, args, sing_coll_noun_inflections) end local function get_singulative_collective_noun_params(defgender, otherinfl) local params = {} add_gender_params(params, defgender) add_infl_list_params(params, noun_basic_inflections) add_infl_params(params, otherinfl) add_infl_list_params(params, noun_shared_inflections) add_infl_params(params, "pauc") return params end pos_functions["collective nouns"] = { params = function() return get_singulative_collective_noun_params("m", "sing") end, func = function(data, args) data.pos_category = "Kata nama" insert(data.categories, langname .. " collective nouns") m_headword_utilities.insert_fixed_inflection { headdata = data, label = "<<collective>>", } handle_gender(data, args) handle_infl_list_args(data, args, noun_basic_inflections) handle_infl(data, args, noun_field_singulative) handle_infl_list_args(data, args, noun_shared_inflections) handle_infl(data, args, noun_field_paucal) end } pos_functions["singulative nouns"] = { params = function() return get_singulative_collective_noun_params("f", "coll") end, func = function(data, args) data.pos_category = "Kata nama" insert(data.categories, langname .. " singulative nouns") m_headword_utilities.insert_fixed_inflection { headdata = data, label = "<<singulative>>", } handle_gender(data, args) handle_infl_list_args(data, args, noun_basic_inflections) handle_infl(data, args, noun_field_collective) handle_infl_list_args(data, args, noun_shared_inflections) handle_infl(data, args, noun_field_paucal) end } -- FIXME: Do numerals really behave almost as nouns? They vary by masc/fem. pos_functions["numerals"] = { params = get_noun_params, func = function(data, args) insert(data.categories, langname .. " cardinal numbers") handle_noun_args(data, args) end } pos_functions["Kata nama khas"] = { params = get_noun_params, func = handle_noun_args, } local function get_pronoun_params() local params = {} add_gender_params(params, defgender) add_infl_list_params(params, noun_basic_inflections) add_infl_list_params(params, noun_shared_inflections) add_infl_params(params, "f") return params end pos_functions["pronouns"] = { params = get_pronoun_params, func = function(data, args) handle_gender(data, args) handle_infl_list_args(data, args, noun_basic_inflections) handle_infl_list_args(data, args, noun_shared_inflections) handle_infl(data, args, noun_field_feminine) end } ----------------------------------------------------------------------------------------- -- Non-lemma forms -- ----------------------------------------------------------------------------------------- local valid_forms = list_to_set( { "I", "II", "III", "IV", "V", "VI", "VII", "VIII", "IX", "X", "XI", "XII", "XIII", "XIV", "XV", "Iq", "IIq", "IIIq", "IVq" }) -- FIXME: Partly duplicated in [[Module:ar-inflections]]. local function handle_conj_form(data, args) local form = args[2] if form then if not valid_forms[form] then error("Invalid verb conjugation form " .. form) end insert(data.inflections, { label = "[[Appendix:Arabic verbs#Form " .. form .. "|form " .. form .. "]]" }) end end pos_functions["verb forms"] = { params = function() return { [2] = {}, } end, func = function(data, args) handle_conj_form(data, args) end } local function get_participle_params() local params = get_adj_params() params[2] = {} return params end pos_functions["active participles"] = { params = get_participle_params, func = function(data, args) data.pos_category = "participles" insert(data.categories, langname .. " active participles") handle_conj_form(data, args) handle_infl_list_args(data, args, adj_inflections) end } pos_functions["passive participles"] = { params = get_participle_params, func = function(data, args) data.pos_category = "participles" insert(data.categories, langname .. " passive participles") handle_conj_form(data, args) handle_infl_list_args(data, args, adj_inflections) end } ----------------------------------------------------------------------------------------- -- Verbs -- ----------------------------------------------------------------------------------------- pos_functions["Kata kerja"] = { head_is_not_1 = true, params = function() return { [1] = {}, -- Comma-separated lists with possible inline modifiers ["past"] = {}, ["past1s"] = {}, ["nonpast"] = {}, ["vn"] = {}, ["noautolinktext"] = {type = "boolean"}, ["noautolinkverb"] = {type = "boolean"}, } end, func = function(data, args) local ar_verb = require(ar_verb_module) local alternant_multiword_spec = args[1] ~= "-" and ar_verb.do_generate_forms(args, "ar-verb", data.pagename) or nil local function do_slot(slots_to_check, override, label, slot_is_headword) -- Do this even with an override so we can return the correct filled slot. local slot, slotval if alternant_multiword_spec then for _, potential_slot in ipairs(slots_to_check) do slotval = alternant_multiword_spec.forms[potential_slot] if slotval then slot = potential_slot break end end end local function get_slot_values() local terms = {} for _, form in ipairs(slotval) do local term = { term = form.form, id = form.id, genders = form.genders, pos = form.pos, lit = form.lit, } term.tr = form.translit if form.footnotes then local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(form.footnotes) term.q = quals term.refs = refs end insert(terms, term) end return terms end if override then local override_param_mods = { alt = {}, t = { -- [[Module:headword]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = {}, g = { -- [[Module:headword]] expects the genders in "genders". item_dest = "genders", type = "genders", }, pos = {}, lit = {}, id = {}, -- Qualifiers and labels q = { type = "qualifier", }, qq = { type = "qualifier", }, l = { type = "labels", }, ll = { type = "labels", }, ref = { -- [[Module:headword]] expects the references in "refs". item_dest = "refs", type = "references", }, } local function generate_obj(formval, parse_err) if formval == "+" then return {term = "+", underlying_terms = get_slot_values()} end local val, uncertain = formval:match("^(.*)(%?)$") val = val or formval uncertain = not not uncertain local ar, translit = val:match("^(.*)//(.*)$") if not ar then ar = formval end local retval = {term = ar, uncertain = uncertain} retval.tr = translit end local terms if override:find("<") then terms = require(parse_utilities_module).parse_inline_modifiers(override, { paramname = paramname, param_mods = override_param_mods, generate_obj = generate_obj, splitchar = "[,،]", escape_fun = escape_comma_whitespace, unescape_fun = unescape_comma_whitespace, }) else terms = split_on_comma(override) for i, split in ipairs(terms) do terms[i] = generate_obj(split) end end -- See if + was supplied and we have to potentially flatten multiple default terms and harmonize -- default properties with override properties. local saw_underlying_terms = false for _, term in ipairs(terms) do if term.underlying_terms then saw_underlying_terms = true break end end if saw_underlying_terms then -- Flatten any default terms, copying the corresponding override properties over the default -- properties. Non-default terms get inserted directly. local flattened = {} for _, term in ipairs(terms) do if term.underlying_terms then for _, underlying in ipairs(term.underlying_terms) do for k, v in pairs(term) do if k ~= "term" and k ~= "underlying_terms" then if k == "uncertain" then underlying.uncertain = underlying.uncertain or v elseif type(v) ~= "table" or v[1] then -- Don't copy empty lists (which are the default) over possibly non-empty -- lists. underlying[k] = v end end end insert(flattened, underlying) end else insert(flattened, term) end end terms = flattened end if not slot_is_headword then terms.label = label end return terms, slot elseif not alternant_multiword_spec then return nil, slot else if not slotval then if slot_is_headword then -- FIXME, put "uncertain" as qualifier? Does this ever happen? return nil, slot elseif alternant_multiword_spec.slot_uncertain[slot] then return {label = label .. " uncertain"}, slot elseif alternant_multiword_spec.slot_explicitly_missing[slot] then return {label = "no " .. label}, slot else -- just say nothing about this slot return nil, slot end end local terms = get_slot_values() if not slot_is_headword then terms.label = label end return terms, slot end end local gloss_parts = {} for _, vform in ipairs(alternant_multiword_spec.verb_forms) do insert(gloss_parts, "[[Appendix:Arabic verbs#Form " .. vform .. "|" .. vform .. "]]") end if gloss_parts[1] then data.gloss = concat(gloss_parts, ", ") end if data.heads[1] and args.past then error("Can't specify both head= and past= to {{ar-verb}}; prefer past=") end if not alternant_multiword_spec.has_active then insert(data.inflections, {label = "passive-only"}) end -- Do this always so `past_slot` is correctly filled. local past, past_slot = do_slot(ar_verb.potential_lemma_slots, args.past, "-", "slot is headword") if data.heads[1] then -- user specified head=; don't override with past= or slot 'past_3sm' etc. else if past then data.heads = past end end local should_do_past1s = not not args.past1s if not should_do_past1s then local is_form_I = false for _, vform in ipairs(alternant_multiword_spec.verb_forms) do if vform == "I" then is_form_I = true break end end if is_form_I then require(inflection_utilities_module).map_word_specs(alternant_multiword_spec, function(base) if base.verb_form == "I" then for _, vowel_spec in ipairs(base.conj_vowels) do -- For form-I geminate verbs, the final vowel of the past is elided in the citation form. -- We want to display it for all cases other than active a~u and a~i (the most common -- cases). if vowel_spec.weakness == "geminate" then if ar_verb.is_passive_only(base.passive) then should_do_past1s = true break end local past_vowel = ar_verb.rget(vowel_spec.past) local nonpast_vowel = ar_verb.rget(vowel_spec.nonpast) if not (past_vowel == ar.A and (nonpast_vowel == ar.U or nonpast_vowel == ar.I)) then should_do_past1s = true break end end end -- FIXME, provide way of breaking early from map_word_specs(). end end) end end local past1s if should_do_past1s then past1s, _ = do_slot({"past_1s", "past_pass_1s"}, args.past1s, "first-person singular past") if past1s then insert(data.inflections, past1s) end end local nonpast_slots if not past_slot or past_slot:find("^past_") then nonpast_slots = {"ind_3ms", "ind_pass_3ms", "imp_2ms"} else nonpast_slots = {} end local nonpast, _ = do_slot(nonpast_slots, args.nonpast, "non-past") if nonpast then insert(data.inflections, nonpast) end local vn, _ = do_slot({"vn"}, args.vn, "verbal noun") if vn then insert(data.inflections, vn) end -- FIXME: Should we insert categories? Conjugation also does it and is more likely to be accurate. --for _, cat in ipairs(alternant_multiword_spec.categories) do -- insert(data.categories, cat) --end --[=[ -- FIXME: Review this to see if we need to port it. -- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to -- incorporate any links in that head into the 1= specification, use the infinitive generated by -- [[Module:pt-verb]] in place of the user-specified or auto-generated head. This was copied from -- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on -- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the -- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian -- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Portuguese equivalent). if not data.user_specified_heads[1] or ( not data.user_specified_heads[2] and alternant_multiword_spec.incorporated_headword_head_into_lemma ) then data.heads = {} for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do local quals, refs = require(inflection_utilities_module). convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes) insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs}) end end ]=] end } ----------------------------------------------------------------------------------------- -- Generic parts of speech -- ----------------------------------------------------------------------------------------- pos_functions.head_with_gender = { params = function() return { [3] = {type = "genders"}, } end, func = function(data, args) handle_gender(data, args, "nonlemma", 3) end, } return export 1mc21n9hd7rlm8q5tnhtlhefbmz6j3l Modul:rhymes 828 11249 375346 228342 2026-09-22T03:11:54Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708593|92708593]]) 375346 Scribunto text/plain local export = {} local force_cat = false -- for testing local decorations_module = "Module:decorations" local IPA_module = "Module:IPA" local parameters_module = "Module:parameters" local parameter_utilities_module = "Module:parameter utilities" local rhymes_styles_css_module = "Module:rhymes/styles.css" local TemplateStyles_module = "Module:TemplateStyles" local utilities_module = "Module:utilities" local rhymes_data = require("Module:rhymes/data") local concat = table.concat local insert = table.insert local function track(page) require("Module:debug/track")("rhymes/" .. page) return true end local function tag_rhyme(rhyme, lang) local formatted_rhyme, cats, err formatted_rhyme, cats, err = require(IPA_module).format_IPA(lang, rhyme, "raw") return formatted_rhyme, cats, err end local function make_rhyme_link(lang, link_rhyme, display_rhyme) local retval, cats local prefix = "[[Rima:Bahasa " if rhymes_data.link_to_category_langs[lang:getCode()] then prefix = "[[:Kategori:Rima:Bahasa " end if not link_rhyme then retval = concat{prefix, lang:getCanonicalName(), "|", lang:getCanonicalName(), "]]"} cats = {} else local formatted_rhyme, err formatted_rhyme, cats, err = tag_rhyme(display_rhyme or link_rhyme, lang) retval = concat{prefix, lang:getCanonicalName(), "/", link_rhyme, "|", formatted_rhyme, "]]", err} end return retval, cats end --[==[ Implementation of {{tl|rhymes row}}. ]==] function export.show_row(frame) local args = require(parameters_module).process( frame.getParent and frame:getParent().args or frame, { [1] = {required = true, type = "full language"}, [2] = {required = true}, [3] = {}, } ) if not args[1] then return "[[Rhymes:English/aɪmz|<span class=\"IPA\">-aɪmz</span>]]" end -- Discard cleanup categories from make_rhyme_link(). return (make_rhyme_link(args[1], args[2], "-" .. args[2])) .. (args[3] and (" (''" .. args[3] .. "'')") or "") end do local function add_syllable_categories(categories, lang, rhyme, num_syl) local prefix = "Rima:Bahasa " .. lang .. "/" .. rhyme insert(categories, prefix) if num_syl then for _, n in ipairs(num_syl) do local c if n > 1 then c = prefix .. "/" .. n .. " suku kata" else c = prefix .. "/1 suku kata" end insert(categories, c) end end end --[==[ Meant to be called from a module. `data` is a table containing the following fields: * `lang`: language object for the rhymes; * `rhymes`: a list of rhymes, each described by an object which specifies the rhyme, optional number of syllables, and optional decoration fields: ** `rhyme`: the rhyme itself; ** `num_syl`: {nil} or a list of numbers, specifying the number of syllables of the word with this rhyme; optional and currently used only for categorization; if omitted, defaults to the top-level `num_syl`; ** `separator`: {nil} or the string used to separate this rhyme from the preceding one when displayed; defaults to the top-level `separator`; ** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the rhyme in question; ** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the rhyme in question; ** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed directly before the rhyme in question; ** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the rhyme in question; ** `refs`: {nil} or a list of references or reference specs to add directly after the rhyme; the value of a list item is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}}) and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the actual reference, located in the {{cd|<nowiki><references /></nowiki>}} section; ** `nocat`: if {true}, suppress categorization for this rhyme only; * `num_syl`: {nil} or a list of numbers, specifying the number of syllables for all rhymes; optional and currently used only for categorization; overridable at the individual rhyme level; * `separator`: {nil} or a string, specifying the separator displayed before all rhymes but the first; by default, {", "}; overridable at the individual rhyme level; * `q`: {nil} or a list of overall left regular qualifier strings, displayed before the initial caption; * `qq`: {nil} or a list of overall right regular qualifier strings, displayed after all rhymes; * `a`: {nil} or a list of overall left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the initial caption; * `aa`: {nil} or a list of right accent qualifier strings, displayed after all rhymes; * `sort`: {nil} or sort key; * `caption`: {nil} or string specifying the caption to use, in place of {"Rhymes"}; a colon and space is automatically added after the caption; * `nocaption`: if {true}, suppress the caption display; * `nocat`: if {true}, suppress categorization; * `force_cat`: if {true}, force categorization even on non-mainspace pages. If both regular and accent qualifiers on the same side and at the same level are specified, the accent qualifiers precede the regular qualifiers on both left and right. '''WARNING''': Destructively modifies the objects inside the `rhymes` field. Note that the number of syllables is currently used only for categorization; if present, an extra category will be added such as [[:Category:Rhymes:Italian/ino/3 syllables]] in addition to [[:Category:Rhymes:Italian/ino]]. ]==] function export.format_rhymes(data) local langname = data.lang:getFullName() local parts = {} local categories = {} local overall_sep = data.separator or ", " for i, r in ipairs(data.rhymes) do local rhyme = r.rhyme local link, link_cats = make_rhyme_link(data.lang, rhyme, "-" .. rhyme) if not r.nocat and not data.nocat then for _, cat in ipairs(link_cats) do insert(categories, cat) end end if r.qualifiers then -- FIXME; Added 2026-09-18; remove in a month error("qualifiers= not allowed here; use q=") end if r.q and r.q[1] or r.qq and r.qq[1] or r.a and r.a[1] or r.aa and r.aa[1] or r.refs and r.refs[1] then link = require(decorations_module).format_decorations { lang = data.lang, text = link, q = r.q, qq = r.qq, a = r.a, aa = r.aa, refs = r.refs, } end insert(parts, r.separator or i > 1 and overall_sep or "") insert(parts, link) if not r.nocat and not data.nocat then add_syllable_categories(categories, langname, rhyme, r.num_syl or data.num_syl) end end local text = concat(parts) if not data.nocaption then text = (data.caption or "Rima") .. ": " .. text end if data.qualifiers then -- FIXME; Added 2026-09-18; remove in a month error("overall qualifiers= not allowed here; use q=") end if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then text = require(decorations_module).format_decorations { lang = data.lang, text = text, q = data.q, qq = data.qq, a = data.a, aa = data.aa, } end if categories[1] then local cat_text = require(utilities_module).format_categories(categories, data.lang, data.sort, nil, force_cat or data.force_cat) text = text .. cat_text end return text end end --[==[ Implementation of {{tl|rhymes}}. ]==] function export.show(frame) local parent_args = frame:getParent().args local compat = parent_args.lang local offset = compat and 0 or 1 local lang_param = compat and "lang" or 1 local plain = {} local boolean = {type = "boolean"} local params = { [lang_param] = {required = true, type = "language", default = "ms"}, [1 + offset] = {list = true, required = true, disallow_holes = true, default = "aɪmz"}, ["caption"] = plain, ["nocaption"] = boolean, ["nocat"] = boolean, ["sort"] = plain, } local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { { param = "s", item_dest = "num_syl", separate_no_index = true, type = "number", sublist = true, }, {group = {"q", "a", "ref"}}, } local rhymes, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 1 + offset, term_dest = "rhyme", track_module = "rhymes", } local lang = args[lang_param] local data = { lang = lang, rhymes = rhymes, num_syl = args.s.default, caption = args.caption, nocaption = args.nocaption, nocat = args.nocat, sort = args.sort, q = args.q.default, qq = args.qq.default, a = args.a.default, aa = args.aa.default, } return export.format_rhymes(data) end --[==[ Implementation of {{tl|rhymes nav}}. ]==] function export.show_nav(frame) local args = require(parameters_module).process( frame:getParent().args, { [1] = {required = true, type = "full language", default = "und"}, [2] = {list = true, allow_holes = true}, ["nocat"] = {type = "boolean"}, } ) local lang = args[1] local langname = lang:getCanonicalName() local parts = args[2] -- Create steps -- FIXME: We should probably use format_categories() in [[Module:utilities]] rather than constructing categories -- manually. local categories = {} -- Here and below, we ignore any cleanup categories coming out of make_rhyme_link() by adding an extra set of parens -- around the call to make_rhyme_link() to cause the second argument (the categories) to be ignored. {{rhymes nav}} -- is run on a rhymes page so it's not clear we want the page to be added to any such categories, if they exist. local steps = {"[[Wikikamus:Rima|Rima]]", (make_rhyme_link(lang))} if #parts > 0 then local last = parts[#parts] parts[#parts] = nil local prefix = "" for i, part in ipairs(parts) do prefix = prefix .. part parts[i] = prefix end for _, part in ipairs(parts) do insert(steps, (make_rhyme_link(lang, part .. "-", "-" .. part .. "-"))) end if last == "-" then insert(steps, (make_rhyme_link(lang, prefix, "-" .. prefix))) insert(categories, "[[Kategori:Rima bahasa " .. langname .. (prefix == "" and "" or "/" .. prefix .. "-") .. "| ]]") elseif mw.title.getCurrentTitle().text == langname .. "/" .. prefix .. last .. "-" then -- DO NOT replace with mw.loadData("Module:headword/data").pagename as we need the root portion insert(steps, (make_rhyme_link(lang, prefix .. last .. "-", "-" .. prefix .. last .. "-"))) insert(categories, "[[Kategori:Rima bahasa " .. langname .. "/" .. prefix .. last .. "-|-]]") else insert(steps, (make_rhyme_link(lang, prefix .. last, "-" .. prefix .. last))) insert(categories, "[[Kategori:Rima bahasa " .. langname .. (prefix == "" and "" or "/" .. prefix .. "-") .. "|" .. last .. "]]") end elseif lang:getCode() ~= "und" then insert(categories, "[[Kategori:Rima bahasa " .. langname .. "| ]]") end if mw.title.getCurrentTitle().nsText == "Rima" then frame:callParserFunction("DISPLAYTITLE", mw.title.getCurrentTitle().fullText:gsub( "/(.+)$", function (rhyme) return "/" .. (tag_rhyme(rhyme, lang)) -- ignore cleanup categories end)) end local templateStyles = require(TemplateStyles_module)(rhymes_styles_css_module) local ol = mw.html.create("ol") for _, step in ipairs(steps) do ol:node(mw.html.create("li"):wikitext(step)) end local div = mw.html.create("div") :attr("role", "navigation") :attr("aria-label", "Breadcrumb") :addClass("ts-rhymesBreadcrumbs") :node(ol) local formatted_cats = args.nocat and "" or concat(categories) return templateStyles .. tostring(div) .. formatted_cats end return export ht9j9li2mw2ejqi4l9yfaule57m75kd Modul:etymology 828 11468 375371 373823 2026-09-22T05:09:38Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708527|92708527]]) 375371 Scribunto text/plain local export = {} -- For testing local force_cat = false local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local insert = table.insert local new_title = mw.title.new local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_language_data_module_name(...) get_language_data_module_name = require(languages_module).getDataModuleName return get_language_data_module_name(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function serial_comma_join(...) serial_comma_join = require(table_module).serialCommaJoin return serial_comma_join(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function track(page, code) local tracking_page = "etymology/" .. page debug_track(tracking_page) if code then debug_track(tracking_page .. "/" .. code) end end local function join_segs(segs, conj) if not segs[2] then return segs[1] elseif conj == "and" or conj == "or" then return serial_comma_join(segs, {conj = conj}) end local sep if conj == "," or conj == ";" then sep = conj .. " " elseif conj == "/" then sep = "/" elseif conj == "~" then sep = " ~ " elseif conj then error(("Internal error: Unrecognized conjunction \"%s\""):format(conj)) else error(("Internal error: No value supplied for conjunction"):format(conj)) end return concat(segs, sep) end -- Returns true if `lang` is the same as `source`, or a variety of it. local function lang_is_source(lang, source) return lang:getCode() == source:getCode() or lang:hasParent(source) end --[==[ Format one or more links as specified in `termobjs`, a list of term objects of the format accepted by `full_link()` in [[Module:links]], including decorations (qualifiers, labels and references). `conj` is used to join multiple terms and must be specified if there is more than one term. `template_name` is the template name used in debug tracking and must be specified. Optional `sourcetext` is text to prepend to the concatenated terms, separated by a space if the concatenated terms are non-empty (which is always the case unless there is a single term with the value "-"). If `decorations_on_outside` is given, any decorations specified in the first term go on the outside of (i.e before) `sourcetext`; otherwise they will end up on the inside. ]==] function export.format_links(termobjs, conj, template_name, sourcetext, decorations_on_outside) if not template_name then error("Internal error: Must specify `template_name` to format_links()") end for i, termobj in ipairs(termobjs) do if termobj.lang:hasType("family") or termobj.lang:getFamilyCode() == "qfa-sub" then if termobj.term and termobj.term ~= "-" then debug_track(template_name .. "/family-with-term") end termobj.term = "-" end if termobj.term == "-" then --[=[ [[Special:WhatLinksHere/Wiktionary:Tracking/cognate/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/derived/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/borrowed/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/calque/no-term]] ]=] debug_track(template_name .. "/no-term") termobjs[i] = i == 1 and sourcetext or "" else if i == 1 and decorations_on_outside and sourcetext then termobj.pretext = sourcetext .. " " sourcetext = nil end termobjs[i] = (i == 1 and sourcetext and sourcetext .. " " or "") .. full_link(termobj, "term") end end return join_segs(termobjs, conj) end function export.get_display_and_cat_name(source, raw) local display, cat_name if source:getCode() == "und" then display = "tidak ditentukan" cat_name = "bahasa lain" elseif source:getCode() == "mul" then display = raw and "rentas bahasa" or "[[w:Translingualisme|rentas bahasa]]" cat_name = "Rentas bahasa" elseif source:getCode() == "mul-tax" then display = raw and "taxonomic name" or "[[w:Tatanama biologi|nama taksonomi]]" cat_name = "Nama taksonomi" else display = raw and source:getCanonicalName() or source:makeWikipediaLink() cat_name = source:getDisplayForm() end return display, cat_name end function export.insert_source_cat_get_display(data) local categories, lang, source = data.categories, data.lang, data.source local display, cat_name = export.get_display_and_cat_name(source, data.raw) if lang and not data.nocat then -- Add the category, but only if there is a current language if not categories then categories = {} end local langname = lang:getFullName() -- If `lang` is an etym-only language, we need to check both it and its parent full language against `source`. -- Otherwise if e.g. `lang` is Medieval Latin and `source` is Latin, we'll end up wrongly constructing a -- category 'Latin terms derived from Latin'. insert(categories, "Perkataan bahasa " .. langname .. ( lang_is_source(lang, source) and " dipinjam balik ke dalam bahasa " .. cat_name or " " .. (data.borrowing_type or "diterbitkan") .. " daripada bahasa " .. cat_name )) end return display, categories end function export.format_source(data) local lang, sort_key = data.lang, data.sort_key -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/sortkey]] if sort_key then track("sortkey") end local display, categories = export.insert_source_cat_get_display(data) if lang and not data.nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories categories = format_categories(categories, lang, sort_key, nil, data.force_cat or force_cat) else categories = "" end return "<span class=\"etyl\">" .. display .. categories .. "</span>" end --[==[ Format sources for etymology templates such as {{tl|bor}}, {{tl|der}}, {{tl|inh}} and {{tl|cog}}. There may potentially be more than one source language (except currently {{tl|inh}}, which doesn't support it because it doesn't really make sense). In that case, all but the last source language is linked to the first term, but only if there is such a term and this linking makes sense, i.e. either (1) the term page exists after stripping diacritics according to the source language in question, or (2) the result of stripping diacritics according to the source language in question results in a different page from the same process applied with the last source language. For example, {{m|ru|соля́нка}} will link to [[солянка]] but {{m|en|соля́нка}} will link to [[соля́нка]] with an accent, and since they are different pages, the use of English as a non-final source with term 'соля́нка' will link to [[соля́нка]] even though it doesn't exist, on the assumption that it is merely a redlink that might exist. If none of the above criteria apply, a non-final source language will be linked to the Wikipedia entry for the language, just as final source languages always are. `data` contains the following fields: * `lang`: The destination language object into which the terms were borrowed, inherited or otherwise derived. Used for categorization and can be nil, as with {{tl|cog}}. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see above. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_sources(data) local lang, sources, terms, borrowing_type, sort_key, categories, nocat = data.lang, data.sources, data.terms, data.borrowing_type, data.sort_key, data.categories, data.nocat local term1, sources_n, source_segs = terms[1], #sources, {} local final_link_page local term1_term, term1_sc = term1.term, term1.sc if sources_n > 1 and term1_term and term1_term ~= "-" then final_link_page = get_link_page(term1_term, sources[sources_n], term1_sc) end for i, source in ipairs(sources) do local seg, display_term if i < sources_n and term1_term and term1_term ~= "-" then local link_page = get_link_page(term1_term, source, term1_sc) display_term = (link_page ~= final_link_page) or (link_page and not not new_title(link_page):getContent()) end -- TODO: if the display forms or transliterations are different, display the terms separately. if display_term then local display, this_cats = export.insert_source_cat_get_display{ lang = lang, source = source, borrowing_type = borrowing_type, raw = true, categories = categories, nocat = nocat, } seg = language_link { lang = source, term = term1_term, alt = display, tr = "-", } if lang and not nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories this_cats = format_categories(this_cats, lang, sort_key, nil, data.force_cat or force_cat) else this_cats = "" end seg = "<span class=\"etyl\">" .. seg .. this_cats .. "</span>" else seg = export.format_source{ lang = lang, source = source, borrowing_type = borrowing_type, sort_key = sort_key, categories = categories, nocat = nocat, } end insert(source_segs, seg) end return join_segs(source_segs, data.sourceconj or "and") end -- Internal implementation of {{cognate}}/{{cog}} template. function export.format_cognate(data) return export.format_derived { sources = data.sources, terms = data.terms, sort_key = data.sort_key, sourceconj = data.sourceconj, conj = data.conj, template_name = "cognate", force_cat = data.force_cat, } end --[==[ Internal implementation of {{derived}}/{{der}} template. This is called externally from [[Module:affix]], [[Module:affixusex]] and [[Module:see]] and needs to support decorations (qualifiers, labels and references) on the outside of the sources for use by those modules. `data` contains the following fields: * `lang`: The destination language object into which the terms were derived. Used for categorization and can be nil, as with {{tl|cog}}; in this case, no categories are added. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see `format_sources()`. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `conj`: Conjunction used to separate multiple terms. '''Required'''. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are as for `conj` above. * `decorations_on_outside`: If specified, any decorations (qualifiers, labels or references) in the first term in `terms` will be displayed on the outside of (before) the source language(s) in `sources`. Normally this should be specified if there is only one term possible in `terms`. * `template_name`: Name of the template invoking this function. Must be specified. Only used for tracking pages. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_derived(data) local terms = data.terms local sourcetext = export.format_sources(data) return export.format_links(terms, data.conj, data.template_name, sourcetext, data.decorations_on_outside) end function export.insert_borrowed_cat(categories, lang, source) if lang_is_source(lang, source) then return end -- If both are the same, we want e.g. [[:Category:English terms borrowed back into English]] not -- [[:Category:English terms borrowed from English]]; the former is inserted automatically by format_source(). -- The second parameter here doesn't matter as it only affects `display`, which we don't use. insert(categories, "Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada " .. select(2, export.get_display_and_cat_name(source, "raw"))) end -- Internal implementation of {{borrowed}}/{{bor}} template. function export.format_borrowed(data) local categories = {} if not data.nocat then local lang = data.lang for _, source in ipairs(data.sources) do export.insert_borrowed_cat(categories, lang, source) end end data = shallow_copy(data) data.categories = categories return export.format_links(data.terms, data.conj, "borrowed", export.format_sources(data)) end do -- Generate the non-ancestor error message. local function show_language(lang) local retval = ("%s (%s)"):format(lang:makeCategoryLink(), lang:getCode()) if lang:hasType("etymology-only") then retval = retval .. (" (an etymology-only language whose regular parent is %s)"):format( show_language(lang:getParent())) end return retval end -- Check that `lang` has `otherlang` (which may be an etymology-only language) as an ancestor. Throw an error if -- not. When `lang` is a family, verifies that `otherlang` is a language in that family. function export.check_ancestor(lang, otherlang) -- When `lang` is a family, verify `otherlang` is in that family or in its parent family. if lang.hasType and lang:hasType("family") then local family_code = lang:getCode() local function in_family_code(fcode, other) if not fcode or fcode == "" then return false end if other.inFamily and other:inFamily(fcode) then return true end if other.getFamilyCode and other:getFamilyCode() == fcode then return true end return false end local in_family = in_family_code(family_code, otherlang) if not in_family then local parent_code if lang.getParent then local parent_family = lang:getParent() if parent_family and parent_family.getCode then parent_code = parent_family:getCode() end end if not parent_code and family_code:find("-", 1, true) then parent_code = family_code:match("^(.+)-[^-]+$") end if parent_code then in_family = in_family_code(parent_code, otherlang) end end if not in_family then local other_display = (otherlang.getCanonicalName and otherlang:getCanonicalName()) or (otherlang.getCode and otherlang:getCode()) or tostring(otherlang) local fam_display = (lang.getCanonicalName and lang:getCanonicalName()) or family_code error(("%s bukan dalam keluarga %s; leluhur diwarisi di bawah keluarga mestilah bahasa dalam keluarga tersebut atau keluarga induknya.") :format(other_display, fam_display)) end return end -- FIXME: I don't know if this function works correctly with etym-only languages in `lang`. I have fixed up -- the module link code appropriately (June 2024) but the remaining logic is untouched. if lang:hasAncestor(otherlang) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/variety]] -- Track inheritance from varieties of Latin that shouldn't have any descendants (everything except Old Latin, Classical Latin and Vulgar Latin). if otherlang:getFullCode() == "la" then otherlang = otherlang:getCode() if not (otherlang == "itc-ola" or otherlang == "la-cla" or otherlang == "la-vul") then track("bad ancestor", otherlang) end end return end local ancestors = lang:getAncestors() local postscript local etym_module_link = lang:hasType("etymology-only") and "[[Module:etymology languages/data]] or " or "" local module_link = "[[" .. get_language_data_module_name(lang:getFullCode()) .. "]]" if not ancestors[1] then postscript = show_language(lang) .. " tidak mempunyai leluhur." else local ancestor_list = {} for _, ancestor in ipairs(ancestors) do insert(ancestor_list, show_language(ancestor)) end postscript = ("Leluhur bahasa%s kepada %s %s %s."):format( ancestors[2] and "" or "", lang:getCanonicalName(), ancestors[2] and "adalah" or "adalah", concat(ancestor_list, " dan ")) end error(("%s tidak ditetapkan sebagai luluhur kepada %s dalam %s%s. %s") :format(show_language(otherlang), show_language(lang), etym_module_link, module_link, postscript)) end end -- Internal implementation of {{inherited}}/{{inh}} template. function export.format_inherited(data) local lang, terms, nocat = data.lang, data.terms, data.nocat local source = terms[1].lang local categories = {} if not nocat then insert(categories,"Perkataan bahasa " .. lang:getFullName() .. " diwariskan daripada bahasa " .. source:getCanonicalName()) end export.check_ancestor(lang, source) data = shallow_copy(data) data.categories = categories data.source = source return export.format_links(terms, data.conj, "inherited", export.format_source(data)) end -- Internal implementation of "misc variant" templates such as {{abbrev}}, {{clipping}}, {{reduplication}} and the like. function export.format_misc_variant(data) local lang, notext, terms, cats, parts = data.lang, data.notext, data.terms, data.cats, {} if not notext then insert(parts, data.text) end if terms[1] then if not notext then -- FIXME: If term is given as '-', we should consider displaying just "Clipping" not "Clipping of". insert(parts, " " .. (data.oftext or "bagi")) end local termparts = {} -- Make links out of all the parts. for _, termobj in ipairs(terms) do local result if termobj.lang then result = export.format_derived { lang = lang, terms = {termobj}, sources = termobj.termlangs or {termobj.lang}, template_name = "misc_variant", decorations_on_outside = true, force_cat = data.force_cat, } else termobj.lang = lang result = export.format_links({termobj}, nil, "misc_variant") end table.insert(termparts, result) end local linktext = join_segs(termparts, data.conj) if not notext and linktext ~= "" then insert(parts, " ") end insert(parts, linktext) end local categories = {} if not data.nocat and cats then for _, cat in ipairs(cats) do insert(categories, cat .. " bahasa " .. lang:getFullName()) end end if categories[1] then insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end -- Implementation of miscellaneous templates such as {{unknown}} and {{onomatopoeia}} that have no associated terms. function export.format_misc_variant_no_term(data) local parts = {} if not data.notext then insert(parts, data.title) end if not data.nocat and data.cat then local lang, categories = data.lang, {} insert(categories, data.cat .. " bahasa " .. lang:getFullName()) insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end return export sp30xyxjrn9ihshbbbmeckrd2czv8nv 375373 375371 2026-09-22T05:18:56Z Hakimi97 2668 tambah "bahasa" untuk kategori perkataan bahasa A dipinjam daripada "bahasa" B 375373 Scribunto text/plain local export = {} -- For testing local force_cat = false local debug_track_module = "Module:debug/track" local languages_module = "Module:languages" local links_module = "Module:links" local table_module = "Module:table" local utilities_module = "Module:utilities" local concat = table.concat local insert = table.insert local new_title = mw.title.new local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function format_categories(...) format_categories = require(utilities_module).format_categories return format_categories(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_language_data_module_name(...) get_language_data_module_name = require(languages_module).getDataModuleName return get_language_data_module_name(...) end local function get_link_page(...) get_link_page = require(links_module).get_link_page return get_link_page(...) end local function language_link(...) language_link = require(links_module).language_link return language_link(...) end local function serial_comma_join(...) serial_comma_join = require(table_module).serialCommaJoin return serial_comma_join(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function track(page, code) local tracking_page = "etymology/" .. page debug_track(tracking_page) if code then debug_track(tracking_page .. "/" .. code) end end local function join_segs(segs, conj) if not segs[2] then return segs[1] elseif conj == "and" or conj == "or" then return serial_comma_join(segs, {conj = conj}) end local sep if conj == "," or conj == ";" then sep = conj .. " " elseif conj == "/" then sep = "/" elseif conj == "~" then sep = " ~ " elseif conj then error(("Internal error: Unrecognized conjunction \"%s\""):format(conj)) else error(("Internal error: No value supplied for conjunction"):format(conj)) end return concat(segs, sep) end -- Returns true if `lang` is the same as `source`, or a variety of it. local function lang_is_source(lang, source) return lang:getCode() == source:getCode() or lang:hasParent(source) end --[==[ Format one or more links as specified in `termobjs`, a list of term objects of the format accepted by `full_link()` in [[Module:links]], including decorations (qualifiers, labels and references). `conj` is used to join multiple terms and must be specified if there is more than one term. `template_name` is the template name used in debug tracking and must be specified. Optional `sourcetext` is text to prepend to the concatenated terms, separated by a space if the concatenated terms are non-empty (which is always the case unless there is a single term with the value "-"). If `decorations_on_outside` is given, any decorations specified in the first term go on the outside of (i.e before) `sourcetext`; otherwise they will end up on the inside. ]==] function export.format_links(termobjs, conj, template_name, sourcetext, decorations_on_outside) if not template_name then error("Internal error: Must specify `template_name` to format_links()") end for i, termobj in ipairs(termobjs) do if termobj.lang:hasType("family") or termobj.lang:getFamilyCode() == "qfa-sub" then if termobj.term and termobj.term ~= "-" then debug_track(template_name .. "/family-with-term") end termobj.term = "-" end if termobj.term == "-" then --[=[ [[Special:WhatLinksHere/Wiktionary:Tracking/cognate/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/derived/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/borrowed/no-term]] [[Special:WhatLinksHere/Wiktionary:Tracking/calque/no-term]] ]=] debug_track(template_name .. "/no-term") termobjs[i] = i == 1 and sourcetext or "" else if i == 1 and decorations_on_outside and sourcetext then termobj.pretext = sourcetext .. " " sourcetext = nil end termobjs[i] = (i == 1 and sourcetext and sourcetext .. " " or "") .. full_link(termobj, "term") end end return join_segs(termobjs, conj) end function export.get_display_and_cat_name(source, raw) local display, cat_name if source:getCode() == "und" then display = "tidak ditentukan" cat_name = "bahasa lain" elseif source:getCode() == "mul" then display = raw and "rentas bahasa" or "[[w:Translingualisme|rentas bahasa]]" cat_name = "Rentas bahasa" elseif source:getCode() == "mul-tax" then display = raw and "taxonomic name" or "[[w:Tatanama biologi|nama taksonomi]]" cat_name = "Nama taksonomi" else display = raw and source:getCanonicalName() or source:makeWikipediaLink() cat_name = source:getDisplayForm() end return display, cat_name end function export.insert_source_cat_get_display(data) local categories, lang, source = data.categories, data.lang, data.source local display, cat_name = export.get_display_and_cat_name(source, data.raw) if lang and not data.nocat then -- Add the category, but only if there is a current language if not categories then categories = {} end local langname = lang:getFullName() -- If `lang` is an etym-only language, we need to check both it and its parent full language against `source`. -- Otherwise if e.g. `lang` is Medieval Latin and `source` is Latin, we'll end up wrongly constructing a -- category 'Latin terms derived from Latin'. insert(categories, "Perkataan bahasa " .. langname .. ( lang_is_source(lang, source) and " dipinjam balik ke dalam bahasa " .. cat_name or " " .. (data.borrowing_type or "diterbitkan") .. " daripada bahasa " .. cat_name )) end return display, categories end function export.format_source(data) local lang, sort_key = data.lang, data.sort_key -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/sortkey]] if sort_key then track("sortkey") end local display, categories = export.insert_source_cat_get_display(data) if lang and not data.nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories categories = format_categories(categories, lang, sort_key, nil, data.force_cat or force_cat) else categories = "" end return "<span class=\"etyl\">" .. display .. categories .. "</span>" end --[==[ Format sources for etymology templates such as {{tl|bor}}, {{tl|der}}, {{tl|inh}} and {{tl|cog}}. There may potentially be more than one source language (except currently {{tl|inh}}, which doesn't support it because it doesn't really make sense). In that case, all but the last source language is linked to the first term, but only if there is such a term and this linking makes sense, i.e. either (1) the term page exists after stripping diacritics according to the source language in question, or (2) the result of stripping diacritics according to the source language in question results in a different page from the same process applied with the last source language. For example, {{m|ru|соля́нка}} will link to [[солянка]] but {{m|en|соля́нка}} will link to [[соля́нка]] with an accent, and since they are different pages, the use of English as a non-final source with term 'соля́нка' will link to [[соля́нка]] even though it doesn't exist, on the assumption that it is merely a redlink that might exist. If none of the above criteria apply, a non-final source language will be linked to the Wikipedia entry for the language, just as final source languages always are. `data` contains the following fields: * `lang`: The destination language object into which the terms were borrowed, inherited or otherwise derived. Used for categorization and can be nil, as with {{tl|cog}}. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see above. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_sources(data) local lang, sources, terms, borrowing_type, sort_key, categories, nocat = data.lang, data.sources, data.terms, data.borrowing_type, data.sort_key, data.categories, data.nocat local term1, sources_n, source_segs = terms[1], #sources, {} local final_link_page local term1_term, term1_sc = term1.term, term1.sc if sources_n > 1 and term1_term and term1_term ~= "-" then final_link_page = get_link_page(term1_term, sources[sources_n], term1_sc) end for i, source in ipairs(sources) do local seg, display_term if i < sources_n and term1_term and term1_term ~= "-" then local link_page = get_link_page(term1_term, source, term1_sc) display_term = (link_page ~= final_link_page) or (link_page and not not new_title(link_page):getContent()) end -- TODO: if the display forms or transliterations are different, display the terms separately. if display_term then local display, this_cats = export.insert_source_cat_get_display{ lang = lang, source = source, borrowing_type = borrowing_type, raw = true, categories = categories, nocat = nocat, } seg = language_link { lang = source, term = term1_term, alt = display, tr = "-", } if lang and not nocat then -- Format categories, but only if there is a current language; {{cog}} currently gets no categories this_cats = format_categories(this_cats, lang, sort_key, nil, data.force_cat or force_cat) else this_cats = "" end seg = "<span class=\"etyl\">" .. seg .. this_cats .. "</span>" else seg = export.format_source{ lang = lang, source = source, borrowing_type = borrowing_type, sort_key = sort_key, categories = categories, nocat = nocat, } end insert(source_segs, seg) end return join_segs(source_segs, data.sourceconj or "and") end -- Internal implementation of {{cognate}}/{{cog}} template. function export.format_cognate(data) return export.format_derived { sources = data.sources, terms = data.terms, sort_key = data.sort_key, sourceconj = data.sourceconj, conj = data.conj, template_name = "cognate", force_cat = data.force_cat, } end --[==[ Internal implementation of {{derived}}/{{der}} template. This is called externally from [[Module:affix]], [[Module:affixusex]] and [[Module:see]] and needs to support decorations (qualifiers, labels and references) on the outside of the sources for use by those modules. `data` contains the following fields: * `lang`: The destination language object into which the terms were derived. Used for categorization and can be nil, as with {{tl|cog}}; in this case, no categories are added. * `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are handled specially; see `format_sources()`. * `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as multiple term objects, the non-final source objects link to the first term object. * `conj`: Conjunction used to separate multiple terms. '''Required'''. Currently recognized values are `and`, `or`, `,`, `;`, `/` and `~`. * `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized values are as for `conj` above. * `decorations_on_outside`: If specified, any decorations (qualifiers, labels or references) in the first term in `terms` will be displayed on the outside of (before) the source language(s) in `sources`. Normally this should be specified if there is only one term possible in `terms`. * `template_name`: Name of the template invoking this function. Must be specified. Only used for tracking pages. * `sort_key`: Sort key for categories. Usually nil. * `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be added. * `nocat`: Don't add any categories to the page. * `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}. * `force_cat`: Force category generation on non-mainspace pages. ]==] function export.format_derived(data) local terms = data.terms local sourcetext = export.format_sources(data) return export.format_links(terms, data.conj, data.template_name, sourcetext, data.decorations_on_outside) end function export.insert_borrowed_cat(categories, lang, source) if lang_is_source(lang, source) then return end -- If both are the same, we want e.g. [[:Category:English terms borrowed back into English]] not -- [[:Category:English terms borrowed from English]]; the former is inserted automatically by format_source(). -- The second parameter here doesn't matter as it only affects `display`, which we don't use. insert(categories, "Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada bahasa " .. select(2, export.get_display_and_cat_name(source, "raw"))) end -- Internal implementation of {{borrowed}}/{{bor}} template. function export.format_borrowed(data) local categories = {} if not data.nocat then local lang = data.lang for _, source in ipairs(data.sources) do export.insert_borrowed_cat(categories, lang, source) end end data = shallow_copy(data) data.categories = categories return export.format_links(data.terms, data.conj, "borrowed", export.format_sources(data)) end do -- Generate the non-ancestor error message. local function show_language(lang) local retval = ("%s (%s)"):format(lang:makeCategoryLink(), lang:getCode()) if lang:hasType("etymology-only") then retval = retval .. (" (an etymology-only language whose regular parent is %s)"):format( show_language(lang:getParent())) end return retval end -- Check that `lang` has `otherlang` (which may be an etymology-only language) as an ancestor. Throw an error if -- not. When `lang` is a family, verifies that `otherlang` is a language in that family. function export.check_ancestor(lang, otherlang) -- When `lang` is a family, verify `otherlang` is in that family or in its parent family. if lang.hasType and lang:hasType("family") then local family_code = lang:getCode() local function in_family_code(fcode, other) if not fcode or fcode == "" then return false end if other.inFamily and other:inFamily(fcode) then return true end if other.getFamilyCode and other:getFamilyCode() == fcode then return true end return false end local in_family = in_family_code(family_code, otherlang) if not in_family then local parent_code if lang.getParent then local parent_family = lang:getParent() if parent_family and parent_family.getCode then parent_code = parent_family:getCode() end end if not parent_code and family_code:find("-", 1, true) then parent_code = family_code:match("^(.+)-[^-]+$") end if parent_code then in_family = in_family_code(parent_code, otherlang) end end if not in_family then local other_display = (otherlang.getCanonicalName and otherlang:getCanonicalName()) or (otherlang.getCode and otherlang:getCode()) or tostring(otherlang) local fam_display = (lang.getCanonicalName and lang:getCanonicalName()) or family_code error(("%s bukan dalam keluarga %s; leluhur diwarisi di bawah keluarga mestilah bahasa dalam keluarga tersebut atau keluarga induknya.") :format(other_display, fam_display)) end return end -- FIXME: I don't know if this function works correctly with etym-only languages in `lang`. I have fixed up -- the module link code appropriately (June 2024) but the remaining logic is untouched. if lang:hasAncestor(otherlang) then -- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/variety]] -- Track inheritance from varieties of Latin that shouldn't have any descendants (everything except Old Latin, Classical Latin and Vulgar Latin). if otherlang:getFullCode() == "la" then otherlang = otherlang:getCode() if not (otherlang == "itc-ola" or otherlang == "la-cla" or otherlang == "la-vul") then track("bad ancestor", otherlang) end end return end local ancestors = lang:getAncestors() local postscript local etym_module_link = lang:hasType("etymology-only") and "[[Module:etymology languages/data]] or " or "" local module_link = "[[" .. get_language_data_module_name(lang:getFullCode()) .. "]]" if not ancestors[1] then postscript = show_language(lang) .. " tidak mempunyai leluhur." else local ancestor_list = {} for _, ancestor in ipairs(ancestors) do insert(ancestor_list, show_language(ancestor)) end postscript = ("Leluhur bahasa%s kepada %s %s %s."):format( ancestors[2] and "" or "", lang:getCanonicalName(), ancestors[2] and "adalah" or "adalah", concat(ancestor_list, " dan ")) end error(("%s tidak ditetapkan sebagai luluhur kepada %s dalam %s%s. %s") :format(show_language(otherlang), show_language(lang), etym_module_link, module_link, postscript)) end end -- Internal implementation of {{inherited}}/{{inh}} template. function export.format_inherited(data) local lang, terms, nocat = data.lang, data.terms, data.nocat local source = terms[1].lang local categories = {} if not nocat then insert(categories,"Perkataan bahasa " .. lang:getFullName() .. " diwariskan daripada bahasa " .. source:getCanonicalName()) end export.check_ancestor(lang, source) data = shallow_copy(data) data.categories = categories data.source = source return export.format_links(terms, data.conj, "inherited", export.format_source(data)) end -- Internal implementation of "misc variant" templates such as {{abbrev}}, {{clipping}}, {{reduplication}} and the like. function export.format_misc_variant(data) local lang, notext, terms, cats, parts = data.lang, data.notext, data.terms, data.cats, {} if not notext then insert(parts, data.text) end if terms[1] then if not notext then -- FIXME: If term is given as '-', we should consider displaying just "Clipping" not "Clipping of". insert(parts, " " .. (data.oftext or "bagi")) end local termparts = {} -- Make links out of all the parts. for _, termobj in ipairs(terms) do local result if termobj.lang then result = export.format_derived { lang = lang, terms = {termobj}, sources = termobj.termlangs or {termobj.lang}, template_name = "misc_variant", decorations_on_outside = true, force_cat = data.force_cat, } else termobj.lang = lang result = export.format_links({termobj}, nil, "misc_variant") end table.insert(termparts, result) end local linktext = join_segs(termparts, data.conj) if not notext and linktext ~= "" then insert(parts, " ") end insert(parts, linktext) end local categories = {} if not data.nocat and cats then for _, cat in ipairs(cats) do insert(categories, cat .. " bahasa " .. lang:getFullName()) end end if categories[1] then insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end -- Implementation of miscellaneous templates such as {{unknown}} and {{onomatopoeia}} that have no associated terms. function export.format_misc_variant_no_term(data) local parts = {} if not data.notext then insert(parts, data.title) end if not data.nocat and data.cat then local lang, categories = data.lang, {} insert(categories, data.cat .. " bahasa " .. lang:getFullName()) insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat)) end return concat(parts) end return export bbu26wo843f49ze9uihvwyyik4ah21g Modul:syllables 828 11549 375369 134546 2026-09-22T04:51:22Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91137329|91137329]]) 375369 Scribunto text/plain local export = {} local m_str_utils = require("Module:string utilities") local gsub = m_str_utils.gsub local match = m_str_utils.match local toNFD = mw.ustring.toNFD local U = m_str_utils.char local diphthongs = mw.loadData("Module:IPA/data").diphthongs local vowels = mw.loadData("Module:IPA/data/symbols").vowels .. "ᵻ" .. "ᵿ" --[[ No use for this at the moment, though it is an interesting catalogue. It might be usable for phonetic transcriptions. Diacritics added to vowels: inverted breve above, inverted breve below, up tack, down tack, left tack, right tack, diaeresis (above), diaeresis below, right half ring, left half ring, plus sign below, minus sign below, combining x above, rhotic hook, tilde (above), tilde below ligature tie (combining double breve), ligature tie below ]] local diacritics = U( 0x311, 0x32F, 0x31D, 0x31E, 0x318, 0x319, 0x308, 0x324, 0x339, 0x31C, 0x31F, 0x320, 0x33D, 0x2DE, 0x303, 0x330, 0x361, 0x35C ) --[[ combining acute and grave tone marks, circumflex ]]-- local tone = "[" .. U(0x341, 0x340, 0x302) .. "]" local nonsyllabicDiacritics = U(0x311, 0x32F) local syllabicDiacritics = U(0x0329, 0x030D) local ties = U(0x361, 0x35C) -- long, half-long, extra short local lengthDiacritics = U(0x2D0, 0x2D1, 0x306) local vowel = "[" .. vowels .. "]" .. tone .. "?" local tie = "[" .. ties .. "]" local nonsyllabicDiacritic = "[" .. nonsyllabicDiacritics .. "]" local syllabicDiacritic = "[" .. syllabicDiacritics .. "]" local UTF8Char = "[%z\1-\127\194-\244][\128-\191]*" function export.getVowels(remainder, lang) if string.find(remainder, "^[%[/]?%-") or string.find(remainder, "%-[%[/]?$") then return nil end -- If a hyphen is at the beginning or end of the transcription, do not count syllables. local count = 0 local diphs = diphthongs[lang:getCode()] or {} remainder = toNFD(remainder) remainder = string.gsub(remainder, "%((.*)%)", "%1") -- Remove parentheses. while remainder ~= "" do -- Ignore nonsyllabic vowels remainder = gsub(remainder, "^" .. vowel .. nonsyllabicDiacritic, "") local m = match(remainder, "^." .. syllabicDiacritic) or -- Syllabic consonant match(remainder, "^" .. vowel .. tie .. vowel) -- Tie bar -- Starts with a recognised diphthong? for _, diph in ipairs(diphs) do if m then break end m = m or match(remainder, "^" .. diph) end -- If we haven't found anything yet, just match on a single vowel m = m or match(remainder, "^" .. vowel) if m then -- Found a vowel, add it count = count + 1 remainder = string.sub(remainder, #m + 1) else -- Found a non-vowel, skip it remainder = string.gsub(remainder, "^" .. UTF8Char, "") end end if count ~= 0 then return count end return nil end function export.countVowels2Test(frame) local params = { [1] = {required = true}, [2] = {default = ""}, } local args = require("Module:parameters").process(frame.args, params) local lang = require("Module:languages").getByCode(args[1]) or require("Module:languages").err(args[1], 1) local count = export.getVowels(args[2], lang) return 'The text "' .. args[2] .. '" contains ' .. count .. ' vowels.' end local function countVowels(text) text = toNFD(text) or error("Invalid UTF-8") local _, count = gsub(text, vowel, "") local _, sequenceCount = gsub(text, vowel.."+", "") local _, nonsyllabicCount = gsub(text, vowel .. nonsyllabicDiacritic, "") local _, tieCount = gsub(text, vowel .. tie .. vowel, "") local diphthongCount = count - (nonsyllabicCount + tieCount) return count, sequenceCount, diphthongCount end local function countDiphthongs(text, lang) text = toNFD(text) or error("Invalid UTF-8") local diphthongs = diphthongs[lang:getCode()] or {} local _, count local total = 0 if diphthongs then for i, diphthong in pairs(diphthongs) do _, count = gsub(text, diphthong, "") total = total + count end end return total end function export.countVowels(frame) local params = { [1] = {default = ""}, } local args = require("Module:parameters").process(frame.args, params) local count, sequenceCount, diphthongCount = countVowels(args[1]) local outputs = {} table.insert(outputs, (count or 'an unknown number of') .. ' vowels') table.insert(outputs, (sequenceCount or 'an unknown number of') .. ' vowel sequences') table.insert(outputs, (diphthongCount or 'an unknown number of') .. ' vowels or vowels and diphthongs') return 'The text "' .. args[1] .. '" contains ' .. mw.text.listToText(outputs) .. "." end function export.countVowelsDiphthongs(frame) local params = { [1] = {required = true}, [2] = {default = ""}, } local args = require("Module:parameters").process(frame.args, params) local lang = require("Module:languages").getByCode(args[1]) or require("Module:languages").err(args[1], 1) local vowels = countVowels(args[2]) local count = vowels - countDiphthongs(args[2], lang) or 0 local out = 'The text "' .. args[2] .. '" contains ' .. (count or 'an unknown number of') if count == 1 then out = out .. ' vowel or diphthong.' else out = out .. ' vowels or diphthongs.' end return out end return export 2vcs9nzyxjbyq8j9l53t5ul8l7zkkz7 Modul:IPA/data/symbols 828 11596 375368 218693 2026-09-22T04:49:53Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91727113|91727113]]) 375368 Scribunto text/plain local data = {} --[=[ Valid IPA symbols. Currently almost all values of "title" and "link" keys are just the comments that were used in [[Module:IPA]]. The "link" fields should be checked (those that start with an uppercase letter are checked). ]=] --[=[ local phones = {} -- Vowels. phones["i"] = { close = true, front = true, unrounded = true, vowel = true, } phones["e"] = { ["close-mid"] = true, front = true, unrounded = true, vowel = true, } phones["ɛ"] = { ["open-mid"] = true, front = true, unrounded = true, vowel = true, } phones["æ"] = { ["near-open"] = true, front = true, unrounded = true, vowel = true, } phones["a"] = { open = true, front = true, unrounded = true, vowel = true, } phones["y"] = { close = true, front = true, rounded = true, vowel = true, } phones["ø"] = { ["close-mid"] = true, front = true, rounded = true, vowel = true, } phones["œ"] = { ["open-mid"] = true, front = true, rounded = true, vowel = true, } phones["ɶ"] = { open = true, front = true, rounded = true, vowel = true, } phones["ɪ"] = { ["near-close"] = true, ["near-front"] = true, unrounded = true, vowel = true, } phones["ʏ"] = { ["near-close"] = true, ["near-front"] = true, rounded = true, vowel = true, } phones["ɨ"] = { close = true, central = true, unrounded = true, vowel = true, } phones["ᵻ"] = { ["near-close"] = true, central = true, unrounded = true, vowel = true, } phones["ɘ"] = { ["close-mid"] = true, central = true, unrounded = true, vowel = true, } phones["ɜ"] = { ["open-mid"] = true, central = true, unrounded = true, vowel = true, } phones["ɝ"] = { rhotic = true, ["open-mid"] = true, central = true, unrounded = true, vowel = true, } phones["ə"] = { mid = true, central = true, vowel = true, } phones["ɚ"] = { rhotic = true, mid = true, central = true, vowel = true, } phones["ɐ"] = { ["near-open"] = true, central = true, vowel = true, } phones["ʉ"] = { close = true, central = true, rounded = true, vowel = true, } phones["ᵿ"] = { ["near-close"] = true, central = true, rounded = true, vowel = true, } phones["ɵ"] = { ["close-mid"] = true, central = true, rounded = true, vowel = true, } phones["ɞ"] = { ["open-mid"] = true, central = true, rounded = true, vowel = true, } phones["ʊ"] = { ["near-close"] = true, ["near-back"] = true, rounded = true, vowel = true, } phones["ɯ"] = { close = true, back = true, unrounded = true, vowel = true, } phones["ɤ"] = { ["close-mid"] = true, back = true, unrounded = true, vowel = true, } phones["ʌ"] = { ["open-mid"] = true, back = true, unrounded = true, vowel = true, } phones["ɑ"] = { open = true, back = true, unrounded = true, vowel = true, } phones["u"] = { close = true, back = true, rounded = true, vowel = true, } phones["o"] = { ["close-mid"] = true, back = true, rounded = true, vowel = true, } phones["ɔ"] = { ["open-mid"] = true, back = true, rounded = true, vowel = true, } phones["ɒ"] = { open = true, back = true, rounded = true, vowel = true, } -- Nasals. phones["m"] = { voiced = true, bilabial = true, nasal = true, } phones["ɱ"] = { voiced = true, labiodental = true, nasal = true, } phones["n"] = { voiced = true, alveolar = true, nasal = true, } phones["ɳ"] = { voiced = true, retroflex = true, nasal = true, } phones["ɲ"] = { voiced = true, palatal = true, nasal = true, } phones["ŋ"] = { voiced = true, velar = true, nasal = true, } phones["𝼇"] = { voiced = true, velodorsal = true, nasal = true, } phones["ɴ"] = { voiced = true, uvular = true, nasal = true, } -- Plosives. phones["p"] = { voiceless = true, bilabial = true, plosive = true, } phones["b"] = { voiced = true, bilabial = true, plosive = true, } phones["t"] = { voiceless = true, alveolar = true, plosive = true, } phones["d"] = { voiced = true, alveolar = true, plosive = true, } phones["ʈ"] = { voiceless = true, retroflex = true, plosive = true, } phones["ɖ"] = { voiced = true, retroflex = true, plosive = true, } phones["c"] = { voiceless = true, palatal = true, plosive = true, } phones["ɟ"] = { voiced = true, palatal = true, plosive = true, } phones["k"] = { voiceless = true, velar = true, plosive = true, } phones["ɡ"] = { voiced = true, velar = true, plosive = true, } phones["𝼃"] = { voiceless = true, velodorsal = true, plosive = true, } phones["𝼁"] = { voiced = true, velodorsal = true, plosive = true, } phones["q"] = { voiceless = true, uvular = true, plosive = true, } phones["ɢ"] = { voiced = true, uvular = true, plosive = true, } phones["ꞯ"] = { voiceless = true, ["upper-pharyngeal"] = true, plosive = true, } phones["𝼂"] = { voiced = true, ["upper-pharyngeal"] = true, plosive = true, } phones["ʡ"] = { epiglottal = true, plosive = true, } phones["ʔ"] = { glottal = true, plosive = true, } -- Fricatives. phones["ɸ"] = { voiceless = true, bilabial = true, fricative = true, } phones["β"] = { voiced = true, bilabial = true, fricative = true, } phones["ʍ"] = { voiceless = true, ["labial-velar"] = true, fricative = true, } phones["f"] = { voiceless = true, labiodental = true, fricative = true, } phones["v"] = { voiced = true, labiodental = true, fricative = true, } phones["θ"] = { voiceless = true, dental = true, ["non-sibilant"] = true, fricative = true, } phones["ð"] = { voiced = true, dental = true, ["non-sibilant"] = true, fricative = true, } phones["s"] = { voiceless = true, alveolar = true, sibilant = true, fricative = true, } phones["z"] = { voiced = true, alveolar = true, sibilant = true, fricative = true, } phones["ɬ"] = { voiceless = true, alveolar = true, lateral = true, fricative = true, } phones["ɮ"] = { voiced = true, alveolar = true, lateral = true, fricative = true, } phones["ʃ"] = { voiceless = true, postalveolar = true, sibilant = true, fricative = true, } phones["ʒ"] = { voiced = true, postalveolar = true, sibilant = true, fricative = true, } phones["ʂ"] = { voiceless = true, retroflex = true, sibilant = true, fricative = true, } phones["ʐ"] = { voiced = true, retroflex = true, sibilant = true, fricative = true, } phones["ꞎ"] = { voiceless = true, retroflex = true, lateral = true, fricative = true, } phones["𝼅"] = { voiced = true, retroflex = true, lateral = true, fricative = true, } phones["ɕ"] = { voiceless = true, ["alveolo-palatal"] = true, sibilant = true, fricative = true, } phones["ʑ"] = { voiced = true, ["alveolo-palatal"] = true, sibilant = true, fricative = true, } phones["ç"] = { voiceless = true, palatal = true, fricative = true, } phones["ʝ"] = { voiced = true, palatal = true, fricative = true, } phones["𝼆"] = { voiceless = true, palatal = true, lateral = true, fricative = true, } phones["ɧ"] = { voiceless = true, ["palatal-velar"] = true, fricative = true, } phones["x"] = { voiceless = true, velar = true, fricative = true, } phones["ɣ"] = { voiced = true, velar = true, fricative = true, } phones["𝼄"] = { voiceless = true, velar = true, lateral = true, fricative = true, } phones["ʩ"] = { voiceless = true, velopharyngeal = true, fricative = true, } phones["χ"] = { voiceless = true, uvular = true, fricative = true, } phones["ʁ"] = { voiced = true, uvular = true, fricative = true, } phones["ħ"] = { voiceless = true, pharyngeal = true, fricative = true, } phones["ʕ"] = { voiced = true, pharyngeal = true, fricative = true, } phones["ʜ"] = { voiceless = true, epiglottal = true, fricative = true, } phones["ʢ"] = { voiced = true, epiglottal = true, fricative = true, } phones["h"] = { voiceless = true, glottal = true, fricative = true, } phones["ɦ"] = { voiced = true, glottal = true, fricative = true, } -- Approximants. phones["ʋ"] = { voiced = true, labiodental = true, approximant = true, } phones["ɥ"] = { voiced = true, ["labial–palatal"] = true, approximant = true, } phones["w"] = { voiced = true, ["labial–velar"] = true, approximant = true, } phones["ɹ"] = { voiced = true, alveolar = true, approximant = true, } phones["ꭨ"] = { ["velarized or pharyngealized"] = true, voiced = true, alveolar = true, approximant = true, } phones["l"] = { voiced = true, alveolar = true, lateral = true, approximant = true, } phones["ɫ"] = { ["velarized or pharyngealized"] = true, voiced = true, alveolar = true, lateral = true, approximant = true, } phones["ɻ"] = { voiced = true, retroflex = true, approximant = true, } phones["ɭ"] = { voiced = true, retroflex = true, lateral = true, approximant = true, } phones["j"] = { voiced = true, palatal = true, approximant = true, } phones["ʎ"] = { voiced = true, palatal = true, lateral = true, approximant = true, } phones["ɰ"] = { voiced = true, velar = true, approximant = true, } phones["ʟ"] = { voiced = true, velar = true, lateral = true, approximant = true, } -- Flaps. phones["ⱱ"] = { voiced = true, labiodental = true, flap = true, } phones["ɾ"] = { voiced = true, alveolar = true, flap = true, } phones["ɺ"] = { voiced = true, alveolar = true, lateral = true, flap = true, } phones["ɽ"] = { voiced = true, retroflex = true, flap = true, } phones["𝼈"] = { voiced = true, retroflex = true, lateral = true, flap = true, } -- Trills. phones["ʙ"] = { voiced = true, bilabial = true, trill = true, } phones["r"] = { voiced = true, alveolar = true, trill = true, } phones["𝼀"] = { voiceless = true, velopharyngeal = true, trill = true, } phones["ʀ"] = { voiced = true, uvular = true, trill = true, } phones["ᴙ"] = { voiced = true, pharyngeal = true, trill = true, } -- Clicks. phones["ʘ"] = { bilabial = true, click = true, } phones["ǀ"] = { dental = true, click = true, } phones["ǃ"] = { alveolar = true, click = true, } phones["𝼊"] = { retroflex = true, click = true, } phones["ǂ"] = { palatal = true, click = true, } phones["ʞ"] = { velar = true, click = true, } phones["ǁ"] = { lateral = true, click = true, } -- Implosives. phones["ɓ"] = { voiced = true, bilabial = true, implosive = true, } phones["ɗ"] = { voiced = true, alveolar = true, implosive = true, } phones["ᶑ"] = { voiced = true, retroflex = true, implosive = true, } phones["ʄ"] = { voiced = true, palatal = true, implosive = true, } phones["ɠ"] = { voiced = true, velar = true, implosive = true, } phones["ʛ"] = { voiced = true, uvular = true, implosive = true, } -- Percussives. phones["ʬ"] = { bilabial = true, percussive = true, } phones["ʭ"] = { bidental = true, percussive = true, } phones["¡"] = { sublaminal = true, ["lower-alveolar"] = true, percussive = true, } ]=] local u = require("Module:string/char") data[1] = { -- PULMONIC CONSONANTS -- nasal ["m"] = { title = "bilabial nasal", link = "w:Bilabial nasal", }, ["ɱ"] = { title = "labiodental nasal", link = "w:Labiodental nasal", }, ["n"] = { title = "alveolar nasal", link = "w:Alveolar nasal", }, ["ɳ"] = { title = "retroflex nasal", link = "w:Retroflex nasal", }, ["ɲ"] = { title = "palatal nasal", link = "w:Palatal nasal", }, ["ŋ"] = { title = "velar nasal", link = "w:Velar nasal", }, ["ɴ"] = { title = "uvular nasal", link = "w:Uvular nasal", }, -- plosive ["p"] = { title = "voiceless bilabial plosive", link = "w:Voiceless bilabial stop", }, ["b"] = { title = "voiced bilabial plosive", link = "w:Voiced bilabial stop", }, ["t"] = { title = "voiceless alveolar plosive", link = "w:Voiceless alveolar stop", }, ["d"] = { title = "voiced alveolar plosive", link = "w:Voiced alveolar stop", }, ["ʈ"] = { title = "voiceless retroflex plosive", link = "w:Voiceless retroflex stop", }, ["ɖ"] = { title = "voiced retroflex plosive", link = "w:Voiced retroflex stop", }, ["c"] = { title = "voiceless palatal plosive", link = "w:Voiceless palatal stop", }, ["ɟ"] = { title = "voiced palatal plosive", link = "w:Voiced palatal stop", }, ["k"] = { title = "voiceless velar plosive", link = "w:Voiceless velar stop", }, ["ɡ"] = { title = "voiced velar plosive", link = "w:Voiced velar stop", }, ["q"] = { title = "voiceless uvular plosive", link = "w:Voiceless uvular stop", }, ["ɢ"] = { title = "voiced uvular plosive", link = "w:Voiced uvular stop", }, ["ʡ"] = { title = "epiglottal plosive", link = "w:Epiglottal stop", }, ["ʔ"] = { title = "glottal stop", link = "w:Glottal stop", }, -- fricative ["ɸ"] = { title = "voiceless bilabial fricative", link = "w:Voiceless bilabial fricative", }, ["β"] = { title = "voiced bilabial fricative", link = "w:Voiced bilabial fricative", }, ["f"] = { title = "voiceless labiodental fricative", link = "w:Voiceless labiodental fricative", }, ["v"] = { title = "voiced labiodental fricative", link = "w:Voiced labiodental fricative", }, ["θ"] = { title = "voiceless dental fricative", link = "w:Voiceless dental fricative", }, ["ð"] = { title = "voiced dental fricative", link = "w:Voiced dental fricative", }, ["s"] = { title = "voiceless alveolar fricative", link = "w:Voiceless alveolar fricative", }, ["z"] = { title = "voiced alveolar fricative", link = "w:Voiced alveolar fricative", }, ["ʃ"] = { title = "voiceless postalveolar fricative", link = "w:Voiceless palato-alveolar sibilant", }, ["ʒ"] = { title = "voiced postalveolar fricative", link = "w:Voiced palato-alveolar sibilant", }, ["ʂ"] = { title = "voiceless retroflex fricative", link = "w:Voiceless retroflex sibilant", }, ["ʐ"] = { title = "voiced retroflex fricative", link = "w:Voiced retroflex sibilant", }, ["ɕ"] = { title = "voiceless alveolo-palatal fricative", link = "w:Voiceless alveolo-palatal sibilant", }, ["ʑ"] = { title = "voiced alveolo-palatal fricative", link = "w:Voiced alveolo-palatal sibilant", }, ["ç"] = { title = "voiceless palatal fricative", link = "w:Voiceless palatal fricative", }, ["ʝ"] = { title = "voiced palatal fricative", link = "w:Voiced palatal fricative", }, ["x"] = { title = "voiceless velar fricative", link = "w:Voiceless velar fricative", }, ["ɣ"] = { title = "voiced velar fricative", link = "w:Voiced velar fricative", }, ["χ"] = { title = "voiceless uvular fricative", link = "w:Voiceless uvular fricative", }, ["ʁ"] = { title = "voiced uvular fricative", link = "w:Voiced uvular fricative", }, ["ħ"] = { title = "voiceless pharyngeal fricative", link = "w:Voiceless pharyngeal fricative", }, ["ʕ"] = { title = "voiced pharyngeal fricative", link = "w:Voiced pharyngeal fricative", }, ["ʜ"] = { title = "voiceless epiglottal fricative", link = "w:Voiceless epiglottal fricative", }, ["ʢ"] = { title = "voiced epiglottal fricative", link = "w:Voiced epiglottal fricative", }, ["h"] = { title = "voiceless glottal fricative", link = "w:Voiceless glottal fricative", }, ["ɦ"] = { title = "voiced glottal fricative", link = "w:Voiced glottal fricative", }, -- approximant ["ʋ"] = { title = "labiodental approximant", link = "w:Labiodental approximant", }, ["ɹ"] = { title = "alveolar approximant", link = "w:Alveolar approximant", }, ["ɻ"] = { title = "retroflex approximant", link = "w:Retroflex approximant", }, ["j"] = { title = "palatal approximant", link = "w:Palatal approximant", }, ["ɰ"] = { title = "velar approximant", link = "w:Velar approximant", }, -- tap, flap ["ⱱ"] = { title = "labiodental tap", link = "w:Labiodental flap", }, ["ɾ"] = { title = "alveolar flap", link = "w:Alveolar flap", }, ["ɽ"] = { title = "retroflex flap", link = "w:Retroflex flap", }, -- trill ["ʙ"] = { title = "bilabial trill", link = "w:Bilabial trill", }, ["r"] = { title = "alveolar trill", link = "w:Alveolar trill", }, ["ʀ"] = { title = "uvular trill", link = "w:Uvular trill", }, ["ᴙ"] = { title = "epiglottal trill", link = "w:Epiglottal trill", }, -- lateral fricative ["ɬ"] = { title = "voiceless alveolar lateral fricative", link = "w:Voiceless alveolar lateral fricative", }, ["ɮ"] = { title = "voiced alveolar lateral fricative", link = "w:Voiced alveolar lateral fricative", }, -- no precomposed Unicode character --TOMOVE --["ɬ̢"] = {title = "voiceless retroflex lateral fricative", link = "w:voiceless retroflex lateral fricative"}, -- no precomposed Unicode character --TOMOVE:3 --["ʎ̝̊"] = {title = "voiceless palatal lateral fricative", link = "w:voiceless palatal lateral fricative"}, -- no precomposed Unicode character --TOMOVE:3 --["ʟ̝̊"] = {title = "voiceless velar lateral fricative", link = "w:voiceless velar lateral fricative"}, -- no precomposed Unicode character --TOMOVE --["ʟ̝"] = {title = "voiced velar lateral fricative", link = "w:voiced velar lateral fricative"}, -- lateral approximant ["l"] = { title = "alveolar lateral approximant", link = "w:Alveolar lateral approximant", }, ["ɭ"] = { title = "retroflex lateral approximant", link = "w:Retroflex lateral approximant", }, ["ʎ"] = { title = "palatal lateral approximant", link = "w:Palatal lateral approximant", }, ["ʟ"] = { title = "velar lateral approximant", link = "w:Velar lateral approximant", }, -- lateral flap ["ɺ"] = { title = "alveolar lateral flap", link = "w:Alveolar lateral flap", }, --["ɭ̆"] = {title = "retroflex lateral flap", link = "w:retroflex lateral flap"}, -- no precomposed Unicode character --TOMOVE --["ɺ˞"] = {title = "retroflex lateral flap", link = "w:retroflex lateral flap"}, -- no precomposed Unicode character --TOMOVE -- NON-PULMONIC CONSONANTS -- clicks ["ʘ"] = { title = "bilabial click", link = "w:Bilabial clicks", }, ["ǀ"] = { title = "dental click", link = "w:Dental clicks", }, ["ǃ"] = { title = "postalveolar click", link = "w:Alveolar clicks", }, ["𝼊"] = { title = "subapical retroflex", link = "w:Retroflex clicks", }, -- NOT IN X-SAMPA ["ǂ"] = { title = "palatal click", link = "w:Palatal clicks", }, ["ǁ"] = { title = "alveolar lateral click", link = "w:Lateral clicks", }, -- implosives ["ɓ"] = { title = "voiced bilabial implosive", link = "w:Voiced bilabial implosive", }, ["ɗ"] = { title = "voiced alveolar implosive", link = "w:Voiced alveolar implosive", }, -- NOT IN X-SAMPA ["ᶑ"] = { title = "retroflex implosive", link = "w:Voiced retroflex implosive", }, ["ʄ"] = { title = "voiced palatal implosive", link = "w:Voiced palatal implosive", }, ["ɠ"] = { title = "voiced velar implosive", link = "w:Voiced velar implosive", }, ["ʛ"] = { title = "voiced uvular implosive", link = "w:Voiced uvular implosive", }, -- ejectives ["ʼ"] = { title = "ejective", link = "w:Ejective consonant", }, -- CO-ARTICULATED CONSONANTS ["ʍ"] = { title = "voiceless labial-velar fricative", link = "w:Voiceless labio-velar approximant", }, ["w"] = { title = "labial-velar approximant", link = "w:Labio-velar approximant", }, ["ɥ"] = { title = "labial-palatal approximant", link = "w:Labialized palatal approximant", }, ["ɧ"] = { title = "voiceless palatal-velar fricative", link = "w:Sj-sound", }, -- should be handled in [[Module:IPA]] and not through this table -- BRACKETS --[[ -- ["//"] = { title = "morphophonemic", link = "w:morphophonemic", }, ["/"] = { title = "phonemic", link = "w:phonemic", }, ["["] = { title = "phonetic", link = "w:phonetic", }, ["["] = { title = "phonetic", link = "w:phonetic", }, ["〈"] = { title = "orthographic", link = "w:orthographic", }, ["〉"] = { title = "orthographic", link = "w:orthographic", }, ["⟨"] = { title = "orthographic", link = "w:orthographic", }, ["⟩"] = { title = "orthographic", link = "w:orthographic", }, ]] -- VOWELS -- close ["i"] = { title = "close front unrounded vowel", link = "w:Close front unrounded vowel", }, ["y"] = { title = "close front rounded vowel", link = "w:Close front rounded vowel", }, ["ɨ"] = { title = "close central unrounded vowel", link = "w:Close central unrounded vowel", }, ["ʉ"] = { title = "close central rounded vowel", link = "w:Close central rounded vowel", }, ["ɯ"] = { title = "close back unrounded vowel", link = "w:Close back unrounded vowel", }, ["u"] = { title = "close back rounded vowel", link = "w:Close back rounded vowel", }, -- near close ["ɪ"] = { title = "near-close near-front unrounded vowel", link = "w:Near-close near-front unrounded vowel", }, ["ʏ"] = { title = "near-close near-front rounded vowel", link = "w:Near-close near-front rounded vowel", }, ["ᵻ"] = { title = "near-close central unrounded vowel", link = "w:Near-close central unrounded vowel", }, -- (alternative) --TOMOVE --[[ ["ɪ̈"] = { title = "near-close central unrounded vowel", link = "w:near-close central unrounded vowel", }, ]] ["ᵿ"] = { title = "near-close central rounded vowel", link = "w:Near-close central rounded vowel", }, --[[ (alternative) TOMOVE ["ʊ̈"] = { title = "near-close central rounded vowel", link = "w:near-close central rounded vowel", }, ]] ["ʊ"] = { title = "near-close near-back rounded vowel", link = "w:Near-close near-back rounded vowel", }, --close mid ["e"] = { title = "close-mid front unrounded vowel", link = "w:Close-mid front unrounded vowel", }, ["ø"] = { title = "close-mid front rounded vowel", link = "w:Close-mid front rounded vowel", }, ["ɘ"] = { title = "close-mid central unrounded vowel", link = "w:Close-mid central unrounded vowel", }, ["ɵ"] = { title = "close-mid central rounded vowel", link = "w:Close-mid central rounded vowel", }, ["ɤ"] = { title = "close-mid back unrounded vowel", link = "w:Close-mid back unrounded vowel", }, ["o"] = { title = "close-mid back rounded vowel", link = "w:Close-mid back rounded vowel", }, -- mid ["ə"] = { title = "schwa", link = "w:Schwa", }, ["ɚ"] = { title = "schwa+r", link = "w:R-colored vowel", }, -- open mid ["ɛ"] = { title = "open-mid front unrounded vowel", link = "w:Open-mid front unrounded vowel", }, ["œ"] = { title = "open-mid front rounded vowel", link = "w:Open-mid front rounded vowel", }, ["ɜ"] = { title = "open-mid central unrounded vowel", link = "w:Open-mid central unrounded vowel", }, ["ɝ"] = { title = "open-mid central unrounded vowel+r", link = "w:R-colored vowel", }, ["ɞ"] = { title = "open-mid central rounded vowel", link = "w:Open-mid central rounded vowel", }, ["ʌ"] = { title = "open-mid back unrounded vowel", link = "w:Open-mid back unrounded vowel", }, ["ɔ"] = { title = "open-mid back rounded vowel", link = "w:Open-mid back rounded vowel", }, -- near open ["æ"] = { title = "near-open front unrounded vowel", link = "w:Near-open front unrounded vowel", }, ["ɐ"] = { title = "near-open central vowel", link = "w:Near-open central vowel", }, -- open ["a"] = { title = "open front unrounded vowel", link = "w:Open front unrounded vowel", }, ["ɶ"] = { title = "open front rounded vowel", link = "w:Open front rounded vowel", }, ["ɑ"] = { title = "open back unrounded vowel", link = "w:Open back unrounded vowel", }, ["ɒ"] = { title = "open back rounded vowel", link = "w:Open back rounded vowel", }, -- SUPRASEGMENTALS ["ˈ"] = {title = "primary stress", link = "w:Stress (linguistics)", XSAMPA = "\""}, --[[ ["???"] = { title = "extra stress: no Unicode char; double primary stress instead", link = "w:extra stress: no Unicode char; double primary stress instead", XSAMPA = "" }, --TOMOVE:3 ]] ["ˌ"] = { title = "secondary stress", link = "w:Secondary stress", }, ["ː"] = { title = "long", link = "w:Length (phonetics)", }, ["ˑ"] = { title = "half long", link = "w:Length (phonetics)", }, ["̆"] = { title = "extra-short", link = "w:Length (phonetics)", }, --[[ ["%."] = { title = "syllable break", link = "w:syllable break", }, ]] --TOMOVE ["‿"] = { title = "linking mark (absence of a break)", link = "w:Tie (typography)#International_Phonetic_Alphabet", }, [" "] = { title = "separator", link = "w:separator", }, -- TONE -- level tones ["˥"] = { title = "top", link = "w:Tone letter", }, ["˦"] = { title = "high", link = "w:Tone letter", }, ["˧"] = { title = "mid", link = "w:Tone letter", }, ["˨"] = { title = "low", link = "w:Tone letter", }, ["˩"] = { title = "bottom", link = "w:Tone letter", }, ["̋"] = { title = "extra high tone", link = "w:Tone letter", }, ["́"] = { title = "high tone", link = "w:Tone letter", }, ["̄"] = { title = "mid tone", link = "w:Tone letter", }, ["̀"] = { title = "low tone", link = "w:Tone letter", }, ["̏"] = { title = "extra low tone", link = "w:Tone letter", }, -- tone terracing ["ꜛ"] = { title = "upstep", link = "w:Upstep", }, ["ꜜ"] = { title = "downstep", link = "w:Downstep", }, -- contour tones ["̌"] = { title = "rising tone", link = "w:Tone (linguistics)", }, ["̂"] = { title = "falling tone", link = "w:Tone (linguistics)", }, ["᷄"] = { title = "high rising tone", link = "w:Tone (linguistics)", }, ["᷅"] = { title = "low rising tone", link = "w:Tone (linguistics)", }, ["᷇"] = { title = "high falling tone", link = "w:Tone (linguistics)", }, ["᷆"] = { title = "low falling tone", link = "w:Tone (linguistics)", }, ["᷈"] = { title = "rising falling tone (peaking)", link = "w:Tone (linguistics)", }, ["᷉"] = { title = "dipping", link = "w:Tone (linguistics)", }, -- [extrapolated from the chart -- please confirm] -- intonation ["|"] = { title = "minor (foot) group", link = "w:Prosodic unit", }, ["‖"] = { title = "major (intonation) group", link = "w:Prosodic unit", }, ["↗"] = { title = "global rise", link = "w:Intonation (linguistics)", }, ["↘"] = { title = "global fall", link = "w:Intonation (linguistics)", }, -- DIACRITICS -- syllabicity & releases ["̩"] = { title = "syllabi ", link = "w:Syllabic consonant", withdescender = "̍" }, -- (or "_=" ["̯"] = { title = "non-syllabic", link = "w:Semivowel", withdescender = "̑" }, ["ʰ"] = { title = "aspirated", link = "w:Aspirated consonant", }, ["ⁿ"] = { title = "nasal release", link = "w:Nasal release", }, ["ˡ"] = { title = "lateral release", link = "w:Lateral release (phonetics)", }, ["̚"] = { title = "no audible release", link = "w:No audible release", }, -- phonation ["̥"] = { title = "voiceless", link = "w:Voicelessness", withdescender = "̊" }, ["̬"] = { title = "voiced", link = "w:Voice (phonetics)", }, ["̤"] = { title = "breathy voice", link = "w:Breathy voice", }, ["̰"] = { title = "creaky voice", link = "w:Creaky voice", }, ["᷽"] = { title = "strident", link = "w:Strident vowel", }, -- primary articulation ["̪"] = { title = "dental", link = "w:Dental consonant", }, ["̺"] = { title = "apical", link = "w:Apical consonant", }, ["̻"] = { title = "laminal", link = "w:Laminal consonant", }, ["̟"] = { title = "advanced", link = "w:Relative articulation#Advanced_and_retracted", withdescender = "˖" }, ["̠"] = { title = "retracted", link = "w:Relative articulation#Retracted", withdescender = "˗" }, ["̼"] = { title = "linguolabial", link = "w:Linguolabial consonant", }, ["̈"] = { title = "centralized", link = "w:Relative articulation#Centralized_vowels", XSAMPA = "_\"" }, ["̽"] = { title = "mid-centralized", link = "Relative articulation#Mid-centralized_vowel", }, ["̞"] = { title = "lowered", link = "w:Relative articulation#Raised_and_lowered", withdescender = "˕" }, ["̝"] = { title = "raised", link = "w:Relative articulation#Raised_and_lowered", withdescender = "˔" }, ["͡"] = { title = "coarticulated", link = "w:Co-articulated consonant", }, ["͈"] = { title = "strong articulation", link = "w:Fortis and lenis", }, -- secondary articulation ["ʷ"] = { title = "labialized", link = "w:Labialization", }, ["ʲ"] = { title = "palatalized", link = "w:Palatalization (phonetics)", }, ["ˠ"] = { title = "velarized", link = "w:Velarization", }, ["ˤ"] = { title = "pharyngealized", link = "w:Pharyngealization", }, -- also see _e ["ɫ"] = { title = "velarized alveolar lateral approximant", link = "w:Alveolar lateral approximant", }, ["̴"] = { title = "velarized or pharyngealized; also see 5", link = "w:Velarization", }, ["̹"] = { title = "more rounded", link = "w:Roundedness", }, ["̜"] = { title = "less rounded", link = "w:Roundedness", }, ["̃"] = { title = "nasalization", link = "w:Nasalization", }, ["˞"] = { title = "rhotacization in vowels, retroflexion in consonants", link = "w:R-colored vowel", }, ["̘"] = { title = "advanced tongue root", link = "w:Advanced and retracted tongue root", }, ["̙"] = { title = "retracted tongue root", link = "w:Advanced and retracted tongue root", }, } data[2] = { -- TODO --["%("] = {}, --["%)"] = {}, ["ːː"] = { title = "extra long", link = "w:Length (phonetics)", }, ["r̥"] = {title = "voiceless alveolar trill", link = "w:Voiceless alveolar trill"}, ["ɬ’"] = {title = "alveolar lateral ejective fricative", link = "w:Alveolar lateral ejective fricative"}, } data[3] = { ["t͡s"] = {title = "voiceless alveolar sibilant affricate", link = "w:Voiceless alveolar affricate"}, ["d͡z"] = {title = "voiced alveolar sibilant affricate", link = "w:Voiced alveolar affricate"}, ["t͡ʃ"] = {title = "voiceless palato-alveolar affricate", link = "w:Voiceless palato-alveolar affricate", descender = true}, ["d͡ʒ"] = {title = "voiced palato-alveolar affricate", link = "w:Voiced palato-alveolar affricate"}, ["ʈ͡ʂ"] = {title = "voiceless retroflex affricate", link = "w:Voiceless retroflex affricate", descender = true}, ["ɖ͡ʐ"] = {title = "voiced retroflex affricate", link = "w:Voiced retroflex affricate, descender = true"}, ["t͡ɕ"] = {title = "voiceless alveolo-palatal affricate", link = "w:Voiceless alveolo-palatal affricate"}, ["d͡ʑ"] = {title = "voiced alveolo-palatal affricate", link = "w:Voiced alveolo-palatal affricate"}, ["c͡ç"] = {title = "voiceless palatal affricate", link = "w:Voiceless palatal affricate, descender = true"}, ["ɟ͡ʝ"] = {title = "voiced palatal affricate", link = "w:Voiced palatal affricate, descender = true"}, ["k͡x"] = {title = "voiceless velar affricate", link = "w:Voiceless velar affricate"}, ["ɡ͡ɣ"] = {title = "voiced velar affricate", link = "w:Voiced velar affricate, descender = true"}, } data[4] = { ["ǃ͡qʼ"] = {title = "alveolar linguo-glottalic stop", link = "w:Ejective-contour clicks, descender = true"}, ["ǁ͡χʼ"] = {title = "lateral linguo-glottalic affricate (homorganic)", link = "w:Ejective-contour clicks", descender = true}, } data[5] = { ["k͡ʟ̝̊"] = {title = "voiceless velar lateral affricate", link = "w:Voiceless velar lateral affricate"}, ["ᶢǀ͡qʼ"] = {title = "voiced dental linguo-glottalic stop", link = "w:Ejective-contour clicks"}, ["ǂ͡kxʼ"] = {title = "palatal linguo-glottalic affricate (heterorganic)", link = "w:Ejective-contour clicks"}, } data[6] = { ["k͡ʟ̝̊ʼ"] = {title = "velar lateral ejective affricate", link = "w:Velar lateral ejective affricate"}, ["ᶢʘ͡kxʼ"] = {title = "voiced labial linguo-glottalic affricate", link = "w:Ejective-contour clicks"}, } data.separator_escapes = { ["⁽"] = "(", ["⁾"] = ")", ["₍"] = "(", ["₎"] = ")", ["ˈ"] = "\1", ["ˌ"] = "\2", ["ː"] = ":", ["ˑ"] = ";", } -- acute and grave tone marks local diacritics = u( -- grave, acute, circumflex, tilde, macron, breve 0x300, 0x301, 0x302, 0x303, 0x304, 0x306, -- diaeresis, ring above, double acute, caron, vertical line above, double grave, left tack 0x308, 0x30A, 0x30B, 0x30C, 0x30D, 0x30F, 0x318, -- right tack, left angle, left half ring below, up tack below, down tack below, plus sign below 0x319, 0x31A, 0x31C, 0x31D, 0x31E, 0x31F, -- minus sign below, rhotic hook below, dot below, diaeresis below, ring below, vertical line below, bridge below 0x320, 0x322, 0x323, 0x324, 0x325, 0x329, 0x32A, -- caron below, inverted breve below 0x32C, 0x32F, -- tilde below, combining tilde overlay, right half ring below, inverted bridge below, square below, seagull below, x above 0x330, 0x334, 0x339, 0x33A, 0x33B, 0x33C, 0x33D, -- grave tone mark, acute tone mark, bridge above, equals sign below, double vertical line below 0x340, 0x341, 0x346, 0x347, 0x348, -- left angle below, not tilde above, homothetic above, almost equal above, left right arrow below 0x349, 0x34A, 0x34B, 0x34C, 0x34D, -- upwards arrow below, left arrowhead below, right arrowhead below 0x34E, 0x354, 0x355, -- double rightwards arrow below, combining Latin small letter a 0x362, 0x361, -- macron–acute, grave–macron, macron–grave, acute–macron, grave–acute–grave, acute–grave–acute 0x1DC4, 0x1DC5, 0x1DC6, 0x1DC7, 0x1DC8, 0x1DC9) data.diacritics = diacritics data.vowels = "iyɨʉɯuɪᵻʏʊᵿeøɘɵɤoəɚɛœɜɝɞʌɔæɐaɶɑɒäëïöüÿ" local tones = "˥˦˧˨˩꜒꜓꜔꜕꜖꜈꜉꜊꜋꜌꜍꜎꜏꜐꜑¹²³⁴⁵⁶⁷⁸⁹⁰" data.tones = tones local superscripts = u(0xA0) .. " ⁰¹²³⁴⁵⁶⁷⁸⁹ᵃ𐞃ᵄᵅᶛᵇ𐞄𐞅ᶜᶝᵈᶞ𐞋𐞌𐞍ᵉᵊᵋ𐞎ᶟᵌ𐞏𐞑ᶠᶢ𐞒𐞓𐞔ˠʰ𐞕𐞖ʱ𐞗ⁱᶦᶤʲᶨᶡ𐞘ᵏˡᶫꭞ𐞛ᶩ𐞞𐞠𐞡ᵐᶬⁿᶰᶮᶯᵑᵒ𐞢ꟹ𐞣ᵓᶱᵖᶲ𐞥ʳ𐞪ʴ𐞦𐞧ʵ𐞨𐞩ʶˢᶳᶴᵗ𐞯ᵘᶶᶣᵚᶭᶷᵛᶹ𐞰ᶺʷꭩˣʸ𐞲ᶻᶼᶽᶾˀˤ𐞳𐞴𐞶𐞷𐞸𐞹𐞵ᵝᶿᵡ˞⁻𐞁𐞂" data.superscripts = superscripts -- An array of patterns of valid character sequences. data.valid = { "⁽[" .. superscripts .. "]+⁾", "[ %(%)%%<>{|}%-→~⁓%.◌abcdefhijklmnopqrstuvwxyz¡àáâãāăēäæçèéêëĕěħìíîïĩīĭĺḿǹńňðòóôõöōŏőœøŕùúûüũūŭűýÿŷŋ" .. "ǀǁǂǃǎǐǒǔřǖǘǚǜǟǣǽǿȁȅȉȍȕȫȭȳɐɑɒɓɔɕɖɗɘəɚɛɜɝɞɟɠɡɢɣɤɥɦɧɨɪᵻɫɬɭɮɯɰɱɲɳɴɵɶɸɹɺ𝼈ɻɽɾʀʁʂʃʄʈʉʊᵿʋṽʌʍʎ𝼆ʏʐʑʒʔʕʘʞʙʛʜʝʟʡʢ𝼊ʬʭ" .. "ʼˈˌːˑˣ˔˕ˬ͗˭ˇ˖β͜θχᴙᶑ᷽ḁḛḭḯṍṏṳṵṹṻạẹẽịọụỳỵỹ‖․‥…‿↑↓↗↘ⱱꜛꜜꟸ𝆏𝆑˗ˋˊ–⸨⸩⁽⁾" .. diacritics .. tones .. superscripts .. "]+" } -- Character sequences which are valid only in a particular language. -- These can be either a single pattern (as a string), or an array of patterns (as a table). data.per_lang_valid = { ["egy"] = "V+", -- V for uncertain vowel ["okm"] = "[LHR!WT]+", -- irregular verb morphophonemes } -- Characters to add VARIATION SELECTOR-15 (U+FE0E) after. -- These are characters with emoji variants that are used by default by some clients. -- Adding VS15 after them instructs them to draw the characters as text instead. data.add_vs15 = "↗↘" data.invalid = { ["!"] = "ǃ", ["ꜝ"] = "ꜜ", ["ꜞ"] = "ꜛ", ["ꜟ"] = "ꜛ", ["'"] = "ˈ", ["’"] = "ʼ", [":"] = "ː", -- Confusable Latin letters ["B"] = "ʙ", ["g"] = "ɡ", ["G"] = "ɢ", ["Ɠ"] = "ʛ", ["H"] = "ʜ", ["ı"] = "ɪ", ["I"] = "ɪ", ["L"] = "ʟ", ["N"] = "ɴ", ["Œ"] = "ɶ", ["Q"] = "ꞯ", ["R"] = "ʀ", ["∫"] = "ʃ", ["⨎"] = "ǂ", -- due to confusion with obsolete 𝼋 below ["ß"] = "β", ["ẞ"] = "β", ["Y"] = "ʏ", ["Ə"] = "ə", ["ǝ"] = "ə", ["Ɂ"] = "ʔ", ["ɂ"] = "ʔ", ["ˁ"] = "ˤ", -- Confusable Greek letters ["α"] = "ɑ", ["γ"] = "ɣ", ["δ"] = "ð", ["ε"] = "ɛ", ["Η"] = "ʜ", ["η"] = "ŋ", ["ι"] = "ɪ", ["λ"] = "ʎ", ["υ"] = "ʋ", ["Ψ"] = "𝼊", ["ψ"] = "𝼊", ["Φ"] = "ɸ", ["ϕ"] = "ɸ", ["ꭓ"] = "χ", -- Actually Latin, since IPA uses the Greek letter(!) -- Confusable Cyrillic letters ["ӕ"] = "æ", ["Ә"] = "ə", ["ә"] = "ə", ["В"] = "ʙ", ["в"] = "ʙ", ["е"] = "e", ["З"] = "ɜ", ["з"] = "ɜ", ["Ѕ"] = "s", ["ѕ"] = "s", ["і"] = "i", ["ј"] = "j", ["Н"] = "ʜ", ["н"] = "ʜ", ["О"] = "o", ["о"] = "o", ["р"] = "p", ["с"] = "c", ["у"] = "y", ["Ү"] = "ʏ", ["ү"] = "ʏ", ["Ф"] = "ɸ", ["ф"] = "ɸ", ["х"] = "x", ["Һ"] = "h", ["һ"] = "h", ["Я"] = "ᴙ", ["я"] = "ᴙ", ["Ѱ"] = "𝼊", ["ѱ"] = "𝼊", ["Ѵ"] = "ⱱ", ["ѵ"] = "ⱱ", ["Ҁ"] = "ʕ", ["ҁ"] = "ʕ", -- Palatalization ["ᶀ"] = "bʲ", ["ꞔ"] = "cʲ", ["ᶁ"] = "dʲ", ["ȡ"] = "d̠ʲ", ["d̂"] = "d̠ʲ", ["ᶂ"] = "fʲ", ["ᶃ"] = "ɡʲ", ["ꞕ"] = "hʲ", ["ᶄ"] = "kʲ", ["ᶅ"] = "lʲ", ["ȴ"] = "l̠ʲ", ["l̂"] = "l̠ʲ", ["𝼓"] = "ɬʲ", ["ᶆ"] = "mʲ", ["ᶇ"] = "nʲ", ["ȵ"] = "n̠ʲ", ["n̂"] = "n̠ʲ", ["𝼔"] = "ŋʲ", ["ᶈ"] = "pʲ", ["ᶉ"] = "rʲ", ["𝼕"] = "ɹʲ", ["𝼖"] = "ɾʲ", ["ᶊ"] = "sʲ", ["𝼞"] = "ɕ", ["𐞺"] = "ᶝ", ["ᶋ"] = "ʃʲ", ["ʆ"] = "ʃʲ", ["ƫ"] = "tʲ", ["ȶ"] = "t̠ʲ", ["t̂"] = "t̠ʲ", ["ᶌ"] = "vʲ", ["ᶍ"] = "xʲ", ["ᶎ"] = "zʲ", ["𝼘"] = "ʒʲ", ["ʓ"] = "ʒʲ", -- Retroflex ["𝼝"] = "ʈ͡ʂ", ["𝼥"] = "ɖ", ["𝼦"] = "ɭ", ["𝼧"] = "ɳ", ["𝼨"] = "ɽ", ["𝼩"] = "ʂ", ["𝼪"] = "ʈ", -- Rhotic vowels ["ᶏ"] = "a˞", ["ᶐ"] = "ɑ˞", ["ᶒ"] = "e˞", ["ə˞"] = "ɚ", ["ᶕ"] = "ɚ", ["ᶓ"] = "ɛ˞", ["ɜ˞"] = "ɝ", ["ᶔ"] = "ɝ", ["ᶖ"] = "i˞", ["𝼚"] = "ɨ˞", ["𝼛"] = "o˞", ["ᶗ"] = "ɔ˞", ["ᶙ"] = "u˞", -- Syllabic approximants ["ɿ"] = "ɹ̩", ["ʅ"] = "ɻ̩", ["ʮ"] = "ɹ̩ʷ", ["ʯ"] = "ɻ̩ʷ", -- Clicks ["ʗ"] = "ǃ", ["𝼋"] = "ǂ", ["ʇ"] = "ǀ", ["ʖ"] = "ǁ", ["‼"] = "𝼊", -- Voiceless implosives ["ƈ"] = "ʄ̊", ["ƙ"] = "ɠ̊", ["ƥ"] = "ɓ̥", ["ʠ"] = "ʛ̥", ["ƭ"] = "ɗ̥", ["𝼉"] = "ᶑ̥", -- Monographs ["ꜰ"] = "ɸ", ["ɩ"] = "ɪ", ["ɼ"] = "r̝", ["ᴜ"] = "ʊ", ["ɷ"] = "ʊ", ["𐞤"] = "ᶷ", ["ƛ"] = "t͡ɬ", ["ƻ"] = "d͡z", ["ƾ"] = "t͡s", -- Digraphs ["ȸ"] = "b̪", ["ʣ"] = "d͡z", ["ʥ"] = "d͡ʑ", ["ꭦ"] = "ɖ͡ʐ", ["ʤ"] = "d͡ʒ", ["𝼒"] = "d͡ʒʲ", ["𝼙"] = "d͡ᶚ", ["ʪ"] = "ɬ͡s", ["ʫ"] = "ɮ͡z", ["ȹ"] = "p̪", ["ʦ"] = "t͡s", ["ʨ"] = "t͡ɕ", ["ꭧ"] = "ʈ͡ʂ", ["ʧ"] = "t͡ʃ", ["𝼗"] = "t͡ʃʲ", ["𝼜"] = "t͡ᶘ", -- Deprecated or confusable diacritics ["̫"] = "ʷ", ["͂"] = "̃", ["᫇"] = "ʷ", ["⸋"] = "̚", ["̱"] = "̠", -- COMBINING MACRON BELOW (U+0331) -> COMBINING MINUS SIGN BELOW (U+0320) -- Precomposed characters with deprecated or confusable diacritics; the left is a precomposed -- version of a lowercase letter with COMBINING MACRON BELOW and the right is the equivalent -- using COMBINING MINUS SIGN BELOW ["ḇ"] = "b̠", ["ḏ"] = "d̠", ["ẖ"] = "h̠", ["ḵ"] = "k̠", ["ḻ"] = "l̠", ["ṉ"] = "n̠", ["ṟ"] = "r̠", ["ṯ"] = "t̠", ["ẕ"] = "z̠", } return data mjopki8sklky9m3kv63kzgka1usuiki Modul:nyms 828 12091 375347 363508 2026-09-22T03:12:08Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708739|92708739]]) 375347 Scribunto text/plain local export = {} local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local labels_module = "Module:labels" local links_module = "Module:links" local parameter_utilities_module = "Module:parameter utilities" local parse_utilities_module = "Module:parse utilities" local function track(page) return require(debug_track_module)("nyms/" .. page) end local function wrap_span(text, lang, sc) return '<span class="' .. sc .. '" lang="' .. lang .. '">' .. text .. '</span>' end local function term_already_linked(term) -- optimization to avoid unnecessarily loading [[Module:parse utilities]] return term:find("[<{]") and require(parse_utilities_module).term_already_linked(term) end function export.nyms(frame) local parent_args = frame:getParent().args -- FIXME: Temporary error message and tracking. for arg, _ in pairs(parent_args) do if type(arg) == "string" and arg:find("^lb[0-9]*$") then local llarg = arg:gsub("^lb", "ll") error(("%s= is deprecated; use %s= instead, per the documentation"):format(arg, llarg)) end if arg == "q" or arg == "qq" then track(arg) end if type(arg) == "string" and arg:find("^tag[0-9]*$") then local larg = arg:gsub("^tag", "l") local llarg = arg:gsub("^tag", "ll") error(("Use %s= (on the left) or %s= (on the right) instead of %s="):format(larg, llarg, arg)) end end local params = { [1] = {required = true, type = "language", default = "und"}, [2] = {list = true, allow_holes = true, required = true}, } local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l"}}, -- For compatibility, we don't distinguish q= from q1= and qq= from q1=. FIXME: Maybe we should change this. {group = "q", separate_no_index = false}, {param = "lb", deprecated = true}, } local special_separators = mw.clone(m_param_utils.default_special_separators) special_separators["<"] = " < " local items, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, parse_lang_prefix = true, track_module = "nyms", lang = 1, special_separators = special_separators, sc = "sc.default", -- because we show them ourselves below so e.g. we can handle already-linked terms no_show_decorations = true, } local nym_type = frame.args[1] local nym_type_class = string.gsub(nym_type, "%s", "-") local lang = args[1] local langcode = lang:getCode() local data = { lang = lang, items = items, sc = args.sc.default, l = args.l.default, ll = args.ll.default, q = args.q.default, qq = args.qq.default, } local parts = {} local thesaurus_parts = {} for i, item in ipairs(data.items) do if item.lb then error("Inline modifier <lb:...> is deprecated; use <ll:...> per the documentation") end local explicit_item_lang = item.lang item.lang = item.lang or data.lang item.sc = item.sc or data.sc local text local is_thesaurus if item.term and item.term:find("^Tesaurus:") then is_thesaurus = true for k, _ in pairs(item) do if m_param_utils.item_key_is_property(k) and k ~= "lang" and k ~= "sc" and k ~= "q" and k ~= "qq" and k ~= "l" and k ~= "ll" and k ~= "refs" then error(("You cannot use most named parameters and inline modifiers with Thesaurus links, but saw %s%s= or its equivalent inline modifier <%s:...>"):format( k, item.itemno, k)) end end local term = item.term:match("^Tesaurus:(.*)$") -- Chop off fragment term = term:match("^(.-)#.*$") or term local lang = item.lang local sccode = (item.sc or lang:findBestScript(term)):getCode() -- FIXME: I assume it's better to include full-language codes in the CSS rather than etym-language codes, -- which are generally specific to Wiktionary. However, we should probably instead be using the functions -- from [[Module:script utilities]] in preference to rolling our own. text = "[[" .. item.term .. "#Bahasa " .. lang:getFullName() .. "|Tesaurus:" .. wrap_span(term, lang:getFullCode(), sccode) .. "]]" else if thesaurus_parts[1] then error("Links to the Thesaurus must follow all non-Thesaurus links") end local raw_term = item.alt or item.term if raw_term and term_already_linked(raw_term) then text = raw_term else text = require(links_module).full_link(item) end end local qq = item.qq -- If a separate language code was given for the term, display the language name as a right qualifier. -- Otherwise it may not be obvious that the term is in a separate language (e.g. if the main language is 'zh' -- and the term language is a Chinese lect such as Min Nan). But don't do this for Translingual terms, which -- are often added to the list of English and other-language terms. if explicit_item_lang then local explicit_code = explicit_item_lang:getCode() if explicit_code ~= langcode and explicit_code ~= "mul" then qq = mw.clone(qq) or {} table.insert(qq, 1, explicit_item_lang:getCanonicalName()) end end if item.q and item.q[1] or qq and qq[1] or item.l and item.l[1] or item.ll and item.ll[1] or item.refs and item.refs[1] then text = require(decorations_module).format_decorations { lang = item.lang, text = text, q = item.q, qq = qq, l = item.l, ll = item.ll, refs = item.refs, } end local insert_place = is_thesaurus and thesaurus_parts or parts -- Don't include the separator if this is the first item of this class that we're inserting. table.insert(insert_place, insert_place[1] and item.separator or "") table.insert(insert_place, text) end local text = table.concat(parts) local thesaurus_text = table.concat(thesaurus_parts) local caption = "<span style=\"font-size: smaller\">" .. mw.getContentLanguage():ucfirst(nym_type) .. ((#items > 1 or thesaurus_text ~= "") and "s" or "") .. ":</span> " text = caption .. text local function decoration_error_if_no_terms() if not parts[1] then error("Cannot specify overall decorations if no non-Thesaurus terms given") end end if data.q and data.q[1] or data.qq and data.qq[1] or data.l and data.l[1] then decoration_error_if_no_terms() text = require(decorations_module).format_decorations { lang = data.lang, text = text, q = data.q, qq = data.qq, l = data.l, -- ll handled specially for compatibility's sake } end if data.ll and data.ll[1] then decoration_error_if_no_terms() text = text .. " &mdash; " .. require(labels_module).show_labels { lang = data.lang, labels = data.ll, nocat = true, open = false, close = false, no_track_already_seen = true, } end if thesaurus_text ~= "" then local thesaurus_intro = parts[1] and "; ''lihat juga'' " or "''lihat'' " text = text .. thesaurus_intro .. thesaurus_text end return "<span class=\"nyms " .. nym_type_class .. "\">" .. text .. "</span>" end return export izsdelhiyeoailt2eqmxaihby4w3q4o Modul:ar-pronunciation 828 12378 375390 365593 2026-09-22T07:32:09Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92717545|92717545]]) 375390 Scribunto text/plain local export = {} local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local audio_module = "Module:audio" local parse_utilities_module = "Module:parse utilities" local rfind = m_str_utils.find local rsplit = m_str_utils.split local ugsub = m_str_utils.gsub local ulen = m_str_utils.len local ulower = m_str_utils.lower local usub = m_str_utils.sub local concat = table.concat local insert = table.insert local lang = require("Module:languages").getByCode("ar") local sc = require("Module:scripts").getByCode("Arab") local correspondences = { ["ʾ"] = "ʔ", ["ṯ"] = "θ", ["j"] = "d͡ʒ", ["ḥ"] = "ħ", ["ḵ"] = "x", ["ḏ"] = "ð", ["š"] = "ʃ", ["ṣ"] = "sˤ", ["ḍ"] = "dˤ", ["ṭ"] = "tˤ", ["ẓ"] = "ðˤ", ["ž"] = "ʒ", ["ʿ"] = "ʕ", ["ḡ"] = "ɣ", ["ḷ"] = "lˤ", ["ū"] = "uː", ["ī"] = "iː", ["ā"] = "aː", ["y"] = "j", ["g"] = "ɡ", ["ē"] = "eː", ["ō"] = "oː", [""] = "", } local vowels = "aāeēiīoōuū" local vowel = "[" .. vowels .. "]" local long_vowels = "āēīōū" local long_vowel = "[" .. long_vowels .. "]" local consonant = "[^" .. vowels .. ". -]" local syllabify_pattern = "(" .. vowel .. ")(" .. consonant .. "?)(" .. consonant .. "?)(" .. vowel .. ")" local tie = "‿" local closed_syllable_shortening_pattern = "(" .. long_vowel .. ")(" .. tie .. ")" .. "(" .. consonant .. ")" local function rsub(term, foo, bar) local retval = ugsub(term, foo, bar) return retval end local function generate_obj(respelling) return { respelling = respelling } end local function split_on_comma(term) if not term then return nil end if term:find(",%s") or term:find("\\") then return require(parse_utilities_module).split_on_comma(term) else return rsplit(term, ",") end end local function parse_respellings_with_modifiers(respelling, paramname) if respelling:find("[<%[]") then local put = require(parse_utilities_module) local segments = put.parse_multi_delimiter_balanced_segment_run(respelling, { { "<", ">" }, { "[", "]" } }) local comma_separated_groups = put.split_alternating_runs_on_comma(segments) local retval = {} for _, group in ipairs(comma_separated_groups) do local j = 2 while j <= #group do if not group[j]:find("^<.*>$") then group[j - 1] = group[j - 1] .. group[j] .. group[j + 1] table.remove(group, j) table.remove(group, j) else j = j + 2 end end local param_mods = { q = { type = "qualifier" }, qq = { type = "qualifier" }, a = { type = "labels" }, aa = { type = "labels" }, ref = { item_dest = "refs", type = "references" }, } table.insert(retval, put.parse_inline_modifiers_from_segments { group = group, arg = respelling, props = { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, }, }) end return retval else local retval = {} for _, item in ipairs(split_on_comma(respelling)) do table.insert(retval, generate_obj(item)) end return retval end end local function parse_pron_modifier(arg, paramname, generate_obj, param_mods, splitchar) splitchar = splitchar or "," if arg:find("<") then param_mods.q = { type = "qualifier" } param_mods.qq = { type = "qualifier" } param_mods.a = { type = "labels" } param_mods.aa = { type = "labels" } param_mods.ref = { item_dest = "refs", type = "references" } return require(parse_utilities_module).parse_inline_modifiers(arg, { param_mods = param_mods, generate_obj = generate_obj, paramname = paramname, splitchar = splitchar, }) else local retval = {} local split_arg = splitchar == "," and split_on_comma(arg) or rsplit(arg, splitchar) for _, term in ipairs(split_arg) do table.insert(retval, generate_obj(term)) end return retval end end local function parse_audio(lang, arg, pagename, paramname) local param_mods = { IPA = { sublist = true }, text = {}, t = { item_dest = "gloss" }, gloss = {}, pos = {}, lit = {}, g = { item_dest = "genders", sublist = true }, bad = {}, cap = { item_dest = "caption" }, } local function process_special_chars(val) if not val then return val end return (val:gsub("#", pagename)) end local function generate_audio_obj(arg) return { file = process_special_chars(arg) } end local retvals = parse_pron_modifier(arg, paramname, generate_audio_obj, param_mods, "%s*;%s*") for _, retval in ipairs(retvals) do retval.lang = lang retval.text = process_special_chars(retval.text) retval.caption = process_special_chars(retval.caption) local textobj = require(audio_module).construct_audio_textobj(retval) retval.text = textobj retval.gloss = nil retval.pos = nil retval.lit = nil retval.genders = nil end return retvals end local function parse_regional_phonetics(ph_arg, pagename) if not ph_arg or ph_arg == "" then return {} end local regionals = {} for _, item in ipairs(rsplit(ph_arg, "%s*;%s*")) do local audio = nil local item_no_mod = item:gsub("<a:([^>]+)>", function(a) audio = a:gsub("#", pagename) return "" end) local region, ipa = item_no_mod:match("^([^:]+):(.+)$") if region and ipa then local regions = rsplit(region, "%s*,%s*") table.insert(regionals, { regions = regions, ipa = ipa, audio = audio }) end end return regionals end local function syllabify(text) text = ugsub(text, "%-(" .. consonant .. ")%-(" .. consonant .. ")", "%1.%2") text = ugsub(text, "%-", ".") for _ = 1, 2 do text = ugsub( text, syllabify_pattern, function(a, b, c, d) if c == "" and b ~= "" then c, b = b, "" end return a .. b .. "." .. c .. d end ) end text = ugsub(text, "(" .. vowel .. ") (" .. consonant .. ")%.?(" .. consonant .. ")", "%1" .. tie .. "%2.%3") return text end local function closed_syllable_shortening(text) local shorten = { ["ā"] = "a", ["ē"] = "e", ["ī"] = "i", ["ō"] = "o", ["ū"] = "u", } text = ugsub(text, closed_syllable_shortening_pattern, function(vowel, tie, consonant) return shorten[vowel] .. tie .. consonant end) return text end function export.link(term) return require("Module:links").full_link { term = term, lang = lang, sc = sc } end function export.toIPA(list, silent_error) local translit if list.tr then translit = list.tr elseif list.term then require("Module:script utilities").checkScript(list.term, "Arab") translit = lang:transliterate(list.term) if not translit then if silent_error then return '' else error('Module:ar-translit failed to generate a transliteration from "' .. list.term .. '".') end end else if silent_error then return '' else error('No Arabic text or transliteration was provided to the function "toIPA".') end end translit = ugsub(translit, "llāh", "ḷḷāh") translit = ugsub(translit, "([iī] ?)ḷḷ", "%1ll") translit = ugsub(translit, "%(t%)", "") translit = ugsub(translit, "(" .. vowel .. ") " .. vowel, "%1 ") translit = ugsub(translit, "%-?l%-?", "l") translit = syllabify(translit) translit = closed_syllable_shortening(translit) local output = ugsub(translit, ".", correspondences) output = ugsub(output, "%-", "") return output end function export.get_pron_info(terms, pagename, paramname) if #terms == 1 and terms[1].respelling == "-" then return { pron_list = nil } end local pron_list = {} local brackets = "/%s/" for _, term in ipairs(terms) do local respelling = term.respelling local ar_term, tr if not respelling or respelling == "" or respelling == "#" then ar_term = pagename elseif rfind(respelling, "[a-zA-Z]") then tr = respelling elseif respelling:find("[ء-ي]") then ar_term = respelling else tr = respelling end local pron = export.toIPA({ term = ar_term, tr = tr }, false) if pron and pron ~= "" then local bracketed_pron = brackets:format(pron) table.insert(pron_list, { pron = bracketed_pron, q = term.q, qq = term.qq, a = term.a, aa = term.aa, refs = term.refs, }) end end return { pron_list = pron_list } end function export.show_old(frame) local params = { [1] = { list = true, allow_holes = true }, ["tr"] = { list = true, allow_holes = true }, ["qual"] = { list = true, allow_holes = true }, ["nl"] = { type = "boolean" }, ["ann"] = {}, } local args = require("Module:parameters").process(frame:getParent().args, params) local ar_terms = args[1] local transliterations = args.tr local qualifiers = args.qual local nl = args.nl if not (ar_terms.maxindex > 0 or transliterations.maxindex > 0) then if mw.title.getCurrentTitle().nsText == "Template" then ar_terms[1] = "كَلِمَة" ar_terms.maxindex = 1 else error( 'Please provide vocalized Arabic in the first parameter of {{[[Template:ar-IPA|ar-IPA]]}}, or transliteration in the "tr" parameter.') end end local pronunciations = {} for i = 1, math.max(ar_terms.maxindex, transliterations.maxindex) do local ar_term = ar_terms[i] local tr = transliterations[i] local qual = qualifiers[i] if not (ar_term or tr) then error("There is a gap in the parameters. Provide either |" .. i .. "= or |tr" .. i .. "=.") elseif ar_term and tr then mw.logObject("Duplicate parameters |" .. i .. "= and |tr" .. i .. "= in {{ar-IPA}},") end local pron = export.toIPA { term = ar_term, tr = tr } table.insert(pronunciations, { pron = "/" .. pron .. "/", q = qual and { qual } or nil }) end local anntext = "" if args.ann then anntext = args.ann if args.ann:find("%+") then local anndefs = {} for i = 1, ar_terms.maxindex do local ar_term = ar_terms[i] if ar_term then table.insert(anndefs, "'''" .. ar_term .. "'''") end end if anndefs[1] then anndefs = table.concat(anndefs, ", ") anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs)) end end anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ":&#32;" end if nl then return anntext .. require("Module:IPA").format_IPA_multiple(lang, pronunciations) else return anntext .. require("Module:IPA").format_IPA_full { lang = lang, items = pronunciations } end end function export.show(frame) local parent_args = frame:getParent().args local process = require("Module:parameters").process local params = { [1] = {}, ["audios"] = {}, ["a"] = { alias_of = "audios" }, ["ph"] = {}, ["pagename"] = {}, ["indent"] = {}, ["ann"] = {}, } local args = process(parent_args, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local indent = args.indent or "*" local termspec = args[1] or "#" local terms = parse_respellings_with_modifiers(termspec, 1) local pronobj = export.get_pron_info(terms, pagename, 1) local regional_phonetics = parse_regional_phonetics(args.ph, pagename) local parts = {} local function ins(text) table.insert(parts, text) end local anntext = "" if args.ann then anntext = args.ann if args.ann:find("%+") then local anndefs = {} for _, term in ipairs(terms) do local respelling = term.respelling if respelling and respelling:find("[ء-ي]") then table.insert(anndefs, "'''" .. respelling .. "'''") end end if anndefs[1] then anndefs = table.concat(anndefs, ", ") anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs)) end end anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ":&#32;" end if pronobj.pron_list and #pronobj.pron_list > 0 then local formatted = require("Module:IPA").format_IPA_full { lang = lang, items = pronobj.pron_list } ins(indent .. anntext .. mw.ustring.toNFC(formatted)) end if args.audios then local format_audio = require("Module:audio").format_audio local audio_objs = parse_audio(lang, args.audios, pagename, "audios") for i, audio_obj in ipairs(audio_objs) do if #audio_objs > 1 and not audio_obj.caption then audio_obj.caption = "Audio " .. i end ins("\n" .. indent .. " " .. format_audio(audio_obj)) end end if #regional_phonetics > 0 then local m_IPA = require("Module:IPA") local m_accent = require("Module:accent qualifier") for _, regional in ipairs(regional_phonetics) do local regions = regional.regions local ipa = regional.ipa local pron_item = { pron = "[" .. ipa .. "]" } local formatted_ipa = m_IPA.format_IPA_full { lang = lang, items = { pron_item } } local formatted_region = m_accent.format_qualifiers(lang, regions) local line = "\n" .. indent .. indent .. " " .. formatted_region .. " " .. mw.ustring.toNFC(formatted_ipa) if regional.audio then local audio_obj = { lang = lang, file = regional.audio, } local textobj = require(audio_module).construct_audio_textobj(audio_obj) audio_obj.text = textobj line = line .. " " .. require("Module:audio").format_audio(audio_obj) end ins(line) end end return concat(parts) end return export 5fgfg64x6pqtsaczm7x9u3q3t4t78ua 375392 375390 2026-09-22T10:18:09Z Hakimi97 2668 Minor correction 375392 Scribunto text/plain local export = {} local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local audio_module = "Module:audio" local parse_utilities_module = "Module:parse utilities" local rfind = m_str_utils.find local rsplit = m_str_utils.split local ugsub = m_str_utils.gsub local ulen = m_str_utils.len local ulower = m_str_utils.lower local usub = m_str_utils.sub local concat = table.concat local insert = table.insert local lang = require("Module:languages").getByCode("ar") local sc = require("Module:scripts").getByCode("Arab") local correspondences = { ["ʾ"] = "ʔ", ["ṯ"] = "θ", ["j"] = "d͡ʒ", ["ḥ"] = "ħ", ["ḵ"] = "x", ["ḏ"] = "ð", ["š"] = "ʃ", ["ṣ"] = "sˤ", ["ḍ"] = "dˤ", ["ṭ"] = "tˤ", ["ẓ"] = "ðˤ", ["ž"] = "ʒ", ["ʿ"] = "ʕ", ["ḡ"] = "ɣ", ["ḷ"] = "lˤ", ["ū"] = "uː", ["ī"] = "iː", ["ā"] = "aː", ["y"] = "j", ["g"] = "ɡ", ["ē"] = "eː", ["ō"] = "oː", [""] = "", } local vowels = "aāeēiīoōuū" local vowel = "[" .. vowels .. "]" local long_vowels = "āēīōū" local long_vowel = "[" .. long_vowels .. "]" local consonant = "[^" .. vowels .. ". -]" local syllabify_pattern = "(" .. vowel .. ")(" .. consonant .. "?)(" .. consonant .. "?)(" .. vowel .. ")" local tie = "‿" local closed_syllable_shortening_pattern = "(" .. long_vowel .. ")(" .. tie .. ")" .. "(" .. consonant .. ")" local function rsub(term, foo, bar) local retval = ugsub(term, foo, bar) return retval end local function generate_obj(respelling) return { respelling = respelling } end local function split_on_comma(term) if not term then return nil end if term:find(",%s") or term:find("\\") then return require(parse_utilities_module).split_on_comma(term) else return rsplit(term, ",") end end local function parse_respellings_with_modifiers(respelling, paramname) if respelling:find("[<%[]") then local put = require(parse_utilities_module) local segments = put.parse_multi_delimiter_balanced_segment_run(respelling, { { "<", ">" }, { "[", "]" } }) local comma_separated_groups = put.split_alternating_runs_on_comma(segments) local retval = {} for _, group in ipairs(comma_separated_groups) do local j = 2 while j <= #group do if not group[j]:find("^<.*>$") then group[j - 1] = group[j - 1] .. group[j] .. group[j + 1] table.remove(group, j) table.remove(group, j) else j = j + 2 end end local param_mods = { q = { type = "qualifier" }, qq = { type = "qualifier" }, a = { type = "labels" }, aa = { type = "labels" }, ref = { item_dest = "refs", type = "references" }, } table.insert(retval, put.parse_inline_modifiers_from_segments { group = group, arg = respelling, props = { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, }, }) end return retval else local retval = {} for _, item in ipairs(split_on_comma(respelling)) do table.insert(retval, generate_obj(item)) end return retval end end local function parse_pron_modifier(arg, paramname, generate_obj, param_mods, splitchar) splitchar = splitchar or "," if arg:find("<") then param_mods.q = { type = "qualifier" } param_mods.qq = { type = "qualifier" } param_mods.a = { type = "labels" } param_mods.aa = { type = "labels" } param_mods.ref = { item_dest = "refs", type = "references" } return require(parse_utilities_module).parse_inline_modifiers(arg, { param_mods = param_mods, generate_obj = generate_obj, paramname = paramname, splitchar = splitchar, }) else local retval = {} local split_arg = splitchar == "," and split_on_comma(arg) or rsplit(arg, splitchar) for _, term in ipairs(split_arg) do table.insert(retval, generate_obj(term)) end return retval end end local function parse_audio(lang, arg, pagename, paramname) local param_mods = { IPA = { sublist = true }, text = {}, t = { item_dest = "gloss" }, gloss = {}, pos = {}, lit = {}, g = { item_dest = "genders", sublist = true }, bad = {}, cap = { item_dest = "caption" }, } local function process_special_chars(val) if not val then return val end return (val:gsub("#", pagename)) end local function generate_audio_obj(arg) return { file = process_special_chars(arg) } end local retvals = parse_pron_modifier(arg, paramname, generate_audio_obj, param_mods, "%s*;%s*") for _, retval in ipairs(retvals) do retval.lang = lang retval.text = process_special_chars(retval.text) retval.caption = process_special_chars(retval.caption) local textobj = require(audio_module).construct_audio_textobj(retval) retval.text = textobj retval.gloss = nil retval.pos = nil retval.lit = nil retval.genders = nil end return retvals end local function parse_regional_phonetics(ph_arg, pagename) if not ph_arg or ph_arg == "" then return {} end local regionals = {} for _, item in ipairs(rsplit(ph_arg, "%s*;%s*")) do local audio = nil local item_no_mod = item:gsub("<a:([^>]+)>", function(a) audio = a:gsub("#", pagename) return "" end) local region, ipa = item_no_mod:match("^([^:]+):(.+)$") if region and ipa then local regions = rsplit(region, "%s*,%s*") table.insert(regionals, { regions = regions, ipa = ipa, audio = audio }) end end return regionals end local function syllabify(text) text = ugsub(text, "%-(" .. consonant .. ")%-(" .. consonant .. ")", "%1.%2") text = ugsub(text, "%-", ".") for _ = 1, 2 do text = ugsub( text, syllabify_pattern, function(a, b, c, d) if c == "" and b ~= "" then c, b = b, "" end return a .. b .. "." .. c .. d end ) end text = ugsub(text, "(" .. vowel .. ") (" .. consonant .. ")%.?(" .. consonant .. ")", "%1" .. tie .. "%2.%3") return text end local function closed_syllable_shortening(text) local shorten = { ["ā"] = "a", ["ē"] = "e", ["ī"] = "i", ["ō"] = "o", ["ū"] = "u", } text = ugsub(text, closed_syllable_shortening_pattern, function(vowel, tie, consonant) return shorten[vowel] .. tie .. consonant end) return text end function export.link(term) return require("Module:links").full_link { term = term, lang = lang, sc = sc } end function export.toIPA(list, silent_error) local translit if list.tr then translit = list.tr elseif list.term then require("Module:script utilities").checkScript(list.term, "Arab") translit = lang:transliterate(list.term) if not translit then if silent_error then return '' else error('Module:ar-translit failed to generate a transliteration from "' .. list.term .. '".') end end else if silent_error then return '' else error('No Arabic text or transliteration was provided to the function "toIPA".') end end translit = ugsub(translit, "llāh", "ḷḷāh") translit = ugsub(translit, "([iī] ?)ḷḷ", "%1ll") translit = ugsub(translit, "%(t%)", "") translit = ugsub(translit, "(" .. vowel .. ") " .. vowel, "%1 ") translit = ugsub(translit, "%-?l%-?", "l") translit = syllabify(translit) translit = closed_syllable_shortening(translit) local output = ugsub(translit, ".", correspondences) output = ugsub(output, "%-", "") return output end function export.get_pron_info(terms, pagename, paramname) if #terms == 1 and terms[1].respelling == "-" then return { pron_list = nil } end local pron_list = {} local brackets = "/%s/" for _, term in ipairs(terms) do local respelling = term.respelling local ar_term, tr if not respelling or respelling == "" or respelling == "#" then ar_term = pagename elseif rfind(respelling, "[a-zA-Z]") then tr = respelling elseif respelling:find("[ء-ي]") then ar_term = respelling else tr = respelling end local pron = export.toIPA({ term = ar_term, tr = tr }, false) if pron and pron ~= "" then local bracketed_pron = brackets:format(pron) table.insert(pron_list, { pron = bracketed_pron, q = term.q, qq = term.qq, a = term.a, aa = term.aa, refs = term.refs, }) end end return { pron_list = pron_list } end function export.show_old(frame) local params = { [1] = { list = true, allow_holes = true }, ["tr"] = { list = true, allow_holes = true }, ["qual"] = { list = true, allow_holes = true }, ["nl"] = { type = "boolean" }, ["ann"] = {}, } local args = require("Module:parameters").process(frame:getParent().args, params) local ar_terms = args[1] local transliterations = args.tr local qualifiers = args.qual local nl = args.nl if not (ar_terms.maxindex > 0 or transliterations.maxindex > 0) then if mw.title.getCurrentTitle().nsText == "Templat" then ar_terms[1] = "كَلِمَة" ar_terms.maxindex = 1 else error( 'Please provide vocalized Arabic in the first parameter of {{[[Templat:ar-IPA|ar-IPA]]}}, or transliteration in the "tr" parameter.') end end local pronunciations = {} for i = 1, math.max(ar_terms.maxindex, transliterations.maxindex) do local ar_term = ar_terms[i] local tr = transliterations[i] local qual = qualifiers[i] if not (ar_term or tr) then error("There is a gap in the parameters. Provide either |" .. i .. "= or |tr" .. i .. "=.") elseif ar_term and tr then mw.logObject("Duplicate parameters |" .. i .. "= and |tr" .. i .. "= in {{ar-IPA}},") end local pron = export.toIPA { term = ar_term, tr = tr } table.insert(pronunciations, { pron = "/" .. pron .. "/", q = qual and { qual } or nil }) end local anntext = "" if args.ann then anntext = args.ann if args.ann:find("%+") then local anndefs = {} for i = 1, ar_terms.maxindex do local ar_term = ar_terms[i] if ar_term then table.insert(anndefs, "'''" .. ar_term .. "'''") end end if anndefs[1] then anndefs = table.concat(anndefs, ", ") anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs)) end end anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ":&#32;" end if nl then return anntext .. require("Module:IPA").format_IPA_multiple(lang, pronunciations) else return anntext .. require("Module:IPA").format_IPA_full { lang = lang, items = pronunciations } end end function export.show(frame) local parent_args = frame:getParent().args local process = require("Module:parameters").process local params = { [1] = {}, ["audios"] = {}, ["a"] = { alias_of = "audios" }, ["ph"] = {}, ["pagename"] = {}, ["indent"] = {}, ["ann"] = {}, } local args = process(parent_args, params) local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local indent = args.indent or "*" local termspec = args[1] or "#" local terms = parse_respellings_with_modifiers(termspec, 1) local pronobj = export.get_pron_info(terms, pagename, 1) local regional_phonetics = parse_regional_phonetics(args.ph, pagename) local parts = {} local function ins(text) table.insert(parts, text) end local anntext = "" if args.ann then anntext = args.ann if args.ann:find("%+") then local anndefs = {} for _, term in ipairs(terms) do local respelling = term.respelling if respelling and respelling:find("[ء-ي]") then table.insert(anndefs, "'''" .. respelling .. "'''") end end if anndefs[1] then anndefs = table.concat(anndefs, ", ") anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs)) end end anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ":&#32;" end if pronobj.pron_list and #pronobj.pron_list > 0 then local formatted = require("Module:IPA").format_IPA_full { lang = lang, items = pronobj.pron_list } ins(indent .. anntext .. mw.ustring.toNFC(formatted)) end if args.audios then local format_audio = require("Module:audio").format_audio local audio_objs = parse_audio(lang, args.audios, pagename, "audios") for i, audio_obj in ipairs(audio_objs) do if #audio_objs > 1 and not audio_obj.caption then audio_obj.caption = "Audio " .. i end ins("\n" .. indent .. " " .. format_audio(audio_obj)) end end if #regional_phonetics > 0 then local m_IPA = require("Module:IPA") local m_accent = require("Module:accent qualifier") for _, regional in ipairs(regional_phonetics) do local regions = regional.regions local ipa = regional.ipa local pron_item = { pron = "[" .. ipa .. "]" } local formatted_ipa = m_IPA.format_IPA_full { lang = lang, items = { pron_item } } local formatted_region = m_accent.format_qualifiers(lang, regions) local line = "\n" .. indent .. indent .. " " .. formatted_region .. " " .. mw.ustring.toNFC(formatted_ipa) if regional.audio then local audio_obj = { lang = lang, file = regional.audio, } local textobj = require(audio_module).construct_audio_textobj(audio_obj) audio_obj.text = textobj line = line .. " " .. require("Module:audio").format_audio(audio_obj) end ins(line) end end return concat(parts) end return export psx2d0cilm9hbxy6gfbf7esbqrthwdf Templat:ar-AFA 10 12379 375389 109891 2026-09-22T07:31:04Z Hakimi97 2668 Kemas kini 375389 wikitext text/x-wiki {{#invoke:ar-pronunciation|show_old}}<noinclude>{{documentation}}</noinclude> qeclh9o7e49x3yzowzjt9uk3z1pwafp Australia 0 12875 375381 334113 2026-09-22T06:59:37Z Hakimi97 2668 Pembetulan 375381 wikitext text/x-wiki ==Bahasa Melayu== {{wikipedia}} [[Fail:Australia satellite plane.jpg|thumb|Gambar satelit Australia]] ===Kata nama khas=== {{ms-knk}} # Sebuah negara di [[Oceania]] dengan ibu negaranya ialah [[Canberra]]. ===Etimologi=== Daripada {{der|ms|en|Australia|}}. ===Sebutan=== * {{a|baku}} {{AFA|ms|/ɔ.stra.li.ja/}} * {{a|Johor-Selangor}} {{AFA|ms|/au̯straliə/}} * {{a|Riau-Lingga}} {{AFA|ms|/au̯stralia/}} * {{penyempangan|ms|Au|stra|lia}} ===Tulisan Jawi=== {{ARchar|اوستراليا}} ===Sinonim=== * [[Komanwel Australia]], nama rasmi. [[Kategori:ms:Negara]] ey6m4e4opemvk21nfi74rb8334yt07m Templat:ms-jawi 10 13235 375384 228927 2026-09-22T07:08:51Z Hakimi97 2668 Kemas kini templat 375384 wikitext text/x-wiki {{#invoke:form of/templates|form_of_t|lang=ms|Ejaan [[w:Tulisan Jawi|Jawi]] bagi|withdot=true}}‎<noinclude>{{documentation}}</noinclude> a33znw28lum5b9r9mse6rv40vwiycbm saro 0 14248 375344 190601 2026-09-21T16:04:52Z Muhamad Izzul Fiqih 7922 /* Bahasa Batak Mandailing */ 375344 wikitext text/x-wiki == Bahasa Batak Mandailing == === Takrifan === ==== Kata nama ==== {{head|btm|kata nama}} # [[bahasa]] === Sebutan === * {{penyempangan|btm|sa|ro}} == Bahasa Melayu Jambi == ===Takrifan=== ====Kata nama==== {{inti|jax|kata sifat}} # [[payah]] guy3762r3vyymg4ihwctfftjkk99r38 375345 375344 2026-09-21T16:05:16Z Muhamad Izzul Fiqih 7922 /* Kata nama */ 375345 wikitext text/x-wiki == Bahasa Batak Mandailing == === Takrifan === ==== Kata nama ==== {{head|btm|kata nama}} # [[bahasa]] === Sebutan === * {{penyempangan|btm|sa|ro}} == Bahasa Melayu Jambi == ===Takrifan=== ====Kata nama==== {{inti|jax|kata sifat}} # [[susah]] es07wyeaoc32vv6bhkljvvssw7ycbhn Modul:hyphenation 828 14282 375350 236146 2026-09-22T03:12:42Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708593|92708593]]) 375350 Scribunto text/plain local export = {} local decorations_module = "Module:decorations" local links_module = "Module:links" local parameters_module = "Module:parameters" local function track(page) require("Module:debug/track")("hyphenation/" .. page) return true end --[==[ Meant to be called from a module. `data` is a table containing the following fields: * `lang`: language object for the hyphenations or syllabifications; * `hyphs`: a list of hyphenations/syllabifications, each described by an object which can contain the following fields: ** `hyph`: list of syllables comprising the hyphenation or syllabification, each a string; ** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the hyphenation in question; ** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the hyphenation in question; ** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed directly before the hyphenation in question; ** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the hyphenation in question; ** `refs`: {nil} or a list of references or reference specs to add directly after the hyphenation; the value of a list item is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}}) and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the actual reference, located in the {{cd|<nowiki><references /></nowiki>}} section; ** `sc`: {nil} or script object for this particular hyphenation/syllabification; * `q`: {nil} or a list of overall left regular qualifier strings, displayed before the initial caption; * `qq`: {nil} or a list of overall right regular qualifier strings, displayed after all hyphenations; * `a`: {nil} or a list of overall left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the initial caption; * `aa`: {nil} or a list of right accent qualifier strings, displayed after all hyphenations; * `sc`: {nil} or script object for the hyphenations/syllabifications; * `caption`, {nil} or a string specifying the caption to use, in place of {"Hyphenation"}; e.g. use {"Syllabification"} if what is passed in is actually a syllabification, as is common; a colon and space is automatically added after the caption; * `nocaption`: if true, suppress the caption display. ]==] function export.format_hyphenations(data) local hyphtexts = {} for _, hyph in ipairs(data.hyphs) do if #hyph.hyph == 0 then error("Saw empty hyphenation; use || to separate hyphenations") end if hyph.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("`.qualifiers` is no longer supported; change the code to use `.qq` or `.q`") end local text = require(links_module).full_link { lang = data.lang, sc = hyph.sc or data.sc, alt = table.concat(hyph.hyph, "‧"), tr = "-", q = hyph.q, qq = hyph.qq, a = hyph.a, aa = hyph.aa, refs = hyph.refs, show_decorations = true, } table.insert(hyphtexts, text) end local text = ((data.nocaption and "") or ((data.caption or "Penyempangan") .. ": ")) .. table.concat(hyphtexts, ", ") if data.qualifiers then -- FIXME: added 2026-09-18; consider removing eventually. error("overall `.qualifiers` is no longer supported; change the code to use `.qq` or `.q`") end if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then text = require(decorations_module).format_decorations { lang = data.lang, text = text, q = data.q, qq = data.qq, a = data.a, aa = data.aa, } end return text end --[==[ Entry point for {{tl|hyphenation}} template (also written {{tl|hyph}}). ]==] function export.hyphenation(frame) local parent_args = frame:getParent().args local compat = parent_args.lang local lang_param = compat and "lang" or 1 local offset = compat and 0 or 1 local params = { [lang_param] = {required = true, type = "language", default = "und"}, [1 + offset] = {list = true, required = true, allow_holes = true, default = "{{{2}}}"}, -- FIXME: For compatibility, q= and qq= refer to the first hyphenation rather than overall. Consider tracking and -- changing this. ["q"] = {list = true, allow_holes = true, type = "qualifier"}, ["qq"] = {list = true, allow_holes = true, type = "qualifier"}, ["a"] = {list = true, allow_holes = true, separate_no_index = true, type = "labels"}, ["aa"] = {list = true, allow_holes = true, separate_no_index = true, type = "labels"}, ["ref"] = {list = true, allow_holes = true, type = "references"}, ["caption"] = {}, ["nocaption"] = {type = "boolean"}, ["sc"] = {list = true, allow_holes = true, separate_no_index = true, type = "script"}, } local args = require(parameters_module).process(parent_args, params) local lang = args[lang_param] local sc = args.sc.default local data = { lang = lang, sc = sc, hyphs = {}, caption = args.caption, nocaption = args.nocaption, a = args.a.default, aa = args.aa.default, } local this_hyph = {hyph = {}} local maxindex = args[1 + offset].maxindex local function insert_hyph() local hyphnum = #data.hyphs + 1 this_hyph.q = args.q[hyphnum] this_hyph.qq = args.qq[hyphnum] this_hyph.a = args.a[hyphnum] this_hyph.aa = args.aa[hyphnum] this_hyph.refs = args.ref[hyphnum] table.insert(data.hyphs, this_hyph) end for i=1, maxindex do local syl = args[1 + offset][i] if not syl then insert_hyph() this_hyph = {hyph = {}} else table.insert(this_hyph.hyph, syl) end end insert_hyph() return export.format_hyphenations(data) end return export noi6hiwft43y5yinfixbo8jrkhbmo8m Modul:affixusex 828 22134 375357 223660 2026-09-22T03:14:12Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708532|92708532]]) 375357 Scribunto text/plain local export = {} local decorations_module = "Module:decorations" local etymology_module = "Module:etymology" local links_module = "Module:links" -- main function --[==[ Format the affix usexes in `data`. We more or less simply call full_link() on each item, along with the associated params, to format the link, but need some special-casing for affixes. On input, the `data` object contains the following fields: * `lang` ('''required'''): Overall language object; default for items not specifying their own language. * `sc`: Overall script object; default for items not specifying their own script. * `items`: List of items. Each is an object with the following fields: ** `term`: The term (affix or resulting term). ** `gloss`, `tr`, `ts`, `genders`, `alt`, `id`, `lit`, `pos`, `ng`: The same as for `full_links()` in [[Module:links]]. ** `lang`: Language of the term. Should only be set when the term has its own language, and will cause the language to be displayed before the term. Defaults to the overall `lang`. ** `sc`: Script of the term. Defaults to the overall `sc`. ** `fulljoiner`: Text of the separator appearing before the item, including spaces. Takes precedence over `joiner` and `arrow`. ** `joiner`: Text of the separator appearing before the item, not including spaces. Takes precedence over `arrow`. ** `arrow`: If specified, the separator is a right arrow. If none of `fulljoiner`, `joiner` and `arrow` are given, the separator is a right arrow if it's the last item, otherwise a plus sign if it's not the first item, otherwise there's no displayed separator. ** `l`: Left labels for the term. ** `ll`: Right labels for the term. ** `q`: Left regular qualifier(s) for the term. ** `qq`: Right regular qualifier(s) for the term. ** `refs`: References for the term, in the structure expected by [[Module:references]]. * `lit`: Overall literal meaning. * `l`: Overall left labels. * `ll`: Overall right labels. * `q`: Overall left regular qualifier(s). * `qq`: Overall right regular qualifier(s). '''WARNING:''' This destructively modifies the `items` objects (specifically by adding default values for `lang` and `sc`). ]==] function export.format_affixusex(data) local result = {} for index, item in ipairs(data.items) do if item.fulljoiner then table.insert(result, item.fulljoiner) elseif item.joiner then table.insert(result, " " .. item.joiner .. " ") elseif index == #data.items or item.arrow then table.insert(result, " → ") elseif index > 1 then table.insert(result, " + ") end table.insert(result, "&lrm;") local text local item_lang_specific = item.lang item.lang = item.lang or data.lang item.sc = item.sc or data.sc if item_lang_specific then text = require(etymology_module).format_derived { sources = {item.lang}, terms = {item}, -- Don't need to specify `nocat = true` because we don't pass in `lang`. template_name = "affixusex", decorations_on_outside = true, } else text = require(links_module).full_link(item, "term") end table.insert(result, text) end result = table.concat(result) .. (data.lit and ", secara harfiah " .. require(links_module).mark(data.lit, "gloss") or "") if data.q and data.q[1] or data.qq and data.qq[1] or data.l and data.l[1] or data.ll and data.ll[1] then result = require(decorations_module).format_decorations { lang = data.lang, text = result, q = data.q, qq = data.qq, l = data.l, ll = data.ll, } end return result end return export lyzn6lkv4gd5qxbwwauzu6e1iowi4c8 Modul:pendokumenan 828 24368 375361 335738 2026-09-22T03:50:53Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92487256|92487256]]) 375361 Scribunto text/plain local export = {} local debug_track_module = "Modul:debug/track" local frame_module = "Modul:frame" local fun_is_callable_module = "Modul:fun/isCallable" local languages_module = "Modul:languages" local links_module = "Modul:links" local load_module = "Modul:load" local module_categorization_module = "Modul:module categorization" local number_list_show_module = "Modul:number list/show" local chemical_element_list_show_module = "Modul:chemical element list/show" local pages_module = "Modul:pages" local parameters_module = "Modul:parameters" local scripts_module = "Modul:scripts" local string_endswith_module = "Modul:string/endswith" local string_gline_module = "Modul:string/gline" local string_startswith_module = "Modul:string/startswith" local string_utilities_module = "Modul:string utilities" local template_parser_module = "Modul:template parser" local title_exists_module = "Modul:title/exists" local title_new_title_module = "Modul:title/newTitle" local concat = table.concat local error = error local full_url = mw.uri.fullUrl local get_current_title = mw.title.getCurrentTitle local insert = table.insert local ipairs = ipairs local list_to_text = mw.text.listToText local new_message = mw.message.new local pcall = pcall local require = require local tonumber = tonumber local tostring = tostring local type = type local unpack = unpack or table.unpack -- Lua 5.2 compatibility local function categorize_module(...) categorize_module = require(module_categorization_module).categorize_module return categorize_module(...) end local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function endswith(...) endswith = require(string_endswith_module) return endswith(...) end local function expand_template(...) expand_template = require(frame_module).expandTemplate return expand_template(...) end local function find_templates(...) find_templates = require(template_parser_module).find_templates return find_templates(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function get_pagetype(...) get_pagetype = require(pages_module).get_pagetype return get_pagetype(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function gline(...) gline = require(string_gline_module) return gline(...) end local function is_callable(...) is_callable = require(fun_is_callable_module) return is_callable(...) end local function is_documentation(...) is_documentation = require(pages_module).is_documentation return is_documentation(...) end local function is_sandbox(...) is_sandbox = require(pages_module).is_sandbox return is_sandbox(...) end local function new_title(...) new_title = require(title_new_title_module) return new_title(...) end local function number_list_show_table(...) number_list_show_table = require(number_list_show_module).table return number_list_show_table(...) end local function chemical_element_list_show_table(...) chemical_element_list_show_table = require(chemical_element_list_show_module).table return chemical_element_list_show_table(...) end local function preprocess(...) preprocess = require(frame_module).preprocess return preprocess(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function safe_load_data(...) safe_load_data = require(load_module).safe_load_data return safe_load_data(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function startswith(...) startswith = require(string_startswith_module) return startswith(...) end local function title_exists(...) title_exists = require(title_exists_module) return title_exists(...) end local function ugsub(...) ugsub = require(string_utilities_module).gsub return ugsub(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end local skins = { ["common"] = "", ["vector"] = "Vector", ["monobook"] = "Monobook", ["cologneblue"] = "Cologne Blue", ["modern"] = "Modern", } local function track(page) debug_track("documentation/" .. page) return true end local function compare_pages(page1, page2, text) return "[" .. tostring( full_url("Khas:ComparePages", { page1 = page1, page2 = page2 })) .. " " .. text .. "]" end -- Avoid transcluding [[Modul:languages/cache]] everywhere. local lang_cache = setmetatable({}, { __index = function(self, k) return require("Modul:languages/cache")[k] end }) local function zh_link(word) return full_link { lang = lang_cache.zh, term = word } end local function make_languages_data_documentation(_title, cats, division) local doc_template, module_cat if endswith(division, "/extra") then division = division:sub(1, -7) doc_template = "language extradata documentation" module_cat = "Modul data ekstra bahasa" else doc_template = "language data documentation" module_cat = "Modul data bahasa" end local sort_key if division == "exceptional" then sort_key = "x" else sort_key = division:gsub("/", "") end insert(cats, module_cat .. "|" .. sort_key) return { title = doc_template } end local function make_Unicode_data_documentation(title, _cats) local subpage, first_three_of_code_point = title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$") if subpage == "names" or subpage == "images" or subpage == "emoji images" then local low, high = tonumber(first_three_of_code_point .. "000", 16), tonumber(first_three_of_code_point .. "FFF", 16) local text, text_type if subpage == "names" then text_type = "titles of images" elseif subpage == "images" then text_type = "titles of images" elseif subpage == "emoji images" then text_type = "emoji-style images" end text = string.format( "Modul data ini mengandungi " .. text_type .. " kepada " .. "titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.", low, high) if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then text = text .. " Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" .. first_three_of_code_point .. "]]." elseif subpage == "emoji images" then text = text .. " Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]." end return text end end local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat) local lang = lang_cache[langcode] if lang then local langname if lang._fullCode then langname = lang_cache[lang._fullCode]:getCanonicalName() else langname = lang:getCanonicalName() end insert(cats, overall_data_module_cat .. "|" .. langname) insert(cats, "Modul bahasa " .. langname) insert(cats, "Modul data bahasa " .. langname) return lang, langname end end --[=[ This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't unnecessarily cluttered. It is a list of tables, each of which have the following possible fields: `regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used. Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the `process` function as the third and subsequent parameters. `process` (optional): This may be a function or a string. If it is a function, it is called as follows: `process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)` where: * TITLE is a title object describing the module's title; see [https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects]. * CATS is an list of categories that the module will be added to. * CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field. The return value of `process` should either be a string (which will be used as the module's documentation), or a table specifying the name of a template to expand to get the documentation, along with the arguments to that template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title` field, and any arguments should be in `args; in this case, the template name will be listed above the generated documentation as the source of the documentation, along with an edit button to edit the template's contents. If, however, the return value of the `process` function is a string, any template invocations will be expanded using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation. If `process` itself is a string rather than a function, it should name a submodule under [[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule will be specified as the source of the documentation (unless it returns a table naming a template to expand to get the documentation, as described above). If `process` is omitted entirely, the module will have no documentation. `cat` (optional): A string naming the category into which the module should be placed, or a list of such strings. Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second, etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS list (the second parameter). ]=] local module_regex = { { regex = "^Modul:languages/data/(3/%l/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(3/%l)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:scripts/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:data tables/data..?.?.?$", cat = "Jadual data berpecah modul rujukan", }, { regex = "^Modul:zh/data/dial%-pron/.+$", cat = "Modul data sebutan dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/dial%-syn/.+$", cat = "Modul data sinonim dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/glyph%-data/.+$", cat = "Modul data bentuk aksara Cina bersejarah", process = function(title, _cats) local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)") if character then return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.") :format(zh_link(character)) end end, }, { regex = "^Modul:zh/data/ltc%-pron/(.+)$", cat = "Modul data sebutan bahasa Cina Pertengahan|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-BS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1", process = "zh data", }, { -- capture rest of zh/data submodules regex = "^Modul:zh/data/(.+)$", cat = "Modul data bahasa Cina|%1", }, { regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$", process = "guoxue-data", }, { regex = "^Modul:Unicode data/(.+)$", cat = "Modul data Unicode|%1", process = make_Unicode_data_documentation, }, { regex = "^Modul:number list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor") if lang then return ("This module contains data on various types of numbers in %s.\n%s") :format(lang:makeCategoryLink(), number_list_show_table() or "") end end, }, { regex = "^Modul:chemical element list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia") if lang then return ("This module contains data on chemical elements in %s.\n%s") :format(lang:makeCategoryLink(), chemical_element_list_show_table() or "") end end, }, { regex = "^Modul:accel/(.+)$", process = function(title, cats) local lang_code = title.subpageText local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel") insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName())) return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.") :format(lang:makeCategoryLink()) end end, }, { regex = "^Modul:inc%-ash/dial/data/(.+)$", cat = "Modul Prakrit Ashoka|%1", process = function(title, _cats) local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$") if word then local lang = lang_cache["inc-ash"] return ("This module contains data on the pronunciation of %s in dialects of %s.") :format(full_link({ term = word, lang = lang }, "term"), lang:makeCategoryLink()) end end, }, { regex = "^.+%-translit$", process = function(title, _cats) return require("Modul:documentation/translit-like").documentation { operation = "translit", title_without_namespace = title.text, } end, }, { regex = "^.+%-sortkey$", process = function(title, _cats) return require("Modul:documentation/translit-like").documentation { operation = "sortkey", title_without_namespace = title.text, } end, }, { regex = "^.+%-stripdiacritics$", process = function(title, _cats) return require("Modul:documentation/translit-like").documentation { operation = "strip diacritics", title_without_namespace = title.text, } end, }, { regex = "^Modul:form of/lang%-data/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa") if lang then -- FIXME, display more info. return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " .. langname .. "." end end }, { regex = "^Modul:labels/data/lang/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa") if lang then return { title = "label language-specific data documentation", args = { [1] = lang_code }, } end end }, { regex = "^Modul:category tree/lang/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang") if lang then return "This module handles generating the descriptions and categorization for " .. langname .. " category pages " .. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are " .. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. " .. "This module is part of the category tree system, which is a general framework for generating the " .. "descriptions and categorization of category pages.\n\n" .. "For more information, see [[Modul:category tree/lang/documentation]].\n\n" .. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the " .. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized." end end }, { regex = "^Modul:category tree/topic/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/topic| ") return { title = "topic cat data submodule documentation" } end }, { regex = "^Modul:category tree/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/grammar| ") return { title = "category tree data submodule documentation" } end }, { regex = "^Modul:ja/data/(.+)$", cat = "Modul data bahasa Jepun|%1", }, { regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$", cat = "Modul atlas data dialek Finland|%1", process = function(_title, _cats, shard) return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " .. "''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " .. "It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " .. "''Kotimaisten kielten keskus'' under the CC BY 4.0 license." end }, { regex = "^Modul:fi%-dialects/data/feature/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:fi%-dialects/data/word/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:Swadesh/data/([%l-]+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "." end end }, { regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$", process = function(_title, cats, lang_code, variety) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the " local etym_lang = get_lang(variety, nil, "allow etym") if etym_lang then return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname) end local script = get_script(variety) if script then return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang)) end return ("%s %s variety of %s."):format(prefix, variety, langname) end end }, { regex = "^Modul:typing%-aids", process = function(title, cats) local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$") local sortkey if data_suffix then if data_suffix:find "^[%l-]+$" then local lang = get_lang(data_suffix) if lang then sortkey = lang:getCanonicalName() insert(cats, "Modul data bahasa " .. sortkey) end elseif data_suffix:find "^%u%l%l%l$" then local script = get_script(data_suffix) if script then -- FIXME: no lang to pass here sortkey = script:getCanonicalName() insert(cats, script:getCategoryName()) end end insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix)) end end, }, { regex = "^Modul:R:([%l-]+):(.+)$", process = function(_title, cats, lang_code, refname) local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname) insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName())) return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}." end end, }, { regex = "^Modul:Quotations/([%l-]+)/?(.*)", process = "Quotation", }, { regex = "^Modul:affix/lang%-data/([%l-]+)", process = "affix lang-data", }, { regex = "^Modul:dialect synonyms/([%l-]+)$", process = function(_title, cats, lang_code) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "| ") return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " .. "{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" .. "==== Struktur modul data bahasa ====\n" .. "* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" .. "* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" .. "* <code>export.notes</code> — optional; table of note keys to text.\n" .. "* <code>export.sources</code> — optional; table of source keys to text.\n" .. "* <code>export.note_aliases</code> — optional; alias map for notes.\n" .. "* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" .. expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } }) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$", process = function(_title, cats, lang_code, term) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term } })) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$", process = function(_title, cats, lang_code, term, id) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term, id = id } })) end end, }, { regex = "^Modul:bibliography/data/([%l-]+)$", process = function(title, cats, lang_code) if lang_code == "preload" then return 'Used as a base model for other languages when the button "create new language submodule" is clicked.' end local page = require(title.fullText).bib_page if not page then page = lang_cache[lang_code]:getCanonicalName() if page then insert(cats, "Modul bahasa " .. page) end end insert(cats, "Modul rujukan") return "This module holds bibliographical data for " .. page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''." end, }, } function export.show(frame) local boolean_default_false = { type = "boolean", default = false } local args = process_params(frame.args, { ["hr"] = true, ["for"] = true, ["from"] = true, ["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage. ["notsubpage"] = boolean_default_false, ["nodoc"] = boolean_default_false, ["nolinks"] = boolean_default_false, -- suppress all "Useful links" ["nosandbox"] = boolean_default_false, -- supress sandbox }) local output = {'\n<div class="documentation" style="display:block; clear:both">\n'} local function ins(txt) insert(output, txt) end local cats = {} local function inscat(cat) insert(cats, cat) end local nodoc = args.nodoc if (not args.hr) or (args.hr == "above") then ins("----\n") end local title = args["for"] and new_title(args["for"]) or get_current_title() local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil local contentModel = title.contentModel local pagetype, is_script_or_stylesheet = get_pagetype(title) local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc local doc_content_source = "Modul:pendokumenan" local auto_generated_cat_source local cats_auto_generated = false if not args.allowondoc and is_documentation(title) then -- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type. error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].") elseif is_sandbox(title) then local sandbox_ns = title.nsText preload = ("Templat:pendokumenan/preload%s%sSandbox"):format( sandbox_ns == "Modul" and sandbox_ns or "Templat", title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or "" ) elseif pagetype:match("%f[%w]gadget%f[%W]") then preload = "Templat:pendokumenan/preloadGadget" elseif pagetype:match("%f[%w]script%f[%W]") then -- .js if title.nsText == "MediaWiki" then preload = "Templat:pendokumenan/preloadMediaWikiJavaScript" else preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end end is_script_or_stylesheet = true elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end is_script_or_stylesheet = true elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto. preload = "Templat:pendokumenan/preloadModule" elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then preload = "Templat:pendokumenan/preloadTemplate" end if doc_title and doc_title.isRedirect then old_doc_title = doc_title doc_title = doc_title.redirectTarget end ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">") local function get_module_doc_and_cats(categories_only) cats_auto_generated = true local automatic_cats = nil if user_name then fallback_docs = "pendokumenan/fallback/user module" automatic_cats = { "Modul kotak pasir pengguna" } else for _, data in ipairs(module_regex) do local captures = { umatch(title.fullText, data.regex) } if #captures > 0 then local cat, process_function if is_callable(data.process) then process_function = data.process elseif type(data.process) == "string" then doc_content_source = "Modul:pendokumenan/functions/" .. data.process process_function = require(doc_content_source) end if process_function then doc_content = process_function(title, cats, unpack(captures)) end if type(doc_content) == "table" then doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source doc_content = expand_template(doc_content) elseif doc_content ~= nil then doc_content = preprocess(doc_content) end cat = data.cat if cat then if type(cat) == "string" then cat = { cat } end for _, c in ipairs(cat) do insert(cats, (ugsub(title.fullText, data.regex, c))) end end break end end end if title.subpageText == "templates" then inscat("Modul antara muka templat") end if automatic_cats then for _, c in ipairs(automatic_cats) do inscat(c) end end if #cats == 0 then local auto_cats = categorize_module { return_raw = true, noerror = true, } if #auto_cats > 0 then auto_generated_cat_source = "Modul:module categorization" end for _, category in ipairs(auto_cats) do inscat(category) end end -- meaning module is not in user’s sandbox or one of many datamodule boring series needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs) end -- Override automatic documentation, if present. if doc_title and doc_title.exists then local cats_auto_generated_text = "" if contentModel == "Scribunto" then local doc_page_content = doc_title.content -- Track then do nothing if there are uses of includeonly. The -- pattern is slightly too permissive, but any false-positives are -- obvious typos that should be corrected. if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then track("module-includeonly") else -- Check for uses of {{module cat}}. find_templates treats the -- input as transcluded by default (i.e. it parses the wikitext -- which will be transcluded through to the module page). local module_cat for template in find_templates(doc_page_content) do if Templat:get_name() == "module cat" then module_cat = true break end end if not module_cat then get_module_doc_and_cats("categories only") auto_generated_cat_source = auto_generated_cat_source or doc_content_source cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" .. auto_generated_cat_source .. "]]. <sup>[[" .. new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>" end end end ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" .. doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" .. cats_auto_generated_text .. "</i></dd>") else if contentModel == "Scribunto" then get_module_doc_and_cats(false) elseif title.nsText == "Templat" then --inscat("Uncategorized templates") needs_doc = not (fallback_docs or nodoc) elseif user_name and is_script_or_stylesheet then skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")] if skin_name then fallback_docs = "pendokumenan/fallback/user " .. contentModel end end if doc_content then ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " .. "dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" .. new_title(doc_content_source):fullUrl { action = "edit" } .. " sunting]]</sup> </i></dd>") elseif not nodoc then if doc_title then ins( "<dd><i style=\"font-size: larger;\">Laman " .. pagetype .. " ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " .. (fallback_docs and "Anda boleh " or "Minta tolong ") .. "[" .. doc_title:fullUrl { action = "edit", preload = preload } .. " ciptakan laman pendokumenan tersebut].</i></dd>\n") else ins( "<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " .. "pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n") end end end if startswith(title.fullText, "MediaWiki:Gadget-") then local is_gadget = false for line in gline(new_title("MediaWiki:Gadgets-definition").content) do local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$") if not gadget then gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$") end if gadget then items = split(items, "|") for i, item in ipairs(items) do if title.fullText == ("MediaWiki:Gadget-" .. item) then is_gadget = true ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>") ins(gadget) ins("</code> gajet ([") ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(" sunting takrifan])'' <dl>") ins("<dd> ''Huraian ([") ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" }))) ins(" sunting])'': ") ins(preprocess(new_message('Gadget-' .. gadget):plain())) ins(" </dd>") table.remove(items, i) if #items > 0 then for j, item in ipairs(items) do items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]' end ins("<dd> ''Bahagian lain'': ") ins(list_to_text(items)) ins("</dd>") end ins("</dl></dd>") break end end end end if not is_gadget then ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [") ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" }))) ins(' gajet] ([') ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(' sunting takrifan]).</dd>') -- else -- inscat("Wiktionary gadgets") end end if old_doc_title then ins("<dd> ''Dilencong daripada'' [") ins(old_doc_title:fullUrl { redirect = "no" }) ins(" ") ins(old_doc_title.fullText) ins("] ([") ins(old_doc_title:fullUrl { action = "edit" }) ins(" sunting]).</dd>\n") end if not args.nolinks then local links = {} local function inslinks(txt) insert(links, txt) end if title.isSubpage and not args.notsubpage then inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]") inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]") else inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]") end inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) .. " pautan]") if contentModel ~= "Scribunto" then inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) .. " lencongan]") end if is_script_or_stylesheet then if user_name then inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]") end else inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) .. " transklusi]") end if contentModel == "Scribunto" then local is_testcases = title.isSubpage and title.subpageText == "testcases" local without_subpage = title.nsText .. ":" .. title.baseText if is_testcases then inslinks("[[:" .. without_subpage .. "|modul ujian]]") else inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]") end if user_name then inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]") inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]") inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]") -- If sandbox module, add a link to the module that this is a sandbox of. -- Exclude user sandbox modules like [[User:Dine2016/sandbox]]. elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then inscat("Modul kotak pasir") -- Sandbox modules don’t really need documentation. needs_doc = false -- Don't track user sandbox modules. local text_title = new_title(title.text) if not (text_title and text_title.nsText == "Pengguna") then local diff local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]") if sandbox_of then track("sandbox to be moved") else sandbox_of = title.text:match("^sandbox%d*/(.*)$") end if not sandbox_of then error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'") :format(title.text)) end sandbox_of = title.nsText .. ":" .. sandbox_of if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) end -- If not a sandbox module, add link to sandbox module. -- Sometimes there are multiple sandboxes for a single Modul: -- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]]. else local sandbox_title local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$") if not user_prefix then user_prefix = "" user_rest = title.text end sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks(sandbox_link .. (diff or "")) end end if title.nsText == "Templat" then -- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]] -- apparently same with/without: &profile=advanced&fulltext=1 local errorq = 'searchengineselect=mediawiki&search=all: hasTemplat:\"' .. title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:' local eincategory = "Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors" inslinks( '[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]' .. ' (' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]' .. '/' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]' .. ')' ) if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template. -- At the moment there are no user sandbox templates with subpage -- “/sandbox”. inscat("Templat kotak pasir") -- Sandbox templates don’t really need documentation. needs_doc = false -- Will behave badly if “/sandbox” occurs twice in title! local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "") local diff if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) -- This is a template that can have a sandbox. elseif not args.nosandbox then -- unless we tell it not to local sandbox_title = title.fullText .. "/sandbox" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or "")) end end if #links > 0 then ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>") end end ins("</dl>\n") -- Show error from [[Modul:category tree/topic cat/data]] on its submodules' -- documentation to, for instance, warn about duplicate labels. if startswith(title.fullText, "Modul:category tree/topic/") then local ok, err = pcall(require, "Modul:category tree/topic/data") if not ok then ins('<span class="error">' .. err .. '</span>\n\n') end end if doc_title and doc_title.exists then -- Override automatic documentation, if present. doc_content = expand_template { title = doc_title.fullText } elseif not doc_content and fallback_docs then doc_content = expand_template { title = fallback_docs, args = { ['user'] = user_name, ['page'] = title.fullText, ['skin name'] = skin_name, }, } end if doc_content then ins(doc_content) end ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br")) if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then if contentModel == "Scribunto" then inscat("Modul belum dikategorikan") -- elseif title.nsText == "Templat" then -- inscat("Templat belum dikategorikan") end end if needs_doc then inscat("Templat dan modul yang memerlukan pendokumenan") end for _, cat in ipairs(cats) do ins("[[Kategori:" .. cat .. "]]") end ins("</div>\n") return concat(output) end function export.module_auto_doc_table() local parts = {} local function ins(text) insert(parts, text) end ins('{|class="wikitable"') ins("! Regex !! Kategori !! Modul yang dikendalikan") for _, spec in ipairs(module_regex) do local cat_text local cats = spec.cat if cats then local cat_parts = {} if type(cats) == "string" then cats = { cats } end for _, cat in ipairs(cats) do insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "&#124;")))) end cat_text = concat(cat_parts, ", ") else cat_text = "''(tidak dinyatakan secara khusus)''" end ins("|-") ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text, is_callable(spec.process) and "''(dikendali secara dalaman)''" or type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or "''(tiada penjana pendokumenan)''")) end ins("|}") return concat(parts, "\n") end return export pllus5uhsp0qm036oww1d624tmv9bxr 375362 375361 2026-09-22T04:13:41Z Hakimi97 2668 Pembetulan kecil 375362 Scribunto text/plain local export = {} local debug_track_module = "Modul:debug/track" local frame_module = "Modul:frame" local fun_is_callable_module = "Modul:fun/isCallable" local languages_module = "Modul:languages" local links_module = "Modul:links" local load_module = "Modul:load" local module_categorization_module = "Modul:module categorization" local number_list_show_module = "Modul:number list/show" local chemical_element_list_show_module = "Modul:chemical element list/show" local pages_module = "Modul:pages" local parameters_module = "Modul:parameters" local scripts_module = "Modul:scripts" local string_endswith_module = "Modul:string/endswith" local string_gline_module = "Modul:string/gline" local string_startswith_module = "Modul:string/startswith" local string_utilities_module = "Modul:string utilities" local template_parser_module = "Modul:template parser" local title_exists_module = "Modul:title/exists" local title_new_title_module = "Modul:title/newTitle" local concat = table.concat local error = error local full_url = mw.uri.fullUrl local get_current_title = mw.title.getCurrentTitle local insert = table.insert local ipairs = ipairs local list_to_text = mw.text.listToText local new_message = mw.message.new local pcall = pcall local require = require local tonumber = tonumber local tostring = tostring local type = type local unpack = unpack or table.unpack -- Lua 5.2 compatibility local function categorize_module(...) categorize_module = require(module_categorization_module).categorize_module return categorize_module(...) end local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function endswith(...) endswith = require(string_endswith_module) return endswith(...) end local function expand_template(...) expand_template = require(frame_module).expandTemplate return expand_template(...) end local function find_templates(...) find_templates = require(template_parser_module).find_templates return find_templates(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function get_pagetype(...) get_pagetype = require(pages_module).get_pagetype return get_pagetype(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function gline(...) gline = require(string_gline_module) return gline(...) end local function is_callable(...) is_callable = require(fun_is_callable_module) return is_callable(...) end local function is_documentation(...) is_documentation = require(pages_module).is_documentation return is_documentation(...) end local function is_sandbox(...) is_sandbox = require(pages_module).is_sandbox return is_sandbox(...) end local function new_title(...) new_title = require(title_new_title_module) return new_title(...) end local function number_list_show_table(...) number_list_show_table = require(number_list_show_module).table return number_list_show_table(...) end local function chemical_element_list_show_table(...) chemical_element_list_show_table = require(chemical_element_list_show_module).table return chemical_element_list_show_table(...) end local function preprocess(...) preprocess = require(frame_module).preprocess return preprocess(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function safe_load_data(...) safe_load_data = require(load_module).safe_load_data return safe_load_data(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function startswith(...) startswith = require(string_startswith_module) return startswith(...) end local function title_exists(...) title_exists = require(title_exists_module) return title_exists(...) end local function ugsub(...) ugsub = require(string_utilities_module).gsub return ugsub(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end local skins = { ["common"] = "", ["vector"] = "Vector", ["monobook"] = "Monobook", ["cologneblue"] = "Cologne Blue", ["modern"] = "Modern", } local function track(page) debug_track("pendokumenan/" .. page) return true end local function compare_pages(page1, page2, text) return "[" .. tostring( full_url("Khas:ComparePages", { page1 = page1, page2 = page2 })) .. " " .. text .. "]" end -- Avoid transcluding [[Modul:languages/cache]] everywhere. local lang_cache = setmetatable({}, { __index = function(self, k) return require("Modul:languages/cache")[k] end }) local function zh_link(word) return full_link { lang = lang_cache.zh, term = word } end local function make_languages_data_documentation(_title, cats, division) local doc_template, module_cat if endswith(division, "/extra") then division = division:sub(1, -7) doc_template = "language extradata documentation" module_cat = "Modul data ekstra bahasa" else doc_template = "language data documentation" module_cat = "Modul data bahasa" end local sort_key if division == "exceptional" then sort_key = "x" else sort_key = division:gsub("/", "") end insert(cats, module_cat .. "|" .. sort_key) return { title = doc_template } end local function make_Unicode_data_documentation(title, _cats) local subpage, first_three_of_code_point = title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$") if subpage == "names" or subpage == "images" or subpage == "emoji images" then local low, high = tonumber(first_three_of_code_point .. "000", 16), tonumber(first_three_of_code_point .. "FFF", 16) local text, text_type if subpage == "names" then text_type = "titles of images" elseif subpage == "images" then text_type = "titles of images" elseif subpage == "emoji images" then text_type = "emoji-style images" end text = string.format( "Modul data ini mengandungi " .. text_type .. " kepada " .. "titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.", low, high) if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then text = text .. " Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" .. first_three_of_code_point .. "]]." elseif subpage == "emoji images" then text = text .. " Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]." end return text end end local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat) local lang = lang_cache[langcode] if lang then local langname if lang._fullCode then langname = lang_cache[lang._fullCode]:getCanonicalName() else langname = lang:getCanonicalName() end insert(cats, overall_data_module_cat .. "|" .. langname) insert(cats, "Modul bahasa " .. langname) insert(cats, "Modul data bahasa " .. langname) return lang, langname end end --[=[ This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't unnecessarily cluttered. It is a list of tables, each of which have the following possible fields: `regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used. Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the `process` function as the third and subsequent parameters. `process` (optional): This may be a function or a string. If it is a function, it is called as follows: `process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)` where: * TITLE is a title object describing the module's title; see [https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects]. * CATS is an list of categories that the module will be added to. * CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field. The return value of `process` should either be a string (which will be used as the module's documentation), or a table specifying the name of a template to expand to get the documentation, along with the arguments to that template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title` field, and any arguments should be in `args; in this case, the template name will be listed above the generated documentation as the source of the documentation, along with an edit button to edit the template's contents. If, however, the return value of the `process` function is a string, any template invocations will be expanded using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation. If `process` itself is a string rather than a function, it should name a submodule under [[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule will be specified as the source of the documentation (unless it returns a table naming a template to expand to get the documentation, as described above). If `process` is omitted entirely, the module will have no documentation. `cat` (optional): A string naming the category into which the module should be placed, or a list of such strings. Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second, etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS list (the second parameter). ]=] local module_regex = { { regex = "^Modul:languages/data/(3/%l/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(3/%l)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:scripts/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:data tables/data..?.?.?$", cat = "Jadual data berpecah modul rujukan", }, { regex = "^Modul:zh/data/dial%-pron/.+$", cat = "Modul data sebutan dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/dial%-syn/.+$", cat = "Modul data sinonim dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/glyph%-data/.+$", cat = "Modul data bentuk aksara Cina bersejarah", process = function(title, _cats) local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)") if character then return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.") :format(zh_link(character)) end end, }, { regex = "^Modul:zh/data/ltc%-pron/(.+)$", cat = "Modul data sebutan bahasa Cina Pertengahan|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-BS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1", process = "zh data", }, { -- capture rest of zh/data submodules regex = "^Modul:zh/data/(.+)$", cat = "Modul data bahasa Cina|%1", }, { regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$", process = "guoxue-data", }, { regex = "^Modul:Unicode data/(.+)$", cat = "Modul data Unicode|%1", process = make_Unicode_data_documentation, }, { regex = "^Modul:number list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor") if lang then return ("This module contains data on various types of numbers in %s.\n%s") :format(lang:makeCategoryLink(), number_list_show_table() or "") end end, }, { regex = "^Modul:chemical element list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia") if lang then return ("This module contains data on chemical elements in %s.\n%s") :format(lang:makeCategoryLink(), chemical_element_list_show_table() or "") end end, }, { regex = "^Modul:accel/(.+)$", process = function(title, cats) local lang_code = title.subpageText local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel") insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName())) return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.") :format(lang:makeCategoryLink()) end end, }, { regex = "^Modul:inc%-ash/dial/data/(.+)$", cat = "Modul Prakrit Ashoka|%1", process = function(title, _cats) local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$") if word then local lang = lang_cache["inc-ash"] return ("This module contains data on the pronunciation of %s in dialects of %s.") :format(full_link({ term = word, lang = lang }, "term"), lang:makeCategoryLink()) end end, }, { regex = "^.+%-translit$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "translit", title_without_namespace = title.text, } end, }, { regex = "^.+%-sortkey$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "sortkey", title_without_namespace = title.text, } end, }, { regex = "^.+%-stripdiacritics$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "strip diacritics", title_without_namespace = title.text, } end, }, { regex = "^Modul:form of/lang%-data/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa") if lang then -- FIXME, display more info. return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " .. langname .. "." end end }, { regex = "^Modul:labels/data/lang/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa") if lang then return { title = "label language-specific data documentation", args = { [1] = lang_code }, } end end }, { regex = "^Modul:category tree/lang/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang") if lang then return "This module handles generating the descriptions and categorization for " .. langname .. " category pages " .. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are " .. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. " .. "This module is part of the category tree system, which is a general framework for generating the " .. "descriptions and categorization of category pages.\n\n" .. "For more information, see [[Modul:category tree/lang/documentation]].\n\n" .. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the " .. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized." end end }, { regex = "^Modul:category tree/topic/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/topic| ") return { title = "topic cat data submodule documentation" } end }, { regex = "^Modul:category tree/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/grammar| ") return { title = "category tree data submodule documentation" } end }, { regex = "^Modul:ja/data/(.+)$", cat = "Modul data bahasa Jepun|%1", }, { regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$", cat = "Modul atlas data dialek Finland|%1", process = function(_title, _cats, shard) return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " .. "''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " .. "It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " .. "''Kotimaisten kielten keskus'' under the CC BY 4.0 license." end }, { regex = "^Modul:fi%-dialects/data/feature/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:fi%-dialects/data/word/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:Swadesh/data/([%l-]+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "." end end }, { regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$", process = function(_title, cats, lang_code, variety) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the " local etym_lang = get_lang(variety, nil, "allow etym") if etym_lang then return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname) end local script = get_script(variety) if script then return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang)) end return ("%s %s variety of %s."):format(prefix, variety, langname) end end }, { regex = "^Modul:typing%-aids", process = function(title, cats) local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$") local sortkey if data_suffix then if data_suffix:find "^[%l-]+$" then local lang = get_lang(data_suffix) if lang then sortkey = lang:getCanonicalName() insert(cats, "Modul data bahasa " .. sortkey) end elseif data_suffix:find "^%u%l%l%l$" then local script = get_script(data_suffix) if script then -- FIXME: no lang to pass here sortkey = script:getCanonicalName() insert(cats, script:getCategoryName()) end end insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix)) end end, }, { regex = "^Modul:R:([%l-]+):(.+)$", process = function(_title, cats, lang_code, refname) local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname) insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName())) return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}." end end, }, { regex = "^Modul:Quotations/([%l-]+)/?(.*)", process = "Quotation", }, { regex = "^Modul:affix/lang%-data/([%l-]+)", process = "affix lang-data", }, { regex = "^Modul:dialect synonyms/([%l-]+)$", process = function(_title, cats, lang_code) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "| ") return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " .. "{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" .. "==== Struktur modul data bahasa ====\n" .. "* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" .. "* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" .. "* <code>export.notes</code> — optional; table of note keys to text.\n" .. "* <code>export.sources</code> — optional; table of source keys to text.\n" .. "* <code>export.note_aliases</code> — optional; alias map for notes.\n" .. "* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" .. expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } }) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$", process = function(_title, cats, lang_code, term) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term } })) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$", process = function(_title, cats, lang_code, term, id) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term, id = id } })) end end, }, { regex = "^Modul:bibliography/data/([%l-]+)$", process = function(title, cats, lang_code) if lang_code == "preload" then return 'Used as a base model for other languages when the button "create new language submodule" is clicked.' end local page = require(title.fullText).bib_page if not page then page = lang_cache[lang_code]:getCanonicalName() if page then insert(cats, "Modul bahasa " .. page) end end insert(cats, "Modul rujukan") return "This module holds bibliographical data for " .. page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''." end, }, } function export.show(frame) local boolean_default_false = { type = "boolean", default = false } local args = process_params(frame.args, { ["hr"] = true, ["for"] = true, ["from"] = true, ["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage. ["notsubpage"] = boolean_default_false, ["nodoc"] = boolean_default_false, ["nolinks"] = boolean_default_false, -- suppress all "Useful links" ["nosandbox"] = boolean_default_false, -- supress sandbox }) local output = {'\n<div class="documentation" style="display:block; clear:both">\n'} local function ins(txt) insert(output, txt) end local cats = {} local function inscat(cat) insert(cats, cat) end local nodoc = args.nodoc if (not args.hr) or (args.hr == "above") then ins("----\n") end local title = args["for"] and new_title(args["for"]) or get_current_title() local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil local contentModel = title.contentModel local pagetype, is_script_or_stylesheet = get_pagetype(title) local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc local doc_content_source = "Modul:pendokumenan" local auto_generated_cat_source local cats_auto_generated = false if not args.allowondoc and is_documentation(title) then -- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type. error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].") elseif is_sandbox(title) then local sandbox_ns = title.nsText preload = ("Templat:pendokumenan/preload%s%sSandbox"):format( sandbox_ns == "Modul" and sandbox_ns or "Templat", title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or "" ) elseif pagetype:match("%f[%w]gadget%f[%W]") then preload = "Templat:pendokumenan/preloadGadget" elseif pagetype:match("%f[%w]script%f[%W]") then -- .js if title.nsText == "MediaWiki" then preload = "Templat:pendokumenan/preloadMediaWikiJavaScript" else preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end end is_script_or_stylesheet = true elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end is_script_or_stylesheet = true elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto. preload = "Templat:pendokumenan/preloadModule" elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then preload = "Templat:pendokumenan/preloadTemplate" end if doc_title and doc_title.isRedirect then old_doc_title = doc_title doc_title = doc_title.redirectTarget end ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">") local function get_module_doc_and_cats(categories_only) cats_auto_generated = true local automatic_cats = nil if user_name then fallback_docs = "pendokumenan/fallback/user module" automatic_cats = { "Modul kotak pasir pengguna" } else for _, data in ipairs(module_regex) do local captures = { umatch(title.fullText, data.regex) } if #captures > 0 then local cat, process_function if is_callable(data.process) then process_function = data.process elseif type(data.process) == "string" then doc_content_source = "Modul:pendokumenan/functions/" .. data.process process_function = require(doc_content_source) end if process_function then doc_content = process_function(title, cats, unpack(captures)) end if type(doc_content) == "table" then doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source doc_content = expand_template(doc_content) elseif doc_content ~= nil then doc_content = preprocess(doc_content) end cat = data.cat if cat then if type(cat) == "string" then cat = { cat } end for _, c in ipairs(cat) do insert(cats, (ugsub(title.fullText, data.regex, c))) end end break end end end if title.subpageText == "templates" then inscat("Modul antara muka templat") end if automatic_cats then for _, c in ipairs(automatic_cats) do inscat(c) end end if #cats == 0 then local auto_cats = categorize_module { return_raw = true, noerror = true, } if #auto_cats > 0 then auto_generated_cat_source = "Modul:module categorization" end for _, category in ipairs(auto_cats) do inscat(category) end end -- meaning module is not in user’s sandbox or one of many datamodule boring series needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs) end -- Override automatic documentation, if present. if doc_title and doc_title.exists then local cats_auto_generated_text = "" if contentModel == "Scribunto" then local doc_page_content = doc_title.content -- Track then do nothing if there are uses of includeonly. The -- pattern is slightly too permissive, but any false-positives are -- obvious typos that should be corrected. if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then track("module-includeonly") else -- Check for uses of {{module cat}}. find_templates treats the -- input as transcluded by default (i.e. it parses the wikitext -- which will be transcluded through to the module page). local module_cat for template in find_templates(doc_page_content) do if Templat:get_name() == "module cat" then module_cat = true break end end if not module_cat then get_module_doc_and_cats("categories only") auto_generated_cat_source = auto_generated_cat_source or doc_content_source cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" .. auto_generated_cat_source .. "]]. <sup>[[" .. new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>" end end end ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" .. doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" .. cats_auto_generated_text .. "</i></dd>") else if contentModel == "Scribunto" then get_module_doc_and_cats(false) elseif title.nsText == "Templat" then --inscat("Uncategorized templates") needs_doc = not (fallback_docs or nodoc) elseif user_name and is_script_or_stylesheet then skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")] if skin_name then fallback_docs = "pendokumenan/fallback/user " .. contentModel end end if doc_content then ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " .. "dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" .. new_title(doc_content_source):fullUrl { action = "edit" } .. " sunting]]</sup> </i></dd>") elseif not nodoc then if doc_title then ins( "<dd><i style=\"font-size: larger;\">Laman " .. pagetype .. " ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " .. (fallback_docs and "Anda boleh " or "Minta tolong ") .. "[" .. doc_title:fullUrl { action = "edit", preload = preload } .. " ciptakan laman pendokumenan tersebut].</i></dd>\n") else ins( "<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " .. "pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n") end end end if startswith(title.fullText, "MediaWiki:Gadget-") then local is_gadget = false for line in gline(new_title("MediaWiki:Gadgets-definition").content) do local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$") if not gadget then gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$") end if gadget then items = split(items, "|") for i, item in ipairs(items) do if title.fullText == ("MediaWiki:Gadget-" .. item) then is_gadget = true ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>") ins(gadget) ins("</code> gajet ([") ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(" sunting takrifan])'' <dl>") ins("<dd> ''Huraian ([") ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" }))) ins(" sunting])'': ") ins(preprocess(new_message('Gadget-' .. gadget):plain())) ins(" </dd>") table.remove(items, i) if #items > 0 then for j, item in ipairs(items) do items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]' end ins("<dd> ''Bahagian lain'': ") ins(list_to_text(items)) ins("</dd>") end ins("</dl></dd>") break end end end end if not is_gadget then ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [") ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" }))) ins(' gajet] ([') ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(' sunting takrifan]).</dd>') -- else -- inscat("Wiktionary gadgets") end end if old_doc_title then ins("<dd> ''Dilencong daripada'' [") ins(old_doc_title:fullUrl { redirect = "no" }) ins(" ") ins(old_doc_title.fullText) ins("] ([") ins(old_doc_title:fullUrl { action = "edit" }) ins(" sunting]).</dd>\n") end if not args.nolinks then local links = {} local function inslinks(txt) insert(links, txt) end if title.isSubpage and not args.notsubpage then inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]") inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]") else inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]") end inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) .. " pautan]") if contentModel ~= "Scribunto" then inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) .. " lencongan]") end if is_script_or_stylesheet then if user_name then inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]") end else inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) .. " transklusi]") end if contentModel == "Scribunto" then local is_testcases = title.isSubpage and title.subpageText == "testcases" local without_subpage = title.nsText .. ":" .. title.baseText if is_testcases then inslinks("[[:" .. without_subpage .. "|modul ujian]]") else inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]") end if user_name then inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]") inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]") inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]") -- If sandbox module, add a link to the module that this is a sandbox of. -- Exclude user sandbox modules like [[User:Dine2016/sandbox]]. elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then inscat("Modul kotak pasir") -- Sandbox modules don’t really need documentation. needs_doc = false -- Don't track user sandbox modules. local text_title = new_title(title.text) if not (text_title and text_title.nsText == "Pengguna") then local diff local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]") if sandbox_of then track("sandbox to be moved") else sandbox_of = title.text:match("^sandbox%d*/(.*)$") end if not sandbox_of then error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'") :format(title.text)) end sandbox_of = title.nsText .. ":" .. sandbox_of if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) end -- If not a sandbox module, add link to sandbox module. -- Sometimes there are multiple sandboxes for a single Modul: -- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]]. else local sandbox_title local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$") if not user_prefix then user_prefix = "" user_rest = title.text end sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks(sandbox_link .. (diff or "")) end end if title.nsText == "Templat" then -- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]] -- apparently same with/without: &profile=advanced&fulltext=1 local errorq = 'searchengineselect=mediawiki&search=all: hasTemplat:\"' .. title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:' local eincategory = "Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors" inslinks( '[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]' .. ' (' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]' .. '/' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]' .. ')' ) if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template. -- At the moment there are no user sandbox templates with subpage -- “/sandbox”. inscat("Templat kotak pasir") -- Sandbox templates don’t really need documentation. needs_doc = false -- Will behave badly if “/sandbox” occurs twice in title! local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "") local diff if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) -- This is a template that can have a sandbox. elseif not args.nosandbox then -- unless we tell it not to local sandbox_title = title.fullText .. "/sandbox" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or "")) end end if #links > 0 then ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>") end end ins("</dl>\n") -- Show error from [[Modul:category tree/topic cat/data]] on its submodules' -- documentation to, for instance, warn about duplicate labels. if startswith(title.fullText, "Modul:category tree/topic/") then local ok, err = pcall(require, "Modul:category tree/topic/data") if not ok then ins('<span class="error">' .. err .. '</span>\n\n') end end if doc_title and doc_title.exists then -- Override automatic documentation, if present. doc_content = expand_template { title = doc_title.fullText } elseif not doc_content and fallback_docs then doc_content = expand_template { title = fallback_docs, args = { ['user'] = user_name, ['page'] = title.fullText, ['skin name'] = skin_name, }, } end if doc_content then ins(doc_content) end ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br")) if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then if contentModel == "Scribunto" then inscat("Modul belum dikategorikan") -- elseif title.nsText == "Templat" then -- inscat("Templat belum dikategorikan") end end if needs_doc then inscat("Templat dan modul yang memerlukan pendokumenan") end for _, cat in ipairs(cats) do ins("[[Kategori:" .. cat .. "]]") end ins("</div>\n") return concat(output) end function export.module_auto_doc_table() local parts = {} local function ins(text) insert(parts, text) end ins('{|class="wikitable"') ins("! Regex !! Kategori !! Modul yang dikendalikan") for _, spec in ipairs(module_regex) do local cat_text local cats = spec.cat if cats then local cat_parts = {} if type(cats) == "string" then cats = { cats } end for _, cat in ipairs(cats) do insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "&#124;")))) end cat_text = concat(cat_parts, ", ") else cat_text = "''(tidak dinyatakan secara khusus)''" end ins("|-") ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text, is_callable(spec.process) and "''(dikendali secara dalaman)''" or type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or "''(tiada penjana pendokumenan)''")) end ins("|}") return concat(parts, "\n") end return export a6872eb4t58l8znugkpylbbqxanopxl 375366 375362 2026-09-22T04:42:04Z Hakimi97 2668 Baiki ralat 375366 Scribunto text/plain local export = {} local debug_track_module = "Modul:debug/track" local frame_module = "Modul:frame" local fun_is_callable_module = "Modul:fun/isCallable" local languages_module = "Modul:languages" local links_module = "Modul:links" local load_module = "Modul:load" local module_categorization_module = "Modul:module categorization" local number_list_show_module = "Modul:number list/show" local chemical_element_list_show_module = "Modul:chemical element list/show" local pages_module = "Modul:pages" local parameters_module = "Modul:parameters" local scripts_module = "Modul:scripts" local string_endswith_module = "Modul:string/endswith" local string_gline_module = "Modul:string/gline" local string_startswith_module = "Modul:string/startswith" local string_utilities_module = "Modul:string utilities" local template_parser_module = "Modul:template parser" local title_exists_module = "Modul:title/exists" local title_new_title_module = "Modul:title/newTitle" local concat = table.concat local error = error local full_url = mw.uri.fullUrl local get_current_title = mw.title.getCurrentTitle local insert = table.insert local ipairs = ipairs local list_to_text = mw.text.listToText local new_message = mw.message.new local pcall = pcall local require = require local tonumber = tonumber local tostring = tostring local type = type local unpack = unpack or table.unpack -- Lua 5.2 compatibility local function categorize_module(...) categorize_module = require(module_categorization_module).categorize_module return categorize_module(...) end local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function endswith(...) endswith = require(string_endswith_module) return endswith(...) end local function expand_template(...) expand_template = require(frame_module).expandTemplate return expand_template(...) end local function find_templates(...) find_templates = require(template_parser_module).find_templates return find_templates(...) end local function full_link(...) full_link = require(links_module).full_link return full_link(...) end local function get_lang(...) get_lang = require(languages_module).getByCode return get_lang(...) end local function get_pagetype(...) get_pagetype = require(pages_module).get_pagetype return get_pagetype(...) end local function get_script(...) get_script = require(scripts_module).getByCode return get_script(...) end local function gline(...) gline = require(string_gline_module) return gline(...) end local function is_callable(...) is_callable = require(fun_is_callable_module) return is_callable(...) end local function is_documentation(...) is_documentation = require(pages_module).is_documentation return is_documentation(...) end local function is_sandbox(...) is_sandbox = require(pages_module).is_sandbox return is_sandbox(...) end local function new_title(...) new_title = require(title_new_title_module) return new_title(...) end local function number_list_show_table(...) number_list_show_table = require(number_list_show_module).table return number_list_show_table(...) end local function chemical_element_list_show_table(...) chemical_element_list_show_table = require(chemical_element_list_show_module).table return chemical_element_list_show_table(...) end local function preprocess(...) preprocess = require(frame_module).preprocess return preprocess(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function safe_load_data(...) safe_load_data = require(load_module).safe_load_data return safe_load_data(...) end local function split(...) split = require(string_utilities_module).split return split(...) end local function startswith(...) startswith = require(string_startswith_module) return startswith(...) end local function title_exists(...) title_exists = require(title_exists_module) return title_exists(...) end local function ugsub(...) ugsub = require(string_utilities_module).gsub return ugsub(...) end local function umatch(...) umatch = require(string_utilities_module).match return umatch(...) end local skins = { ["common"] = "", ["vector"] = "Vector", ["monobook"] = "Monobook", ["cologneblue"] = "Cologne Blue", ["modern"] = "Modern", } local function track(page) debug_track("pendokumenan/" .. page) return true end local function compare_pages(page1, page2, text) return "[" .. tostring( full_url("Khas:ComparePages", { page1 = page1, page2 = page2 })) .. " " .. text .. "]" end -- Avoid transcluding [[Modul:languages/cache]] everywhere. local lang_cache = setmetatable({}, { __index = function(self, k) return require("Modul:languages/cache")[k] end }) local function zh_link(word) return full_link { lang = lang_cache.zh, term = word } end local function make_languages_data_documentation(_title, cats, division) local doc_template, module_cat if endswith(division, "/extra") then division = division:sub(1, -7) doc_template = "language extradata documentation" module_cat = "Modul data ekstra bahasa" else doc_template = "language data documentation" module_cat = "Modul data bahasa" end local sort_key if division == "exceptional" then sort_key = "x" else sort_key = division:gsub("/", "") end insert(cats, module_cat .. "|" .. sort_key) return { title = doc_template } end local function make_Unicode_data_documentation(title, _cats) local subpage, first_three_of_code_point = title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$") if subpage == "names" or subpage == "images" or subpage == "emoji images" then local low, high = tonumber(first_three_of_code_point .. "000", 16), tonumber(first_three_of_code_point .. "FFF", 16) local text, text_type if subpage == "names" then text_type = "titles of images" elseif subpage == "images" then text_type = "titles of images" elseif subpage == "emoji images" then text_type = "emoji-style images" end text = string.format( "Modul data ini mengandungi " .. text_type .. " kepada " .. "titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.", low, high) if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then text = text .. " Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" .. first_three_of_code_point .. "]]." elseif subpage == "emoji images" then text = text .. " Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]." end return text end end local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat) local lang = lang_cache[langcode] if lang then local langname if lang._fullCode then langname = lang_cache[lang._fullCode]:getCanonicalName() else langname = lang:getCanonicalName() end insert(cats, overall_data_module_cat .. "|" .. langname) insert(cats, "Modul bahasa " .. langname) insert(cats, "Modul data bahasa " .. langname) return lang, langname end end --[=[ This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't unnecessarily cluttered. It is a list of tables, each of which have the following possible fields: `regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used. Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the `process` function as the third and subsequent parameters. `process` (optional): This may be a function or a string. If it is a function, it is called as follows: `process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)` where: * TITLE is a title object describing the module's title; see [https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects]. * CATS is an list of categories that the module will be added to. * CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field. The return value of `process` should either be a string (which will be used as the module's documentation), or a table specifying the name of a template to expand to get the documentation, along with the arguments to that template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title` field, and any arguments should be in `args; in this case, the template name will be listed above the generated documentation as the source of the documentation, along with an edit button to edit the template's contents. If, however, the return value of the `process` function is a string, any template invocations will be expanded using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation. If `process` itself is a string rather than a function, it should name a submodule under [[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule will be specified as the source of the documentation (unless it returns a table naming a template to expand to get the documentation, as described above). If `process` is omitted entirely, the module will have no documentation. `cat` (optional): A string naming the category into which the module should be placed, or a list of such strings. Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second, etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS list (the second parameter). ]=] local module_regex = { { regex = "^Modul:languages/data/(3/%l/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(3/%l)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(2)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional/extra)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/data/(exceptional)$", process = make_languages_data_documentation, }, { regex = "^Modul:languages/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:scripts/.+$", cat = "Modul bahasa dan tulisan", }, { regex = "^Modul:data tables/data..?.?.?$", cat = "Jadual data berpecah modul rujukan", }, { regex = "^Modul:zh/data/dial%-pron/.+$", cat = "Modul data sebutan dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/dial%-syn/.+$", cat = "Modul data sinonim dialek bahasa Cina", process = "zh dial or syn", }, { regex = "^Modul:zh/data/glyph%-data/.+$", cat = "Modul data bentuk aksara Cina bersejarah", process = function(title, _cats) local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)") if character then return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.") :format(zh_link(character)) end end, }, { regex = "^Modul:zh/data/ltc%-pron/(.+)$", cat = "Modul data sebutan bahasa Cina Pertengahan|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-BS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1", process = "zh data", }, { regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$", cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1", process = "zh data", }, { -- capture rest of zh/data submodules regex = "^Modul:zh/data/(.+)$", cat = "Modul data bahasa Cina|%1", }, { regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$", process = "guoxue-data", }, { regex = "^Modul:Unicode data/(.+)$", cat = "Modul data Unicode|%1", process = make_Unicode_data_documentation, }, { regex = "^Modul:number list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor") if lang then return ("This module contains data on various types of numbers in %s.\n%s") :format(lang:makeCategoryLink(), number_list_show_table() or "") end end, }, { regex = "^Modul:chemical element list/data/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia") if lang then return ("This module contains data on chemical elements in %s.\n%s") :format(lang:makeCategoryLink(), chemical_element_list_show_table() or "") end end, }, { regex = "^Modul:accel/(.+)$", process = function(title, cats) local lang_code = title.subpageText local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel") insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName())) return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.") :format(lang:makeCategoryLink()) end end, }, { regex = "^Modul:inc%-ash/dial/data/(.+)$", cat = "Modul Prakrit Ashoka|%1", process = function(title, _cats) local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$") if word then local lang = lang_cache["inc-ash"] return ("This module contains data on the pronunciation of %s in dialects of %s.") :format(full_link({ term = word, lang = lang }, "term"), lang:makeCategoryLink()) end end, }, { regex = "^.+%-translit$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "translit", title_without_namespace = title.text, } end, }, { regex = "^.+%-sortkey$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "sortkey", title_without_namespace = title.text, } end, }, { regex = "^.+%-stripdiacritics$", process = function(title, _cats) return require("Modul:pendokumenan/translit-like").documentation { operation = "strip diacritics", title_without_namespace = title.text, } end, }, { regex = "^Modul:form of/lang%-data/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa") if lang then -- FIXME, display more info. return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " .. langname .. "." end end }, { regex = "^Modul:labels/data/lang/(.+)$", process = function(_title, cats, lang_code) local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa") if lang then return { title = "label language-specific data documentation", args = { [1] = lang_code }, } end end }, { regex = "^Modul:category tree/lang/(.+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang") if lang then return "This module handles generating the descriptions and categorization for " .. langname .. " category pages " .. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are " .. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. " .. "This module is part of the category tree system, which is a general framework for generating the " .. "descriptions and categorization of category pages.\n\n" .. "For more information, see [[Modul:category tree/lang/documentation]].\n\n" .. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the " .. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized." end end }, { regex = "^Modul:category tree/topic/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/topic| ") return { title = "topic cat data submodule documentation" } end }, { regex = "^Modul:category tree/(.+)$", process = function(_title, cats, _submodule) insert(cats, "Modul data category tree/grammar| ") return { title = "category tree data submodule documentation" } end }, { regex = "^Modul:ja/data/(.+)$", cat = "Modul data bahasa Jepun|%1", }, { regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$", cat = "Modul atlas data dialek Finland|%1", process = function(_title, _cats, shard) return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " .. "''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " .. "It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " .. "''Kotimaisten kielten keskus'' under the CC BY 4.0 license." end }, { regex = "^Modul:fi%-dialects/data/feature/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:fi%-dialects/data/word/(.+)", cat = "Modul data dialek Finland|%1", }, { regex = "^Modul:Swadesh/data/([%l-]+)$", process = function(_title, cats, lang_code) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "." end end }, { regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$", process = function(_title, cats, lang_code, variety) local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh") if lang then local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the " local etym_lang = get_lang(variety, nil, "allow etym") if etym_lang then return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname) end local script = get_script(variety) if script then return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang)) end return ("%s %s variety of %s."):format(prefix, variety, langname) end end }, { regex = "^Modul:typing%-aids", process = function(title, cats) local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$") local sortkey if data_suffix then if data_suffix:find "^[%l-]+$" then local lang = get_lang(data_suffix) if lang then sortkey = lang:getCanonicalName() insert(cats, "Modul data bahasa " .. sortkey) end elseif data_suffix:find "^%u%l%l%l$" then local script = get_script(data_suffix) if script then -- FIXME: no lang to pass here sortkey = script:getCanonicalName() insert(cats, script:getCategoryName()) end end insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix)) end end, }, { regex = "^Modul:R:([%l-]+):(.+)$", process = function(_title, cats, lang_code, refname) local lang = lang_cache[lang_code] if lang then insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname) insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName())) return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}." end end, }, { regex = "^Modul:Quotations/([%l-]+)/?(.*)", process = "Quotation", }, { regex = "^Modul:affix/lang%-data/([%l-]+)", process = "affix lang-data", }, { regex = "^Modul:dialect synonyms/([%l-]+)$", process = function(_title, cats, lang_code) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "| ") return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " .. "{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" .. "==== Struktur modul data bahasa ====\n" .. "* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" .. "* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" .. "* <code>export.notes</code> — optional; table of note keys to text.\n" .. "* <code>export.sources</code> — optional; table of source keys to text.\n" .. "* <code>export.note_aliases</code> — optional; alias map for notes.\n" .. "* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" .. expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } }) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$", process = function(_title, cats, lang_code, term) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term } })) end end, }, { regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$", process = function(_title, cats, lang_code, term, id) local lang = lang_cache[lang_code] if lang then local langname = "bahasa " .. lang:getCanonicalName() insert(cats, "Modul data sinonim dialek|" .. langname) insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term) return ("%s\n\n%s"):format( "==== Term/sense module structure ====\n" .. "* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" .. "* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" .. "* <code>export.gloss</code> — optional; short meaning for the table.\n" .. "* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" .. "* <code>export.notes</code> — optional; list of note keys.\n" .. "* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" .. "* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" .. "* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" .. "Example (custom title and data column, IPA realizations):\n" .. "<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" .. "export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" .. "export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n", expand_template({ title = 'dial syn', args = { lang_code, term, id = id } })) end end, }, { regex = "^Modul:bibliography/data/([%l-]+)$", process = function(title, cats, lang_code) if lang_code == "preload" then return 'Used as a base model for other languages when the button "create new language submodule" is clicked.' end local page = require(title.fullText).bib_page if not page then page = lang_cache[lang_code]:getCanonicalName() if page then insert(cats, "Modul bahasa " .. page) end end insert(cats, "Modul rujukan") return "This module holds bibliographical data for " .. page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''." end, }, } function export.show(frame) local boolean_default_false = { type = "boolean", default = false } local args = process_params(frame.args, { ["hr"] = true, ["for"] = true, ["from"] = true, ["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage. ["notsubpage"] = boolean_default_false, ["nodoc"] = boolean_default_false, ["nolinks"] = boolean_default_false, -- suppress all "Useful links" ["nosandbox"] = boolean_default_false, -- supress sandbox }) local output = {'\n<div class="documentation" style="display:block; clear:both">\n'} local function ins(txt) insert(output, txt) end local cats = {} local function inscat(cat) insert(cats, cat) end local nodoc = args.nodoc if (not args.hr) or (args.hr == "above") then ins("----\n") end local title = args["for"] and new_title(args["for"]) or get_current_title() local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil local contentModel = title.contentModel local pagetype, is_script_or_stylesheet = get_pagetype(title) local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc local doc_content_source = "Modul:pendokumenan" local auto_generated_cat_source local cats_auto_generated = false if not args.allowondoc and is_documentation(title) then -- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type. error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].") elseif is_sandbox(title) then local sandbox_ns = title.nsText preload = ("Templat:pendokumenan/preload%s%sSandbox"):format( sandbox_ns == "Modul" and sandbox_ns or "Templat", title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or "" ) elseif pagetype:match("%f[%w]gadget%f[%W]") then preload = "Templat:pendokumenan/preloadGadget" elseif pagetype:match("%f[%w]script%f[%W]") then -- .js if title.nsText == "MediaWiki" then preload = "Templat:pendokumenan/preloadMediaWikiJavaScript" else preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end end is_script_or_stylesheet = true elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css preload = "Templat:pendokumenan/preloadTemplate" -- XXX if title.nsText == "Pengguna" then user_name = title.rootText end is_script_or_stylesheet = true elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto. preload = "Templat:pendokumenan/preloadModule" elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then preload = "Templat:pendokumenan/preloadTemplate" end if doc_title and doc_title.isRedirect then old_doc_title = doc_title doc_title = doc_title.redirectTarget end ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">") local function get_module_doc_and_cats(categories_only) cats_auto_generated = true local automatic_cats = nil if user_name then fallback_docs = "pendokumenan/fallback/user module" automatic_cats = { "Modul kotak pasir pengguna" } else for _, data in ipairs(module_regex) do local captures = { umatch(title.fullText, data.regex) } if #captures > 0 then local cat, process_function if is_callable(data.process) then process_function = data.process elseif type(data.process) == "string" then doc_content_source = "Modul:pendokumenan/functions/" .. data.process process_function = require(doc_content_source) end if process_function then doc_content = process_function(title, cats, unpack(captures)) end if type(doc_content) == "table" then doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source doc_content = expand_template(doc_content) elseif doc_content ~= nil then doc_content = preprocess(doc_content) end cat = data.cat if cat then if type(cat) == "string" then cat = { cat } end for _, c in ipairs(cat) do insert(cats, (ugsub(title.fullText, data.regex, c))) end end break end end end if title.subpageText == "templates" then inscat("Modul antara muka templat") end if automatic_cats then for _, c in ipairs(automatic_cats) do inscat(c) end end if #cats == 0 then local auto_cats = categorize_module { return_raw = true, noerror = true, } if #auto_cats > 0 then auto_generated_cat_source = "Modul:module categorization" end for _, category in ipairs(auto_cats) do inscat(category) end end -- meaning module is not in user’s sandbox or one of many datamodule boring series needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs) end -- Override automatic documentation, if present. if doc_title and doc_title.exists then local cats_auto_generated_text = "" if contentModel == "Scribunto" then local doc_page_content = doc_title.content -- Track then do nothing if there are uses of includeonly. The -- pattern is slightly too permissive, but any false-positives are -- obvious typos that should be corrected. if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then track("module-includeonly") else -- Check for uses of {{module cat}}. find_templates treats the -- input as transcluded by default (i.e. it parses the wikitext -- which will be transcluded through to the module page). local module_cat for template in find_templates(doc_page_content) do if template:get_name() == "module cat" then module_cat = true break end end if not module_cat then get_module_doc_and_cats("categories only") auto_generated_cat_source = auto_generated_cat_source or doc_content_source cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" .. auto_generated_cat_source .. "]]. <sup>[[" .. new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>" end end end ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" .. doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" .. cats_auto_generated_text .. "</i></dd>") else if contentModel == "Scribunto" then get_module_doc_and_cats(false) elseif title.nsText == "Templat" then --inscat("Uncategorized templates") needs_doc = not (fallback_docs or nodoc) elseif user_name and is_script_or_stylesheet then skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")] if skin_name then fallback_docs = "pendokumenan/fallback/user " .. contentModel end end if doc_content then ins( "<dd><i style=\"font-size: larger;\">Berikut merupakan " .. "[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " .. "dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" .. new_title(doc_content_source):fullUrl { action = "edit" } .. " sunting]]</sup> </i></dd>") elseif not nodoc then if doc_title then ins( "<dd><i style=\"font-size: larger;\">Laman " .. pagetype .. " ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " .. (fallback_docs and "Anda boleh " or "Minta tolong ") .. "[" .. doc_title:fullUrl { action = "edit", preload = preload } .. " ciptakan laman pendokumenan tersebut].</i></dd>\n") else ins( "<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " .. "pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n") end end end if startswith(title.fullText, "MediaWiki:Gadget-") then local is_gadget = false for line in gline(new_title("MediaWiki:Gadgets-definition").content) do local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$") if not gadget then gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$") end if gadget then items = split(items, "|") for i, item in ipairs(items) do if title.fullText == ("MediaWiki:Gadget-" .. item) then is_gadget = true ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>") ins(gadget) ins("</code> gajet ([") ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(" sunting takrifan])'' <dl>") ins("<dd> ''Huraian ([") ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" }))) ins(" sunting])'': ") ins(preprocess(new_message('Gadget-' .. gadget):plain())) ins(" </dd>") table.remove(items, i) if #items > 0 then for j, item in ipairs(items) do items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]' end ins("<dd> ''Bahagian lain'': ") ins(list_to_text(items)) ins("</dd>") end ins("</dl></dd>") break end end end end if not is_gadget then ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [") ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" }))) ins(' gajet] ([') ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" }))) ins(' sunting takrifan]).</dd>') -- else -- inscat("Wiktionary gadgets") end end if old_doc_title then ins("<dd> ''Dilencong daripada'' [") ins(old_doc_title:fullUrl { redirect = "no" }) ins(" ") ins(old_doc_title.fullText) ins("] ([") ins(old_doc_title:fullUrl { action = "edit" }) ins(" sunting]).</dd>\n") end if not args.nolinks then local links = {} local function inslinks(txt) insert(links, txt) end if title.isSubpage and not args.notsubpage then inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]") inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]") else inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]") end inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) .. " pautan]") if contentModel ~= "Scribunto" then inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) .. " lencongan]") end if is_script_or_stylesheet then if user_name then inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]") end else inslinks( "[" .. tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) .. " transklusi]") end if contentModel == "Scribunto" then local is_testcases = title.isSubpage and title.subpageText == "testcases" local without_subpage = title.nsText .. ":" .. title.baseText if is_testcases then inslinks("[[:" .. without_subpage .. "|modul ujian]]") else inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]") end if user_name then inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]") inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]") inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]") -- If sandbox module, add a link to the module that this is a sandbox of. -- Exclude user sandbox modules like [[User:Dine2016/sandbox]]. elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then inscat("Modul kotak pasir") -- Sandbox modules don’t really need documentation. needs_doc = false -- Don't track user sandbox modules. local text_title = new_title(title.text) if not (text_title and text_title.nsText == "Pengguna") then local diff local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]") if sandbox_of then track("sandbox to be moved") else sandbox_of = title.text:match("^sandbox%d*/(.*)$") end if not sandbox_of then error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'") :format(title.text)) end sandbox_of = title.nsText .. ":" .. sandbox_of if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) end -- If not a sandbox module, add link to sandbox module. -- Sometimes there are multiple sandboxes for a single Modul: -- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]]. else local sandbox_title local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$") if not user_prefix then user_prefix = "" user_rest = title.text end sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks(sandbox_link .. (diff or "")) end end if title.nsText == "Templat" then -- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]] -- apparently same with/without: &profile=advanced&fulltext=1 local errorq = 'searchengineselect=mediawiki&search=all: hastemplate:\"' .. title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:' local eincategory = "Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors" inslinks( '[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]' .. ' (' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]' .. '/' .. '[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]' .. ')' ) if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template. -- At the moment there are no user sandbox templates with subpage -- “/sandbox”. inscat("Templat kotak pasir") -- Sandbox templates don’t really need documentation. needs_doc = false -- Will behave badly if “/sandbox” occurs twice in title! local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "") local diff if title_exists(sandbox_of) then diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")" else track("no sandbox of") end inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or "")) -- This is a template that can have a sandbox. elseif not args.nosandbox then -- unless we tell it not to local sandbox_title = title.fullText .. "/sandbox" local diff if title_exists(sandbox_title) then diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")" end inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or "")) end end if #links > 0 then ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>") end end ins("</dl>\n") -- Show error from [[Modul:category tree/topic cat/data]] on its submodules' -- documentation to, for instance, warn about duplicate labels. if startswith(title.fullText, "Modul:category tree/topic/") then local ok, err = pcall(require, "Modul:category tree/topic/data") if not ok then ins('<span class="error">' .. err .. '</span>\n\n') end end if doc_title and doc_title.exists then -- Override automatic documentation, if present. doc_content = expand_template { title = doc_title.fullText } elseif not doc_content and fallback_docs then doc_content = expand_template { title = fallback_docs, args = { ['user'] = user_name, ['page'] = title.fullText, ['skin name'] = skin_name, }, } end if doc_content then ins(doc_content) end ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br")) if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then if contentModel == "Scribunto" then inscat("Modul belum dikategorikan") -- elseif title.nsText == "Templat" then -- inscat("Templat belum dikategorikan") end end if needs_doc then inscat("Templat dan modul yang memerlukan pendokumenan") end for _, cat in ipairs(cats) do ins("[[Kategori:" .. cat .. "]]") end ins("</div>\n") return concat(output) end function export.module_auto_doc_table() local parts = {} local function ins(text) insert(parts, text) end ins('{|class="wikitable"') ins("! Regex !! Kategori !! Modul yang dikendalikan") for _, spec in ipairs(module_regex) do local cat_text local cats = spec.cat if cats then local cat_parts = {} if type(cats) == "string" then cats = { cats } end for _, cat in ipairs(cats) do insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "&#124;")))) end cat_text = concat(cat_parts, ", ") else cat_text = "''(tidak dinyatakan secara khusus)''" end ins("|-") ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text, is_callable(spec.process) and "''(dikendali secara dalaman)''" or type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or "''(tiada penjana pendokumenan)''")) end ins("|}") return concat(parts, "\n") end return export iqaaks3ujjv11d34e9oefx3iq8b1c5i Modul:names 828 27469 375348 244788 2026-09-22T03:12:17Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708781|92708781]]) 375348 Scribunto text/plain local export = {} local m_languages = require("Module:languages") local m_links = require("Module:links") local m_utilities = require("Module:utilities") local m_str_utils = require("Module:string utilities") local m_table = require("Module:table") local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local parameter_utilities_module = "Module:parameter utilities" local parse_interface_module = "Module:parse interface" local parse_utilities_module = "Module:parse utilities" local enlang = m_languages.getByCode("ms") local rsubn = m_str_utils.gsub local rsplit = m_str_utils.split local u = m_str_utils.char local function rsub(str, from, to) return (rsubn(str, from, to)) end local TEMP_LESS_THAN = u(0xFFF2) local force_cat = false -- for testing --[=[ FIXME: 1. from=the Bible (DONE) 2. origin=18th century [DONE] 3. popular= (DONE) 4. varoftype= (DONE) 5. eqtype= [DONE] 6. dimoftype= [DONE] 7. from=de:Elisabeth (same language) (DONE) 8. blendof=, blendof2= [DONE] 9. varform, dimform [DONE] 10. from=English < Latin [DONE] 11. usage=rare -> categorize as rare? 12. dimeq= (also vareq=?) [DONE] 13. fromtype= [DONE] 14. <tr:...> and similar params [DONE] ]=] -- Used in category code; name types which are full-word end-matching substrings of longer name types (e.g. "surnames" -- of "male surnames", but not "male surnames" of "female surnames" because "male" only matches a part of the word -- "female") should follow the longer name. export.personal_name_types = { "nama keluarga lelaki", "nama keluarga perempuan", "nama keluarga bebas jantina", "nama keluarga", "patronimik", "matronimik", } export.personal_name_type_set = m_table.listToSet(export.personal_name_types) export.given_name_genders = { lelaki = {type = "human"}, perempuan = {type = "human"}, uniseks = {type = "human", cat = {"nama diri lelaki", "nama diri perempuan", "nama diri uniseks"}, article = ""}, ["jantina tidak diketahui"] = {type = "human", cat = {}, track = true}, haiwan = {type = "animal", track = true}, kucing = {type = "animal"}, lembu = {type = "animal"}, anjing = {type = "animal"}, kuda = {type = "animal"}, khinzir = {type = "animal"}, } local function get_given_name_cats(gender, props) local cats = props.cat if not cats then if props.type == "animal" then cats = {"nama " .. gender} else cats = {"nama diri " .. gender} end end return cats end do local function do_cat(cat) if not export.personal_name_type_set[cat] then export.personal_name_type_set[cat] = true table.insert(export.personal_name_types, cat) end end for gender, props in pairs(export.given_name_genders) do local cats = get_given_name_cats(gender, props) for _, cat in ipairs(cats) do do_cat("bentuk singkat " .. cat) do_cat("bentuk agam " .. cat) do_cat(cat) end end do_cat("nama diri") end local translit_name_type_list = { "nama keluarga", "nama diri lelaki", "nama diri perempuan", "nama diri uniseks", "patronimik" } local function track(page) require("Module:debug").track("names/" .. page) end -- Get raw text, for use in computing the indefinite article. Use get_plaintext() in [[Module:utilities]] and also -- remove parens that may surround qualifier or label text preceding a term. local function get_rawtext(text) text = m_utilities.get_plaintext(text) text = text:gsub("[()%[%]]", "") return text end --[=[ Parse a term and associated properties. This works with parameters of the form 'Karlheinz' or 'Kunigunde<q:medieval, now rare>' or 'non:Óláfr' or 'ru:Фру́нзе<tr:Frúnzɛ><q:rare>' where the modifying properties are contained in <...> specifications after the term. `term` is the full parameter value including any angle brackets and colons; `paramname` is the name of the parameter that this value comes from, for error purposes; `deflang` is a language object used in the return value when the language isn't specified (e.g. in the examples 'Karlheinz' and 'Kunigunde<q:medieval, now rare>' above); `allow_explicit_lang` indicates whether the language can be explicitly given (e.g. in the examples 'non:Óláfr' or 'ru:Фру́нзе<tr:Frúnzɛ><q:rare>' above). Normally the return value is a terminfo object that can be passed to full_link() in [[Module:links]]), additionally with optional fields `.q`, `.qq`, `.l`, `.ll`, `.refs` and `.eq` (a list of objects of the same form as the returned terminfo object. However, if `allow_multiple_terms` is given, multiple comma-separated names can be given in `term`, and the return value is a list of objects of the form described just above. ]=] local function parse_term_with_annotations(term, paramname, deflang, allow_explicit_lang, allow_multiple_terms) local param_mods = require(parameter_utilities_module).construct_param_mods { {group = {"link", "l", "q", "ref"}}, {param = "eq", convert = function(eqval, parse_err) return parse_term_with_annotations(eqval, paramname .. ".eq", enlang, false, "allow multiple terms") end}, } local function generate_obj(term, parse_err) local termlang if allow_explicit_lang then local actual_term actual_term, termlang = require(parse_interface_module).parse_term_with_lang { term = term, parse_err = parse_err, paramname = paramname, } term = actual_term or term end return { term = term, lang = termlang or deflang, } end return require(parse_interface_module).parse_inline_modifiers(term, { param_mods = param_mods, paramname = paramname, generate_obj = generate_obj, splitchar = allow_multiple_terms and "," or nil, }) end --[=[ Link a single term. If `do_language_link` is given and a given term's language is English, the link will be constructed using language_link() in [[Module:links]]; otherwise, with full_link(). `termobj` is an object as returned by parse_term_with_annotations(), i.e. it is suitable for passing to [[Module:links]] and additionally contains optional fields `.q`, `.qq`, `.l`, `.ll`, `.refs` and `.eq` (a list of objects of the same form as `termobj`). ]=] local function link_one_term(termobj, do_language_link) local link if do_language_link and termobj.lang:getCode() == "ms" then link = m_links.language_link(termobj) else link = m_links.full_link(termobj) end if termobj.q and termobj.q[1] or termobj.qq and termobj.qq[1] or termobj.l and termobj.l[1] or termobj.ll and termobj.ll[1] or termobj.refs and termobj.refs[1] then link = require(decorations_module).format_decorations { lang = termobj.lang, text = link, q = termobj.q, qq = termobj.qq, l = termobj.l, ll = termobj.ll, refs = termobj.refs, } end if termobj.eq then local eqtext = {} for _, eqobj in ipairs(termobj.eq) do table.insert(eqtext, link_one_term(eqobj, true)) end link = link .. " [=" .. m_table.serialCommaJoin(eqtext, {conj = "atau"}) .. "]" end return link end --[=[ Link the terms in `terms`, and join them using the conjunction in `conj` (defaulting to "or"). Joining is done using serialCommaJoin() in [[Module:table]], so that e.g. two terms are joined as "TERM or TERM" while three terms are joined as "TERM, TERM or TERM" with special CSS spans before the final "or" to allow an "Oxford comma" to appear if configured appropriately. (However, if `conj` is the special value ", ", joining is done directly using that value.) If `include_langname` is given, the language of the first term will be prepended to the joined terms. If `do_language_link` is given and a given term's language is English, the link will be constructed using language_link() in [[Module:links]]; otherwise, with full_link(). Each term in `terms` is an object as returned by parse_term_with_annotations(). ]=] local function join_terms(terms, include_langname, do_language_link, conj) local links = {} local langnametext for _, termobj in ipairs(terms) do if include_langname and not langnametext then langnametext = termobj.lang:getCanonicalName() .. " " end table.insert(links, link_one_term(termobj, do_language_link)) end local joined_terms if conj == ", " then joined_terms = table.concat(links, conj) else joined_terms = m_table.serialCommaJoin(links, {conj = conj or "atau"}) end return (langnametext or "") .. joined_terms end --[=[ Gather the parameters for multiple names and link each name using full_link() (for foreign names) or language_link() (for English names), joining the names using serialCommaJoin() in [[Module:table]] with the conjunction `conj` (defaulting to "or"). (However, if `conj` is the special value ", ", joining is done directly using that value.) This can be used, for example, to fetch and join all the masculine equivalent names for a feminine given name. Each name is specified using parameters beginning with `pname` in `args`, e.g. "m", "m2", "m3", etc. `lang` is a language object specifying the language of the names (defaulting to English), for use in linking them. If `allow_explicit_lang` is given, the language of the terms can be specified explicitly by prefixing a term with a language code, e.g. 'sv:Björn' or 'la:[[Nicolaus|Nīcolāī]]'. This function assumes that the parameters have already been parsed by [[Module:parameters]] and gathered into lists, so that e.g. all "mN" parameters are in a list in args["m"]. ]=] local function join_names(lang, args, pname, conj, allow_explicit_lang) local termobjs = {} local do_language_link = false if not lang then lang = enlang do_language_link = true end local function process_one_term(term, i) for _, termobj in ipairs(parse_term_with_annotations(term, pname .. (i == 1 and "" or i), lang, allow_explicit_lang, "allow multiple terms")) do table.insert(termobjs, termobj) end end if not args[pname] then return "", 0 elseif type(args[pname]) == "table" then for i, term in ipairs(args[pname]) do process_one_term(term, i) end else process_one_term(args[pname], 1) end return join_terms(termobjs, nil, do_language_link, conj or "dan"), #termobjs end local function get_eqtext(args) local eqsegs = {} local lastlang = nil local last_eqseg = {} local function process_one_term(term, i) for _, termobj in ipairs(parse_term_with_annotations(term, "eq" .. (i == 1 and "" or i), enlang, "allow explicit lang", "allow multiple terms")) do local termlang = termobj.lang:getCode() if lastlang and lastlang ~= termlang then if #last_eqseg > 0 then table.insert(eqsegs, last_eqseg) end last_eqseg = {} end lastlang = termlang table.insert(last_eqseg, termobj) end end if type(args.eq) == "table" then for i, term in ipairs(args.eq) do process_one_term(term, i) end elseif type(args.eq) == "string" then process_one_term(args.eq, 1) end if #last_eqseg > 0 then table.insert(eqsegs, last_eqseg) end local eqtextsegs = {} for _, eqseg in ipairs(eqsegs) do table.insert(eqtextsegs, join_terms(eqseg, "include langname")) end return m_table.serialCommaJoin(eqtextsegs, {conj = "atau"}) end local function get_fromtext(lang, args) local catparts = {} local fromsegs = {} local i = 1 local function parse_from(from) local unrecognized = false local prefix, suffix if from == "nama keluarga" or from == "nama diri" or from == "nama panggilan" or from == "nama tempat" or from == "kata nama am" or from == "nama bulan" then prefix = "dipindahkan daripada " suffix = from table.insert(catparts, from) elseif from == "patronimik" or from == "matronimik" or from == "ciptaan baharu" then prefix = "berasal " suffix = "sebagai " .. from table.insert(catparts, from) elseif from == "pekerjaan" or from == "etnonim" then prefix = "berasal " suffix = "sebagai " .. from table.insert(catparts, from) elseif from == "Alkitab" then prefix = "berasal " suffix = "daripada Alkitab" table.insert(catparts, from) else prefix = "daripada " if from:find(":") then local termobj = parse_term_with_annotations(from, "from" .. (i == 1 and "" or i), lang, "allow explicit lang") local fromlangname = "" if termobj.lang:getCode() ~= lang:getCode() then -- If name is derived from another name in the same language, don't include lang name after text -- "from " or create a category like "German male given names derived from German". local canonical_name = termobj.lang:getCanonicalName() fromlangname = "bahasa " .. canonical_name .. " " table.insert(catparts, canonical_name) end suffix = fromlangname .. link_one_term(termobj) else local family = from:match("^[Bb]ahasa%-bahasa (.+)$") if family then if require("Module:families").getByCanonicalName(family) then table.insert(catparts, from) else unrecognized = true end suffix = from else if m_languages.getByCanonicalName(from, nil, "allow etym") then table.insert(catparts, from) else unrecognized = true end suffix = from end end end if unrecognized then track("unrecognized from") track("unrecognized from/" .. from) end return prefix, suffix end local last_fromseg = nil local put = require(parse_utilities_module) local from_args = args.from or {} if type(from_args) == "string" then from_args = {from_args} end while from_args[i] do -- We may have multiple comma-separated items, each of which may have multiple items separated by a -- space-delimited < sign, each of which may have inline modifiers with embedded commas in them. To handle -- this correctly, first replace space-delimited < signs with a special character, then split on balanced -- <...> and [...] signs, then split on comma, then rejoin the stuff between commas. We will then split on -- TEMP_LESS_THAN (the replacement for space-delimited < signs) and reparse. local rawfroms = rsub(from_args[i], "%s+<%s+", TEMP_LESS_THAN) local segments = put.parse_multi_delimiter_balanced_segment_run(rawfroms, {{"<", ">"}, {"[", "]"}}) local comma_separated_groups = put.split_alternating_runs_on_comma(segments) for j, comma_separated_group in ipairs(comma_separated_groups) do comma_separated_groups[j] = table.concat(comma_separated_group) end for _, rawfrom in ipairs(comma_separated_groups) do local froms = rsplit(rawfrom, TEMP_LESS_THAN) if #froms == 1 then local prefix, suffix = parse_from(froms[1]) if last_fromseg and (last_fromseg.has_multiple_froms or last_fromseg.prefix ~= prefix) then table.insert(fromsegs, last_fromseg) last_fromseg = nil end if not last_fromseg then last_fromseg = {prefix = prefix, suffixes = {}} end table.insert(last_fromseg.suffixes, suffix) else if last_fromseg then table.insert(fromsegs, last_fromseg) last_fromseg = nil end local first_suffixpart = "" local rest_suffixparts = {} for j, from in ipairs(froms) do local prefix, suffix = parse_from(from) if j == 1 then first_suffixpart = prefix .. suffix else table.insert(rest_suffixparts, prefix .. suffix) end end local full_suffix = first_suffixpart .. " [yang seterusnya " .. table.concat(rest_suffixparts, ", yang seterusnya ") .. "]" last_fromseg = {prefix = "", has_multiple_froms = true, suffixes = {full_suffix}} end end i = i + 1 end table.insert(fromsegs, last_fromseg) local fromtextsegs = {} for _, fromseg in ipairs(fromsegs) do table.insert(fromtextsegs, fromseg.prefix .. m_table.serialCommaJoin(fromseg.suffixes, {conj = "atau"})) end return m_table.serialCommaJoin(fromtextsegs, {conj = "atau"}), catparts end local function parse_given_name_genders(genderspec) if export.given_name_genders[genderspec] then -- optimization return {{ type = genderspec, props = export.given_name_genders[genderspec], }}, export.given_name_genders[genderspec].type == "animal" end local is_animal = nil local param_mods = require(parameter_utilities_module).construct_param_mods { {group = {"l", "q", "ref"}}, {param = {"text", "article"}}, } local function generate_obj(term, parse_err) if not export.given_name_genders[term] then local valid_genders = {} for k, _ in pairs(export.given_name_genders) do table.insert(valid_genders, k) end table.sort(valid_genders) parse_err(("Jantina '%s' tidak dikenali: jantina yang sah ialah %s"):format( term, table.concat(valid_genders, ", "))) end return { type = term, props = export.given_name_genders[term], } end local retval = require(parse_interface_module).parse_inline_modifiers(genderspec, { param_mods = param_mods, paramname = "2", generate_obj = generate_obj, splitchar = ",", }) for _, spec in ipairs(retval) do local this_is_animal = spec.props.type == "animal" if is_animal == nil then is_animal = this_is_animal elseif is_animal ~= this_is_animal then error("Jenis nama haiwan dan jantina manusia tidak boleh dicampurkan") end end return retval, is_animal end local function generate_given_name_genders(lang, genders) local parts = {} for _, spec in ipairs(genders) do local text if spec.text then -- NOTE: This assumes no % sign in the gender type, which seems safe. text = spec.text:gsub("%+", spec.type) else if spec.props.type == "animal" then text = "[[" .. spec.type .. "]]" else text = spec.type end end if spec.q and spec.q[1] or spec.qq and spec.qq[1] or spec.l and spec.l[1] or spec.ll and spec.ll[1] or spec.refs and spec.refs[1] then text = require(decorations_module).format_decorations { lang = lang, text = text, q = spec.q, qq = spec.qq, l = spec.l, ll = spec.ll, refs = spec.refs, raw = true, } end table.insert(parts, text) end local retval = m_table.serialCommaJoin(parts, {conj = "atau"}) local article = genders[1].article if not article and not genders[1].text and not genders[1].q and not genders[1].l then article = genders[1].props.article end if not article then article = "" end return retval, article end -- The entry point for {{given name}}. function export.given_name(frame) local parent_args = frame:getParent().args local compat = parent_args.lang local offset = compat and 0 or 1 local lang_index = compat and "lang" or 1 local boolean = {type = "boolean"} local list = {list = true} local args = require("Module:parameters").process(parent_args, { [lang_index] = {required = true, type = "language", default = "und"}, ["gender"] = {default = "jantina tidak diketahui"}, [1 + offset] = {alias_of = "gender"}, ["usage"] = true, ["origin"] = true, ["popular"] = true, ["populartype"] = true, ["meaning"] = list, ["meaningtype"] = true, ["addl"] = true, ["nocap"] = boolean, -- initial article: A or An ["A"] = true, ["sort"] = true, ["from"] = true, [2 + offset] = {alias_of = "from"}, ["fromtype"] = true, ["xlit"] = true, ["eq"] = true, ["eqtype"] = true, ["varof"] = true, ["varoftype"] = true, ["var"] = {alias_of = "varof"}, ["vartype"] = {alias_of = "varoftype"}, ["varform"] = true, ["varformtype"] = true, ["dimof"] = true, ["dimoftype"] = true, ["dim"] = {alias_of = "dimof"}, ["dimtype"] = {alias_of = "dimoftype"}, ["dimform"] = true, ["dimformtype"] = true, ["augof"] = true, ["augoftype"] = true, ["aug"] = {alias_of = "augof"}, ["augtype"] = {alias_of = "augoftype"}, ["augform"] = true, ["augformtype"] = true, ["clipof"] = true, ["clipoftype"] = true, ["blend"] = true, ["blendtype"] = true, ["m"] = true, ["mtype"] = true, ["f"] = true, ["ftype"] = true, ["nocat"] = boolean, }) local textsegs = {} local lang = args[lang_index] local langcode = lang:getCode() local function fetch_typetext(param) return args[param] and args[param] .. " " or "" end local genders, is_animal = parse_given_name_genders(args.gender) local dimoftext, numdimofs = join_names(lang, args, "dimof") local augoftext, numaugofs = join_names(lang, args, "augof") local xlittext = join_names(nil, args, "xlit") local blendtext = join_names(lang, args, "blend") -- formerly the only one that used "and" instead of "or" local varoftext = join_names(lang, args, "varof") local clipoftext = join_names(lang, args, "clipof") local mtext = join_names(lang, args, "m") local ftext = join_names(lang, args, "f") local varformtext, numvarforms = join_names(lang, args, "varform", ", ") local dimformtext, numdimforms = join_names(lang, args, "dimform", ", ") local augformtext, numaugforms = join_names(lang, args, "augform", ", ") local meaningsegs = {} for _, meaning in ipairs(args.meaning) do table.insert(meaningsegs, '“' .. meaning .. '”') end local meaningtext = m_table.serialCommaJoin(meaningsegs, {conj = "atau"}) local eqtext = get_eqtext(args) local function ins(txt) table.insert(textsegs, txt) end local dimoftype = args.dimoftype local augoftype = args.augoftype local added_text = nil if numdimofs > 0 then added_text = (dimoftype and dimoftype .. " " or "") .. "[[bentuk singkat]]" .. (xlittext ~= "" and ", " .. xlittext .. "," or "") .. " daripada " elseif numaugofs > 0 then added_text = (augoftype and augoftype .. " " or "") .. "[[bentuk agam]]" .. (xlittext ~= "" and ", " .. xlittext .. "," or "") .. " daripada " end local force_plural = false if added_text ~= nil then if args.dimof == "-" then dimoftext = "" force_plural = true else added_text = added_text .. "" end ins(added_text) end local article = args.A if not article and textsegs[1] then article = "" end ins("[[nama diri]]") if not is_animal then local gendertext, gender_article = generate_given_name_genders(lang, genders) article = article or gender_article ins(" ") ins(gendertext) end article = article or "" -- if no article set yet, it's "a" based on "given name" if langcode == "ms" and not args.nocap then article = mw.getContentLanguage():ucfirst(article) end local need_comma = false if numdimofs > 0 then ins(" " .. dimoftext) need_comma = not is_animal elseif numaugofs > 0 then ins(" " .. augoftext) need_comma = not is_animal elseif xlittext ~= "" then ins(", " .. xlittext) need_comma = true end if is_animal then if need_comma then ins(",") end need_comma = true ins(" untuk ") local gendertext, gender_article = generate_given_name_genders(lang, genders) ins(gender_article ~= "" and gender_article .. " " or "") ins(gendertext) end local from_catparts = {} if args.from then if need_comma then ins(",") end need_comma = true ins(" " .. fetch_typetext("fromtype")) local textseg, this_catparts = get_fromtext(lang, args) for _, catpart in ipairs(this_catparts) do m_table.insertIfNot(from_catparts, catpart) end ins(textseg) end if meaningtext ~= "" then if need_comma then ins(",") end need_comma = true ins(" " .. fetch_typetext("meaningtype") .. "bermaksud " .. meaningtext) end if args.origin then if need_comma then ins(",") end need_comma = true ins(" berasal daripada " .. args.origin) end if args.usage then if need_comma then ins(",") end ins(" dengan penggunaan " .. args.usage) end if varoftext ~= "" then ins(", " ..fetch_typetext("varoftype") .. "bentuk variasi daripada " .. varoftext) end if clipoftext ~= "" then ins(", " .. fetch_typetext("clipoftype") .. "kependekan daripada " .. clipoftext) end if blendtext ~= "" then ins(", " .. fetch_typetext("blendtype") .. "lakuran daripada " .. blendtext) end if args.popular then ins(", " .. fetch_typetext("populartype") .. "popular " .. args.popular) end if mtext ~= "" then ins(", " .. fetch_typetext("mtype") .. "padanan maskulin " .. mtext) end if ftext ~= "" then ins(", " .. fetch_typetext("ftype") .. "padanan feminin " .. ftext) end if eqtext ~= "" then ins(", " .. fetch_typetext("eqtype") .. "berpadanan dengan " .. eqtext) end if args.addl then if args.addl:find("^;") then ins(args.addl) elseif args.addl:find("^_") then ins(" " .. args.addl:sub(2)) else ins(", " .. args.addl) end end if varformtext ~= "" then ins("; " .. fetch_typetext("varformtype") .. "bentuk variasi" .. " " .. varformtext) end if dimformtext ~= "" then ins("; " .. fetch_typetext("dimformtype") .. "bentuk singkat" .. " " .. dimformtext) end if augformtext ~= "" then ins("; " .. fetch_typetext("augformtype") .. "bentuk agam" .. " " .. augformtext) end local text = "<span class='use-with-mention'>" .. (article ~= "" and article .. " " or "") .. table.concat(textsegs) .. "</span>" if args.nocat then return text end local categories = {} local langname = " bahasa " .. lang:getCanonicalName() local function insert_cats(dimaugof) if dimaugof == "" and genders[1].props.type == "human" then -- No category such as "English diminutives of given names" table.insert(categories, "nama diri" .. langname) end local function insert_cat(cat) table.insert(categories, dimaugof .. cat .. langname) for _, catpart in ipairs(from_catparts) do table.insert(categories, dimaugof .. cat .. langname .. " daripada " .. catpart) end end for _, spec in ipairs(genders) do local typ = spec.type if spec.props.track then track(typ) end local cats = get_given_name_cats(spec.type, spec.props) for _, cat in ipairs(cats) do insert_cat(cat) end end end insert_cats("") if numdimofs > 0 then insert_cats("bentuk singkat ") elseif numaugofs > 0 then insert_cats("bentuk agam ") end return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat) end -- The entry point for {{surname}}, {{patronymic}} and {{matronymic}}. function export.surname(frame) local iargs = require("Module:parameters").process(frame.args, { ["type"] = {required = true, set = {"nama keluarga", "patronimik", "matronimik"}}, }) local parent_args = frame:getParent().args local compat = parent_args.lang local offset = compat and 0 or 1 if parent_args.dot or parent_args.nodot then error("dot= dan nodot= tidak lagi disokong dalam [[Template:" .. iargs.type .. "]] kerana tanda " .. "noktah tidak lagi ditambah secara lalai; tambahkannya sendiri selepas templat jika diperlukan") end local lang_index = compat and "lang" or 1 local boolean = {type = "boolean"} local list = {list = true} local gender_arg = iargs.type == "nama keluarga" and "g" or 1 + offset local adj_arg = iargs.type == "nama keluarga" and 1 + offset or 2 + offset local args = require("Module:parameters").process(parent_args, { [lang_index] = {required = true, type = "language", template_default = "und"}, [gender_arg] = iargs.type == "nama keluarga" and true or {required = true, template_default = "tidak diketahui"}, -- gender(s) [adj_arg] = true, -- adjective/qualifier ["usage"] = true, ["origin"] = true, ["popular"] = true, ["populartype"] = true, ["meaning"] = list, ["meaningtype"] = true, ["parent"] = true, ["addl"] = true, ["nocap"] = boolean, -- initial article: by default A or An (English), a or an (otherwise) ["A"] = true, ["sort"] = true, ["from"] = true, ["fromtype"] = true, ["xlit"] = true, ["eq"] = true, ["eqtype"] = true, ["varof"] = true, ["varoftype"] = true, ["var"] = {alias_of = "varof"}, ["vartype"] = {alias_of = "varoftype"}, ["varform"] = true, ["varformtype"] = true, ["clipof"] = true, ["clipoftype"] = true, ["blend"] = true, ["blendtype"] = true, ["m"] = true, ["mtype"] = true, ["f"] = true, ["ftype"] = true, ["nocat"] = boolean, }) local textsegs = {} local lang = args[lang_index] local langcode = lang:getCode() local function fetch_typetext(param) return args[param] and args[param] .. " " or "" end local saw_male = false local saw_female = false local genders = {} if args[gender_arg] then for _, g in ipairs(require(parse_interface_module).split_on_comma(args[gender_arg])) do if g == "tidak diketahui" or g == "jantina tidak diketahui" or g == "?" then g = "jantina tidak diketahui" track("jantina tidak diketahui") elseif g == "uniseks" or g == "bebas jantina" or g == "c" then g = "bebas jantina" saw_male = true saw_female = true elseif g == "m" or g == "lelaki" then g = "lelaki" saw_male = true elseif g == "f" or g == "perempuan" then g = "perempuan" saw_female = true else error("Jantina tidak dikenali: " .. g) end table.insert(genders, g) end end local adj = args[adj_arg] local xlittext = join_names(nil, args, "xlit") local blendtext = join_names(lang, args, "blend", "dan") local varoftext = join_names(lang, args, "varof") local clipoftext = join_names(lang, args, "clipof") local mtext = join_names(lang, args, "m") local ftext = join_names(lang, args, "f") local parenttext = join_names(lang, args, "parent", nil, "allow explicit lang") local varformtext, numvarforms = join_names(lang, args, "varform", ", ") local meaningsegs = {} for _, meaning in ipairs(args.meaning) do table.insert(meaningsegs, '“' .. meaning .. '”') end if parenttext ~= "" then local child = saw_male and not saw_female and "anak lelaki" or saw_female and not saw_male and "anak perempuan" or "anak lelaki/perempuan" table.insert(meaningsegs, ("“%s kepada %s”"):format(child, parenttext)) end local meaningtext = m_table.serialCommaJoin(meaningsegs, {conj = "atau"}) local eqtext = get_eqtext(args) local function ins(txt) table.insert(textsegs, txt) end ins("<span class='use-with-mention'>") -- If gender is supplied, it goes before the specified adjective in adj=. The only value of gender that uses "an" is -- "unknown-gender" (note that "unisex" wouldn't use it but in any case we map "unisex" to "common-gender"). If gender -- isn't supplied, look at the first letter of the value of adj= if supplied; otherwise, the article is always "a" -- because the word "surname", "patronymic" or "matronymic" follows. Capitalize "A"/"An" if English. local article if args.A then article = args.A else article = #genders > 0 and genders[1] == "jantina tidak diketahui" and "" or #genders == 0 and adj and "" or "" if langcode == "ms" and not args.nocap then article = mw.getContentLanguage():ucfirst(article) end end ins(article ~= "" and article .. " " or "") ins("[[" .. iargs.type .. "]]") if #genders > 0 then ins(" " .. table.concat(genders, " atau ")) end if adj then ins(" " .. adj) end local need_comma = false if xlittext ~= "" then ins(", " .. xlittext) need_comma = true end local from_catparts = {} if args.from then if need_comma then ins(",") end need_comma = true ins(" " .. fetch_typetext("fromtype")) local textseg, this_catparts = get_fromtext(lang, args) for _, catpart in ipairs(this_catparts) do m_table.insertIfNot(from_catparts, catpart) end ins(textseg) end if meaningtext ~= "" then if need_comma then ins(",") end need_comma = true ins(" " .. fetch_typetext("meaningtype") .. "bermaksud " .. meaningtext) end if args.origin then if need_comma then ins(",") end need_comma = true ins(" berasal daripada " .. args.origin) end if args.usage then if need_comma then ins(",") end need_comma = true ins(" dengan penggunaan " .. args.usage) end if varoftext ~= "" then ins(", " ..fetch_typetext("varoftype") .. "bentuk variasi daripada " .. varoftext) end if clipoftext ~= "" then ins(", " .. fetch_typetext("clipoftype") .. "kependekan daripada " .. clipoftext) end if blendtext ~= "" then ins(", " .. fetch_typetext("blendtype") .. "lakuran daripada " .. blendtext) end if args.popular then ins(", " .. fetch_typetext("populartype") .. "popular " .. args.popular) end if mtext ~= "" then ins(", " .. fetch_typetext("mtype") .. "padanan maskulin " .. mtext) end if ftext ~= "" then ins(", " .. fetch_typetext("ftype") .. "padanan feminin " .. ftext) end if eqtext ~= "" then ins(", " .. fetch_typetext("eqtype") .. "berpadanan dengan " .. eqtext) end if args.addl then if args.addl:find("^;") then ins(args.addl) elseif args.addl:find("^_") then ins(" " .. args.addl:sub(2)) else ins(", " .. args.addl) end end if varformtext ~= "" then ins("; " .. fetch_typetext("varformtype") .. "bentuk variasi" .. " " .. varformtext) end ins("</span>") local text = table.concat(textsegs) if args.nocat then return text end local categories = {} local langname = " bahasa " .. lang:getCanonicalName() local function insert_cats(g) g = g and " " .. g or "" table.insert(categories, iargs.type .. g .. langname) for _, catpart in ipairs(from_catparts) do table.insert(categories, iargs.type .. g .. langname .. " daripada " .. catpart) end end insert_cats(nil) local function insert_cats_gender(g) if g == "jantina tidak diketahui" then return end if g == "bebas jantina" then insert_cats_gender("lelaki") insert_cats_gender("perempuan") end insert_cats(g) end for _, g in ipairs(genders) do insert_cats_gender(g) end return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat) end -- The entry point for {{name translit}}, {{name respelling}}, {{name obor}} and {{foreign name}}. function export.name_translit(frame) local boolean = {type = "boolean"} local iargs = require("Module:parameters").process(frame.args, { ["desctext"] = {required = true}, ["obor"] = boolean, ["foreign_name"] = boolean, }) local parent_args = frame:getParent().args local params = { [1] = {required = true, type = "language", template_default = "ms"}, [2] = {required = true, type = "language", sublist = true, template_default = "ru"}, [3] = {list = true, allow_holes = true}, ["type"] = {required = true, set = translit_name_type_list, sublist = true, default = "patronimik"}, ["dim"] = boolean, ["aug"] = boolean, ["nocap"] = boolean, ["addl"] = true, ["sort"] = true, ["pagename"] = true, ["nocat"] = boolean, } local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "q", "l", "ref"}}, {param = {"xlit", "eq"}}, } local names, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 3, track_module = "names/name translit", disallow_custom_separators = true, -- Use the first source language as the language of the specified names. lang = function(args) return args[2][1] end, sc = "sc.default", -- We display decorations ourselves so we can display them "raw", as we surround everything with -- 'use-with-mention' and want to avoid double-tagging. no_show_decorations = true, } local lang = args[1] local langcode = lang:getCode() local sources = args[2] local pagename = args.pagename or mw.loadData("Module:headword/data").pagename local textsegs = {} local function ins(txt) table.insert(textsegs, txt) end ins("<span class='use-with-mention'>") local desctext = iargs.desctext if langcode == "ms" and not args.nocap then desctext = mw.getContentLanguage():ucfirst(desctext) end ins(desctext .. " ") if not iargs.foreign_name then ins("daripada ") end local langsegs = {} for i, source in ipairs(sources) do local sourcename = source:getCanonicalName() local function get_source_link() local term_to_link = names[1] and names[1].term or pagename -- We link the language name to either the first specified name or the pagename, in the following -- circumstances: -- (1) More than one language was given along with at least one name; or -- (2) We're handling {{foreign name}} or {{name obor}}, and no name was given. -- The reason for (1) is that if more than one language was given, we want a link to the name -- in each language, as the name that's displayed is linked only to the first specified language. -- However, if only one language was given, linking the language to the name is redundant. -- The reason for (2) is that {{foreign name}} is often used when the name in the destination language -- is spelled the same as the name in the source language (e.g. [[Clinton]] or [[Obama]] in Italian), -- and in that case no name will be explicitly specified but we still want a link to the name in the -- source language. The reason we restrict this to {{foreign name}} or {{name obor}}, not to -- {{name translit}} or {{name respelling}}, is that {{name translit}} and {{name respelling}} ought to be -- used for names spelled differently in the destination language (either transliterated or respelled), so -- assuming the pagename is the name in the source language is wrong. if names[1] and #sources > 1 or (iargs.foreign_name or iargs.obor) and not names[1] then return m_links.language_link{ lang = sources[i], term = term_to_link, alt = sourcename, tr = "-" } else return sourcename end end if i == 1 and not iargs.foreign_name then -- If at least one name is given, we say "A transliteration of the LANG surname FOO", linking LANG to FOO. -- Otherwise we say "A transliteration of a LANG surname". if names[1] then table.insert(langsegs, get_source_link()) else table.insert(langsegs, sourcename) end else table.insert(langsegs, get_source_link()) end end local langseg_text = m_table.serialCommaJoin(langsegs, {conj = "atau"}) local augdim_text if args.dim then augdim_text = " [[bentuk singkat]]" elseif args.aug then augdim_text = " [[bentuk agam]]" else augdim_text = "" end local nametype_linked = {} for _, nametype in ipairs(args["type"]) do if nametype == "nama keluarga" or nametype == "patronimik" then table.insert(nametype_linked, "[[" .. nametype .. "]]") elseif nametype == "nama diri lelaki" then table.insert(nametype_linked, "[[nama diri]] lelaki") elseif nametype == "nama diri perempuan" then table.insert(nametype_linked, "[[nama diri]] perempuan") elseif nametype == "nama diri uniseks" then table.insert(nametype_linked, "[[nama diri]] uniseks") else table.insert(nametype_linked, nametype) end end local nametype_text = m_table.serialCommaJoin(nametype_linked, {conj = "atau"}) .. augdim_text if not iargs.foreign_name then ins(nametype_text) ins(" bahasa " .. langseg_text) if names[1] then ins(" ") end else ins(nametype_text) ins(" dalam bahasa " .. langseg_text) if names[1] then ins(", ") end end local linked_names = {} local embedded_comma = false for _, name in ipairs(names) do local linked_name = m_links.full_link(name, "term") if name.q and name.q[1] or name.qq and name.qq[1] or name.l and name.l[1] or name.ll and name.ll[1] or name.refs and name.refs[1] then linked_name = require(decorations_module).format_decorations { lang = name.lang, text = linked_name, q = name.q, qq = name.qq, l = name.l, ll = name.ll, refs = name.refs, raw = true, -- since we put 'use-with-mention' around the entire text } end if name.xlit then embedded_comma = true linked_name = linked_name .. ", " .. m_links.language_link { lang = enlang, term = name.xlit } end if name.eq then embedded_comma = true linked_name = linked_name .. ", berpadanan dengan " .. m_links.language_link { lang = enlang, term = name.eq } end table.insert(linked_names, linked_name) end if embedded_comma then ins(table.concat(linked_names, "; atau daripada ")) else ins(m_table.serialCommaJoin(linked_names, {conj = "atau"})) end if args.addl then if args.addl:find("^;") then ins(args.addl) elseif args.addl:find("^_") then ins(" " .. args.addl:sub(2)) else ins(", " .. args.addl) end end ins("</span>") local text = table.concat(textsegs) if args.nocat then return text end local categories = {} local function inscat(cat) table.insert(categories, cat) end for _, nametype in ipairs(args.type) do local function insert_cats(dimaugof) local function insert_cats_type(ty) if ty == "nama diri uniseks" then insert_cats_type("nama diri lelaki") insert_cats_type("nama diri perempuan") end for _, source in ipairs(sources) do inscat("Kemasan bahasa " .. lang:getFullName() .. " daripada " .. dimaugof .. ty .. " bahasa " .. source:getCanonicalName()) inscat("Perkataan bahasa " .. lang:getFullName() .. " diterbitkan daripada bahasa " .. source:getCanonicalName()) inscat("Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada bahasa " .. source:getCanonicalName()) if iargs.obor then inscat("Pinjaman ortografi bahasa " .. lang:getFullName() .. " daripada bahasa " .. source:getCanonicalName()) end if source:getCode() ~= source:getFullCode() then -- etymology language inscat("Kemasan bahasa " .. lang:getFullName() .. " daripada " .. dimaugof .. ty .. " bahasa " .. source:getFullName()) end end end insert_cats_type(nametype) end insert_cats("") if args.dim then insert_cats("bentuk singkat ") end if args.aug then insert_cats("bentuk agam ") end end return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat) end return export b2vgwn57cpwfm78qqio840nz6cbmyyt Modul:etymology/templates/descendant 828 27509 375370 365964 2026-09-22T05:01:25Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708899|92708899]]) 375370 Scribunto text/plain local export = {} local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local descendants_tree_module = "Module:descendants tree" local etymology_style_css = "Module:etymology/style.css" local labels_module = "Module:labels" local languages_module = "Module:languages" local links_module = "Module:links" local parameter_utilities_module = "Module:parameter utilities" local scripts_module = "Module:scripts" local table_module = "Module:table" local table_module_list_to_set = "Module:table/listToSet" local template_styles_module = "Module:TemplateStyles" local concat = table.concat local insert = table.insert local list_to_set = require(table_module_list_to_set) local error_on_no_descendants = false local function track(page) return require(debug_track_module)("descendant/" .. page) end local function ine(arg) if arg == "" then return nil else return arg end end local function add_tooltip(text, tooltip) return '<span class="desc-arr" title="' .. tooltip .. '">' .. text .. '</span>' end -- Boolean params indicating whether a descendant term (or all terms) are particular sorts of borrowings. local bortypes = {"inh", "bor", "lbor", "slb", "obor", "translit", "der", "clq", "pclq", "sml", "unc"} local bortype_set = list_to_set(bortypes) -- Aliases of clq=. local calque_aliases = {"cal", "calq", "calque"} local calque_alias_set = list_to_set(calque_aliases) -- Aliases of pclq=. local partial_calque_aliases = {"pcal", "pcalq", "pcalque"} local partial_calque_alias_set = list_to_set(partial_calque_aliases) local semi_learned_borrowing_aliases = {"slbor"} local semi_learned_borrowing_alias_set = list_to_set(semi_learned_borrowing_aliases) --- Return a function of one argument `field` (a param name), which fetches `args`[`field`].default if index == 0, else --- `container`[`field`]. local function get_val(container, args, index) return function(field) if index == 0 then return args[field].default else return container[field] end end end local function get_arrow(container, args, index) local val = get_val(container, args, index) local arrow if val("bor") then arrow = add_tooltip("→", "pinjaman") elseif val("lbor") then arrow = add_tooltip("→", "pinjaman terpelajar") elseif val("slb") then arrow = add_tooltip("→", "pinjaman terpelajar separa") elseif val("obor") then arrow = add_tooltip("→", "pinjaman ortografi") elseif val("translit") then arrow = add_tooltip("→", "transliterasi") elseif val("clq") then arrow = add_tooltip("→", "pinjaman terjemahan") elseif val("pclq") then arrow = add_tooltip("→", "pinjaman terjemahan separa") elseif val("sml") then arrow = add_tooltip("→", "pinjaman semantik") elseif val("inh") or (val("unc") and not val("der")) then arrow = add_tooltip(">", "diwariskan") else arrow = "" end -- allow der=1 in conjunction with bor=1 to indicate e.g. English "pars recta" -- derived and borrowed from Latin "pars". if val("der") then arrow = arrow .. add_tooltip("⇒", "reshaped by analogy or addition of morphemes") end if val("unc") then arrow = arrow .. add_tooltip("?", "uncertain") end if arrow ~= "" then arrow = arrow .. " " end return arrow end -- Return the pre-decoration text for the `index`th term, or the overall pre-decoration text if index == 0. local function get_pre_decorations(container, args, index) if index > 0 then -- per term decorations are handled at the subitem level, by full_link(). return nil, nil end local val = get_val(container, args, index) return val("l"), val("q") end -- Return the post-decoration text for the `index`th term, or the overall post-decoration text if index == 0. local function get_post_decorations(container, args, index, lang) local val = get_val(container, args, index) local boolean_labels = {} if val("inh") then insert(boolean_labels, "diwariskan") end if val("lbor") then insert(boolean_labels, "terpelajar") end if val("slb") then insert(boolean_labels, "terpelajar separa") end if val("translit") then insert(boolean_labels, "transliterasi") end if val("clq") then insert(boolean_labels, "pinjaman terjemahan") end if val("pclq") then insert(boolean_labels, "pinjaman terjemahan separa") end if val("sml") then insert(boolean_labels, "pinjaman semantik") end if index > 0 then -- per term decorations are handled at the subitem level, by full_link(). return boolean_labels else local quals, dash_labels quals = val("qq") if val("ll") then local labels = require(labels_module).show_labels { lang = lang, labels = val("ll"), nocat = true, open = false, close = false, no_track_already_seen = true, ok_to_destructively_modify = true, -- doesn't apply to `labels` } if labels ~= "" then dash_labels = " &mdash; " .. labels end end return boolean_labels, quals, dash_labels end end local function desc_or_desc_tree(frame, desc_tree) local params local boolean = {type = "boolean"} if desc_tree then params = { [1] = {required = true, type = "language", family = true, default = "gem-pro"}, [2] = {required = true, list = true, allow_holes = true, default = "*fuhsaz"}, notext = boolean, noalts = boolean, noparent = boolean, } else params = { [1] = {required = true, type = "language", family = true, default = "en"}, [2] = {list = true, allow_holes = true, template_default = "word"}, alts = boolean, } end -- Add other single params. params.sclang = boolean params.sclb = {replaced_by = "sclang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"} params.nolang = boolean params.nolb = {replaced_by = "nolang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"} local parent_args if frame.args[1] then parent_args = frame.args else parent_args = frame:getParent().args end -- Error to catch most uses of old-style parameters. if ine(parent_args[4]) and not ine(parent_args[3]) and not ine(parent_args.tr2) and not ine(parent_args.ts2) and not ine(parent_args.t2) and not ine(parent_args.gloss2) and not ine(parent_args.g2) and not ine(parent_args.alt2) then error("You specified a term in 4= and not one in 3=. You probably meant to use t= to specify a gloss instead. " .. "If you intended to specify two terms, put the second term in 3=.") end if not ine(parent_args[3]) and not ine(parent_args.alt2) and not ine(parent_args.tr2) and not ine(parent_args.ts2) and ine(parent_args.g2) then error("You specified a gender in g2= but no term in 3=. You were probably trying to specify two genders for " .. "a single term. To do that, put both genders in g=, comma-separated.") end local m_param_utils = require(parameter_utilities_module) local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l", "q"}}, {param = "lb", replaced_by = false, instead = "use 'l' for left labels or 'll' for right labels"}, {param = bortypes, type = "boolean", overall = true, separate_no_index = true}, {param = calque_aliases, alias_of = "clq"}, {param = partial_calque_aliases, alias_of = "pclq"}, {param = semi_learned_borrowing_aliases, alias_of = "slb"}, } local groups, args, globalprops = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, -- Need some work to support this. -- parse_lang_prefix = true, track_module = "descendant", -- Due to allowing families as langs and substituting 'und', it's easier to do this later. -- lang = function() ... end sc = "sc.default", splitchar = "[,~]", subitem_separator_map = {[","] = " / ", ["~"] = " ~ "}, pre_normalize_modifiers = function(data) local modtext = data.modtext modtext = modtext:match("^<(.*)>$") if not modtext then error(("Internal error: Passed-in modifier isn't surrounded by angle brackets: %s"):format( data.modtext)) end if bortype_set[modtext] or calque_alias_set[modtext] or partial_calque_alias_set[modtext] or semi_learned_borrowing_alias_set[modtext] then modtext = modtext .. ":1" end return "<" .. modtext .. ">" end, } local lang = args[1] local namespace = mw.title.getCurrentTitle().nsText if (namespace == "" or namespace == "Rekonstruksi") and ( lang:hasType("appendix-constructed") and not lang:hasType("regular")) then error("Istilah dalam bahasa binaan lampiran-sahaja tidak boleh diberikan sebagai keturunan.") end local fetch_alt_forms = desc_tree and not args.noalts or not desc_tree and args.alts local m_desctree if desc_tree or fetch_alt_forms then m_desctree = require(descendants_tree_module) end if lang:getCode() ~= lang:getFullCode() then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/etymological]] track("etymological") track("etymological/" .. lang:getCode()) end local is_family = lang:hasType("family") local proxy_lang if is_family then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/family]] track("family") track("family/" .. lang:getCode()) proxy_lang = require(languages_module).getByCode("und") else proxy_lang = lang end local langname if is_family then -- The display form for families includes the word "languages", which we probably don't want to -- display. langname = lang:getCanonicalName() else langname = lang:getDisplayForm() end local langtag if args.sclang then local sc_to_use = args.sc.default if not sc_to_use then local first_termobj = groups[1] and groups[1].terms[1] if not first_termobj then error("sclang= diberikan tetapi tiada istilah untuk memaparkan nama skrip") end sc_to_use = first_termobj.sc if not sc_to_use then local first_term = first_termobj.term or first_termobj.alt if not first_term then error("sclang= diberikan tetapi item pertama tiada istilah atau bentuk paparan untuk memaparkan nama skrip") end if first_termobj.lang then sc_to_use = first_termobj.lang:findBestScript(first_term) elseif is_family then sc_to_use = require(scripts_module).findBestScriptWithoutLang(first_term, "none is last resort") else sc_to_use = lang:findBestScript(first_term) end end end langtag = sc_to_use:getDisplayForm(lang) else langtag = langname end local terms_for_descendant_trees = {} -- Keep track of descendants whose descendant tree we fetch. Don't fetch the same descendant tree twice (which -- can happen especially with Arabic-script terms with the same unvocalized spelling but differing vocalization). -- This happens e.g. with Ottoman Turkish [[پورتقال]], which has {{desctree|fa-cls|پُرْتُقَال|پُرْتِقَال|bor=1}}, with -- two terms that have the same unvocalized spelling. local terms_and_ids_fetched = {} local descendant_terms_seen = {} local parts = {} for i, group in ipairs(groups) do local group_parts = {} local terms_for_alt_forms = {} for _, item in ipairs(group.terms) do local link = "" item.lang = item.lang or proxy_lang item.track_sc = true -- Construct a link out of `item`. Also add the term to the list of descendant trees and/or alternative -- forms to fetch, if the page+ID combination hasn't already been seen. if item.term ~= "-" then -- including term == nil link = require(links_module).full_link(item, nil, true) if item.term and (desc_tree or fetch_alt_forms) then local m_links = require(links_module) -- Fetches information under entry. If term is of type A//B, it checks A. local entry_name = m_links.get_link_page(m_links.remove_links(mw.ustring.gsub(item.term, "//.+$", "")), lang, item.sc) -- NOTE: We use the term and ID as the key, but not the language. This is OK currently because -- all terms have the same language; but if we ever add support for a term-specific language, -- we need to fix this. local term_and_id = item.id and entry_name .. "!!!" .. item.id or entry_name if not terms_and_ids_fetched[term_and_id] then terms_and_ids_fetched[term_and_id] = true local term_for_fetching = { lang = lang, entry_name = entry_name, id = item.id } if desc_tree then if is_family then error("Tiada sokongan pada masa ini (dan mungkin tidak akan ada) untuk mengambil pokok keturunan apabila kod keluarga diberikan sebagai ganti kod bahasa") end if error_on_no_descendants then require(table_module).insertIfNot(descendant_terms_seen, { term = item.term, id = item.id }) end table.insert(terms_for_descendant_trees, term_for_fetching) end if fetch_alt_forms then if is_family then error("Tiada sokongan pada masa ini (dan mungkin tidak akan ada) untuk mengambil bentuk alternatif apabila kod keluarga diberikan sebagai ganti kod bahasa") end -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/alts]] track("alts") table.insert(terms_for_alt_forms, term_for_fetching) end end end elseif item.tr or item.ts or item.gloss or item.genders then -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term]] track("no term") item.term = nil item.show_decorations = true link = require(links_module).full_link(item, nil, true) link = link :gsub("<small>%[Istilah%?%]</small> ", "") :gsub("<small>%[Istilah%?%]</small>&nbsp;", "") :gsub("%[%[Category:[^%[%]]+ term requests%]%]", "") else -- display no link at all -- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term or annotations]] track("no term or annotations") end if link ~= "" then insert(group_parts, item.separator) insert(group_parts, link) end end if group_parts[1] then for _, altterm in ipairs(terms_for_alt_forms) do local altform = m_desctree.get_alternative_forms(altterm.lang, altterm.entry_name, altterm.id, globalprops.use_semicolon and "; " or ", ") if altform ~= "" then insert(group_parts, globalprops.use_semicolon and "; " or ", ") insert(group_parts, altform) end end local group_link = concat(group_parts) insert(parts, group.separator) if not args.notext then insert(parts, get_arrow(group, args, i)) end -- no pre-qualifiers/labels and no post-qualifiers/dash-labels local post_boolean_labels = get_post_decorations(group, args, i, proxy_lang) if post_boolean_labels and post_boolean_labels[1] then group_link = require(decorations_module).format_decorations { lang = proxy_lang, text = group_link, ll = post_boolean_labels, } end insert(parts, group_link) end end local descendant_trees = {} for _, descterm in ipairs(terms_for_descendant_trees) do -- When I ([[User:Benwing2]]) first implemented this in Nov 2020, I had `maxmaxindex > 1` as the last argument. -- Since then, [[User:Fytcha]] changed the last param to `true`. local descendant_tree = m_desctree.get_descendants(descterm.lang, descterm.entry_name, descterm.id, true) if descendant_tree and descendant_tree ~= "" then insert(descendant_trees, descendant_tree) end end if error_on_no_descendants and desc_tree and not descendant_trees[1] then local function format_term_seen(term_seen) if term_seen.id then return ("[[%s]] dengan ID '%s'"):format(term_seen.term, term_seen.id) else return ("[[%s]]"):format(term_seen.term) end end if #descendant_terms_seen == 0 then error("[[Template:desctree]] dipanggil tetapi tiada istilah untuk mendapatkan keturunan") elseif #descendant_terms_seen == 1 then error(("Tiada bahagian Keturunan ditemui dalam entri %s di bawah pengepala untuk %s"):format( format_term_seen(descendant_terms_seen[1]), lang:getFullName())) else for i, term_seen in ipairs(descendant_terms_seen) do descendant_terms_seen[i] = format_term_seen(term_seen) end error(("Tiada bahagian Keturunan ditemui dalam mana-mana entri %s di bawah pengepala untuk %s"):format( concat(descendant_terms_seen, ", "), lang:getFullName())) end end local descendants = concat(descendant_trees) if args.noparent then return descendants end local initial_labels, initial_quals = get_pre_decorations(nil, args, 0) local final_boolean_labels, final_quals, final_dash_labels = get_post_decorations(nil, args, 0, proxy_lang) local all_linktext = concat(parts) if initial_labels and initial_labels[1] or initial_quals and initial_quals[1] or final_boolean_labels and final_boolean_labels[1] or final_quals and final_quals[1] then all_linktext = require(decorations_module).format_decorations { lang = proxy_lang, text = all_linktext, l = initial_labels, q = initial_quals, ll = final_boolean_labels, qq = final_quals, } end if final_dash_labels then all_linktext = all_linktext .. final_dash_labels end all_linktext = all_linktext .. descendants if args.notext then return all_linktext end local initial_arrow = get_arrow(nil, args, 0) if args.nolang then return initial_arrow .. all_linktext else return concat { initial_arrow, langtag, ":", all_linktext ~= "" and " " or "", all_linktext } end end function export.descendant(frame) return desc_or_desc_tree(frame, false) .. require(template_styles_module)(etymology_style_css) end function export.descendants_tree(frame) return desc_or_desc_tree(frame, true) end return export 4kh1nl61q5k5lsm5uswzc33ooz8wsiy Modul:languages/data/exceptional 828 33718 375374 375287 2026-09-22T05:25:54Z Hakimi97 2668 Betulkan keluarga bahasa bahasa Temuan terus ke bahasa Melayik Purba 375374 Scribunto text/plain local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared local m = {} m["aav-khs-pro"] = { "Khasi Purba", 116773216, "aav-khs", "Latn", type = "reconstructed", } m["aav-nic-pro"] = { "Nicobar Purba", 116773793, "aav-nic", "Latn", type = "reconstructed", } m["aav-pkl-pro"] = { "Pnar-Khasi-Lyngngam Purba", 116773259, "aav-pkl", "Latn", type = "reconstructed", } m["aav-pro"] = { -- mkh-pro akan digabungkan ke dalam ini "Austroasia Purba", 116773186, "aav", "Latn", type = "reconstructed", } m["afa-pro"] = { "Afroasia Purba", 269125, "afa", "Latn", type = "reconstructed", } m["alg-aga"] = { "Agawam", nil, "alg-eas", "Latn", } m["alg-pro"] = { "Algonquian Purba", 7251834, "alg", "Latn", type = "reconstructed", sort_key = {remove_diacritics = "·"}, } m["alv-ama"] = { "Amasi", 4740400, "nic-grs", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron}, } m["alv-bgu"] = { "Bainouk Gubeeher", 17002646, "alv-bny", "Latn", } m["alv-bua-pro"] = { "Bua Purba", 116773723, "alv-bua", "Latn", type = "reconstructed", } m["alv-cng-pro"] = { "Cangin Purba", 116773726, "alv-cng", "Latn", type = "reconstructed", } m["alv-edo-pro"] = { "Edoid Purba", 116773206, "alv-edo", "Latn", type = "reconstructed", } m["alv-fli-pro"] = { "Fali Purba", 116773754, "alv-fli", "Latn", type = "reconstructed", } m["alv-gbe-pro"] = { "Gbe Purba", 116773208, "alv-gbe", "Latn", type = "reconstructed", } m["alv-gng-pro"] = { "Guang Purba", 116773757, "alv-gng", "Latn", type = "reconstructed", } m["alv-gtm-pro"] = { "Togo Tengah Purba", 116773732, "alv-gtm", "Latn", type = "reconstructed", } m["alv-gwa"] = { "Gwara", 16945580, "nic-pla", "Latn", } m["alv-hei-pro"] = { "Heiban Purba", 116773760, "alv-hei", "Latn", type = "reconstructed", } m["alv-ido-pro"] = { "Idomoid Purba", 116773764, "alv-ido", "Latn", type = "reconstructed", } m["alv-igb-pro"] = { "Igboid Purba", 116773765, "alv-igb", "Latn", type = "reconstructed", } m["alv-kwa-pro"] = { "Kwa Purba", 116773780, "alv-kwa", "Latn", type = "reconstructed", } m["alv-mum-pro"] = { "Mumuye Purba", 116773791, "alv-mum", "Latn", type = "reconstructed", } m["alv-nup-pro"] = { "Nupoid Purba", 116773795, "alv-nup", "Latn", type = "reconstructed", } m["alv-pro"] = { "Atlantik-Congo Purba", 116732838, "alv", "Latn", type = "reconstructed", } m["alv-edk-pro"] = { "Edekiri Purba", nil, "alv-edk", "Latn", type = "reconstructed", } m["alv-yor-pro"] = { "Yoruba Purba", nil, "alv-yor", "Latn", type = "reconstructed", } m["alv-yrd-pro"] = { "Yoruboid Purba", 116773824, "alv-yrd", "Latn", type = "reconstructed", } m["alv-von-pro"] = { "Volta-Niger Purba", 116773820, "alv-von", "Latn", type = "reconstructed", } m["apa-pro"] = { "Apache Purba", 116773135, "apa", "Latn", type = "reconstructed", } m["aql-pro"] = { "Algik Purba", 18389588, "aql", "Latn", type = "reconstructed", sort_key = {remove_diacritics = "·"}, } m["art-adu"] = { "Adûni", 1232159, "art", "Latn", type = "appendix-constructed", } m["art-bel"] = { "Kreol Belter", 108055510, "art", "Latn", type = "appendix-constructed", sort_key = { remove_diacritics = c.acute, from = {"ɒ"}, to = {"a"}, }, } m["art-blk"] = { "Bolak", 2909283, "art", "Latn", type = "appendix-constructed", } m["art-bsp"] = { "Bahasa Hitam", 686210, "art", "Latn, Teng", type = "appendix-constructed", } m["art-com"] = { "Communicationssprache", 35227, "art", "Latn", type = "appendix-constructed", } m["art-dtk"] = { "Dothraki", 2914733, "art", "Latn", type = "appendix-constructed", } m["art-elo"] = { "Eloi", nil, "art", "Latn", type = "appendix-constructed", } m["art-gld"] = { "Goa'uld", 19823, "art", "Latn, Egyp, Mero", type = "appendix-constructed", } m["art-lap"] = { "Lapine", 6488195, "art", "Latn", type = "appendix-constructed", } m["art-man"] = { "Mandalorian", 54289, "art", "Latn", type = "appendix-constructed", } m["art-mun"] = { "Mundolinco", 851355, "art", "Latn", type = "appendix-constructed", } m["art-nav"] = { "Naʼvi", 316939, "art", "Latn", type = "appendix-constructed", } m["art-vlh"] = { "Valyria Tinggi", 64483808, "art", "Latn", type = "appendix-constructed", } m["ath-nic"] = { "Nicola", 20609, "ath-nor", "Latn", } m["ath-pro"] = { "Athabaska Purba", 104841722, "ath", "Latn", type = "reconstructed", } m["auf-pro"] = { "Arawa Purba", 116773706, "auf", "Latn", type = "reconstructed", } m["aus-alu"] = { "Alungul", 16827670, "aus-pmn", "Latn", } m["aus-and"] = { "Andjingith", 4754509, "aus-pmn", "Latn", } m["aus-ang"] = { "Angkula", 16828520, "aus-pmn", "Latn", } m["aus-arn-pro"] = { "Arnhem Purba", 116773720, "aus-arn", "Latn", type = "reconstructed", } m["aus-bra"] = { "Barranbinya", 4863220, "aus-pmn", "Latn", } m["aus-brm"] = { "Barunggam", 4865914, "aus-pmn", "Latn", } m["aus-cww-pro"] = { "New South Wales Tengah Purba", 116773199, "aus-cww", "Latn", type = "reconstructed", } m["aus-dal-pro"] = { "Daly Purba", 116773743, "aus-dal", "Latn", type = "reconstructed", } m["aus-guw"] = { "Guwar", 6652138, "aus-pam", "Latn", } m["aus-lsw"] = { "Little Swanport", 6652138, "qfa-unc", "Latn", } m["aus-mbi"] = { "Mbiywom", 6799701, "aus-pmn", "Latn", } m["aus-ngk"] = { "Ngkoth", 7022405, "aus-pmn", "Latn", } m["aus-nyu-pro"] = { "Nyulnyulan Purba", 116773797, "aus-nyu", "Latn", type = "reconstructed", } m["aus-pam-pro"] = { "Pama-Nyunga Purba", 33942, "aus-pam", "Latn", type = "reconstructed", } m["aus-tul"] = { "Tulua", 16938541, "aus-pam", "Latn", } m["aus-uwi"] = { "Uwinymil", 7903995, "aus-arn", "Latn", } m["aus-wdj-pro"] = { "Iwaidjan Purba", 116773767, "aus-wdj", "Latn", type = "reconstructed", } m["aus-won"] = { "Wong-gie", nil, "aus-pam", "Latn", } m["aus-wul"] = { "Wulguru", 8039196, "aus-dyb", "Latn", } m["aus-ynk"] = { -- kontras nny "Yangkaal", 3913770, "aus-tnk", "Latn", } m["awd-amc-pro"] = { "Amuesha-Chamicuro Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-kmp-pro"] = { "Kampa Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-prw-pro"] = { "Paresi-Waura Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-ama"] = { "Amarizana", 16827787, "awd", "Latn", } m["awd-ana"] = { "Anauyá", 16828252, "awd", "Latn", } m["awd-apo"] = { "Apolista", 16916645, "awd", "Latn", } m["awd-cab"] = { "Cabre", 16850160, "awd", "Latn", } m["awd-gnu"] = { "Guinau", 3504087, "awd", "Latn", } m["awd-kar"] = { "Cariay", 16920253, "awd", "Latn", } m["awd-kaw"] = { "Kawishana", 6379993, "awd-nwk", "Latn", } m["awd-kus"] = { "Kustenau", 5196293, "awd", "Latn", } m["awd-man"] = { "Manao", 6746920, "awd", "Latn", } m["awd-mar"] = { "Marawan", 6755108, "awd", "Latn", } m["awd-mpr"] = { "Maipure", 6736872, "awd", "Latn", } m["awd-mrt"] = { "Mariaté", 16910017, "awd-nwk", "Latn", } m["awd-nwk-pro"] = { "Nawiki Purba", 116773234, "awd-nwk", "Latn", type = "reconstructed", } m["awd-pai"] = { "Paikoneka", 128807835, "awd", "Latn", } m["awd-pas"] = { "Pasé", 7143168, "awd-nwk", "Latn", } m["awd-pro"] = { "Arawak Purba", 97573478, "awd", "Latn", type = "reconstructed", } m["awd-she"] = { "Shebayo", 7492248, "awd", "Latn", } m["awd-taa-pro"] = { "Ta-Arawak Purba", 116773282, "awd-taa", "Latn", type = "reconstructed", } m["awd-wai"] = { "Wainumá", 16910017, "awd-nwk", "Latn", } m["awd-war"] = { "Warekena Kuno", 105320180, "awd-nwk", "Latn", } m["awd-yum"] = { "Yumana", 8061062, "awd-nwk", "Latn", } m["azc-caz"] = { "Cazcan", 5055514, "azc", "Latn", } m["azc-cup-pro"] = { "Cupan Purba", 116773738, "azc-cup", "Latn", type = "reconstructed", } m["azc-ktn"] = { "Kitanemuk", 3197558, "azc-tak", "Latn", } m["azc-nah-pro"] = { "Nahua Purba", 7251860, "azc-nah", "Latn", type = "reconstructed", } m["azc-nic"] = { "Nicoleño", 50241488, "azc", "Latn", } m["azc-num-pro"] = { "Numik Purba", 116773247, "azc-num", "Latn", type = "reconstructed", } m["azc-pro"] = { "Uto-Aztek Purba", 96400333, "azc", "Latn", type = "reconstructed", } m["azc-tak-pro"] = { "Takik Purba", 116773283, "azc-tak", "Latn", type = "reconstructed", } m["azc-tat"] = { "Tataviam", 743736, "azc", "Latn", } m["ber-pro"] = { "Berber Purba", 2855698, "ber", "Latn", type = "reconstructed", } m["ber-fog"] = { "Fogaha", 107610173, "ber", "Latn", } m["ber-zuw"] = { "Zuwara", 4117169, "ber", "Latn", } m["bnt-bal"] = { "Balong", 93935237, "bnt-bbo", "Latn", } m["bnt-bon"] = { "Boma Nkuu", nil, "bnt", "Latn", } m["bnt-boy"] = { "Boma Yumu", nil, "bnt", "Latn", } m["bnt-bwa"] = { "Bwala", 128810345, "bnt-tek", "Latn", } m["bnt-cmw"] = { "Chimwiini", 4958328, "bnt-swh", "Latn", } m["bnt-ind"] = { "Indanga", 51412803, "bnt", "Latn", } m["bnt-lal"] = { "Lala (Afrika Selatan)", 6480154, "bnt-ngu", "Latn", } m["bnt-mpi"] = { "Mpiin", 93937013, "bnt-bdz", "Latn", } m["bnt-mpu"] = { "Mpuono", -- jangan dikelirukan dengan Mbuun zmp 36056, "bnt", "Latn", } m["bnt-ngu-pro"] = { "Nguni Purba", 961559, "bnt-ngu", "Latn", type = "reconstructed", sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.caron}, } m["bnt-phu"] = { "Phuthi", 33796, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute}, } m["bnt-pro"] = { "Bantu Purba", 3408025, "bnt", "Latn", type = "reconstructed", sort_key = "bnt-pro-sortkey", } m["bnt-sab-pro"] = { "Sabaki Purba", nil, -- Q2209395 ialah kod untuk keluarga Sabaki "bnt-sab", "Latn", type = "reconstructed", } m["bnt-sbo"] = { "Boma Selatan", nil, "bnt", "Latn", } m["bnt-sts-pro"] = { "Sotho-Tswana Purba", 116773278, "bnt-sts", "Latn", type = "reconstructed", } m["btk-pro"] = { "Batak Purba", 116773191, "btk", "Latn", type = "reconstructed", } m["cau-abz-pro"] = { "Abkhaz-Abaza Purba", 7251831, "cau-abz", "Latn", type = "reconstructed", } m["cau-and-pro"] = { "Andi Purba", nil, "cau-and", "Latn", type = "reconstructed", } m["cau-ava-pro"] = { "Avar-Andi Purba", 116773187, "cau-ava", "Latn", type = "reconstructed", } m["cau-cir-pro"] = { "Circassia Purba", 7251838, "cau-cir", "Latn", type = "reconstructed", } m["cau-drg-pro"] = { "Dargwa Purba", 116773205, "cau-drg", "Latn", type = "reconstructed", } m["cau-lzg-pro"] = { "Lezghi Purba", 116773223, "cau-lzg", "Latn", type = "reconstructed", } m["cau-nec-pro"] = { "Kaukasia Timur Laut Purba", 116773244, "cau-nec", "Latn", type = "reconstructed", } m["cau-nkh-pro"] = { "Nakh Purba", 108032840, "cau-nkh", "Latn", type = "reconstructed", } m["cau-nwc-pro"] = { "Kaukasia Barat Laut Purba", 7251861, "cau-nwc", "Latn", type = "reconstructed", } m["cau-tsz-pro"] = { "Tsez Purba", 116773287, "cau-tsz", "Latn", type = "reconstructed", } m["cba-ata"] = { "Atanques", 4812783, "cba", "Latn", } m["cba-cat"] = { "Catío Chibcha", 7083619, "cba", "Latn", } m["cba-dor"] = { "Dorasque", 5297532, "cba", "Latn", } m["cba-dui"] = { "Duit", 3041061, "cba", "Latn", } m["cba-hue"] = { "Huetar", 35514, "cba", "Latn", } m["cba-nut"] = { "Nutabe", 7070405, "cba", "Latn", } m["cba-pro"] = { "Chibchan Purba", 116773203, "cba", "Latn", type = "reconstructed", } m["ccs-pro"] = { "Kartvelia Purba", 2608203, "ccs", "Latn", type = "reconstructed", strip_diacritics = { from = {"q̣", "p̣", "ʓ", "ċ"}, to = {"q̇", "ṗ", "ʒ", "c̣"} }, } m["ccs-gzn-pro"] = { "Georgia-Zan Purba", 23808119, "ccs-gzn", "Latn", type = "reconstructed", strip_diacritics = { from = {"q̣", "p̣", "ʓ", "ċ"}, to = {"q̇", "ṗ", "ʒ", "c̣"} }, } m["cdc-cbm-pro"] = { "Chadik Tengah Purba", 116773197, "cdc-cbm", "Latn", type = "reconstructed", } m["cdc-mas-pro"] = { "Masa Purba", 116773789, "cdc-mas", "Latn", type = "reconstructed", } m["cdc-pro"] = { "Chadik Purba", 116773201, "cdc", "Latn", type = "reconstructed", } m["cdd-pro"] = { "Caddoan Purba", 116773725, "cdd", "Latn", type = "reconstructed", } m["cel-bry-pro"] = { "Britonik Purba", 1248800, "cel-bry", "Latn, Polyt", sort_key = { Latn = "cel-bry-pro-sortkey", }, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["cel-gal"] = { "Gallaecia", 3094789, "cel-his", } m["cel-gau"] = { "Gaul", 29977, "cel", "Latn, Polyt, Ital", strip_diacritics = { Latn = {remove_diacritics = c.macron .. c.breve .. c.diaer}, }, sort_key = { Latn = "cel-bry-pro-sortkey", }, -- translit Ital dalam [[Module:scripts/data]] -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["cel-pro"] = { "Keltik Purba", 653649, "cel", "Latn", type = "reconstructed", sort_key = "cel-pro-sortkey", } m["chi-pro"] = { "Chimakuan Purba", 116773734, "chi", "Latn", type = "reconstructed", } m["chm-pro"] = { "Mari Purba", 116773788, "chm", "Latn", type = "reconstructed", } m["cmc-pro"] = { "Chamik Purba", 114793834, "cmc", "Latn", type = "reconstructed", } m["crp-bip"] = { "Pijin Basque-Iceland", 810378, "crp", "Latn", ancestors = "eu", } m["crp-cpr"] = { "Pijin Rusia-China", nil, "crp", "Hani, Cyrl, Latn", ancestors = "ru, zh", translit = {Cyrl = "ru-translit"}, strip_diacritics = { Cyrl = {remove_diacritics = c.acute .. c.grave .. c.macron}, }, } m["crp-gep"] = { "Pijin Greenland Barat", 17036301, "crp", "Latn", ancestors = "kl", } m["crp-kia"] = { "Pijin Jerman Kiautschou", 108314615, "crp", "Latn", ancestors = "de", } m["crp-mar"] = { "Bahasa Roh Maroon", 1093206, "crp", "Latn", ancestors = "en", } m["crp-mpp"] = { "Pijin Portugis Macau", 128804537, "crp", "Hant, Latn", ancestors = "pt", sort_key = {Hant = "Hani-sortkey"}, } m["crp-rsn"] = { "Russenorsk", 505125, "crp", "Cyrl, Latn", ancestors = "nn, ru", translit = {Cyrl = "ru-translit"}, } m["crp-spp"] = { "Pijin Ladang Samoa", 7409948, "crp", "Latn", ancestors = "en", } m["crp-slb"] = { "Inggeris Solombala", 7558525, "crp", "Cyrl, Latn", ancestors = "en, ru", translit = {Cyrl = "ru-translit"}, } m["crp-tpr"] = { "Pijin Rusia Taimyr", 16930506, "crp", "Cyrl", ancestors = "ru", translit = "ru-translit", } m["csu-bba-pro"] = { "Bongo-Bagirmi Purba", 116773722, "csu-bba", "Latn", type = "reconstructed", } m["csu-maa-pro"] = { "Mangbetu Purba", 116773786, "csu-maa", "Latn", type = "reconstructed", } m["csu-pro"] = { "Sudan Tengah Purba", 116773730, "csu", "Latn", type = "reconstructed", } m["csu-sar-pro"] = { "Sara Purba", 116773809, "csu-sar", "Latn", type = "reconstructed", } m["cus-ash"] = { "Ashraaf", 4805855, "cus-som", "Latn", } m["cus-hec-pro"] = { "Kusyi Timur Tanah Tinggi Purba", 116773761, "cus-hec", "Latn", type = "reconstructed", } m["cus-som-pro"] = { "Somaloid Purba", nil, "cus-som", "Latn", type = "reconstructed", } m["cus-sou-pro"] = { "Kusyi Selatan Purba", 126081567, "cus-sou", "Latn", type = "reconstructed", } m["cus-pro"] = { "Kusyi Purba", 116773204, "cus", "Latn", type = "reconstructed", } m["dmn-dam"] = { "Dama (Sierra Leone)", 19601574, "dmn", "Latn", } m["dra-bry"] = { "Beary", 1089116, "qfa-mix", "Mlym, Knda", ancestors = "ml, tcy", -- translit Knda dalam [[Module:scripts/data]] -- translit Mlym dalam [[Module:scripts/data]] } m["dra-cen-pro"] = { "Dravidia Tengah Purba", nil, "dra-cen", "Latn", type = "reconstructed", } m["dra-mkn"] = { "Kannada Pertengahan", 128810572, "dra-kan", "Knda", -- translit Knda dalam [[Module:scripts/data]] } m["dra-nor-pro"] = { "Dravidia Utara Purba", 124433593, "dra-nor", "Latn", type = "reconstructed", } m["dra-okn"] = { "Kannada Kuno", 15723156, "dra-kan", "Knda", -- translit Knda dalam [[Module:scripts/data]] } m["dra-ote"] = { "Telugu Kuno", 126720868, "dra-tel", "Telu", translit = "te-translit", } m["dra-pro"] = { "Dravidia Purba", 1702853, "dra", "Latn", type = "reconstructed", } m["dra-sdo-pro"] = { "Dravidia Selatan I Purba", 104847952, -- "Proto-Dravidia Selatan" Wikipedia ialah Proto-Dravidia Selatan I dalam skema ini. "dra-sdo", "Latn", type = "reconstructed", } m["dra-sdt-pro"] = { "Dravidia Selatan II Purba", 128885257, "dra-sdt", "Latn", type = "reconstructed", } m["dra-sou-pro"] = { "Dravidia Selatan Purba", 128886121, "dra-sou", "Latn", type = "reconstructed", } m["egx-dem"] = { "Mesir Demotik", 36765, "egx", "Latn, Egyd, Polyt", sort_key = { Latn = { remove_diacritics = "'%-%s", from = {"ꜣ", "j", "e", "ꜥ", "y", "w", "b", "p", "f", "m", "n", "r", "l", "ḥ", "ḫ", "h̭", "ẖ", "h", "š", "s", "q", "k", "g", "ṱ", "ṯ", "t", "ḏ", "%.", "⸗"}, to = {p[1], p[2], p[3], p[4], p[5], p[6], p[7], p[8], p[9], p[10], p[11], p[12], p[13], p[15], p[16], p[16], p[17], p[14], p[19], p[18], p[20], p[21], p[22], p[23], p[24], p[23], p[25], p[26], p[26]} }, }, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["dmn-pro"] = { "Mande Purba", 116773785, "dmn", "Latn", type = "reconstructed", } m["dmn-mdw-pro"] = { "Mande Barat Purba", 116773822, "dmn-mdw", "Latn", type = "reconstructed", } m["dru-pro"] = { "Rukai Purba", 116773807, "map", "Latn", type = "reconstructed", } m["ero-gsz"] = { "Geshiza", nil, "ero", "Latn", } m["ero-nya"] = { "Nyagrong Minyag", nil, "ero", "Latn", } m["ero-tau"] = { "Stau", nil, "ero", "Latn", } m["esx-esk-pro"] = { "Eskimo Purba", 7251842, "esx-esk", "Latn", type = "reconstructed", } m["esx-ink"] = { "Inuktun", 1671647, "esx-inu", "Latn", } m["esx-inq"] = { "Inuinnaqtun", 28070, "esx-inu", "Latn", } m["esx-inu-pro"] = { "Inuit Purba", 60785588, "esx-inu", "Latn", type = "reconstructed", } m["esx-pro"] = { "Eskimo-Aleut Purba", 7251843, "esx", "Latn", type = "reconstructed", } m["esx-tut"] = { "Tunumiisut", 15665389, "esx-inu", "Latn", } m["euq-pro"] = { "Basque Purba", 938011, "euq", "Latn", type = "reconstructed", } m["gba-pro"] = { "Gbaya Purba", nil, "gba", "Latn", type = "reconstructed", } m["gem-pro"] = { "Jermanik Purba", 669623, "gem", "Latn", type = "reconstructed", sort_key = "gem-pro-sortkey", } m["gme-bur"] = { "Burgundia", 47625, "gme", "Latn", } m["gme-cgo"] = { "Goth Crimea", 36211, "gme", "Latn", } m["gmq-gut"] = { "Gutnish", 1256646, "gmq", "Latn", ancestors = "gmq-ogt", } m["gmq-jmk"] = { "Jamtish", 35512, "gmq-eas", "Latn", } m["gmq-mno"] = { "Norway Pertengahan", 3417070, "gmq-wes", "Latn", } m["gmq-oda"] = { "Denmark Kuno", 12330003, "gmq-eas", "Latn, Runr", strip_diacritics = {remove_diacritics = c.macron}, } m["gmq-ogt"] = { "Gutnish Kuno", 1133488, "gmq", "Latn, Runr", ancestors = "non", } m["gmq-osw"] = { "Sweden Kuno", 2417210, "gmq-eas", "Latn, Runr", strip_diacritics = {remove_diacritics = c.macron}, } m["gmq-pro"] = { "Norse Purba", 1671294, "gmq", "Runr", translit = "Runr-translit", } m["gmq-scy"] = { "Scanian", 768017, "gmq-eas", "Latn", } m["gmw-bgh"] = { "Bergish", 329030, "gmw-frk", "Latn", } m["gmw-cfr"] = { "Franconia Tengah", 572197, "gmw-hgm", "Latn", ancestors = "gmh", wikimedia_codes = "ksh", } m["gmw-ecg"] = { "Jerman Tengah Timur", 499344, -- merangkumi Q699284, Q152965 "gmw-hgm", "Latn", ancestors = "gmh", } m["gmw-fin"] = { "Fingallian", 3072588, "gmw-ian", "Latn", } m["gmw-gts"] = { "Gottscheerish", 533109, "gmw-hgm", "Latn", ancestors = "bar", } m["gmw-jdt"] = { "Belanda Jersey", 1687911, "gmw-frk", "Latn", ancestors = "nl", } m["gmw-msc"] = { "Scots Pertengahan", 3327000, "gmw-ang", "Latn", ancestors = "enm-esc", } m["gmw-pro"] = { "Jermanik Barat Purba", 78079021, "gmw", "Latn, Runr", -- type = "reconstructed", -- sebahagian besarnya tetapi tidak sepenuhnya direkonstruksi (seperti Proto-Norse); lihat BP Apr '24, tetapkan kembali kepada direkonstruksi (?) jika 'anti-asterisk' ditambah sort_key = "gmw-pro-sortkey", } m["gmw-rfr"] = { "Franconia Rhine", 707007, "gmw-hgm", "Latn", ancestors = "gmh", } m["gmw-stm"] = { "Schwaben Szatmár", 2223059, "gmw-hgm", "Latn", ancestors = "swg", } m["gmw-tsx"] = { "Saxon Transylvania", 260942, "gmw-hgm", "Latn", ancestors = "gmw-cfr", } m["gmw-vog"] = { "Jerman Volga", 312574, "gmw-hgm", "Latn", ancestors = "gmw-rfr", } m["gmw-zps"] = { "Jerman Zipser", 205548, "gmw-hgm", "Latn", ancestors = "gmh", } m["gn-cls"] = { "Guarani Klasik", 17478065, "gn", "Latn", } m["grk-cal"] = { "Yunani Calabria", 1146398, "grk", "Latn, Grek", ancestors = "grk-ita", translit = { Grek = "el-translit", }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-ita"] = { "Yunani Italiot", 19720507, "grk", "Latn, Grek", ancestors = "gkm", translit = { Grek = "el-translit", }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-mar"] = { "Yunani Mariupol", 4400023, "grk", "Cyrl, Latn, Grek", ancestors = "gkm", translit = { Cyrl = "grk-mar-translit", Grek = "grk-mar-translit", }, override_translit = true, strip_diacritics = { Cyrl = {remove_diacritics = c.acute}, }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-pro"] = { "Hellenik Purba", 1231805, "grk", "Latn, Polyt", type = "reconstructed", sort_key = {Latn = { from = {"ʰ", "ʷ"}, to = {"h", "w"}, remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.caron .. c.CGJ }}, display_text = {Latn = { from = {"([dlLt])" .. c.caron}, to = {"%1" .. c.CGJ .. c.caron}, }}, strip_diacritics = {Latn = { from = {"([dlLt])" .. c.caron}, to = {"%1" .. c.CGJ .. c.caron}, }}, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] -- NOTA: dahulunya tiada translit ditentukan untuk Polyt; mungkin ketinggalan secara tidak sengaja; jika tidak, tetapkan Polyt = false dalam -- bahagian translit } m["hmn-pro"] = { "Hmongik Purba", 116773210, "hmn", "Latn", type = "reconstructed", } m["hmx-mie-pro"] = { "Mienik Purba", 116773229, "hmx-mie", "Latn", type = "reconstructed", } m["hmx-pro"] = { "Hmong-Mien Purba", 7251846, "hmx", "Latn", type = "reconstructed", } m["hyx-pro"] = { "Armenia Purba", 3848498, "hyx", "Latn", type = "reconstructed", } m["iir-nur-pro"] = { "Nuristani Purba", 116773248, "iir-nur", "Latn", type = "reconstructed", } m["iir-pro"] = { "Indo-Iran Purba", 966439, "iir", "Latn", type = "reconstructed", } m["ijo-pro"] = { "Ijoid Purba", 116773766, "ijo", "Latn", type = "reconstructed", } m["inc-apa"] = { "Apabhramsa", 616419, "inc-mid", "Deva, Shrd, Sidd", ancestors = "pra", translit = { Deva = "sa-translit", -- translit Shrd dalam [[Module:scripts/data]] -- translit Sidd dalam [[Module:scripts/data]] }, } m["inc-ash"] = { "Prakrit Ashoka", 104854379, "inc-mid", "Brah, Khar", ancestors = "sa", translit = { -- translit Brah dalam [[Module:scripts/data]] Khar = "Khar-translit", }, } m["inc-dng-pro"] = { "Dangari Purba", nil, "inc-dng", "Latn", type = "reconstructed", } m["inc-kam"] = { "Prakrit Kamarupi", 6356097, "inc-bas", "Brah, Sidd", -- translit Brah, Sidd dalam [[Module:scripts/data]] } m["inc-kho"] = { "Kholosi", 24952008, "inc-snd", "Latn", } m["inc-khr"] = { "Khortha", 13406670, "inc-sad", "Deva, Kthi", translit = { Deva = "bho-translit", Kthi = "bho-Kthi-translit", }, } m["inc-krd-pro"] = { "Kamta Purba", 128816843, "inc-bas", "Latn", ancestors = "inc-kam", type = "reconstructed", } m["inc-mas"] = { "Assam Pertengahan", 128806836, "inc-bas", "as-Beng", ancestors = "inc-oas", translit = "inc-mas-translit", } m["inc-mbn"] = { "Benggali Pertengahan", 113559927, "inc-bas", "Beng", ancestors = "inc-obn", translit = "inc-mbn-translit", } m["inc-mgu"] = { "Gujarati Pertengahan", 24907429, "inc-wes", "Deva", ancestors = "inc-ogu", } m["inc-mor"] = { "Odia Pertengahan", 128810882, "inc-eas", "Orya", ancestors = "inc-oor", } m["inc-oas"] = { "Assam Awal", 85758237, "inc-bas", "as-Beng", ancestors = "inc-kam", translit = "inc-oas-translit", } m["inc-oaw"] = { "Awadhi Kuno", nil, "inc-hie", "Deva, Kthi, Aran", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Deva = "sa-translit", Kthi = "sa-Kthi-translit", Aran = "inc-ohi-translit", }, } m["inc-obn"] = { "Benggali Kuno", 113559926, "inc-bas", "Beng", } m["inc-ogu"] = { "Gujarati Kuno", 24907427, "inc-wes", "Deva", translit = "sa-translit", } m["inc-ohi"] = { "Hindi Kuno", 48767781, "inc-hiw", "Deva, Aran", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Deva = "sa-translit", Aran = "inc-ohi-translit", }, } m["inc-oor"] = { "Odia Kuno", 128807801, "inc-eas", "Orya", } m["inc-opa"] = { "Punjabi Kuno", 115270971, "inc-pan", "Guru, Aran", translit = { Guru = "inc-opa-Guru-translit", Aran = "pa-Aran-translit", }, strip_diacritics = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun}, } m["inc-pro"] = { "Indo-Arya Purba", 23808344, "inc", "Latn", type = "reconstructed", } m["inc-sar"] = { "Sarazi", 85799728, "him", "Aran, Deva, Takr", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Aran = "ur-translit", Deva = "hi-translit", -- Takr = "Takr-translit", }, } m["ine-ana-pro"] = { "Anatolia Purba", 7251833, "ine-ana", "Latn", type = "reconstructed", } m["ine-bsl-pro"] = { "Balto-Slavik Purba", 1703347, "ine-bsl", "Latn", type = "reconstructed", sort_key = { from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", c.acute, c.macron, "ˀ"}, to = {"a", "e", "i", "o", "u"} }, } m["ine-kal"] = { "Kalašma", 122770439, "ine-ana", "Xsux", } m["ine-pae"] = { "Paeonia", 2705672, "ine", "Polyt", -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["ine-pro"] = { "Indo-Eropah Purba", 37178, "ine", "Latn", type = "reconstructed", sort_key = { from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", "ĺ", "ḿ", "ń", "ŕ", "ǵ", "ḱ", "ʰ", "ʷ", "₁", "₂", "₃", c.ringbelow, c.acute, c.macron}, to = {"a", "e", "i", "o", "u", "l", "m", "n", "r", "g'", "k'", "¯h", "¯w", "1", "2", "3"} }, } m["ine-toc-pro"] = { "Tocharia Purba", 104841462, "ine-toc", "Latn", type = "reconstructed", } m["xme-old"] = { "Median Kuno", 36461, "xme", "Polyt, Latn", -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["xme-mid"] = { "Median Pertengahan", 12836150, "xme", "Latn", } m["xme-ker"] = { "Kerman", 129850, "xme", "Arab, Latn, Hebr", ancestors = "xme-mid", -- display_text, strip_diacritics, sort_key Hebr dalam [[Module:scripts/data]] } m["xme-taf"] = { "Tafreshi", nil, "xme", "Arab, Latn", ancestors = "xme-mid", } m["xme-ttc-pro"] = { "Tatik Purba", 122973870, "xme-ttc", "Latn", ancestors = "xme-mid", } m["xme-kls"] = { "Kalasuri", nil, "xme-ttc", ancestors = "xme-ttc-nor", } m["xme-klt"] = { "Kilit", 3612452, "xme-ttc", "Cyrl", -- dan Arab? } m["xme-ott"] = { "Tati Kuno", 434697, "xme-ttc", "Arab, Latn", } m["ira-kms-pro"] = { "Komisenia Purba", 116773777, "ira-kms", "Latn", type = "reconstructed", } m["ira-mpr-pro"] = { "Medo-Parthia Purba", 116773227, "ira-mpr", "Latn", type = "reconstructed", } m["ira-pat-pro"] = { "Pathan Purba", 116773255, "ira-pat", "Latn", type = "reconstructed", } m["ira-pro"] = { "Iran Purba", 4167865, "ira", "Latn", type = "reconstructed", } m["ira-zgr-pro"] = { "Zaza-Gorani Purba", 116775031, "ira-zgr", "Latn", type = "reconstructed", } m["xsc-pro"] = { "Scythia Purba", 116773273, "xsc", "Latn", type = "reconstructed", } m["xsc-sar-pro"] = { "Sarmatia Purba", 116773249, "xsc-sar", "Latn", type = "reconstructed", } m["xsc-skw-pro"] = { "Saka-Wakhi Purba", 116773267, "xsc-skw", "Latn", type = "reconstructed", } m["xsc-sak-pro"] = { "Saka Purba", 116773264, "xsc-sak", "Latn", type = "reconstructed", } m["ira-sym-pro"] = { "Shughni-Yazghulami-Munji Purba", 116773813, "ira-sym", "Latn", type = "reconstructed", } m["ira-sgi-pro"] = { "Sanglechi-Ishkashimi Purba", 116773808, "ira-sgi", "Latn", type = "reconstructed", } m["ira-mny-pro"] = { "Munji-Yidgha Purba", 116773792, "ira-mny", "Latn", type = "reconstructed", } m["ira-shy-pro"] = { "Shughni-Yazghulami Purba", 116773812, "ira-shy", "Latn", type = "reconstructed", } m["ira-shr-pro"] = { "Shughni-Roshani Purba", 116773811, "ira-shr", "Latn", type = "reconstructed", } m["ira-sgc-pro"] = { "Sogdia Purba", 116773276, "ira-sgc", "Latn", type = "reconstructed", } m["ira-wnj"] = { "Vanji", 3398419, "ira-shy", "Latn", } m["iro-ere"] = { "Erie", 5388365, "iro-nor", "Latn", } m["iro-min"] = { "Mingo", 128531, "iro-nor", "Latn", ietf_subtag = "i-mingo", -- tag IETF yang diwarisi } m["iro-nor-pro"] = { "Iroquois Utara Purba", 116773242, "iro-nor", "Latn", type = "reconstructed", } m["iro-pro"] = { "Iroquois Purba", 7251852, "iro", "Latn", type = "reconstructed", } m["itc-pro"] = { "Italik Purba", 17102720, "itc", "Latn", type = "reconstructed", } m["itc-psa"] = { "Pra-Samnit", 7239186, "itc-sbl", "Ital, Polyt, Latn", -- translit Ital dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja) -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["jpx-hcj"] = { "Hachijō", 5637049, "jpx", "Jpan", ancestors = "ojp-eas", translit = s["jpx-translit"], display_text = s["jpx-displaytext"], strip_diacritics = s["jpx-stripdiacritics"], sort_key = s["jpx-sortkey"], } m["jpx-pro"] = { "Jepunik Purba", 3924309, "jpx", "Latn", type = "reconstructed", } m["jpx-ryu-pro"] = { "Ryukyu Purba", 56349069, "jpx-ryu", "Latn", type = "reconstructed", } m["kar-pro"] = { "Karen Purba", 85794783, "kar", "Latn", type = "reconstructed", } m["kca-eas"] = { "Khanty Timur", 30304622, "kca", "Cyrl", translit = "kca-translit", override_translit = true, -- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka) sort_key = { Cyrl = { from = {"ᲊ"}, to = {"Ᲊ"} } }, } m["kca-nor"] = { "Khanty Utara", 30304527, "kca", "Cyrl", translit = "kca-translit", override_translit = true, -- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka) sort_key = { Cyrl = { from = {"ᲊ"}, to = {"Ᲊ"} } }, } m["kca-pro"] = { "Khanty Purba", 127505171, "kca", "Latn", type = "reconstructed", } m["kca-sou"] = { "Khanty Selatan", 30304618, "kca", "Cyrl", translit = "kca-translit", override_translit = true, } m["khi-kho-pro"] = { "Khoe Purba", 116773218, "khi-kho", "Latn", type = "reconstructed", } m["khi-kun"] = { "ǃKung", 32904, "khi-kxa", "Latn", } m["ko-ear"] = { "Korea Moden Awal", 756014, "qfa-kor", "Kore", ancestors = "okm", translit = "okm-translit", -- strip_diacritics Kore dalam [[Module:scripts/data]] } m["kro-pro"] = { "Kru Purba", 116773778, "kro", "Latn", type = "reconstructed", } m["ku-pro"] = { "Kurdi Purba", 116773221, "ku", "Latn", type = "reconstructed", } m["map-ata-pro"] = { "Atayalik Purba", 116773151, "map-ata", "Latn", type = "reconstructed", } m["map-bms"] = { "Banyumasan", 33219, "map", "Latn, Java", } m["map-pro"] = { "Austronesia Purba", 49230, "map", "Latn", type = "reconstructed", } m["mis-hkl"] = { "Hokkien Peranakan Kelantan", 108794818, "qfa-mix", ancestors = "nan-hbl, sou, mfa", } m["mis-idn"] = { "Idiom Neutral", 35847, "art", "Latn", type = "appendix-constructed", } m["mis-isa"] = { "Isauria", 16956868, nil, -- "Xsux, Hluw, Latn", } m["mis-jie"] = { "Jie", 124424186, nil, "Hani", sort_key = "Hani-sortkey", } m["mis-jzh"] = { "Jizhao", 45242758, "qfa-bej", "Latn", } m["mis-kas"] = { "Kassite", 35612, nil, "Xsux", } m["mis-mmd"] = { "Mimi Decorse", 6862206, nil, "Latn", } m["mis-mmn"] = { "Mimi Nachtigal", 6862207, nil, "Latn", } m["mis-phi"] = { "Filistin", 2230924, nil, "Phnx", -- translit Phnx dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja) } m["mis-rou"] = { "Rouran", 48816637, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-tdl"] = { "Turdulia", 133176492, } m["mis-tdt"] = { "Turdetania", 133176461, } m["mis-tnw"] = { "Tangwang", 7683179, "qfa-mix", "Latn", ancestors = "cmn, sce", } m["mis-tuh"] = { "Tuyuhun", 48816625, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-tuo"] = { "Tuoba", 48816629, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-wuh"] = { "Wuhuan", 118976867, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-xbi"] = { "Xianbei", 4448647, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-xnu"] = { "Xiongnu", 10901674, nil, "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mjg-mgl"] = { "Mongghul", 53765528, "mjg", "Latn", -- juga Mong, Cyrl? } m["mjg-mgr"] = { "Mangghuer", 56285392, "mjg", "Latn", -- juga Mong, Cyrl? } m["mkh-asl-pro"] = { "Asli Purba", 55630680, "mkh-asl", "Latn", type = "reconstructed", } m["mkh-ban-pro"] = { "Bahnar Purba", 116773189, "mkh-ban", "Latn", type = "reconstructed", } m["mkh-kat-pro"] = { "Katuik Purba", 116773772, "mkh-kat", "Latn", type = "reconstructed", } m["mkh-khm-pro"] = { "Khmuik Purba", 116773774, "mkh-khm", "Latn", type = "reconstructed", } m["mkh-kmr-pro"] = { "Khmer Purba", 55630684, "mkh-kmr", "Latn", type = "reconstructed", } m["mkh-mmn"] = { "Mon Pertengahan", 121337926, "mkh-mnc", "Latn, Mymr", --dan juga Pallava ancestors = "omx", } m["mkh-mnc-pro"] = { "Monik Purba", 116773231, "mkh-mnc", "Latn", type = "reconstructed", } m["mkh-mvi"] = { "Vietnam Pertengahan", 9199, "mkh-vie", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mkh-pal-pro"] = { "Palaungik Purba", 104847372, "mkh-pal", "Latn", type = "reconstructed", } m["mkh-pea-pro"] = { "Pearik Purba", 116773804, "mkh-pea", "Latn", type = "reconstructed", } m["mkh-pkn-pro"] = { "Pakanik Purba", 116773803, "mkh-pkn", "Latn", type = "reconstructed", } m["mkh-pro"] = { --Ini akan digabungkan ke dalam aav-pro 2015. "Mon-Khmer Purba", 7251859, "mkh", "Latn", type = "reconstructed", } m["mnw-tha"] = { -- Untuk dibuang. "Mon Thailand", nil, "mkh-mnc", "Mymr, Thai", ancestors = "mkh-mmn", sort_key = { from = {"[%p]", "ျ", "ြ", "ွ", "ှ", "ၞ", "ၟ", "ၠ", "ၚ", "ဿ", "[็-๎]", "([เแโใไ])([ก-ฮ])ฺ?"}, to = {"", "္ယ", "္ရ", "္ဝ", "္ဟ", "္န", "္မ", "္လ", "င", "သ္သ", "", "%2%1"} }, } m["mkh-vie-pro"] = { "Vietik Purba", 109432616, "mkh-vie", "Latn", type = "reconstructed", } m["mns-cen"] = { "Mansi Tengah", 128810384, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mns-nor"] = { "Mansi Utara", 30304537, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mns-pro"] = { "Mansi Purba", 128883093, "mns", "Latn", type = "reconstructed", } m["mns-sou"] = { "Mansi Selatan", 30304629, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mun-pro"] = { "Munda Purba", 105102373, "mun", "Latn", type = "reconstructed", } m["myn-chl"] = { -- peringkat selepas ''emy'' "Ch'olti'", 873995, "myn", "Latn", } m["myn-pro"] = { "Maya Purba", 3321532, "myn", "Latn", type = "reconstructed", } m["nai-ala"] = { "Alazapa", 128810233, nil, "Latn", } m["nai-bay"] = { "Bayogoula", 1563704, nil, "Latn", } m["nai-cal"] = { "Calusa", 51782, nil, "Latn", } m["nai-chi"] = { "Chiquimulilla", 25339627, "nai-xin", "Latn", } m["nai-chu-pro"] = { "Chumash Purba", 116773736, "nai-chu", "Latn", type = "reconstructed", } m["nai-cig"] = { "Ciguayo", 20741700, nil, "Latn", } m["nai-ckn-pro"] = { "Chinook Purba", 116773735, "nai-ckn", "Latn", type = "reconstructed", } m["nai-guz"] = { "Guazacapán", 19572028, "nai-xin", "Latn", } m["nai-hit"] = { "Hitchiti", 1542882, "nai-mus", "Latn", } m["nai-ipa"] = { "Ipai", 3027474, "nai-yuc", "Latn", } m["nai-jtp"] = { "Jutiapa", nil, "nai-xin", "Latn", } m["nai-jum"] = { "Jumaytepeque", 25339626, "nai-xin", "Latn", } m["nai-kat"] = { "Kathlamet", 6376639, "nai-ckn", "Latn", } m["nai-klp-pro"] = { "Kalapuya Purba", 116773771, "nai-klp", "Latn", type = "reconstructed", } m["nai-knm"] = { "Konomihu", 3198734, "nai-shs", "Latn", } m["nai-kum"] = { "Kumeyaay", 4910139, "nai-yuc", "Latn", } m["nai-mac"] = { "Macoris", 21070851, nil, "Latn", } m["nai-mdu-pro"] = { "Maidu Purba", 116773784, "nai-mdu", "Latn", type = "reconstructed", } m["nai-miz-pro"] = { "Mixe-Zoque Purba", 7251858, "nai-miz", "Latn", type = "reconstructed", } m["nai-mus-pro"] = { "Muskogi Purba", 116775368, "nai-mus", "Latn", type = "reconstructed", } m["nai-nao"] = { "Naolan", 6964594, nil, "Latn", } m["nai-nrs"] = { "Shasta Sungai Baru", 7011254, "nai-shs", "Latn", } m["nai-okw"] = { "Okwanuchu", 3350126, "nai-shs", "Latn", } m["nai-per"] = { "Pericú", 3375369, nil, "Latn", } m["nai-pic"] = { "Picuris", 7191257, "nai-kta", "Latn", } m["nai-plp-pro"] = { "Penuti Penara Purba", 116773806, "nai-plp", "Latn", type = "reconstructed", } m["nai-pom-pro"] = { "Pomo Purba", 116773262, "nai-pom", "Latn", type = "reconstructed", } m["nai-qng"] = { "Quinigua", 36360, nil, "Latn", } m["nai-sca-pro"] = { -- PERHATIAN 'sio-pro' "Proto-Siouan" iaitu Proto-Sioux Barat "Siouan-Catawba Purba", 116773275, "nai-sca", "Latn", type = "reconstructed", } m["nai-sin"] = { "Sinacantán", 24190249, "nai-xin", "Latn", } m["nai-sln"] = { "Lenca Salvador", 3229434, "nai-len", "Latn", } m["nai-spt"] = { "Sahaptin", 3833015, "nai-shp", "Latn", } m["nai-tap"] = { "Tapachultec", 7684401, "nai-miz", "Latn", } m["nai-taw"] = { "Tawasa", 7689233, nil, "Latn", } m["nai-teq"] = { "Tequistlatec", 2964454, "nai-tqn", "Latn", } m["nai-tip"] = { "Tipai", 3027471, "nai-yuc", "Latn", } m["nai-tot-pro"] = { "Totozoquean Purba", 116773285, "nai-tot", "Latn", type = "reconstructed", } m["nai-tsi-pro"] = { "Tsimshianik Purba", nil, "nai-tsi", "Latn", type = "reconstructed", } m["nai-utn-pro"] = { "Utik Purba", 116773290, "nai-utn", "Latn", type = "reconstructed", } m["nai-wai"] = { "Waikuri", 3118702, nil, "Latn", } m["nai-wji"] = { "Jicaque Barat", 3178610, "nai-jcq", "Latn", } m["nai-yup"] = { "Yupiltepeque", 25339628, "nai-xin", "Latn", } m["nan-dat"] = { "Min Datian", 19855572, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-hbl"] = { "Hokkien", 1624231, "zhx-nan", "Hants, Latn, Bopo, Kana", wikimedia_codes = "zh-min-nan", generate_forms = "zh-generateforms", sort_key = { Hani = "Hani-sortkey", Kana = "Kana-sortkey" }, } m["nan-hlh"] = { "Min Hailufeng", 120755728, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-lnx"] = { "Min Longyan", 6674568, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-tws"] = { "Teochew", 36759, "zhx-nan", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["nan-zhe"] = { "Min Zhenan", 3846710, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-zsh"] = { "Min Sanxiang", 7420769, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["ngf-bin-pro"] = { "Binandere Purba", 137881672, "ngf-bin", "Latn", type = "reconstructed", } m["ngf-pro"] = { "Trans-New Guinea Purba", 85794785, "ngf", "Latn", type = "reconstructed", } m["nic-bco-pro"] = { "Benue-Congo Purba", 116773194, "nic-bco", "Latn", type = "reconstructed", } m["nic-bod-pro"] = { "Bantoid Purba", 116773190, "nic-bod", "Latn", type = "reconstructed", } m["nic-eov-pro"] = { "Oti-Volta Timur Purba", 116773753, "nic-eov", "Latn", type = "reconstructed", } m["nic-gns-pro"] = { "Gurunsi Purba", 116773759, "nic-gns", "Latn", type = "reconstructed", } m["nic-grf-pro"] = { "Grassfields Purba", 116773755, "nic-grf", "Latn", type = "reconstructed", } m["nic-gur-pro"] = { "Gur Purba", 116773758, "nic-gur", "Latn", type = "reconstructed", } m["nic-jkn-pro"] = { "Jukunoid Purba", 116773769, "nic-jkn", "Latn", type = "reconstructed", } m["nic-lcr-pro"] = { "Cross River Hilir Purba", 116773782, "nic-lcr", "Latn", type = "reconstructed", } m["nic-ogo-pro"] = { "Ogoni Purba", 116773799, "nic-ogo", "Latn", type = "reconstructed", } m["nic-ovo-pro"] = { "Oti-Volta Purba", 116773802, "nic-ovo", "Latn", type = "reconstructed", } m["nic-plt-pro"] = { "Plateau Purba", 116773805, "nic-plt", "Latn", type = "reconstructed", } m["nic-pro"] = { "Niger-Congo Purba", 108000748, "nic", "Latn", type = "reconstructed", } m["nic-ubg-pro"] = { "Ubangi Purba", 116773818, "nic-ubg", "Latn", type = "reconstructed", } m["nic-ucr-pro"] = { "Cross River Hulu Purba", 116773819, "nic-ucr", "Latn", type = "reconstructed", } m["nic-vco-pro"] = { "Volta-Congo Purba", 116773293, "nic-vco", "Latn", type = "reconstructed", } m["njo-jgl"] = { "Ao Chungli", 55607615, "njo", "Latn", } m["njo-mng"] = { "Ao Mongsen", 85383221, "njo", "Latn", } m["nub-har"] = { "Haraza", 19572059, "nub", "Arab, Latn", } m["nub-pro"] = { "Nubia Purba", 116773246, "nub", "Latn", type = "reconstructed", } m["omq-cha-pro"] = { "Chatino Purba", 116773202, "omq-cha", "Latn", type = "reconstructed", } m["omq-maz-pro"] = { "Mazatec Purba", 116773790, "omq-maz", "Latn", type = "reconstructed", } m["omq-mix-pro"] = { "Mixtecan Purba", 21573423, "omq-mix", "Latn", type = "reconstructed", } m["omq-mxt-pro"] = { "Mixtec Purba", 21573424, "omq-mxt", "Latn", type = "reconstructed", } m["omq-otp-pro"] = { "Oto-Pamean Purba", 116773251, "omq-otp", "Latn", type = "reconstructed", } m["omq-pro"] = { "Oto-Manguean Purba", 33669, "omq", "Latn", type = "reconstructed", } m["omq-sjq"] = { "Chatino San Juan Quiahije", 138330751, "omq-cha", "Latn", } m["omq-tel"] = { "Mixtec Teposcolula", nil, "omq-mxt", "Latn", } m["omq-teo"] = { "Chatino Teojomulco", 25340451, "omq-cha", "Latn", } m["omq-tri-pro"] = { "Triqui Purba", 116773817, "omq-tri", "Latn", type = "reconstructed", } m["omq-zap-pro"] = { "Zapotecan Purba", 116773297, "omq-zap", "Latn", type = "reconstructed", } m["omq-zpc-pro"] = { "Zapotec Purba", 116773296, "omq-zpc", "Latn", type = "reconstructed", } m["omv-aro-pro"] = { "Aroid Purba", 116773721, "omv-aro", "Latn", type = "reconstructed", } m["omv-diz-pro"] = { "Dizoid Purba", 116773750, "omv-diz", "Latn", type = "reconstructed", } m["omv-pro"] = { "Omotik Purba", 116773800, "omv", "Latn", type = "reconstructed", } m["oto-otm-pro"] = { "Otomi Purba", 5908710, "oto-otm", "Latn", type = "reconstructed", } m["oto-pro"] = { "Otomian Purba", 116773252, "oto", "Latn", type = "reconstructed", } m["paa-kmn"] = { "Kómnzo", 18344310, "paa-wko", "Latn", } m["paa-kwn"] = { "Kuwani", 6449056, "qfa-unc", -- kurang dibuktikan, mungkin sama dengan atau berkaitan dengan Kalabra "Latn", } m["paa-lei"] = { "Leitre", 85776228, "paa-isk", } m["paa-nha-pro"] = { "Halmahera Utara Purba", 116773241, "paa-nha", "Latn", type = "reconstructed" } m["paa-nun"] = { "Nungon", 128807788, "ngf-ynu", "Latn", } m["phi-din"] = { "Agta Dinapigue", 16945774, "phi", "Latn", } m["phi-kal-pro"] = { "Kalamian Purba", 116773213, "phi-kal", "Latn", type = "reconstructed", } m["phi-nag"] = { "Agta Nagtipunan", 16966111, "phi", "Latn", } m["phi-pro"] = { "Filipina Purba", 18204898, "phi", "Latn", type = "reconstructed", } m["poz-abi"] = { "Abai", 19570729, "poz-san", "Latn", } m["poz-bal"] = { "Baliledo", 4850912, "poz", "Latn", } m["poz-btk-pro"] = { "Bungku-Tolaki Purba", 116773724, "poz-btk", "Latn", type = "reconstructed", } m["poz-cet-pro"] = { "Melayu-Polinesia Tengah-Timur Purba", 2269883, "poz-cet", "Latn", type = "reconstructed", } m["poz-hce-pro"] = { "Halmahera-Cenderawasih Purba", 116773209, "poz-hce", "Latn", type = "reconstructed", } m["poz-lgx-pro"] = { "Lampung Purba", 116773222, "poz-lgx", "Latn", type = "reconstructed", } m["poz-mcm-pro"] = { "Melayu-Chamik Purba", 116773225, "poz-mcm", "Latn", type = "reconstructed", } m["poz-mic-pro"] = { "Mikronesia Purba", 111939079, "poz-mic", "Latn", type = "reconstructed", } m["poz-mly-pro"] = { "Melayik Purba", 98057728, "poz-mly", "Latn", type = "reconstructed", } m["poz-msa-pro"] = { "Melayu-Sumbawa Purba", 116773226, "poz-msa", "Latn", type = "reconstructed", } m["poz-nes"] = { "Nese", 2157412, "poz-vnc", "Latn", } m["poz-oce-pro"] = { "Oceania Purba", 141741, "poz-oce", "Latn", type = "reconstructed", } m["poz-pcc-pro"] = { "Pasifik Tengah Purba", 111962726, "poz-pcc", "Latn", type = "reconstructed", } m["poz-pep-pro"] = { "Polinesia Timur Purba", 113988745, "poz-pep", "Latn", type = "reconstructed", } m["poz-pnp-pro"] = { "Polinesia Teras Purba", 113988746, "poz-pnp", "Latn", type = "reconstructed", } m["poz-pol-pro"] = { "Polinesia Purba", 1658709, "poz-pol", "Latn", type = "reconstructed", } m["poz-pro"] = { "Melayu-Polinesia Purba", 3832960, "poz", "Latn", type = "reconstructed", } m["poz-sml"] = { "Melayu Sarawak", 4251702, "poz-mly", "Latn, Arab", } m["poz-ssw-pro"] = { "Sulawesi Selatan Purba", 116773279, "poz-ssw", "Latn", type = "reconstructed", } m["poz-swa-pro"] = { "Sarawak Utara Purba", 116773243, "poz-swa", "Latn", type = "reconstructed", } m["poz-ter"] = { "Melayu Terengganu", 4207412, "poz-mly", "Latn, Arab", } m["pqe-pro"] = { "Melayu-Polinesia Timur Purba", 2269883, "pqe", "Latn", type = "reconstructed", } m["pra-niy"] = { "Prakrit Niya", 11991601, "inc-mid", "Khar", ancestors = "inc-ash", translit = "Khar-translit", } m["qfa-adm-pro"] = { "Andaman Besar Purba", 116773756, "qfa-adm", "Latn", type = "reconstructed", } m["qfa-bet-pro"] = { "Be-Tai Purba", 116773193, "qfa-bet", "Latn", type = "reconstructed", } m["qfa-cka-pro"] = { "Chukotko-Kamchatka Purba", 7251837, "qfa-cka", "Latn", type = "reconstructed", } m["qfa-hur-pro"] = { "Hurro-Urartia Purba", 116773211, "qfa-hur", "Latn", type = "reconstructed", } m["qfa-kad-pro"] = { "Kadu Purba", 116773770, "qfa-kad", "Latn", type = "reconstructed", } m["qfa-kms-pro"] = { "Kam-Sui Purba", 55630682, "qfa-kms", "Latn", type = "reconstructed", } m["qfa-kor-pro"] = { "Korea Purba", 467883, "qfa-kor", "Latn", type = "reconstructed", } m["qfa-kra-pro"] = { "Kra Purba", 7251854, "qfa-kra", "Latn", type = "reconstructed", } m["qfa-lic-pro"] = { "Hlai Purba", 7251845, "qfa-lic", "Latn", type = "reconstructed", } m["qfa-onb-pro"] = { "Be Purba", 116773192, "qfa-onb", "Latn", type = "reconstructed", } m["qfa-ong-pro"] = { "Onga Purba", 116773801, "qfa-ong", "Latn", type = "reconstructed", } m["qfa-tak-pro"] = { "Kra-Dai Purba", 104901616, "qfa-tak", "Latn", type = "reconstructed", } m["qfa-yen-pro"] = { "Yenisei Purba", 27639, "qfa-yen", "Latn", type = "reconstructed", } m["qfa-yuk-pro"] = { "Yukaghir Purba", 116773294, "qfa-yuk", "Latn", type = "reconstructed", } m["qwe-kch"] = { "Kichwa", 1740805, "qwe", "Latn", ancestors = "qu", } m["qwe-pro"] = { "Quechua Purba", 5575757, "qwe", "Latn", type = "reconstructed", } m["roa-ang"] = { "Angevin", 56782, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-bbn"] = { "Bourbonnais-Berrichon", 2899128, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-brg"] = { "Bourguignon", 508332, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-can"] = { "Cantabrian", 917021, "roa-asl", "Latn", } m["roa-cha"] = { "Champenois", 430018, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-fcm"] = { "Franc-Comtois", 510561, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-gal"] = { "Gallo", 37300, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-gib"] = { "Gallo-Italik Basilicata", 3094838, "roa-git", ancestors = "pms-old", "Latn", } m["roa-gis"] = { "Gallo-Italik Sicily", 2629019, "roa-git", "Latn", ancestors = "pms-old", } m["roa-leo"] = { "Leon", 34108, "roa-asl", "Latn", } m["roa-lor"] = { "Lorrain", 671198, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-oca"] = { "Catalonia Kuno", 15478520, "roa-ocr", "Latn", sort_key = {remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla .. "·"}, } m["roa-ole"] = { "Leon Kuno", 125977465, "roa-asl", "Latn", } m["roa-ona"] = { "Navarro-Aragon Kuno", 2736184, "roa-nar", "Latn", } m["roa-opt"] = { "Galicia-Portugis Kuno", 1072111, "roa-gap", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ}, } m["roa-orl"] = { "Orléanais", 28497058, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-poi"] = { "Poitevin-Saintongeais", 514123, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-tar"] = { "Tarantino", 695526, "roa-itr", "Latn", wikimedia_codes = "roa-tara", } m["sai-all"] = { "Allentiac", 19570789, "sai-hrp", "Latn", } m["sai-and"] = { "Andoquero", 16828359, "sai-wit", "Latn", } m["sai-ayo"] = { "Ayomán", 16937754, "sai-jir", "Latn", } m["sai-bae"] = { "Baenan", 3401998, "qfa-unc", -- pupus, kurang dibuktikan; hanya dikenali melalui 9 perkataan "Latn", } m["sai-bag"] = { "Bagua", 5390321, "qfa-unc", -- pupus, kurang dibuktikan; mungkin bahasa Carib "Latn", } m["sai-bet"] = { "Betoi", 926551, "qfa-iso", "Latn", } m["sai-bor-pro"] = { "Bora Purba", nil, "sai-bor", "Latn", } m["sai-cac"] = { "Cacán", 945482, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan "Latn", } m["sai-caq"] = { "Caranqui", 2937753, "sai-bar", "Latn", } m["sai-car-pro"] = { "Carib Purba", 116773196, "sai-car", "Latn", type = "reconstructed", } m["sai-cat"] = { "Catacao", 5051136, "sai-ctc", "Latn", } m["sai-cer-pro"] = { "Cerrado Purba", 116773200, "sai-cer", "Latn", type = "reconstructed", } m["sai-chi"] = { "Chirino", 5390321, "qfa-unc", -- pupus, hanya empat perkataan diketahui; mungkin berkaitan dengan Candoshi-Shapra (cbu) "Latn", } m["sai-chn"] = { "Chaná", 5072718, "sai-crn", "Latn", } m["sai-chp"] = { "Chapacura", 5072884, "sai-cpc", "Latn", } m["sai-chr"] = { "Charrua", 5086680, "sai-crn", "Latn", } m["sai-chu"] = { "Churuya", 5118339, "sai-guh", "Latn", } m["sai-cje-pro"] = { "Jê Tengah Purba", 116773198, "sai-cje", "Latn", type = "reconstructed", } m["sai-cmg"] = { "Comechingon", 6644203, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan "Latn", } m["sai-cno"] = { "Chono", 5104704, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan, mungkin palsu "Latn", } m["sai-cnr"] = { "Cañari", 5055572, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Chimuan atau Barbacoan "Latn", } m["sai-coe"] = { "Coeruna", 6425639, "sai-wit", "Latn", } m["sai-col"] = { "Colán", 5141893, "sai-ctc", "Latn", } m["sai-cop"] = { "Copallén", 5390321, "qfa-unc", -- pupus, hanya empat perkataan dibuktikan; mungkin Cholonan "Latn", } m["sai-crd"] = { "Coroado Puri", 24191321, "sai-mje", "Latn", } m["sai-ctq"] = { "Catuquinaru", 16858455, "qfa-unc", -- pupus, kurang dibuktikan; kosa kata tidak menyerupai bahasa lain "Latn", } m["sai-cul"] = { "Culli", 2879660, "qfa-unc", -- pupus, kurang dibuktikan; sering dianggap sebagai pencilan "Latn", } m["sai-cva"] = { "Cueva", 5192644, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Chocoan "Latn", } m["sai-esm"] = { "Esmeralda", 3058083, "qfa-unc", -- pupus, kurang dibuktikan; mungkin berkaitan dengan Yaruro "Latn", } m["sai-ewa"] = { "Ewarhuyana", 16898104, nil, "Latn", } m["sai-gam"] = { "Gamela", 5403661, "qfa-unc", -- pupus, kurang dibuktikan; mungkin pencilan "Latn", } m["sai-gay"] = { "Gayón", 5528902, "sai-jir", "Latn", } m["sai-gmo"] = { "Guamo", 5613495, "qfa-unc", -- pupus; "Kaufman (1990) mendapati hubungan dengan bahasa-bahasa Chapacuran meyakinkan." [Wikipedia] Dianggap sebagai pencilan oleh Campbell (2024). "Latn", } m["sai-gua"] = { "Guachí", 5613172, "sai-guc", "Latn", } m["sai-gue"] = { "Güenoa", 5626799, "sai-crn", "Latn", } m["sai-hau"] = { "Haush", 3128376, "sai-cho", "Latn", } m["sai-jee-pro"] = { "Jê Purba", 116773212, "sai-jee", "Latn", type = "reconstructed", } m["sai-jko"] = { "Jeikó", 6176527, "sai-mje", "Latn", } m["sai-jrj"] = { "Jirajara", 6202966, "sai-jir", "Latn", } m["sai-kat"] = { -- kontras xoo, kzw, sai-xoc "Katembri", 6375925, "qfa-unc", -- pupus, kurang dibuktikan; "Kaufman (1990) telah menghubungkannya dengan bahasa Taruma yang hampir pupus, walaupun ini tidak diterima oleh sarjana lain." [Wikipedia] "Latn", } m["sai-mal"] = { "Malalí", 6741212, "sai-mje", -- dianggap sebagai bahasa Maxakalían yang paling divergen (subbahagian kepada Macro-Jê), yang mana kami tiada entri "Latn", } m["sai-mar"] = { "Maratino", 6755055, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Uto-Aztecan "Latn", } m["sai-mat"] = { "Matanawi", 6786047, "qfa-unc", -- pupus; sama ada pencilan atau berkait jauh dengan bahasa-bahasa Muran; Campbell (2024) menyenarainya sebagai pencilan, Glottolog memberikannya sebagai tidak terkelas "Latn", } m["sai-mcn"] = { "Mocana", 3402048, "qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad) "Latn", } m["sai-men"] = { "Menien", 16890110, "sai-mje", "Latn", } m["sai-mil"] = { "Millcayac", 19573012, "sai-hrp", "Latn", } m["sai-mlb"] = { "Malibu", 134374036, "qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad) "Latn", } m["sai-msk"] = { "Masakará", 6782426, "sai-mje", "Latn", } m["sai-muc"] = { "Mucuchí", 6931290, nil, -- lazimnya dianggap sebagai Timotean, yang mana kami tiada entri "Latn", } m["sai-mue"] = { "Muellama", 16886936, "sai-bar", "Latn", } m["sai-muz"] = { "Muzo", 6644203, "qfa-unc", -- bahasa pupus di Colombia, kurang dibuktikan; mungkin Pijao (Cariban) "Latn", } m["sai-mys"] = { "Maynas", 16919393, "sai-cah", -- mengikut Campbell (2024); dahulu dianggap tidak terkelas "Latn", } m["sai-nat"] = { "Natú", 9006749, "qfa-unc", -- pupus, kurang dibuktikan; "hanya Greenberg yang berani mengelaskannya".[Wikipedia, memetik Moseley, Christopher; Asher, R. E.; Tait, Mary (1994), Atlas of the world's languages] "Latn", } m["sai-nje-pro"] = { "Jê Utara Purba", 116773245, "sai-nje", "Latn", type = "reconstructed", } m["sai-opo"] = { "Opón", 7099152, "sai-car", "Latn", } m["sai-oto"] = { "Otomaco", 16879234, "sai-otm", "Latn", } m["sai-pal"] = { "Palta", 3042978, "qfa-unc", -- pupus, tidak terkelas; mungkin Chicham "Latn", } m["sai-pam"] = { "Pamigua", 5908689, "sai-tin", "Latn", } m["sai-par"] = { "Paratió", 16890038, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Xukuruan "Latn", } m["sai-peb"] = { "Peba", 3373890, "sai-pey", "Latn", } m["sai-pnz"] = { "Panzaleo", 3123275, "qfa-unc", -- pupus, tidak terkelas; mungkin Paezan "Latn", } m["sai-prh"] = { "Puruhá", 3410994, "qfa-unc", -- pupus, kurang dibuktikan; mungkin dalam keluarga dengan Cañari "Latn", } m["sai-ptg"] = { "Patagón", 128807870, "sai-tar", -- pupus, hanya diketahui daripada 4 perkataan, yang mencadangkan susur galur Cariban (Campbell 2024) "Latn", } m["sai-pur"] = { "Purukotó", 7261622, "sai-pem", "Latn", } m["sai-pyg"] = { "Payaguá", 7156643, "sai-guc", "Latn", } m["sai-pyk"] = { "Pykobjê", 98113977, "sai-nje", "Latn", } m["sai-qmb"] = { "Quimbaya", 7272043, "qfa-unc", -- pupus, mungkin tidak wujud; sedikit perkataan yang diketahui "Latn", } m["sai-qtm"] = { "Quitemo", 7272651, "sai-cpc", "Latn", } m["sai-rab"] = { "Rabona", 6644203, "qfa-unc", -- pupus, kurang dibuktikan, kebanyakan nama tumbuhan; mungkin Candoshi-Shapra "Latn", } m["sai-ram"] = { "Ramanos", 16902824, "qfa-unc", -- pupus, kurang dibuktikan, mungkin pencilan; mengikut Glottolog: "senarai perkataan yang kerdil ... tidak menunjukkan persamaan yang meyakinkan dengan bahasa sekeliling" "Latn", } m["sai-sac"] = { "Sácata", 5390321, "qfa-unc", -- pupus, hanya 3 perkataan diketahui; mungkin Candoshí atau Arawak "Latn", } m["sai-san"] = { "Sanaviron", 16895999, "qfa-unc", -- pupus, tidak terkelas; tiada konsensus mengenai pengelasan "Latn", } m["sai-sap"] = { "Sapará", 7420922, "sai-car", "Latn", } m["sai-sec"] = { "Sechura", 7442912, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan "Latn", } m["sai-sin"] = { "Sinúfana", 7525275, "qfa-unc", -- hampir pupus, kurang dibuktikan; mungkin Chocoan "Latn", } m["sai-sje-pro"] = { "Jê Selatan Purba", 116773814, "sai-sje", "Latn", type = "reconstructed", } m["sai-tab"] = { "Tabancale", 5390321, "qfa-unc", -- pupus, hanya 5 perkataan diketahui; tiada kaitan yang jelas, mungkin pencilan "Latn", } m["sai-tal"] = { "Tallán", 16910468, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan "Latn", } m["sai-tap"] = { "Tapayuna", 30719984, "sai-nje", "Latn", } m["sai-tar-pro"] = { "Taranoan Purba", 116773816, "sai-tar", "Latn", type = "reconstructed", } m["sai-teu"] = { "Teushen", 3519243, "qfa-unc", -- mungkin pupus menjelang 1950-an; mungkin Chonan "Latn", } m["sai-tim"] = { "Timote", 7806995, nil, -- mungkin dalam keluarga Timote kecil "Latn", } m["sai-tpr"] = { "Taparita", 7684460, "sai-otm", "Latn", } m["sai-trr"] = { "Tarairiú", 7685313, "qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan "Latn", } m["sai-wai"] = { "Waitaká", 16918610, "qfa-unc", -- pupus, mungkin Purian "Latn", } m["sai-way"] = { "Wayumara", 7960726, "sai-car", "Latn", } m["sai-wit-pro"] = { "Witotoan Purba", 116773823, "sai-wit", "Latn", type = "reconstructed", } m["sai-wnm"] = { "Wanham", 16879440, "sai-cpc", "Latn", } m["sai-xoc"] = { -- kontras xoo, kzw, sai-kat "Xocó", 12953620, "qfa-unc", -- pupus dan kurang dibuktikan; tidak jelas sama ada satu atau tiga bahasa "Latn", } m["sai-yao"] = { "Yao (Amerika Selatan)", 16979655, "sai-ven", "Latn", } m["sai-yar"] = { -- bukan keluarga yang sama dengan 'suy' "Yarumá", 3505859, "sai-pek", "Latn", } m["sai-yri"] = { "Yuri", 2669157, "sai-tyu", "Latn", } m["sai-yup"] = { "Yupua", 8061430, "sai-tuc", "Latn", } m["sai-yur"] = { "Yurumanguí", 1281291, "qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan "Latn", } m["sal-pro"] = { "Salish Purba", 116773269, "sal", "Latn", type = "reconstructed", } m["sdv-daj-pro"] = { "Daju Purba", 116773739, "sdv-daj", "Latn", type = "reconstructed", } m["sdv-eje-pro"] = { "Jebel Timur Purba", 116773751, "sdv-eje", "Latn", type = "reconstructed", } m["sdv-nil-pro"] = { "Nilotik Purba", 116773794, "sdv-nil", "Latn", type = "reconstructed", } m["sdv-nyi-pro"] = { "Nyima Purba", 116773796, "sdv-nyi", "Latn", type = "reconstructed", } m["sdv-tmn-pro"] = { "Taman Purba", 116773815, "sdv-tmn", "Latn", type = "reconstructed", } m["sel-nor"] = { "Selkup Utara", 30304565, "sel", "Cyrl", translit = "sel-nor-translit", } m["sel-pro"] = { "Selkup Purba", 128884235, "sel", "Latn", type = "reconstructed", } m["sel-sou"] = { "Selkup Selatan", 30304639, "sel", "Cyrl", translit = "sel-sou-translit", } m["sem-amm"] = { "Ammon", 279181, "sem-can", "Phnx", -- translit Phnx dalam [[Module:scripts/data]] } m["sem-amo"] = { "Amor", 35941, "sem-nwe", "Xsux, Latn", } m["sem-cha"] = { "Chaha", 35543, "sem-eth", "Ethi", translit = "Ethi-translit", } m["sem-dad"] = { "Dadan", 21838040, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-dum"] = { "Dumait", 128810397, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-has"] = { "Hasait", 3541433, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-his"] = { "Hisma", 22948260, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-mhr"] = { "Muher", 33743, "sem-eth", "Latn", } m["sem-pro"] = { "Samiah Purba", 1658554, "sem", "Latn", type = "reconstructed", } m["sem-saf"] = { "Safait", 472586, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-sam"] = { "Samal", 85847147, "sem-nwe", "Phnx", -- translit Phnx dalam [[Module:scripts/data]] } m["sem-srb"] = { "Arab Selatan Kuno", 35025, "sem-osa", "Sarb", -- translit Sarb dalam [[Module:scripts/data]] } m["sem-tay"] = { "Tayman", 24912301, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-tha"] = { "Thamud", 843030, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-wes-pro"] = { "Samiah Barat Purba", 98021726, "sem-wes", "Latn", type = "reconstructed", } m["sio-pro"] = { -- PERHATIAN ini bukan 'nai-sca-pro' "Proto-Siouan-Catawban" iaitu Proto-Sioux Barat "Sioux Purba", 34181, "sio", "Latn", type = "reconstructed", } m["sit-aao-pro"] = { "Naga Tengah Purba", nil, "sit-aao", "Latn", type = "reconstructed", } m["sit-bai-pro"] = { "Bai Purba", nil, "sit-bai", "Latn", type = "reconstructed", } m["sit-ban"] = { "Bangru", 56071779, "sit-hrs", "Latn", } m["sit-bdi-pro"] = { "Bodish Purba", nil, "sit-bdi", "Latn", type = "reconstructed", } m["sit-bok"] = { "Bokar", 4938727, "sit-tan", "Latn, Tibt", override_translit = true, -- translit, display_text, strip_diacritics, sort_key Tibt dalam [[Module:scripts/data]] } m["sit-cai"] = { "Caijia", 5017528, "sit-cln", "Latn" } m["sit-cha"] = { "Chairel", 5068066, "sit-luu", "Latn", } m["sit-ers-pro"] = { "Ersu Purba", nil, "sit-ers", "Latn", type = "reconstructed", } m["sit-hrs-pro"] = { "Hrusish Purba", 116773762, "sit-hrs", "Latn", type = "reconstructed", } m["sit-jap"] = { "Japhug", 3162245, "sit-egy", "Latn", } m["sit-kha-pro"] = { "Kham Purba", 116773773, "sit-kha", "Latn", type = "reconstructed", } m["sit-khb-pro"] = { "Kho-Bwa Purba", nil, "sit-khb", "Latn", type = "reconstructed", } m["sit-khp-pro"] = { "Puroik Purba", nil, "sit-khb", "Latn", type = "reconstructed", } m["sit-khw-pro"] = { "Kho-Bwa Barat Purba", nil, "sit-khw", "Latn", type = "reconstructed", } m["sit-kon-pro"] = { "Naga Utara Purba", nil, "sit-kon", "Latn", type = "reconstructed", } m["sit-liz"] = { "Lizu", 6660653, "sit-ers", "Latn", -- dan Ersu Shaba } m["sit-lnj"] = { "Longjia", 17096251, "sit-cln", "Latn" } m["sit-lrn"] = { "Luren", 16946370, "sit-cln", "Latn" } m["sit-luu-pro"] = { "Luish Purba", 116773783, "sit-luu", "Latn", type = "reconstructed", } m["sit-nas-pro"] = { "Naish Purba", nil, "sit-nas", "Latn", type = "reconstructed", } m["sit-prn"] = { "Puiron", 7259048, "sit-zem", } m["sit-pro"] = { "Sino-Tibet Purba", 24839178, "sit", "Latn", type = "reconstructed", } m["sit-sit"] = { "Situ", 19840830, "sit-egy", "Latn", } m["sit-tam-pro"] = { "Tamang Purba", 117469295, "sit-tam", "Latn", type = "reconstructed", } m["sit-tan-pro"] = { "Tani Purba", 116773284, "sit-tan", "Latn", -- memerlukan pengesahan type = "reconstructed", } m["sit-tgm"] = { "Tangam", 17041370, "sit-tan", "Latn", } m["sit-tng-pro"] = { "Tangkhul Purba", nil, "sit-tng", "Latn", type = "reconstructed", } m["sit-tos"] = { "Tosu", 7827899, "sit-ers", "Latn", -- juga Ersu Shaba } m["sit-tsh"] = { "Tshobdun", 19840950, "sit-egy", "Latn", } m["sit-zbu"] = { "Zbu", 19841106, "sit-egy", "Latn", } m["sla-pro"] = { "Slav Purba", 747537, "sla", "Latn", type = "reconstructed", strip_diacritics = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {'ś'}, }, sort_key = { from = {"č", "ď", "ě", "ę", "ь", "ľ", "ň", "ǫ", "ř", "š", "ś", "ť", "ъ", "ž"}, to = {"c²", "d²", "e²", "e³", "i²", "l²", "nj", "o²", "r²", "s²", "s³", "t²", "u²", "z²"}, } } m["smi-pro"] = { "Sami Purba", 7251862, "smi", "Latn", type = "reconstructed", sort_key = { from = {"ā", "č", "δ", "[ëē]", "ŋ", "ń", "ō", "š", "θ", "%([^()]+%)"}, to = {"a", "c²", "d", "e", "n²", "n³", "o", "s²", "t²"} }, } m["son-pro"] = { "Songhai Purba", 116773277, "son", "Latn", type = "reconstructed", } m["sqj-pro"] = { "Albania Purba", 18210846, "sqj", "Latn", type = "reconstructed", } m["ssa-klk-pro"] = { "Kuliak Purba", 116773779, "ssa-klk", "Latn", type = "reconstructed", } m["ssa-kom-pro"] = { "Koma Purba", 116773775, "ssa-kom", "Latn", type = "reconstructed", } m["ssa-pro"] = { "Nilo-Sahara Purba", 116773236, "ssa", "Latn", type = "reconstructed", } m["syd-pro"] = { "Samoyed Purba", 7251863, "syd", "Latn", type = "reconstructed", } m["tai-pro"] = { "Tai Purba", 6583709, "tai", "Latn", type = "reconstructed", } m["tai-swe-pro"] = { "Tai Barat Daya Purba", 116773280, "tai-swe", "Latn", type = "reconstructed", } m["tbq-bdg-pro"] = { "Bodo-Garo Purba", 116773195, "tbq-bdg", "Latn", type = "reconstructed", } m["tbq-blg"] = { "Bailang", 2879843, "tbq-lob", "Hani", sort_key = "Hani-sortkey", } m["tbq-brm-pro"] = { "Burma Purba", nil, "tbq-brm", "Latn", type = "reconstructed", } m["tbq-gkh"] = { "Gokhy", 5578069, "tbq-sil", "Latn", } m["tbq-kuk-pro"] = { "Kuki-Chin Purba", 116773220, "tbq-kuk", "Latn", type = "reconstructed", } m["tbq-lal-pro"] = { "Lalo Purba", 116773781, "tbq-lal", "Latn", type = "reconstructed", } m["tbq-laz"] = { "Laze", 17007626, "sit-nas", "Latn", } m["tbq-lob-pro"] = { "Lolo-Burma Purba", 116773224, "tbq-lob", "Latn", type = "reconstructed", } m["tbq-lol-pro"] = { "Lolo Purba", 7251855, "tbq-lol", "Latn", type = "reconstructed", } m["tbq-mil"] = { "Milang", 6850761, "sit-gsi", "Deva, Latn", } m["tbq-mor"] = { "Moran", 6909216, "tbq-bdg", "Latn", } m["tbq-ngo"] = { "Ngochang", 56582, "tbq-brm", "Latn", } -- tbq-pro kini khusus etimologi m["trk-dkh"] = { "Dukhan", 12809273, "trk-ssb", "Latn, Cyrl, Mong", -- translit, display_text dan strip_diacritics Mong dalam [[Module:scripts/data]] } -- Seperti yang diuraikan dalam ''Dīwān Lughāt al-Turk'' karya Mahmud al-Kashgari abad ke-11. m["trk-eog"] = { "Oghuz Kuno Awal", nil, "trk-ogz", "Arab", strip_diacritics = {Arab = "ar-stripdiacritics"}, } m["trk-oat"] = { "Turki Anatolia Kuno", 7083390, "trk-ogz", "Arab", strip_diacritics = {Arab = "ar-stripdiacritics"}, ancestors = "trk-eog", } m["trk-pro"] = { "Turkik Purba", 3657773, "trk", "Latn", type = "reconstructed", standard_chars = { Latn = " ()-abdegiklmnoprstuxyzïöüāčēīĺŋōŕšūǖȫẹ" .. c.macron, } } m["tup-gua-pro"] = { "Tupi-Guarani Purba", 116773288, "tup-gua", "Latn", type = "reconstructed", } m["tup-kab"] = { "Kabishiana", 15302988, "tup", "Latn", } m["tup-kaw"] = { "Kawahiva", 6346712, "tup-gua", "Latn", } m["tup-pro"] = { "Tupi Purba", 10354700, "tup", "Latn", type = "reconstructed", } m["tuw-alk"] = { "Alchuka", 113553616, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-bal"] = { "Bala", 86730632, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-kkl"] = { "Kyakala", 118875708, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-kli"] = { "Kili", 6406892, "tuw-ewe", "Cyrl", } m["tuw-pro"] = { "Tungus Purba", 85872335, "tuw", "Latn", type = "reconstructed", } m["tuw-sol"] = { "Solon", 30004, "tuw-ewe", } m["urj-fin-pro"] = { "Finnik Purba", 11883720, "urj-fin", "Latn", type = "reconstructed", } m["urj-koo"] = { "Komi Kuno", 86679962, "kv", "Perm, Cyrs", translit = "urj-koo-translit", -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]; sebelum ini, strip_diacritics Cyrs tidak hadir } m["urj-kuk"] = { "Kukkuzi", 107410460, "urj-fin", "Latn", ancestors = "vot", } m["urj-kya"] = { "Komi-Yazva", 2365210, "kv", "Cyrl", translit = "kv-translit", override_translit = true, strip_diacritics = {remove_diacritics = c.acute}, } m["urj-mdv-pro"] = { "Mordvinik Purba", 116773232, "urj-mdv", "Latn", type = "reconstructed", } m["urj-prm-pro"] = { "Permik Purba", 116773257, "urj-prm", "Latn", type = "reconstructed", } m["urj-pro"] = { "Uralik Purba", 288765, "urj", "Latn", type = "reconstructed", } m["urj-ugr-pro"] = { "Ugrik Purba", 156631, "urj-ugr", "Latn", type = "reconstructed", } m["xnd-pro"] = { "Na-Dene Purba", 116773233, "xnd", "Latn", type = "reconstructed", } m["xgn-pro"] = { "Mongol Purba", 2493677, "xgn", "Latn", type = "reconstructed", sort_key = { from = {"č", "i", "ï", "ǰ", "ŋ", "ö", "š", "ü"}, to = {"c", "i" .. p[1], "i", "j", "n" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1]}, }, } m["yok-bvy"] = { "Yokuts Buena Vista", 4985474, "yok", "Latn", } m["yok-dly"] = { "Yokuts Delta", 70923266, "yok", "Latn", } m["yok-gsy"] = { "Yokuts Gashowu", 3098708, "yok", "Latn", } m["yok-kry"] = { "Yokuts Sungai Kings", 6413014, "yok", "Latn", } m["yok-nvy"] = { "Yokuts Lembah Utara", 85789777, "yok", "Latn", } m["yok-ply"] = { "Yokuts Palewyami", 2387391, "yok", "Latn", } m["yok-svy"] = { "Yokuts Lembah Selatan", 12642473, "yok", "Latn", } m["yok-tky"] = { "Yokuts Tule-Kaweah", 7851988, "yok", "Latn", } m["ypk-pro"] = { "Yupik Purba", 116773295, "ypk", "Latn", type = "reconstructed", } m["yrk-for"] = { "Nenets Hutan", 1295107, "yrk", "Cyrl", translit = "yrk-for-translit", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.dotabove}, } m["yrk-tun"] = { "Nenets Tundra", 36452, "yrk", "Cyrl", strip_diacritics = { from = {"ӑ", "а̄", "э̇", "ӣ", "ы̄", "ӯ", "ю̄", "я̆", "я̄"}, to = {"а", "а", "э", "и", "ы", "у", "ю", "я", "я"}, }, translit = "yrk-tun-translit", } m["zhx-min-pro"] = { "Min Purba", 19646347, "zhx-min", "Latn", type = "reconstructed", } m["zhx-sht"] = { "Tuhua Shaozhou", 1920769, "zhx", "Nshu, Hants", generate_forms = "zh-generateforms", sort_key = {Hani = "Hani-sortkey"}, } m["zhx-sic"] = { "Sichuan", 2278732, "zhx-man", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["zhx-tai"] = { "Taishan", 2208940, "zhx-yue", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["zle-ono"] = { "Novgorod Kuno", 162013, "zle", "Cyrs, Glag", translit = {Cyrs = "Cyrs-translit", Glag = "Glag-translit"}, -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]] } m["zle-ort"] = { "Ruthenia Kuno", 13211, "zle", "Arab, Cyrs, Latn", ancestors = "orv", translit = { Cyrs = "zle-ort-translit", Arab = "zle-ort-Arab-translit", }, strip_diacritics = { Cyrs = { remove_diacritics = m_langdata.chars_substitutions["Cyrs_remove_diacritics"], remove_exceptions = {"Ї", "ї"}, }, Arab = "ar-stripdiacritics", }, -- sort_key Cyrs dalam [[Module:scripts/data]] } m["zls-chs"] = { "Slav Gereja", 33251, "zls", "Cyrs, Glag, Latn", ancestors = "cu", translit = { Cyrs = "Cyrs-translit", Glag = "Glag-translit" }, -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]] } m["zlw-ocs"] = { "Czech Kuno", 593096, "zlw", "Latn", } m["zlw-opl"] = { "Poland Kuno", 149838, "zlw-lch", "Latn", strip_diacritics = {remove_diacritics = c.ringabove}, } m["zlw-osk"] = { "Slovak Kuno", 12776676, "zlw", "Latn", } m["zlw-slv"] = { "Slovincia", 36822, "zlw-pom", "Latn", strip_diacritics = {remove_diacritics = c.macron .. c.breve}, } -- Kod tambahan untuk bahasa-bahasa yang digunakan di Malaysia, yang tidak wujud di Wikikamus Bahasa Inggeris m["zlm-coa"] = { "Melayu Terengganu Pesisir", 4207412, "poz-mly", "Latn, ms-Arab", } m["zlm-pah"] = { "Melayu Pahang", 7310370, "poz-mly", "Latn", } m["tmw"] = { "Temuan", 3025610, "poz-mly", "Latn", } m["kzt"] = { "Dusun Tambunan", 12953514, "poz-san", "Latn", } return require("Module:languages").finalizeData(m, "language") 8hpheldymxlci4iezyguza9t4ttu1az 375376 375374 2026-09-22T05:35:06Z Hakimi97 2668 Move "Temuan" back to Module:languages/data/3/t, and move "Dusun Tambunan" back to Module:languages/data/3/k 375376 Scribunto text/plain local m_langdata = require("Module:languages/data") -- Loaded on demand, as it may not be needed (depending on the data). local function u(...) u = require("Module:string utilities").char return u(...) end local c = m_langdata.chars local p = m_langdata.puaChars local s = m_langdata.shared local m = {} m["aav-khs-pro"] = { "Khasi Purba", 116773216, "aav-khs", "Latn", type = "reconstructed", } m["aav-nic-pro"] = { "Nicobar Purba", 116773793, "aav-nic", "Latn", type = "reconstructed", } m["aav-pkl-pro"] = { "Pnar-Khasi-Lyngngam Purba", 116773259, "aav-pkl", "Latn", type = "reconstructed", } m["aav-pro"] = { -- mkh-pro akan digabungkan ke dalam ini "Austroasia Purba", 116773186, "aav", "Latn", type = "reconstructed", } m["afa-pro"] = { "Afroasia Purba", 269125, "afa", "Latn", type = "reconstructed", } m["alg-aga"] = { "Agawam", nil, "alg-eas", "Latn", } m["alg-pro"] = { "Algonquian Purba", 7251834, "alg", "Latn", type = "reconstructed", sort_key = {remove_diacritics = "·"}, } m["alv-ama"] = { "Amasi", 4740400, "nic-grs", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron}, } m["alv-bgu"] = { "Bainouk Gubeeher", 17002646, "alv-bny", "Latn", } m["alv-bua-pro"] = { "Bua Purba", 116773723, "alv-bua", "Latn", type = "reconstructed", } m["alv-cng-pro"] = { "Cangin Purba", 116773726, "alv-cng", "Latn", type = "reconstructed", } m["alv-edo-pro"] = { "Edoid Purba", 116773206, "alv-edo", "Latn", type = "reconstructed", } m["alv-fli-pro"] = { "Fali Purba", 116773754, "alv-fli", "Latn", type = "reconstructed", } m["alv-gbe-pro"] = { "Gbe Purba", 116773208, "alv-gbe", "Latn", type = "reconstructed", } m["alv-gng-pro"] = { "Guang Purba", 116773757, "alv-gng", "Latn", type = "reconstructed", } m["alv-gtm-pro"] = { "Togo Tengah Purba", 116773732, "alv-gtm", "Latn", type = "reconstructed", } m["alv-gwa"] = { "Gwara", 16945580, "nic-pla", "Latn", } m["alv-hei-pro"] = { "Heiban Purba", 116773760, "alv-hei", "Latn", type = "reconstructed", } m["alv-ido-pro"] = { "Idomoid Purba", 116773764, "alv-ido", "Latn", type = "reconstructed", } m["alv-igb-pro"] = { "Igboid Purba", 116773765, "alv-igb", "Latn", type = "reconstructed", } m["alv-kwa-pro"] = { "Kwa Purba", 116773780, "alv-kwa", "Latn", type = "reconstructed", } m["alv-mum-pro"] = { "Mumuye Purba", 116773791, "alv-mum", "Latn", type = "reconstructed", } m["alv-nup-pro"] = { "Nupoid Purba", 116773795, "alv-nup", "Latn", type = "reconstructed", } m["alv-pro"] = { "Atlantik-Congo Purba", 116732838, "alv", "Latn", type = "reconstructed", } m["alv-edk-pro"] = { "Edekiri Purba", nil, "alv-edk", "Latn", type = "reconstructed", } m["alv-yor-pro"] = { "Yoruba Purba", nil, "alv-yor", "Latn", type = "reconstructed", } m["alv-yrd-pro"] = { "Yoruboid Purba", 116773824, "alv-yrd", "Latn", type = "reconstructed", } m["alv-von-pro"] = { "Volta-Niger Purba", 116773820, "alv-von", "Latn", type = "reconstructed", } m["apa-pro"] = { "Apache Purba", 116773135, "apa", "Latn", type = "reconstructed", } m["aql-pro"] = { "Algik Purba", 18389588, "aql", "Latn", type = "reconstructed", sort_key = {remove_diacritics = "·"}, } m["art-adu"] = { "Adûni", 1232159, "art", "Latn", type = "appendix-constructed", } m["art-bel"] = { "Kreol Belter", 108055510, "art", "Latn", type = "appendix-constructed", sort_key = { remove_diacritics = c.acute, from = {"ɒ"}, to = {"a"}, }, } m["art-blk"] = { "Bolak", 2909283, "art", "Latn", type = "appendix-constructed", } m["art-bsp"] = { "Bahasa Hitam", 686210, "art", "Latn, Teng", type = "appendix-constructed", } m["art-com"] = { "Communicationssprache", 35227, "art", "Latn", type = "appendix-constructed", } m["art-dtk"] = { "Dothraki", 2914733, "art", "Latn", type = "appendix-constructed", } m["art-elo"] = { "Eloi", nil, "art", "Latn", type = "appendix-constructed", } m["art-gld"] = { "Goa'uld", 19823, "art", "Latn, Egyp, Mero", type = "appendix-constructed", } m["art-lap"] = { "Lapine", 6488195, "art", "Latn", type = "appendix-constructed", } m["art-man"] = { "Mandalorian", 54289, "art", "Latn", type = "appendix-constructed", } m["art-mun"] = { "Mundolinco", 851355, "art", "Latn", type = "appendix-constructed", } m["art-nav"] = { "Naʼvi", 316939, "art", "Latn", type = "appendix-constructed", } m["art-vlh"] = { "Valyria Tinggi", 64483808, "art", "Latn", type = "appendix-constructed", } m["ath-nic"] = { "Nicola", 20609, "ath-nor", "Latn", } m["ath-pro"] = { "Athabaska Purba", 104841722, "ath", "Latn", type = "reconstructed", } m["auf-pro"] = { "Arawa Purba", 116773706, "auf", "Latn", type = "reconstructed", } m["aus-alu"] = { "Alungul", 16827670, "aus-pmn", "Latn", } m["aus-and"] = { "Andjingith", 4754509, "aus-pmn", "Latn", } m["aus-ang"] = { "Angkula", 16828520, "aus-pmn", "Latn", } m["aus-arn-pro"] = { "Arnhem Purba", 116773720, "aus-arn", "Latn", type = "reconstructed", } m["aus-bra"] = { "Barranbinya", 4863220, "aus-pmn", "Latn", } m["aus-brm"] = { "Barunggam", 4865914, "aus-pmn", "Latn", } m["aus-cww-pro"] = { "New South Wales Tengah Purba", 116773199, "aus-cww", "Latn", type = "reconstructed", } m["aus-dal-pro"] = { "Daly Purba", 116773743, "aus-dal", "Latn", type = "reconstructed", } m["aus-guw"] = { "Guwar", 6652138, "aus-pam", "Latn", } m["aus-lsw"] = { "Little Swanport", 6652138, "qfa-unc", "Latn", } m["aus-mbi"] = { "Mbiywom", 6799701, "aus-pmn", "Latn", } m["aus-ngk"] = { "Ngkoth", 7022405, "aus-pmn", "Latn", } m["aus-nyu-pro"] = { "Nyulnyulan Purba", 116773797, "aus-nyu", "Latn", type = "reconstructed", } m["aus-pam-pro"] = { "Pama-Nyunga Purba", 33942, "aus-pam", "Latn", type = "reconstructed", } m["aus-tul"] = { "Tulua", 16938541, "aus-pam", "Latn", } m["aus-uwi"] = { "Uwinymil", 7903995, "aus-arn", "Latn", } m["aus-wdj-pro"] = { "Iwaidjan Purba", 116773767, "aus-wdj", "Latn", type = "reconstructed", } m["aus-won"] = { "Wong-gie", nil, "aus-pam", "Latn", } m["aus-wul"] = { "Wulguru", 8039196, "aus-dyb", "Latn", } m["aus-ynk"] = { -- kontras nny "Yangkaal", 3913770, "aus-tnk", "Latn", } m["awd-amc-pro"] = { "Amuesha-Chamicuro Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-kmp-pro"] = { "Kampa Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-prw-pro"] = { "Paresi-Waura Purba", nil, "awd", "Latn", type = "reconstructed", } m["awd-ama"] = { "Amarizana", 16827787, "awd", "Latn", } m["awd-ana"] = { "Anauyá", 16828252, "awd", "Latn", } m["awd-apo"] = { "Apolista", 16916645, "awd", "Latn", } m["awd-cab"] = { "Cabre", 16850160, "awd", "Latn", } m["awd-gnu"] = { "Guinau", 3504087, "awd", "Latn", } m["awd-kar"] = { "Cariay", 16920253, "awd", "Latn", } m["awd-kaw"] = { "Kawishana", 6379993, "awd-nwk", "Latn", } m["awd-kus"] = { "Kustenau", 5196293, "awd", "Latn", } m["awd-man"] = { "Manao", 6746920, "awd", "Latn", } m["awd-mar"] = { "Marawan", 6755108, "awd", "Latn", } m["awd-mpr"] = { "Maipure", 6736872, "awd", "Latn", } m["awd-mrt"] = { "Mariaté", 16910017, "awd-nwk", "Latn", } m["awd-nwk-pro"] = { "Nawiki Purba", 116773234, "awd-nwk", "Latn", type = "reconstructed", } m["awd-pai"] = { "Paikoneka", 128807835, "awd", "Latn", } m["awd-pas"] = { "Pasé", 7143168, "awd-nwk", "Latn", } m["awd-pro"] = { "Arawak Purba", 97573478, "awd", "Latn", type = "reconstructed", } m["awd-she"] = { "Shebayo", 7492248, "awd", "Latn", } m["awd-taa-pro"] = { "Ta-Arawak Purba", 116773282, "awd-taa", "Latn", type = "reconstructed", } m["awd-wai"] = { "Wainumá", 16910017, "awd-nwk", "Latn", } m["awd-war"] = { "Warekena Kuno", 105320180, "awd-nwk", "Latn", } m["awd-yum"] = { "Yumana", 8061062, "awd-nwk", "Latn", } m["azc-caz"] = { "Cazcan", 5055514, "azc", "Latn", } m["azc-cup-pro"] = { "Cupan Purba", 116773738, "azc-cup", "Latn", type = "reconstructed", } m["azc-ktn"] = { "Kitanemuk", 3197558, "azc-tak", "Latn", } m["azc-nah-pro"] = { "Nahua Purba", 7251860, "azc-nah", "Latn", type = "reconstructed", } m["azc-nic"] = { "Nicoleño", 50241488, "azc", "Latn", } m["azc-num-pro"] = { "Numik Purba", 116773247, "azc-num", "Latn", type = "reconstructed", } m["azc-pro"] = { "Uto-Aztek Purba", 96400333, "azc", "Latn", type = "reconstructed", } m["azc-tak-pro"] = { "Takik Purba", 116773283, "azc-tak", "Latn", type = "reconstructed", } m["azc-tat"] = { "Tataviam", 743736, "azc", "Latn", } m["ber-pro"] = { "Berber Purba", 2855698, "ber", "Latn", type = "reconstructed", } m["ber-fog"] = { "Fogaha", 107610173, "ber", "Latn", } m["ber-zuw"] = { "Zuwara", 4117169, "ber", "Latn", } m["bnt-bal"] = { "Balong", 93935237, "bnt-bbo", "Latn", } m["bnt-bon"] = { "Boma Nkuu", nil, "bnt", "Latn", } m["bnt-boy"] = { "Boma Yumu", nil, "bnt", "Latn", } m["bnt-bwa"] = { "Bwala", 128810345, "bnt-tek", "Latn", } m["bnt-cmw"] = { "Chimwiini", 4958328, "bnt-swh", "Latn", } m["bnt-ind"] = { "Indanga", 51412803, "bnt", "Latn", } m["bnt-lal"] = { "Lala (Afrika Selatan)", 6480154, "bnt-ngu", "Latn", } m["bnt-mpi"] = { "Mpiin", 93937013, "bnt-bdz", "Latn", } m["bnt-mpu"] = { "Mpuono", -- jangan dikelirukan dengan Mbuun zmp 36056, "bnt", "Latn", } m["bnt-ngu-pro"] = { "Nguni Purba", 961559, "bnt-ngu", "Latn", type = "reconstructed", sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.caron}, } m["bnt-phu"] = { "Phuthi", 33796, "bnt-ngu", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute}, } m["bnt-pro"] = { "Bantu Purba", 3408025, "bnt", "Latn", type = "reconstructed", sort_key = "bnt-pro-sortkey", } m["bnt-sab-pro"] = { "Sabaki Purba", nil, -- Q2209395 ialah kod untuk keluarga Sabaki "bnt-sab", "Latn", type = "reconstructed", } m["bnt-sbo"] = { "Boma Selatan", nil, "bnt", "Latn", } m["bnt-sts-pro"] = { "Sotho-Tswana Purba", 116773278, "bnt-sts", "Latn", type = "reconstructed", } m["btk-pro"] = { "Batak Purba", 116773191, "btk", "Latn", type = "reconstructed", } m["cau-abz-pro"] = { "Abkhaz-Abaza Purba", 7251831, "cau-abz", "Latn", type = "reconstructed", } m["cau-and-pro"] = { "Andi Purba", nil, "cau-and", "Latn", type = "reconstructed", } m["cau-ava-pro"] = { "Avar-Andi Purba", 116773187, "cau-ava", "Latn", type = "reconstructed", } m["cau-cir-pro"] = { "Circassia Purba", 7251838, "cau-cir", "Latn", type = "reconstructed", } m["cau-drg-pro"] = { "Dargwa Purba", 116773205, "cau-drg", "Latn", type = "reconstructed", } m["cau-lzg-pro"] = { "Lezghi Purba", 116773223, "cau-lzg", "Latn", type = "reconstructed", } m["cau-nec-pro"] = { "Kaukasia Timur Laut Purba", 116773244, "cau-nec", "Latn", type = "reconstructed", } m["cau-nkh-pro"] = { "Nakh Purba", 108032840, "cau-nkh", "Latn", type = "reconstructed", } m["cau-nwc-pro"] = { "Kaukasia Barat Laut Purba", 7251861, "cau-nwc", "Latn", type = "reconstructed", } m["cau-tsz-pro"] = { "Tsez Purba", 116773287, "cau-tsz", "Latn", type = "reconstructed", } m["cba-ata"] = { "Atanques", 4812783, "cba", "Latn", } m["cba-cat"] = { "Catío Chibcha", 7083619, "cba", "Latn", } m["cba-dor"] = { "Dorasque", 5297532, "cba", "Latn", } m["cba-dui"] = { "Duit", 3041061, "cba", "Latn", } m["cba-hue"] = { "Huetar", 35514, "cba", "Latn", } m["cba-nut"] = { "Nutabe", 7070405, "cba", "Latn", } m["cba-pro"] = { "Chibchan Purba", 116773203, "cba", "Latn", type = "reconstructed", } m["ccs-pro"] = { "Kartvelia Purba", 2608203, "ccs", "Latn", type = "reconstructed", strip_diacritics = { from = {"q̣", "p̣", "ʓ", "ċ"}, to = {"q̇", "ṗ", "ʒ", "c̣"} }, } m["ccs-gzn-pro"] = { "Georgia-Zan Purba", 23808119, "ccs-gzn", "Latn", type = "reconstructed", strip_diacritics = { from = {"q̣", "p̣", "ʓ", "ċ"}, to = {"q̇", "ṗ", "ʒ", "c̣"} }, } m["cdc-cbm-pro"] = { "Chadik Tengah Purba", 116773197, "cdc-cbm", "Latn", type = "reconstructed", } m["cdc-mas-pro"] = { "Masa Purba", 116773789, "cdc-mas", "Latn", type = "reconstructed", } m["cdc-pro"] = { "Chadik Purba", 116773201, "cdc", "Latn", type = "reconstructed", } m["cdd-pro"] = { "Caddoan Purba", 116773725, "cdd", "Latn", type = "reconstructed", } m["cel-bry-pro"] = { "Britonik Purba", 1248800, "cel-bry", "Latn, Polyt", sort_key = { Latn = "cel-bry-pro-sortkey", }, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["cel-gal"] = { "Gallaecia", 3094789, "cel-his", } m["cel-gau"] = { "Gaul", 29977, "cel", "Latn, Polyt, Ital", strip_diacritics = { Latn = {remove_diacritics = c.macron .. c.breve .. c.diaer}, }, sort_key = { Latn = "cel-bry-pro-sortkey", }, -- translit Ital dalam [[Module:scripts/data]] -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["cel-pro"] = { "Keltik Purba", 653649, "cel", "Latn", type = "reconstructed", sort_key = "cel-pro-sortkey", } m["chi-pro"] = { "Chimakuan Purba", 116773734, "chi", "Latn", type = "reconstructed", } m["chm-pro"] = { "Mari Purba", 116773788, "chm", "Latn", type = "reconstructed", } m["cmc-pro"] = { "Chamik Purba", 114793834, "cmc", "Latn", type = "reconstructed", } m["crp-bip"] = { "Pijin Basque-Iceland", 810378, "crp", "Latn", ancestors = "eu", } m["crp-cpr"] = { "Pijin Rusia-China", nil, "crp", "Hani, Cyrl, Latn", ancestors = "ru, zh", translit = {Cyrl = "ru-translit"}, strip_diacritics = { Cyrl = {remove_diacritics = c.acute .. c.grave .. c.macron}, }, } m["crp-gep"] = { "Pijin Greenland Barat", 17036301, "crp", "Latn", ancestors = "kl", } m["crp-kia"] = { "Pijin Jerman Kiautschou", 108314615, "crp", "Latn", ancestors = "de", } m["crp-mar"] = { "Bahasa Roh Maroon", 1093206, "crp", "Latn", ancestors = "en", } m["crp-mpp"] = { "Pijin Portugis Macau", 128804537, "crp", "Hant, Latn", ancestors = "pt", sort_key = {Hant = "Hani-sortkey"}, } m["crp-rsn"] = { "Russenorsk", 505125, "crp", "Cyrl, Latn", ancestors = "nn, ru", translit = {Cyrl = "ru-translit"}, } m["crp-spp"] = { "Pijin Ladang Samoa", 7409948, "crp", "Latn", ancestors = "en", } m["crp-slb"] = { "Inggeris Solombala", 7558525, "crp", "Cyrl, Latn", ancestors = "en, ru", translit = {Cyrl = "ru-translit"}, } m["crp-tpr"] = { "Pijin Rusia Taimyr", 16930506, "crp", "Cyrl", ancestors = "ru", translit = "ru-translit", } m["csu-bba-pro"] = { "Bongo-Bagirmi Purba", 116773722, "csu-bba", "Latn", type = "reconstructed", } m["csu-maa-pro"] = { "Mangbetu Purba", 116773786, "csu-maa", "Latn", type = "reconstructed", } m["csu-pro"] = { "Sudan Tengah Purba", 116773730, "csu", "Latn", type = "reconstructed", } m["csu-sar-pro"] = { "Sara Purba", 116773809, "csu-sar", "Latn", type = "reconstructed", } m["cus-ash"] = { "Ashraaf", 4805855, "cus-som", "Latn", } m["cus-hec-pro"] = { "Kusyi Timur Tanah Tinggi Purba", 116773761, "cus-hec", "Latn", type = "reconstructed", } m["cus-som-pro"] = { "Somaloid Purba", nil, "cus-som", "Latn", type = "reconstructed", } m["cus-sou-pro"] = { "Kusyi Selatan Purba", 126081567, "cus-sou", "Latn", type = "reconstructed", } m["cus-pro"] = { "Kusyi Purba", 116773204, "cus", "Latn", type = "reconstructed", } m["dmn-dam"] = { "Dama (Sierra Leone)", 19601574, "dmn", "Latn", } m["dra-bry"] = { "Beary", 1089116, "qfa-mix", "Mlym, Knda", ancestors = "ml, tcy", -- translit Knda dalam [[Module:scripts/data]] -- translit Mlym dalam [[Module:scripts/data]] } m["dra-cen-pro"] = { "Dravidia Tengah Purba", nil, "dra-cen", "Latn", type = "reconstructed", } m["dra-mkn"] = { "Kannada Pertengahan", 128810572, "dra-kan", "Knda", -- translit Knda dalam [[Module:scripts/data]] } m["dra-nor-pro"] = { "Dravidia Utara Purba", 124433593, "dra-nor", "Latn", type = "reconstructed", } m["dra-okn"] = { "Kannada Kuno", 15723156, "dra-kan", "Knda", -- translit Knda dalam [[Module:scripts/data]] } m["dra-ote"] = { "Telugu Kuno", 126720868, "dra-tel", "Telu", translit = "te-translit", } m["dra-pro"] = { "Dravidia Purba", 1702853, "dra", "Latn", type = "reconstructed", } m["dra-sdo-pro"] = { "Dravidia Selatan I Purba", 104847952, -- "Proto-Dravidia Selatan" Wikipedia ialah Proto-Dravidia Selatan I dalam skema ini. "dra-sdo", "Latn", type = "reconstructed", } m["dra-sdt-pro"] = { "Dravidia Selatan II Purba", 128885257, "dra-sdt", "Latn", type = "reconstructed", } m["dra-sou-pro"] = { "Dravidia Selatan Purba", 128886121, "dra-sou", "Latn", type = "reconstructed", } m["egx-dem"] = { "Mesir Demotik", 36765, "egx", "Latn, Egyd, Polyt", sort_key = { Latn = { remove_diacritics = "'%-%s", from = {"ꜣ", "j", "e", "ꜥ", "y", "w", "b", "p", "f", "m", "n", "r", "l", "ḥ", "ḫ", "h̭", "ẖ", "h", "š", "s", "q", "k", "g", "ṱ", "ṯ", "t", "ḏ", "%.", "⸗"}, to = {p[1], p[2], p[3], p[4], p[5], p[6], p[7], p[8], p[9], p[10], p[11], p[12], p[13], p[15], p[16], p[16], p[17], p[14], p[19], p[18], p[20], p[21], p[22], p[23], p[24], p[23], p[25], p[26], p[26]} }, }, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["dmn-pro"] = { "Mande Purba", 116773785, "dmn", "Latn", type = "reconstructed", } m["dmn-mdw-pro"] = { "Mande Barat Purba", 116773822, "dmn-mdw", "Latn", type = "reconstructed", } m["dru-pro"] = { "Rukai Purba", 116773807, "map", "Latn", type = "reconstructed", } m["ero-gsz"] = { "Geshiza", nil, "ero", "Latn", } m["ero-nya"] = { "Nyagrong Minyag", nil, "ero", "Latn", } m["ero-tau"] = { "Stau", nil, "ero", "Latn", } m["esx-esk-pro"] = { "Eskimo Purba", 7251842, "esx-esk", "Latn", type = "reconstructed", } m["esx-ink"] = { "Inuktun", 1671647, "esx-inu", "Latn", } m["esx-inq"] = { "Inuinnaqtun", 28070, "esx-inu", "Latn", } m["esx-inu-pro"] = { "Inuit Purba", 60785588, "esx-inu", "Latn", type = "reconstructed", } m["esx-pro"] = { "Eskimo-Aleut Purba", 7251843, "esx", "Latn", type = "reconstructed", } m["esx-tut"] = { "Tunumiisut", 15665389, "esx-inu", "Latn", } m["euq-pro"] = { "Basque Purba", 938011, "euq", "Latn", type = "reconstructed", } m["gba-pro"] = { "Gbaya Purba", nil, "gba", "Latn", type = "reconstructed", } m["gem-pro"] = { "Jermanik Purba", 669623, "gem", "Latn", type = "reconstructed", sort_key = "gem-pro-sortkey", } m["gme-bur"] = { "Burgundia", 47625, "gme", "Latn", } m["gme-cgo"] = { "Goth Crimea", 36211, "gme", "Latn", } m["gmq-gut"] = { "Gutnish", 1256646, "gmq", "Latn", ancestors = "gmq-ogt", } m["gmq-jmk"] = { "Jamtish", 35512, "gmq-eas", "Latn", } m["gmq-mno"] = { "Norway Pertengahan", 3417070, "gmq-wes", "Latn", } m["gmq-oda"] = { "Denmark Kuno", 12330003, "gmq-eas", "Latn, Runr", strip_diacritics = {remove_diacritics = c.macron}, } m["gmq-ogt"] = { "Gutnish Kuno", 1133488, "gmq", "Latn, Runr", ancestors = "non", } m["gmq-osw"] = { "Sweden Kuno", 2417210, "gmq-eas", "Latn, Runr", strip_diacritics = {remove_diacritics = c.macron}, } m["gmq-pro"] = { "Norse Purba", 1671294, "gmq", "Runr", translit = "Runr-translit", } m["gmq-scy"] = { "Scanian", 768017, "gmq-eas", "Latn", } m["gmw-bgh"] = { "Bergish", 329030, "gmw-frk", "Latn", } m["gmw-cfr"] = { "Franconia Tengah", 572197, "gmw-hgm", "Latn", ancestors = "gmh", wikimedia_codes = "ksh", } m["gmw-ecg"] = { "Jerman Tengah Timur", 499344, -- merangkumi Q699284, Q152965 "gmw-hgm", "Latn", ancestors = "gmh", } m["gmw-fin"] = { "Fingallian", 3072588, "gmw-ian", "Latn", } m["gmw-gts"] = { "Gottscheerish", 533109, "gmw-hgm", "Latn", ancestors = "bar", } m["gmw-jdt"] = { "Belanda Jersey", 1687911, "gmw-frk", "Latn", ancestors = "nl", } m["gmw-msc"] = { "Scots Pertengahan", 3327000, "gmw-ang", "Latn", ancestors = "enm-esc", } m["gmw-pro"] = { "Jermanik Barat Purba", 78079021, "gmw", "Latn, Runr", -- type = "reconstructed", -- sebahagian besarnya tetapi tidak sepenuhnya direkonstruksi (seperti Proto-Norse); lihat BP Apr '24, tetapkan kembali kepada direkonstruksi (?) jika 'anti-asterisk' ditambah sort_key = "gmw-pro-sortkey", } m["gmw-rfr"] = { "Franconia Rhine", 707007, "gmw-hgm", "Latn", ancestors = "gmh", } m["gmw-stm"] = { "Schwaben Szatmár", 2223059, "gmw-hgm", "Latn", ancestors = "swg", } m["gmw-tsx"] = { "Saxon Transylvania", 260942, "gmw-hgm", "Latn", ancestors = "gmw-cfr", } m["gmw-vog"] = { "Jerman Volga", 312574, "gmw-hgm", "Latn", ancestors = "gmw-rfr", } m["gmw-zps"] = { "Jerman Zipser", 205548, "gmw-hgm", "Latn", ancestors = "gmh", } m["gn-cls"] = { "Guarani Klasik", 17478065, "gn", "Latn", } m["grk-cal"] = { "Yunani Calabria", 1146398, "grk", "Latn, Grek", ancestors = "grk-ita", translit = { Grek = "el-translit", }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-ita"] = { "Yunani Italiot", 19720507, "grk", "Latn, Grek", ancestors = "gkm", translit = { Grek = "el-translit", }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-mar"] = { "Yunani Mariupol", 4400023, "grk", "Cyrl, Latn, Grek", ancestors = "gkm", translit = { Cyrl = "grk-mar-translit", Grek = "grk-mar-translit", }, override_translit = true, strip_diacritics = { Cyrl = {remove_diacritics = c.acute}, }, -- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]] } m["grk-pro"] = { "Hellenik Purba", 1231805, "grk", "Latn, Polyt", type = "reconstructed", sort_key = {Latn = { from = {"ʰ", "ʷ"}, to = {"h", "w"}, remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.caron .. c.CGJ }}, display_text = {Latn = { from = {"([dlLt])" .. c.caron}, to = {"%1" .. c.CGJ .. c.caron}, }}, strip_diacritics = {Latn = { from = {"([dlLt])" .. c.caron}, to = {"%1" .. c.CGJ .. c.caron}, }}, -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] -- NOTA: dahulunya tiada translit ditentukan untuk Polyt; mungkin ketinggalan secara tidak sengaja; jika tidak, tetapkan Polyt = false dalam -- bahagian translit } m["hmn-pro"] = { "Hmongik Purba", 116773210, "hmn", "Latn", type = "reconstructed", } m["hmx-mie-pro"] = { "Mienik Purba", 116773229, "hmx-mie", "Latn", type = "reconstructed", } m["hmx-pro"] = { "Hmong-Mien Purba", 7251846, "hmx", "Latn", type = "reconstructed", } m["hyx-pro"] = { "Armenia Purba", 3848498, "hyx", "Latn", type = "reconstructed", } m["iir-nur-pro"] = { "Nuristani Purba", 116773248, "iir-nur", "Latn", type = "reconstructed", } m["iir-pro"] = { "Indo-Iran Purba", 966439, "iir", "Latn", type = "reconstructed", } m["ijo-pro"] = { "Ijoid Purba", 116773766, "ijo", "Latn", type = "reconstructed", } m["inc-apa"] = { "Apabhramsa", 616419, "inc-mid", "Deva, Shrd, Sidd", ancestors = "pra", translit = { Deva = "sa-translit", -- translit Shrd dalam [[Module:scripts/data]] -- translit Sidd dalam [[Module:scripts/data]] }, } m["inc-ash"] = { "Prakrit Ashoka", 104854379, "inc-mid", "Brah, Khar", ancestors = "sa", translit = { -- translit Brah dalam [[Module:scripts/data]] Khar = "Khar-translit", }, } m["inc-dng-pro"] = { "Dangari Purba", nil, "inc-dng", "Latn", type = "reconstructed", } m["inc-kam"] = { "Prakrit Kamarupi", 6356097, "inc-bas", "Brah, Sidd", -- translit Brah, Sidd dalam [[Module:scripts/data]] } m["inc-kho"] = { "Kholosi", 24952008, "inc-snd", "Latn", } m["inc-khr"] = { "Khortha", 13406670, "inc-sad", "Deva, Kthi", translit = { Deva = "bho-translit", Kthi = "bho-Kthi-translit", }, } m["inc-krd-pro"] = { "Kamta Purba", 128816843, "inc-bas", "Latn", ancestors = "inc-kam", type = "reconstructed", } m["inc-mas"] = { "Assam Pertengahan", 128806836, "inc-bas", "as-Beng", ancestors = "inc-oas", translit = "inc-mas-translit", } m["inc-mbn"] = { "Benggali Pertengahan", 113559927, "inc-bas", "Beng", ancestors = "inc-obn", translit = "inc-mbn-translit", } m["inc-mgu"] = { "Gujarati Pertengahan", 24907429, "inc-wes", "Deva", ancestors = "inc-ogu", } m["inc-mor"] = { "Odia Pertengahan", 128810882, "inc-eas", "Orya", ancestors = "inc-oor", } m["inc-oas"] = { "Assam Awal", 85758237, "inc-bas", "as-Beng", ancestors = "inc-kam", translit = "inc-oas-translit", } m["inc-oaw"] = { "Awadhi Kuno", nil, "inc-hie", "Deva, Kthi, Aran", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Deva = "sa-translit", Kthi = "sa-Kthi-translit", Aran = "inc-ohi-translit", }, } m["inc-obn"] = { "Benggali Kuno", 113559926, "inc-bas", "Beng", } m["inc-ogu"] = { "Gujarati Kuno", 24907427, "inc-wes", "Deva", translit = "sa-translit", } m["inc-ohi"] = { "Hindi Kuno", 48767781, "inc-hiw", "Deva, Aran", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Deva = "sa-translit", Aran = "inc-ohi-translit", }, } m["inc-oor"] = { "Odia Kuno", 128807801, "inc-eas", "Orya", } m["inc-opa"] = { "Punjabi Kuno", 115270971, "inc-pan", "Guru, Aran", translit = { Guru = "inc-opa-Guru-translit", Aran = "pa-Aran-translit", }, strip_diacritics = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun}, } m["inc-pro"] = { "Indo-Arya Purba", 23808344, "inc", "Latn", type = "reconstructed", } m["inc-sar"] = { "Sarazi", 85799728, "him", "Aran, Deva, Takr", strip_diacritics = { from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه" to = {"ہ", "ہ"}, remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef }, translit = { Aran = "ur-translit", Deva = "hi-translit", -- Takr = "Takr-translit", }, } m["ine-ana-pro"] = { "Anatolia Purba", 7251833, "ine-ana", "Latn", type = "reconstructed", } m["ine-bsl-pro"] = { "Balto-Slavik Purba", 1703347, "ine-bsl", "Latn", type = "reconstructed", sort_key = { from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", c.acute, c.macron, "ˀ"}, to = {"a", "e", "i", "o", "u"} }, } m["ine-kal"] = { "Kalašma", 122770439, "ine-ana", "Xsux", } m["ine-pae"] = { "Paeonia", 2705672, "ine", "Polyt", -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["ine-pro"] = { "Indo-Eropah Purba", 37178, "ine", "Latn", type = "reconstructed", sort_key = { from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", "ĺ", "ḿ", "ń", "ŕ", "ǵ", "ḱ", "ʰ", "ʷ", "₁", "₂", "₃", c.ringbelow, c.acute, c.macron}, to = {"a", "e", "i", "o", "u", "l", "m", "n", "r", "g'", "k'", "¯h", "¯w", "1", "2", "3"} }, } m["ine-toc-pro"] = { "Tocharia Purba", 104841462, "ine-toc", "Latn", type = "reconstructed", } m["xme-old"] = { "Median Kuno", 36461, "xme", "Polyt, Latn", -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["xme-mid"] = { "Median Pertengahan", 12836150, "xme", "Latn", } m["xme-ker"] = { "Kerman", 129850, "xme", "Arab, Latn, Hebr", ancestors = "xme-mid", -- display_text, strip_diacritics, sort_key Hebr dalam [[Module:scripts/data]] } m["xme-taf"] = { "Tafreshi", nil, "xme", "Arab, Latn", ancestors = "xme-mid", } m["xme-ttc-pro"] = { "Tatik Purba", 122973870, "xme-ttc", "Latn", ancestors = "xme-mid", } m["xme-kls"] = { "Kalasuri", nil, "xme-ttc", ancestors = "xme-ttc-nor", } m["xme-klt"] = { "Kilit", 3612452, "xme-ttc", "Cyrl", -- dan Arab? } m["xme-ott"] = { "Tati Kuno", 434697, "xme-ttc", "Arab, Latn", } m["ira-kms-pro"] = { "Komisenia Purba", 116773777, "ira-kms", "Latn", type = "reconstructed", } m["ira-mpr-pro"] = { "Medo-Parthia Purba", 116773227, "ira-mpr", "Latn", type = "reconstructed", } m["ira-pat-pro"] = { "Pathan Purba", 116773255, "ira-pat", "Latn", type = "reconstructed", } m["ira-pro"] = { "Iran Purba", 4167865, "ira", "Latn", type = "reconstructed", } m["ira-zgr-pro"] = { "Zaza-Gorani Purba", 116775031, "ira-zgr", "Latn", type = "reconstructed", } m["xsc-pro"] = { "Scythia Purba", 116773273, "xsc", "Latn", type = "reconstructed", } m["xsc-sar-pro"] = { "Sarmatia Purba", 116773249, "xsc-sar", "Latn", type = "reconstructed", } m["xsc-skw-pro"] = { "Saka-Wakhi Purba", 116773267, "xsc-skw", "Latn", type = "reconstructed", } m["xsc-sak-pro"] = { "Saka Purba", 116773264, "xsc-sak", "Latn", type = "reconstructed", } m["ira-sym-pro"] = { "Shughni-Yazghulami-Munji Purba", 116773813, "ira-sym", "Latn", type = "reconstructed", } m["ira-sgi-pro"] = { "Sanglechi-Ishkashimi Purba", 116773808, "ira-sgi", "Latn", type = "reconstructed", } m["ira-mny-pro"] = { "Munji-Yidgha Purba", 116773792, "ira-mny", "Latn", type = "reconstructed", } m["ira-shy-pro"] = { "Shughni-Yazghulami Purba", 116773812, "ira-shy", "Latn", type = "reconstructed", } m["ira-shr-pro"] = { "Shughni-Roshani Purba", 116773811, "ira-shr", "Latn", type = "reconstructed", } m["ira-sgc-pro"] = { "Sogdia Purba", 116773276, "ira-sgc", "Latn", type = "reconstructed", } m["ira-wnj"] = { "Vanji", 3398419, "ira-shy", "Latn", } m["iro-ere"] = { "Erie", 5388365, "iro-nor", "Latn", } m["iro-min"] = { "Mingo", 128531, "iro-nor", "Latn", ietf_subtag = "i-mingo", -- tag IETF yang diwarisi } m["iro-nor-pro"] = { "Iroquois Utara Purba", 116773242, "iro-nor", "Latn", type = "reconstructed", } m["iro-pro"] = { "Iroquois Purba", 7251852, "iro", "Latn", type = "reconstructed", } m["itc-pro"] = { "Italik Purba", 17102720, "itc", "Latn", type = "reconstructed", } m["itc-psa"] = { "Pra-Samnit", 7239186, "itc-sbl", "Ital, Polyt, Latn", -- translit Ital dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja) -- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]] } m["jpx-hcj"] = { "Hachijō", 5637049, "jpx", "Jpan", ancestors = "ojp-eas", translit = s["jpx-translit"], display_text = s["jpx-displaytext"], strip_diacritics = s["jpx-stripdiacritics"], sort_key = s["jpx-sortkey"], } m["jpx-pro"] = { "Jepunik Purba", 3924309, "jpx", "Latn", type = "reconstructed", } m["jpx-ryu-pro"] = { "Ryukyu Purba", 56349069, "jpx-ryu", "Latn", type = "reconstructed", } m["kar-pro"] = { "Karen Purba", 85794783, "kar", "Latn", type = "reconstructed", } m["kca-eas"] = { "Khanty Timur", 30304622, "kca", "Cyrl", translit = "kca-translit", override_translit = true, -- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka) sort_key = { Cyrl = { from = {"ᲊ"}, to = {"Ᲊ"} } }, } m["kca-nor"] = { "Khanty Utara", 30304527, "kca", "Cyrl", translit = "kca-translit", override_translit = true, -- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka) sort_key = { Cyrl = { from = {"ᲊ"}, to = {"Ᲊ"} } }, } m["kca-pro"] = { "Khanty Purba", 127505171, "kca", "Latn", type = "reconstructed", } m["kca-sou"] = { "Khanty Selatan", 30304618, "kca", "Cyrl", translit = "kca-translit", override_translit = true, } m["khi-kho-pro"] = { "Khoe Purba", 116773218, "khi-kho", "Latn", type = "reconstructed", } m["khi-kun"] = { "ǃKung", 32904, "khi-kxa", "Latn", } m["ko-ear"] = { "Korea Moden Awal", 756014, "qfa-kor", "Kore", ancestors = "okm", translit = "okm-translit", -- strip_diacritics Kore dalam [[Module:scripts/data]] } m["kro-pro"] = { "Kru Purba", 116773778, "kro", "Latn", type = "reconstructed", } m["ku-pro"] = { "Kurdi Purba", 116773221, "ku", "Latn", type = "reconstructed", } m["map-ata-pro"] = { "Atayalik Purba", 116773151, "map-ata", "Latn", type = "reconstructed", } m["map-bms"] = { "Banyumasan", 33219, "map", "Latn, Java", } m["map-pro"] = { "Austronesia Purba", 49230, "map", "Latn", type = "reconstructed", } m["mis-hkl"] = { "Hokkien Peranakan Kelantan", 108794818, "qfa-mix", ancestors = "nan-hbl, sou, mfa", } m["mis-idn"] = { "Idiom Neutral", 35847, "art", "Latn", type = "appendix-constructed", } m["mis-isa"] = { "Isauria", 16956868, nil, -- "Xsux, Hluw, Latn", } m["mis-jie"] = { "Jie", 124424186, nil, "Hani", sort_key = "Hani-sortkey", } m["mis-jzh"] = { "Jizhao", 45242758, "qfa-bej", "Latn", } m["mis-kas"] = { "Kassite", 35612, nil, "Xsux", } m["mis-mmd"] = { "Mimi Decorse", 6862206, nil, "Latn", } m["mis-mmn"] = { "Mimi Nachtigal", 6862207, nil, "Latn", } m["mis-phi"] = { "Filistin", 2230924, nil, "Phnx", -- translit Phnx dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja) } m["mis-rou"] = { "Rouran", 48816637, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-tdl"] = { "Turdulia", 133176492, } m["mis-tdt"] = { "Turdetania", 133176461, } m["mis-tnw"] = { "Tangwang", 7683179, "qfa-mix", "Latn", ancestors = "cmn, sce", } m["mis-tuh"] = { "Tuyuhun", 48816625, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-tuo"] = { "Tuoba", 48816629, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-wuh"] = { "Wuhuan", 118976867, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-xbi"] = { "Xianbei", 4448647, "qfa-xgx", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mis-xnu"] = { "Xiongnu", 10901674, nil, "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mjg-mgl"] = { "Mongghul", 53765528, "mjg", "Latn", -- juga Mong, Cyrl? } m["mjg-mgr"] = { "Mangghuer", 56285392, "mjg", "Latn", -- juga Mong, Cyrl? } m["mkh-asl-pro"] = { "Asli Purba", 55630680, "mkh-asl", "Latn", type = "reconstructed", } m["mkh-ban-pro"] = { "Bahnar Purba", 116773189, "mkh-ban", "Latn", type = "reconstructed", } m["mkh-kat-pro"] = { "Katuik Purba", 116773772, "mkh-kat", "Latn", type = "reconstructed", } m["mkh-khm-pro"] = { "Khmuik Purba", 116773774, "mkh-khm", "Latn", type = "reconstructed", } m["mkh-kmr-pro"] = { "Khmer Purba", 55630684, "mkh-kmr", "Latn", type = "reconstructed", } m["mkh-mmn"] = { "Mon Pertengahan", 121337926, "mkh-mnc", "Latn, Mymr", --dan juga Pallava ancestors = "omx", } m["mkh-mnc-pro"] = { "Monik Purba", 116773231, "mkh-mnc", "Latn", type = "reconstructed", } m["mkh-mvi"] = { "Vietnam Pertengahan", 9199, "mkh-vie", "Hani, Latn", sort_key = {Hani = "Hani-sortkey"}, } m["mkh-pal-pro"] = { "Palaungik Purba", 104847372, "mkh-pal", "Latn", type = "reconstructed", } m["mkh-pea-pro"] = { "Pearik Purba", 116773804, "mkh-pea", "Latn", type = "reconstructed", } m["mkh-pkn-pro"] = { "Pakanik Purba", 116773803, "mkh-pkn", "Latn", type = "reconstructed", } m["mkh-pro"] = { --Ini akan digabungkan ke dalam aav-pro 2015. "Mon-Khmer Purba", 7251859, "mkh", "Latn", type = "reconstructed", } m["mnw-tha"] = { -- Untuk dibuang. "Mon Thailand", nil, "mkh-mnc", "Mymr, Thai", ancestors = "mkh-mmn", sort_key = { from = {"[%p]", "ျ", "ြ", "ွ", "ှ", "ၞ", "ၟ", "ၠ", "ၚ", "ဿ", "[็-๎]", "([เแโใไ])([ก-ฮ])ฺ?"}, to = {"", "္ယ", "္ရ", "္ဝ", "္ဟ", "္န", "္မ", "္လ", "င", "သ္သ", "", "%2%1"} }, } m["mkh-vie-pro"] = { "Vietik Purba", 109432616, "mkh-vie", "Latn", type = "reconstructed", } m["mns-cen"] = { "Mansi Tengah", 128810384, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mns-nor"] = { "Mansi Utara", 30304537, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mns-pro"] = { "Mansi Purba", 128883093, "mns", "Latn", type = "reconstructed", } m["mns-sou"] = { "Mansi Selatan", 30304629, "mns", "Cyrl", translit = "mns-translit", override_translit = true, } m["mun-pro"] = { "Munda Purba", 105102373, "mun", "Latn", type = "reconstructed", } m["myn-chl"] = { -- peringkat selepas ''emy'' "Ch'olti'", 873995, "myn", "Latn", } m["myn-pro"] = { "Maya Purba", 3321532, "myn", "Latn", type = "reconstructed", } m["nai-ala"] = { "Alazapa", 128810233, nil, "Latn", } m["nai-bay"] = { "Bayogoula", 1563704, nil, "Latn", } m["nai-cal"] = { "Calusa", 51782, nil, "Latn", } m["nai-chi"] = { "Chiquimulilla", 25339627, "nai-xin", "Latn", } m["nai-chu-pro"] = { "Chumash Purba", 116773736, "nai-chu", "Latn", type = "reconstructed", } m["nai-cig"] = { "Ciguayo", 20741700, nil, "Latn", } m["nai-ckn-pro"] = { "Chinook Purba", 116773735, "nai-ckn", "Latn", type = "reconstructed", } m["nai-guz"] = { "Guazacapán", 19572028, "nai-xin", "Latn", } m["nai-hit"] = { "Hitchiti", 1542882, "nai-mus", "Latn", } m["nai-ipa"] = { "Ipai", 3027474, "nai-yuc", "Latn", } m["nai-jtp"] = { "Jutiapa", nil, "nai-xin", "Latn", } m["nai-jum"] = { "Jumaytepeque", 25339626, "nai-xin", "Latn", } m["nai-kat"] = { "Kathlamet", 6376639, "nai-ckn", "Latn", } m["nai-klp-pro"] = { "Kalapuya Purba", 116773771, "nai-klp", "Latn", type = "reconstructed", } m["nai-knm"] = { "Konomihu", 3198734, "nai-shs", "Latn", } m["nai-kum"] = { "Kumeyaay", 4910139, "nai-yuc", "Latn", } m["nai-mac"] = { "Macoris", 21070851, nil, "Latn", } m["nai-mdu-pro"] = { "Maidu Purba", 116773784, "nai-mdu", "Latn", type = "reconstructed", } m["nai-miz-pro"] = { "Mixe-Zoque Purba", 7251858, "nai-miz", "Latn", type = "reconstructed", } m["nai-mus-pro"] = { "Muskogi Purba", 116775368, "nai-mus", "Latn", type = "reconstructed", } m["nai-nao"] = { "Naolan", 6964594, nil, "Latn", } m["nai-nrs"] = { "Shasta Sungai Baru", 7011254, "nai-shs", "Latn", } m["nai-okw"] = { "Okwanuchu", 3350126, "nai-shs", "Latn", } m["nai-per"] = { "Pericú", 3375369, nil, "Latn", } m["nai-pic"] = { "Picuris", 7191257, "nai-kta", "Latn", } m["nai-plp-pro"] = { "Penuti Penara Purba", 116773806, "nai-plp", "Latn", type = "reconstructed", } m["nai-pom-pro"] = { "Pomo Purba", 116773262, "nai-pom", "Latn", type = "reconstructed", } m["nai-qng"] = { "Quinigua", 36360, nil, "Latn", } m["nai-sca-pro"] = { -- PERHATIAN 'sio-pro' "Proto-Siouan" iaitu Proto-Sioux Barat "Siouan-Catawba Purba", 116773275, "nai-sca", "Latn", type = "reconstructed", } m["nai-sin"] = { "Sinacantán", 24190249, "nai-xin", "Latn", } m["nai-sln"] = { "Lenca Salvador", 3229434, "nai-len", "Latn", } m["nai-spt"] = { "Sahaptin", 3833015, "nai-shp", "Latn", } m["nai-tap"] = { "Tapachultec", 7684401, "nai-miz", "Latn", } m["nai-taw"] = { "Tawasa", 7689233, nil, "Latn", } m["nai-teq"] = { "Tequistlatec", 2964454, "nai-tqn", "Latn", } m["nai-tip"] = { "Tipai", 3027471, "nai-yuc", "Latn", } m["nai-tot-pro"] = { "Totozoquean Purba", 116773285, "nai-tot", "Latn", type = "reconstructed", } m["nai-tsi-pro"] = { "Tsimshianik Purba", nil, "nai-tsi", "Latn", type = "reconstructed", } m["nai-utn-pro"] = { "Utik Purba", 116773290, "nai-utn", "Latn", type = "reconstructed", } m["nai-wai"] = { "Waikuri", 3118702, nil, "Latn", } m["nai-wji"] = { "Jicaque Barat", 3178610, "nai-jcq", "Latn", } m["nai-yup"] = { "Yupiltepeque", 25339628, "nai-xin", "Latn", } m["nan-dat"] = { "Min Datian", 19855572, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-hbl"] = { "Hokkien", 1624231, "zhx-nan", "Hants, Latn, Bopo, Kana", wikimedia_codes = "zh-min-nan", generate_forms = "zh-generateforms", sort_key = { Hani = "Hani-sortkey", Kana = "Kana-sortkey" }, } m["nan-hlh"] = { "Min Hailufeng", 120755728, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-lnx"] = { "Min Longyan", 6674568, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-tws"] = { "Teochew", 36759, "zhx-nan", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["nan-zhe"] = { "Min Zhenan", 3846710, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["nan-zsh"] = { "Min Sanxiang", 7420769, "zhx-nan", "Hants", generate_forms = "zh-generateforms", sort_key = "Hani-sortkey", } m["ngf-bin-pro"] = { "Binandere Purba", 137881672, "ngf-bin", "Latn", type = "reconstructed", } m["ngf-pro"] = { "Trans-New Guinea Purba", 85794785, "ngf", "Latn", type = "reconstructed", } m["nic-bco-pro"] = { "Benue-Congo Purba", 116773194, "nic-bco", "Latn", type = "reconstructed", } m["nic-bod-pro"] = { "Bantoid Purba", 116773190, "nic-bod", "Latn", type = "reconstructed", } m["nic-eov-pro"] = { "Oti-Volta Timur Purba", 116773753, "nic-eov", "Latn", type = "reconstructed", } m["nic-gns-pro"] = { "Gurunsi Purba", 116773759, "nic-gns", "Latn", type = "reconstructed", } m["nic-grf-pro"] = { "Grassfields Purba", 116773755, "nic-grf", "Latn", type = "reconstructed", } m["nic-gur-pro"] = { "Gur Purba", 116773758, "nic-gur", "Latn", type = "reconstructed", } m["nic-jkn-pro"] = { "Jukunoid Purba", 116773769, "nic-jkn", "Latn", type = "reconstructed", } m["nic-lcr-pro"] = { "Cross River Hilir Purba", 116773782, "nic-lcr", "Latn", type = "reconstructed", } m["nic-ogo-pro"] = { "Ogoni Purba", 116773799, "nic-ogo", "Latn", type = "reconstructed", } m["nic-ovo-pro"] = { "Oti-Volta Purba", 116773802, "nic-ovo", "Latn", type = "reconstructed", } m["nic-plt-pro"] = { "Plateau Purba", 116773805, "nic-plt", "Latn", type = "reconstructed", } m["nic-pro"] = { "Niger-Congo Purba", 108000748, "nic", "Latn", type = "reconstructed", } m["nic-ubg-pro"] = { "Ubangi Purba", 116773818, "nic-ubg", "Latn", type = "reconstructed", } m["nic-ucr-pro"] = { "Cross River Hulu Purba", 116773819, "nic-ucr", "Latn", type = "reconstructed", } m["nic-vco-pro"] = { "Volta-Congo Purba", 116773293, "nic-vco", "Latn", type = "reconstructed", } m["njo-jgl"] = { "Ao Chungli", 55607615, "njo", "Latn", } m["njo-mng"] = { "Ao Mongsen", 85383221, "njo", "Latn", } m["nub-har"] = { "Haraza", 19572059, "nub", "Arab, Latn", } m["nub-pro"] = { "Nubia Purba", 116773246, "nub", "Latn", type = "reconstructed", } m["omq-cha-pro"] = { "Chatino Purba", 116773202, "omq-cha", "Latn", type = "reconstructed", } m["omq-maz-pro"] = { "Mazatec Purba", 116773790, "omq-maz", "Latn", type = "reconstructed", } m["omq-mix-pro"] = { "Mixtecan Purba", 21573423, "omq-mix", "Latn", type = "reconstructed", } m["omq-mxt-pro"] = { "Mixtec Purba", 21573424, "omq-mxt", "Latn", type = "reconstructed", } m["omq-otp-pro"] = { "Oto-Pamean Purba", 116773251, "omq-otp", "Latn", type = "reconstructed", } m["omq-pro"] = { "Oto-Manguean Purba", 33669, "omq", "Latn", type = "reconstructed", } m["omq-sjq"] = { "Chatino San Juan Quiahije", 138330751, "omq-cha", "Latn", } m["omq-tel"] = { "Mixtec Teposcolula", nil, "omq-mxt", "Latn", } m["omq-teo"] = { "Chatino Teojomulco", 25340451, "omq-cha", "Latn", } m["omq-tri-pro"] = { "Triqui Purba", 116773817, "omq-tri", "Latn", type = "reconstructed", } m["omq-zap-pro"] = { "Zapotecan Purba", 116773297, "omq-zap", "Latn", type = "reconstructed", } m["omq-zpc-pro"] = { "Zapotec Purba", 116773296, "omq-zpc", "Latn", type = "reconstructed", } m["omv-aro-pro"] = { "Aroid Purba", 116773721, "omv-aro", "Latn", type = "reconstructed", } m["omv-diz-pro"] = { "Dizoid Purba", 116773750, "omv-diz", "Latn", type = "reconstructed", } m["omv-pro"] = { "Omotik Purba", 116773800, "omv", "Latn", type = "reconstructed", } m["oto-otm-pro"] = { "Otomi Purba", 5908710, "oto-otm", "Latn", type = "reconstructed", } m["oto-pro"] = { "Otomian Purba", 116773252, "oto", "Latn", type = "reconstructed", } m["paa-kmn"] = { "Kómnzo", 18344310, "paa-wko", "Latn", } m["paa-kwn"] = { "Kuwani", 6449056, "qfa-unc", -- kurang dibuktikan, mungkin sama dengan atau berkaitan dengan Kalabra "Latn", } m["paa-lei"] = { "Leitre", 85776228, "paa-isk", } m["paa-nha-pro"] = { "Halmahera Utara Purba", 116773241, "paa-nha", "Latn", type = "reconstructed" } m["paa-nun"] = { "Nungon", 128807788, "ngf-ynu", "Latn", } m["phi-din"] = { "Agta Dinapigue", 16945774, "phi", "Latn", } m["phi-kal-pro"] = { "Kalamian Purba", 116773213, "phi-kal", "Latn", type = "reconstructed", } m["phi-nag"] = { "Agta Nagtipunan", 16966111, "phi", "Latn", } m["phi-pro"] = { "Filipina Purba", 18204898, "phi", "Latn", type = "reconstructed", } m["poz-abi"] = { "Abai", 19570729, "poz-san", "Latn", } m["poz-bal"] = { "Baliledo", 4850912, "poz", "Latn", } m["poz-btk-pro"] = { "Bungku-Tolaki Purba", 116773724, "poz-btk", "Latn", type = "reconstructed", } m["poz-cet-pro"] = { "Melayu-Polinesia Tengah-Timur Purba", 2269883, "poz-cet", "Latn", type = "reconstructed", } m["poz-hce-pro"] = { "Halmahera-Cenderawasih Purba", 116773209, "poz-hce", "Latn", type = "reconstructed", } m["poz-lgx-pro"] = { "Lampung Purba", 116773222, "poz-lgx", "Latn", type = "reconstructed", } m["poz-mcm-pro"] = { "Melayu-Chamik Purba", 116773225, "poz-mcm", "Latn", type = "reconstructed", } m["poz-mic-pro"] = { "Mikronesia Purba", 111939079, "poz-mic", "Latn", type = "reconstructed", } m["poz-mly-pro"] = { "Melayik Purba", 98057728, "poz-mly", "Latn", type = "reconstructed", } m["poz-msa-pro"] = { "Melayu-Sumbawa Purba", 116773226, "poz-msa", "Latn", type = "reconstructed", } m["poz-nes"] = { "Nese", 2157412, "poz-vnc", "Latn", } m["poz-oce-pro"] = { "Oceania Purba", 141741, "poz-oce", "Latn", type = "reconstructed", } m["poz-pcc-pro"] = { "Pasifik Tengah Purba", 111962726, "poz-pcc", "Latn", type = "reconstructed", } m["poz-pep-pro"] = { "Polinesia Timur Purba", 113988745, "poz-pep", "Latn", type = "reconstructed", } m["poz-pnp-pro"] = { "Polinesia Teras Purba", 113988746, "poz-pnp", "Latn", type = "reconstructed", } m["poz-pol-pro"] = { "Polinesia Purba", 1658709, "poz-pol", "Latn", type = "reconstructed", } m["poz-pro"] = { "Melayu-Polinesia Purba", 3832960, "poz", "Latn", type = "reconstructed", } m["poz-sml"] = { "Melayu Sarawak", 4251702, "poz-mly", "Latn, Arab", } m["poz-ssw-pro"] = { "Sulawesi Selatan Purba", 116773279, "poz-ssw", "Latn", type = "reconstructed", } m["poz-swa-pro"] = { "Sarawak Utara Purba", 116773243, "poz-swa", "Latn", type = "reconstructed", } m["poz-ter"] = { "Melayu Terengganu", 4207412, "poz-mly", "Latn, Arab", } m["pqe-pro"] = { "Melayu-Polinesia Timur Purba", 2269883, "pqe", "Latn", type = "reconstructed", } m["pra-niy"] = { "Prakrit Niya", 11991601, "inc-mid", "Khar", ancestors = "inc-ash", translit = "Khar-translit", } m["qfa-adm-pro"] = { "Andaman Besar Purba", 116773756, "qfa-adm", "Latn", type = "reconstructed", } m["qfa-bet-pro"] = { "Be-Tai Purba", 116773193, "qfa-bet", "Latn", type = "reconstructed", } m["qfa-cka-pro"] = { "Chukotko-Kamchatka Purba", 7251837, "qfa-cka", "Latn", type = "reconstructed", } m["qfa-hur-pro"] = { "Hurro-Urartia Purba", 116773211, "qfa-hur", "Latn", type = "reconstructed", } m["qfa-kad-pro"] = { "Kadu Purba", 116773770, "qfa-kad", "Latn", type = "reconstructed", } m["qfa-kms-pro"] = { "Kam-Sui Purba", 55630682, "qfa-kms", "Latn", type = "reconstructed", } m["qfa-kor-pro"] = { "Korea Purba", 467883, "qfa-kor", "Latn", type = "reconstructed", } m["qfa-kra-pro"] = { "Kra Purba", 7251854, "qfa-kra", "Latn", type = "reconstructed", } m["qfa-lic-pro"] = { "Hlai Purba", 7251845, "qfa-lic", "Latn", type = "reconstructed", } m["qfa-onb-pro"] = { "Be Purba", 116773192, "qfa-onb", "Latn", type = "reconstructed", } m["qfa-ong-pro"] = { "Onga Purba", 116773801, "qfa-ong", "Latn", type = "reconstructed", } m["qfa-tak-pro"] = { "Kra-Dai Purba", 104901616, "qfa-tak", "Latn", type = "reconstructed", } m["qfa-yen-pro"] = { "Yenisei Purba", 27639, "qfa-yen", "Latn", type = "reconstructed", } m["qfa-yuk-pro"] = { "Yukaghir Purba", 116773294, "qfa-yuk", "Latn", type = "reconstructed", } m["qwe-kch"] = { "Kichwa", 1740805, "qwe", "Latn", ancestors = "qu", } m["qwe-pro"] = { "Quechua Purba", 5575757, "qwe", "Latn", type = "reconstructed", } m["roa-ang"] = { "Angevin", 56782, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-bbn"] = { "Bourbonnais-Berrichon", 2899128, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-brg"] = { "Bourguignon", 508332, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-can"] = { "Cantabrian", 917021, "roa-asl", "Latn", } m["roa-cha"] = { "Champenois", 430018, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-fcm"] = { "Franc-Comtois", 510561, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-gal"] = { "Gallo", 37300, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-gib"] = { "Gallo-Italik Basilicata", 3094838, "roa-git", ancestors = "pms-old", "Latn", } m["roa-gis"] = { "Gallo-Italik Sicily", 2629019, "roa-git", "Latn", ancestors = "pms-old", } m["roa-leo"] = { "Leon", 34108, "roa-asl", "Latn", } m["roa-lor"] = { "Lorrain", 671198, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-oca"] = { "Catalonia Kuno", 15478520, "roa-ocr", "Latn", sort_key = {remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla .. "·"}, } m["roa-ole"] = { "Leon Kuno", 125977465, "roa-asl", "Latn", } m["roa-ona"] = { "Navarro-Aragon Kuno", 2736184, "roa-nar", "Latn", } m["roa-opt"] = { "Galicia-Portugis Kuno", 1072111, "roa-gap", "Latn", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ}, } m["roa-orl"] = { "Orléanais", 28497058, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-poi"] = { "Poitevin-Saintongeais", 514123, "roa-oil", "Latn", sort_key = s["roa-oil-sortkey"], } m["roa-tar"] = { "Tarantino", 695526, "roa-itr", "Latn", wikimedia_codes = "roa-tara", } m["sai-all"] = { "Allentiac", 19570789, "sai-hrp", "Latn", } m["sai-and"] = { "Andoquero", 16828359, "sai-wit", "Latn", } m["sai-ayo"] = { "Ayomán", 16937754, "sai-jir", "Latn", } m["sai-bae"] = { "Baenan", 3401998, "qfa-unc", -- pupus, kurang dibuktikan; hanya dikenali melalui 9 perkataan "Latn", } m["sai-bag"] = { "Bagua", 5390321, "qfa-unc", -- pupus, kurang dibuktikan; mungkin bahasa Carib "Latn", } m["sai-bet"] = { "Betoi", 926551, "qfa-iso", "Latn", } m["sai-bor-pro"] = { "Bora Purba", nil, "sai-bor", "Latn", } m["sai-cac"] = { "Cacán", 945482, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan "Latn", } m["sai-caq"] = { "Caranqui", 2937753, "sai-bar", "Latn", } m["sai-car-pro"] = { "Carib Purba", 116773196, "sai-car", "Latn", type = "reconstructed", } m["sai-cat"] = { "Catacao", 5051136, "sai-ctc", "Latn", } m["sai-cer-pro"] = { "Cerrado Purba", 116773200, "sai-cer", "Latn", type = "reconstructed", } m["sai-chi"] = { "Chirino", 5390321, "qfa-unc", -- pupus, hanya empat perkataan diketahui; mungkin berkaitan dengan Candoshi-Shapra (cbu) "Latn", } m["sai-chn"] = { "Chaná", 5072718, "sai-crn", "Latn", } m["sai-chp"] = { "Chapacura", 5072884, "sai-cpc", "Latn", } m["sai-chr"] = { "Charrua", 5086680, "sai-crn", "Latn", } m["sai-chu"] = { "Churuya", 5118339, "sai-guh", "Latn", } m["sai-cje-pro"] = { "Jê Tengah Purba", 116773198, "sai-cje", "Latn", type = "reconstructed", } m["sai-cmg"] = { "Comechingon", 6644203, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan "Latn", } m["sai-cno"] = { "Chono", 5104704, "qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan, mungkin palsu "Latn", } m["sai-cnr"] = { "Cañari", 5055572, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Chimuan atau Barbacoan "Latn", } m["sai-coe"] = { "Coeruna", 6425639, "sai-wit", "Latn", } m["sai-col"] = { "Colán", 5141893, "sai-ctc", "Latn", } m["sai-cop"] = { "Copallén", 5390321, "qfa-unc", -- pupus, hanya empat perkataan dibuktikan; mungkin Cholonan "Latn", } m["sai-crd"] = { "Coroado Puri", 24191321, "sai-mje", "Latn", } m["sai-ctq"] = { "Catuquinaru", 16858455, "qfa-unc", -- pupus, kurang dibuktikan; kosa kata tidak menyerupai bahasa lain "Latn", } m["sai-cul"] = { "Culli", 2879660, "qfa-unc", -- pupus, kurang dibuktikan; sering dianggap sebagai pencilan "Latn", } m["sai-cva"] = { "Cueva", 5192644, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Chocoan "Latn", } m["sai-esm"] = { "Esmeralda", 3058083, "qfa-unc", -- pupus, kurang dibuktikan; mungkin berkaitan dengan Yaruro "Latn", } m["sai-ewa"] = { "Ewarhuyana", 16898104, nil, "Latn", } m["sai-gam"] = { "Gamela", 5403661, "qfa-unc", -- pupus, kurang dibuktikan; mungkin pencilan "Latn", } m["sai-gay"] = { "Gayón", 5528902, "sai-jir", "Latn", } m["sai-gmo"] = { "Guamo", 5613495, "qfa-unc", -- pupus; "Kaufman (1990) mendapati hubungan dengan bahasa-bahasa Chapacuran meyakinkan." [Wikipedia] Dianggap sebagai pencilan oleh Campbell (2024). "Latn", } m["sai-gua"] = { "Guachí", 5613172, "sai-guc", "Latn", } m["sai-gue"] = { "Güenoa", 5626799, "sai-crn", "Latn", } m["sai-hau"] = { "Haush", 3128376, "sai-cho", "Latn", } m["sai-jee-pro"] = { "Jê Purba", 116773212, "sai-jee", "Latn", type = "reconstructed", } m["sai-jko"] = { "Jeikó", 6176527, "sai-mje", "Latn", } m["sai-jrj"] = { "Jirajara", 6202966, "sai-jir", "Latn", } m["sai-kat"] = { -- kontras xoo, kzw, sai-xoc "Katembri", 6375925, "qfa-unc", -- pupus, kurang dibuktikan; "Kaufman (1990) telah menghubungkannya dengan bahasa Taruma yang hampir pupus, walaupun ini tidak diterima oleh sarjana lain." [Wikipedia] "Latn", } m["sai-mal"] = { "Malalí", 6741212, "sai-mje", -- dianggap sebagai bahasa Maxakalían yang paling divergen (subbahagian kepada Macro-Jê), yang mana kami tiada entri "Latn", } m["sai-mar"] = { "Maratino", 6755055, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Uto-Aztecan "Latn", } m["sai-mat"] = { "Matanawi", 6786047, "qfa-unc", -- pupus; sama ada pencilan atau berkait jauh dengan bahasa-bahasa Muran; Campbell (2024) menyenarainya sebagai pencilan, Glottolog memberikannya sebagai tidak terkelas "Latn", } m["sai-mcn"] = { "Mocana", 3402048, "qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad) "Latn", } m["sai-men"] = { "Menien", 16890110, "sai-mje", "Latn", } m["sai-mil"] = { "Millcayac", 19573012, "sai-hrp", "Latn", } m["sai-mlb"] = { "Malibu", 134374036, "qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad) "Latn", } m["sai-msk"] = { "Masakará", 6782426, "sai-mje", "Latn", } m["sai-muc"] = { "Mucuchí", 6931290, nil, -- lazimnya dianggap sebagai Timotean, yang mana kami tiada entri "Latn", } m["sai-mue"] = { "Muellama", 16886936, "sai-bar", "Latn", } m["sai-muz"] = { "Muzo", 6644203, "qfa-unc", -- bahasa pupus di Colombia, kurang dibuktikan; mungkin Pijao (Cariban) "Latn", } m["sai-mys"] = { "Maynas", 16919393, "sai-cah", -- mengikut Campbell (2024); dahulu dianggap tidak terkelas "Latn", } m["sai-nat"] = { "Natú", 9006749, "qfa-unc", -- pupus, kurang dibuktikan; "hanya Greenberg yang berani mengelaskannya".[Wikipedia, memetik Moseley, Christopher; Asher, R. E.; Tait, Mary (1994), Atlas of the world's languages] "Latn", } m["sai-nje-pro"] = { "Jê Utara Purba", 116773245, "sai-nje", "Latn", type = "reconstructed", } m["sai-opo"] = { "Opón", 7099152, "sai-car", "Latn", } m["sai-oto"] = { "Otomaco", 16879234, "sai-otm", "Latn", } m["sai-pal"] = { "Palta", 3042978, "qfa-unc", -- pupus, tidak terkelas; mungkin Chicham "Latn", } m["sai-pam"] = { "Pamigua", 5908689, "sai-tin", "Latn", } m["sai-par"] = { "Paratió", 16890038, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Xukuruan "Latn", } m["sai-peb"] = { "Peba", 3373890, "sai-pey", "Latn", } m["sai-pnz"] = { "Panzaleo", 3123275, "qfa-unc", -- pupus, tidak terkelas; mungkin Paezan "Latn", } m["sai-prh"] = { "Puruhá", 3410994, "qfa-unc", -- pupus, kurang dibuktikan; mungkin dalam keluarga dengan Cañari "Latn", } m["sai-ptg"] = { "Patagón", 128807870, "sai-tar", -- pupus, hanya diketahui daripada 4 perkataan, yang mencadangkan susur galur Cariban (Campbell 2024) "Latn", } m["sai-pur"] = { "Purukotó", 7261622, "sai-pem", "Latn", } m["sai-pyg"] = { "Payaguá", 7156643, "sai-guc", "Latn", } m["sai-pyk"] = { "Pykobjê", 98113977, "sai-nje", "Latn", } m["sai-qmb"] = { "Quimbaya", 7272043, "qfa-unc", -- pupus, mungkin tidak wujud; sedikit perkataan yang diketahui "Latn", } m["sai-qtm"] = { "Quitemo", 7272651, "sai-cpc", "Latn", } m["sai-rab"] = { "Rabona", 6644203, "qfa-unc", -- pupus, kurang dibuktikan, kebanyakan nama tumbuhan; mungkin Candoshi-Shapra "Latn", } m["sai-ram"] = { "Ramanos", 16902824, "qfa-unc", -- pupus, kurang dibuktikan, mungkin pencilan; mengikut Glottolog: "senarai perkataan yang kerdil ... tidak menunjukkan persamaan yang meyakinkan dengan bahasa sekeliling" "Latn", } m["sai-sac"] = { "Sácata", 5390321, "qfa-unc", -- pupus, hanya 3 perkataan diketahui; mungkin Candoshí atau Arawak "Latn", } m["sai-san"] = { "Sanaviron", 16895999, "qfa-unc", -- pupus, tidak terkelas; tiada konsensus mengenai pengelasan "Latn", } m["sai-sap"] = { "Sapará", 7420922, "sai-car", "Latn", } m["sai-sec"] = { "Sechura", 7442912, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan "Latn", } m["sai-sin"] = { "Sinúfana", 7525275, "qfa-unc", -- hampir pupus, kurang dibuktikan; mungkin Chocoan "Latn", } m["sai-sje-pro"] = { "Jê Selatan Purba", 116773814, "sai-sje", "Latn", type = "reconstructed", } m["sai-tab"] = { "Tabancale", 5390321, "qfa-unc", -- pupus, hanya 5 perkataan diketahui; tiada kaitan yang jelas, mungkin pencilan "Latn", } m["sai-tal"] = { "Tallán", 16910468, "qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan "Latn", } m["sai-tap"] = { "Tapayuna", 30719984, "sai-nje", "Latn", } m["sai-tar-pro"] = { "Taranoan Purba", 116773816, "sai-tar", "Latn", type = "reconstructed", } m["sai-teu"] = { "Teushen", 3519243, "qfa-unc", -- mungkin pupus menjelang 1950-an; mungkin Chonan "Latn", } m["sai-tim"] = { "Timote", 7806995, nil, -- mungkin dalam keluarga Timote kecil "Latn", } m["sai-tpr"] = { "Taparita", 7684460, "sai-otm", "Latn", } m["sai-trr"] = { "Tarairiú", 7685313, "qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan "Latn", } m["sai-wai"] = { "Waitaká", 16918610, "qfa-unc", -- pupus, mungkin Purian "Latn", } m["sai-way"] = { "Wayumara", 7960726, "sai-car", "Latn", } m["sai-wit-pro"] = { "Witotoan Purba", 116773823, "sai-wit", "Latn", type = "reconstructed", } m["sai-wnm"] = { "Wanham", 16879440, "sai-cpc", "Latn", } m["sai-xoc"] = { -- kontras xoo, kzw, sai-kat "Xocó", 12953620, "qfa-unc", -- pupus dan kurang dibuktikan; tidak jelas sama ada satu atau tiga bahasa "Latn", } m["sai-yao"] = { "Yao (Amerika Selatan)", 16979655, "sai-ven", "Latn", } m["sai-yar"] = { -- bukan keluarga yang sama dengan 'suy' "Yarumá", 3505859, "sai-pek", "Latn", } m["sai-yri"] = { "Yuri", 2669157, "sai-tyu", "Latn", } m["sai-yup"] = { "Yupua", 8061430, "sai-tuc", "Latn", } m["sai-yur"] = { "Yurumanguí", 1281291, "qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan "Latn", } m["sal-pro"] = { "Salish Purba", 116773269, "sal", "Latn", type = "reconstructed", } m["sdv-daj-pro"] = { "Daju Purba", 116773739, "sdv-daj", "Latn", type = "reconstructed", } m["sdv-eje-pro"] = { "Jebel Timur Purba", 116773751, "sdv-eje", "Latn", type = "reconstructed", } m["sdv-nil-pro"] = { "Nilotik Purba", 116773794, "sdv-nil", "Latn", type = "reconstructed", } m["sdv-nyi-pro"] = { "Nyima Purba", 116773796, "sdv-nyi", "Latn", type = "reconstructed", } m["sdv-tmn-pro"] = { "Taman Purba", 116773815, "sdv-tmn", "Latn", type = "reconstructed", } m["sel-nor"] = { "Selkup Utara", 30304565, "sel", "Cyrl", translit = "sel-nor-translit", } m["sel-pro"] = { "Selkup Purba", 128884235, "sel", "Latn", type = "reconstructed", } m["sel-sou"] = { "Selkup Selatan", 30304639, "sel", "Cyrl", translit = "sel-sou-translit", } m["sem-amm"] = { "Ammon", 279181, "sem-can", "Phnx", -- translit Phnx dalam [[Module:scripts/data]] } m["sem-amo"] = { "Amor", 35941, "sem-nwe", "Xsux, Latn", } m["sem-cha"] = { "Chaha", 35543, "sem-eth", "Ethi", translit = "Ethi-translit", } m["sem-dad"] = { "Dadan", 21838040, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-dum"] = { "Dumait", 128810397, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-has"] = { "Hasait", 3541433, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-his"] = { "Hisma", 22948260, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-mhr"] = { "Muher", 33743, "sem-eth", "Latn", } m["sem-pro"] = { "Samiah Purba", 1658554, "sem", "Latn", type = "reconstructed", } m["sem-saf"] = { "Safait", 472586, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-sam"] = { "Samal", 85847147, "sem-nwe", "Phnx", -- translit Phnx dalam [[Module:scripts/data]] } m["sem-srb"] = { "Arab Selatan Kuno", 35025, "sem-osa", "Sarb", -- translit Sarb dalam [[Module:scripts/data]] } m["sem-tay"] = { "Tayman", 24912301, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-tha"] = { "Thamud", 843030, "sem-cen", "Narb", -- translit Narb dalam [[Module:scripts/data]] } m["sem-wes-pro"] = { "Samiah Barat Purba", 98021726, "sem-wes", "Latn", type = "reconstructed", } m["sio-pro"] = { -- PERHATIAN ini bukan 'nai-sca-pro' "Proto-Siouan-Catawban" iaitu Proto-Sioux Barat "Sioux Purba", 34181, "sio", "Latn", type = "reconstructed", } m["sit-aao-pro"] = { "Naga Tengah Purba", nil, "sit-aao", "Latn", type = "reconstructed", } m["sit-bai-pro"] = { "Bai Purba", nil, "sit-bai", "Latn", type = "reconstructed", } m["sit-ban"] = { "Bangru", 56071779, "sit-hrs", "Latn", } m["sit-bdi-pro"] = { "Bodish Purba", nil, "sit-bdi", "Latn", type = "reconstructed", } m["sit-bok"] = { "Bokar", 4938727, "sit-tan", "Latn, Tibt", override_translit = true, -- translit, display_text, strip_diacritics, sort_key Tibt dalam [[Module:scripts/data]] } m["sit-cai"] = { "Caijia", 5017528, "sit-cln", "Latn" } m["sit-cha"] = { "Chairel", 5068066, "sit-luu", "Latn", } m["sit-ers-pro"] = { "Ersu Purba", nil, "sit-ers", "Latn", type = "reconstructed", } m["sit-hrs-pro"] = { "Hrusish Purba", 116773762, "sit-hrs", "Latn", type = "reconstructed", } m["sit-jap"] = { "Japhug", 3162245, "sit-egy", "Latn", } m["sit-kha-pro"] = { "Kham Purba", 116773773, "sit-kha", "Latn", type = "reconstructed", } m["sit-khb-pro"] = { "Kho-Bwa Purba", nil, "sit-khb", "Latn", type = "reconstructed", } m["sit-khp-pro"] = { "Puroik Purba", nil, "sit-khb", "Latn", type = "reconstructed", } m["sit-khw-pro"] = { "Kho-Bwa Barat Purba", nil, "sit-khw", "Latn", type = "reconstructed", } m["sit-kon-pro"] = { "Naga Utara Purba", nil, "sit-kon", "Latn", type = "reconstructed", } m["sit-liz"] = { "Lizu", 6660653, "sit-ers", "Latn", -- dan Ersu Shaba } m["sit-lnj"] = { "Longjia", 17096251, "sit-cln", "Latn" } m["sit-lrn"] = { "Luren", 16946370, "sit-cln", "Latn" } m["sit-luu-pro"] = { "Luish Purba", 116773783, "sit-luu", "Latn", type = "reconstructed", } m["sit-nas-pro"] = { "Naish Purba", nil, "sit-nas", "Latn", type = "reconstructed", } m["sit-prn"] = { "Puiron", 7259048, "sit-zem", } m["sit-pro"] = { "Sino-Tibet Purba", 24839178, "sit", "Latn", type = "reconstructed", } m["sit-sit"] = { "Situ", 19840830, "sit-egy", "Latn", } m["sit-tam-pro"] = { "Tamang Purba", 117469295, "sit-tam", "Latn", type = "reconstructed", } m["sit-tan-pro"] = { "Tani Purba", 116773284, "sit-tan", "Latn", -- memerlukan pengesahan type = "reconstructed", } m["sit-tgm"] = { "Tangam", 17041370, "sit-tan", "Latn", } m["sit-tng-pro"] = { "Tangkhul Purba", nil, "sit-tng", "Latn", type = "reconstructed", } m["sit-tos"] = { "Tosu", 7827899, "sit-ers", "Latn", -- juga Ersu Shaba } m["sit-tsh"] = { "Tshobdun", 19840950, "sit-egy", "Latn", } m["sit-zbu"] = { "Zbu", 19841106, "sit-egy", "Latn", } m["sla-pro"] = { "Slav Purba", 747537, "sla", "Latn", type = "reconstructed", strip_diacritics = { remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve, remove_exceptions = {'ś'}, }, sort_key = { from = {"č", "ď", "ě", "ę", "ь", "ľ", "ň", "ǫ", "ř", "š", "ś", "ť", "ъ", "ž"}, to = {"c²", "d²", "e²", "e³", "i²", "l²", "nj", "o²", "r²", "s²", "s³", "t²", "u²", "z²"}, } } m["smi-pro"] = { "Sami Purba", 7251862, "smi", "Latn", type = "reconstructed", sort_key = { from = {"ā", "č", "δ", "[ëē]", "ŋ", "ń", "ō", "š", "θ", "%([^()]+%)"}, to = {"a", "c²", "d", "e", "n²", "n³", "o", "s²", "t²"} }, } m["son-pro"] = { "Songhai Purba", 116773277, "son", "Latn", type = "reconstructed", } m["sqj-pro"] = { "Albania Purba", 18210846, "sqj", "Latn", type = "reconstructed", } m["ssa-klk-pro"] = { "Kuliak Purba", 116773779, "ssa-klk", "Latn", type = "reconstructed", } m["ssa-kom-pro"] = { "Koma Purba", 116773775, "ssa-kom", "Latn", type = "reconstructed", } m["ssa-pro"] = { "Nilo-Sahara Purba", 116773236, "ssa", "Latn", type = "reconstructed", } m["syd-pro"] = { "Samoyed Purba", 7251863, "syd", "Latn", type = "reconstructed", } m["tai-pro"] = { "Tai Purba", 6583709, "tai", "Latn", type = "reconstructed", } m["tai-swe-pro"] = { "Tai Barat Daya Purba", 116773280, "tai-swe", "Latn", type = "reconstructed", } m["tbq-bdg-pro"] = { "Bodo-Garo Purba", 116773195, "tbq-bdg", "Latn", type = "reconstructed", } m["tbq-blg"] = { "Bailang", 2879843, "tbq-lob", "Hani", sort_key = "Hani-sortkey", } m["tbq-brm-pro"] = { "Burma Purba", nil, "tbq-brm", "Latn", type = "reconstructed", } m["tbq-gkh"] = { "Gokhy", 5578069, "tbq-sil", "Latn", } m["tbq-kuk-pro"] = { "Kuki-Chin Purba", 116773220, "tbq-kuk", "Latn", type = "reconstructed", } m["tbq-lal-pro"] = { "Lalo Purba", 116773781, "tbq-lal", "Latn", type = "reconstructed", } m["tbq-laz"] = { "Laze", 17007626, "sit-nas", "Latn", } m["tbq-lob-pro"] = { "Lolo-Burma Purba", 116773224, "tbq-lob", "Latn", type = "reconstructed", } m["tbq-lol-pro"] = { "Lolo Purba", 7251855, "tbq-lol", "Latn", type = "reconstructed", } m["tbq-mil"] = { "Milang", 6850761, "sit-gsi", "Deva, Latn", } m["tbq-mor"] = { "Moran", 6909216, "tbq-bdg", "Latn", } m["tbq-ngo"] = { "Ngochang", 56582, "tbq-brm", "Latn", } -- tbq-pro kini khusus etimologi m["trk-dkh"] = { "Dukhan", 12809273, "trk-ssb", "Latn, Cyrl, Mong", -- translit, display_text dan strip_diacritics Mong dalam [[Module:scripts/data]] } -- Seperti yang diuraikan dalam ''Dīwān Lughāt al-Turk'' karya Mahmud al-Kashgari abad ke-11. m["trk-eog"] = { "Oghuz Kuno Awal", nil, "trk-ogz", "Arab", strip_diacritics = {Arab = "ar-stripdiacritics"}, } m["trk-oat"] = { "Turki Anatolia Kuno", 7083390, "trk-ogz", "Arab", strip_diacritics = {Arab = "ar-stripdiacritics"}, ancestors = "trk-eog", } m["trk-pro"] = { "Turkik Purba", 3657773, "trk", "Latn", type = "reconstructed", standard_chars = { Latn = " ()-abdegiklmnoprstuxyzïöüāčēīĺŋōŕšūǖȫẹ" .. c.macron, } } m["tup-gua-pro"] = { "Tupi-Guarani Purba", 116773288, "tup-gua", "Latn", type = "reconstructed", } m["tup-kab"] = { "Kabishiana", 15302988, "tup", "Latn", } m["tup-kaw"] = { "Kawahiva", 6346712, "tup-gua", "Latn", } m["tup-pro"] = { "Tupi Purba", 10354700, "tup", "Latn", type = "reconstructed", } m["tuw-alk"] = { "Alchuka", 113553616, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-bal"] = { "Bala", 86730632, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-kkl"] = { "Kyakala", 118875708, "tuw-jrc", "Latn, Hans", sort_key = {Hans = "Hani-sortkey"}, } m["tuw-kli"] = { "Kili", 6406892, "tuw-ewe", "Cyrl", } m["tuw-pro"] = { "Tungus Purba", 85872335, "tuw", "Latn", type = "reconstructed", } m["tuw-sol"] = { "Solon", 30004, "tuw-ewe", } m["urj-fin-pro"] = { "Finnik Purba", 11883720, "urj-fin", "Latn", type = "reconstructed", } m["urj-koo"] = { "Komi Kuno", 86679962, "kv", "Perm, Cyrs", translit = "urj-koo-translit", -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]; sebelum ini, strip_diacritics Cyrs tidak hadir } m["urj-kuk"] = { "Kukkuzi", 107410460, "urj-fin", "Latn", ancestors = "vot", } m["urj-kya"] = { "Komi-Yazva", 2365210, "kv", "Cyrl", translit = "kv-translit", override_translit = true, strip_diacritics = {remove_diacritics = c.acute}, } m["urj-mdv-pro"] = { "Mordvinik Purba", 116773232, "urj-mdv", "Latn", type = "reconstructed", } m["urj-prm-pro"] = { "Permik Purba", 116773257, "urj-prm", "Latn", type = "reconstructed", } m["urj-pro"] = { "Uralik Purba", 288765, "urj", "Latn", type = "reconstructed", } m["urj-ugr-pro"] = { "Ugrik Purba", 156631, "urj-ugr", "Latn", type = "reconstructed", } m["xnd-pro"] = { "Na-Dene Purba", 116773233, "xnd", "Latn", type = "reconstructed", } m["xgn-pro"] = { "Mongol Purba", 2493677, "xgn", "Latn", type = "reconstructed", sort_key = { from = {"č", "i", "ï", "ǰ", "ŋ", "ö", "š", "ü"}, to = {"c", "i" .. p[1], "i", "j", "n" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1]}, }, } m["yok-bvy"] = { "Yokuts Buena Vista", 4985474, "yok", "Latn", } m["yok-dly"] = { "Yokuts Delta", 70923266, "yok", "Latn", } m["yok-gsy"] = { "Yokuts Gashowu", 3098708, "yok", "Latn", } m["yok-kry"] = { "Yokuts Sungai Kings", 6413014, "yok", "Latn", } m["yok-nvy"] = { "Yokuts Lembah Utara", 85789777, "yok", "Latn", } m["yok-ply"] = { "Yokuts Palewyami", 2387391, "yok", "Latn", } m["yok-svy"] = { "Yokuts Lembah Selatan", 12642473, "yok", "Latn", } m["yok-tky"] = { "Yokuts Tule-Kaweah", 7851988, "yok", "Latn", } m["ypk-pro"] = { "Yupik Purba", 116773295, "ypk", "Latn", type = "reconstructed", } m["yrk-for"] = { "Nenets Hutan", 1295107, "yrk", "Cyrl", translit = "yrk-for-translit", strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.dotabove}, } m["yrk-tun"] = { "Nenets Tundra", 36452, "yrk", "Cyrl", strip_diacritics = { from = {"ӑ", "а̄", "э̇", "ӣ", "ы̄", "ӯ", "ю̄", "я̆", "я̄"}, to = {"а", "а", "э", "и", "ы", "у", "ю", "я", "я"}, }, translit = "yrk-tun-translit", } m["zhx-min-pro"] = { "Min Purba", 19646347, "zhx-min", "Latn", type = "reconstructed", } m["zhx-sht"] = { "Tuhua Shaozhou", 1920769, "zhx", "Nshu, Hants", generate_forms = "zh-generateforms", sort_key = {Hani = "Hani-sortkey"}, } m["zhx-sic"] = { "Sichuan", 2278732, "zhx-man", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["zhx-tai"] = { "Taishan", 2208940, "zhx-yue", "Hants", generate_forms = "zh-generateforms", translit = "zh-translit", sort_key = "Hani-sortkey", } m["zle-ono"] = { "Novgorod Kuno", 162013, "zle", "Cyrs, Glag", translit = {Cyrs = "Cyrs-translit", Glag = "Glag-translit"}, -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]] } m["zle-ort"] = { "Ruthenia Kuno", 13211, "zle", "Arab, Cyrs, Latn", ancestors = "orv", translit = { Cyrs = "zle-ort-translit", Arab = "zle-ort-Arab-translit", }, strip_diacritics = { Cyrs = { remove_diacritics = m_langdata.chars_substitutions["Cyrs_remove_diacritics"], remove_exceptions = {"Ї", "ї"}, }, Arab = "ar-stripdiacritics", }, -- sort_key Cyrs dalam [[Module:scripts/data]] } m["zls-chs"] = { "Slav Gereja", 33251, "zls", "Cyrs, Glag, Latn", ancestors = "cu", translit = { Cyrs = "Cyrs-translit", Glag = "Glag-translit" }, -- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]] } m["zlw-ocs"] = { "Czech Kuno", 593096, "zlw", "Latn", } m["zlw-opl"] = { "Poland Kuno", 149838, "zlw-lch", "Latn", strip_diacritics = {remove_diacritics = c.ringabove}, } m["zlw-osk"] = { "Slovak Kuno", 12776676, "zlw", "Latn", } m["zlw-slv"] = { "Slovincia", 36822, "zlw-pom", "Latn", strip_diacritics = {remove_diacritics = c.macron .. c.breve}, } -- Kod tambahan untuk bahasa-bahasa yang digunakan di Malaysia, yang tidak wujud di Wikikamus Bahasa Inggeris m["zlm-coa"] = { "Melayu Terengganu Pesisir", 4207412, "poz-mly", "Latn, ms-Arab", } m["zlm-pah"] = { "Melayu Pahang", 7310370, "poz-mly", "Latn", } return require("Module:languages").finalizeData(m, "language") dluvfqb7mu3yt1ue6qh0s4xr04usufo Modul:languages/data/3/k/extra 828 33767 375378 373746 2026-09-22T05:36:05Z Hakimi97 2668 [[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]] 375378 Scribunto text/plain local m = {} m["kaa"] = { aliases = {"Qaraqalpaq"}, } m["kab"] = { aliases = {"Kabylian"}, } m["kac"] = { aliases = {"Kachin"}, } m["kad"] = { } m["kae"] = { } m["kaf"] = { otherNames = {"Kazhuo"}, } m["kag"] = { } m["kah"] = { otherNames = {"Kara"}, } m["kai"] = { } m["kaj"] = { } m["kak"] = { } m["kam"] = { otherNames = {"Kikamba", "Kamba (Kenya)"}, } m["kao"] = { otherNames = {"Khasonke", "Kasonke", "Khassonké"}, } m["kap"] = { otherNames = {"Bezheta", "Kapucha", "Bezhita"}, } m["kaq"] = { otherNames = {"Kapanawa"}, } m["kaw"] = { aliases = {"Kawi"}, } m["kax"] = { } m["kay"] = { } m["kba"] = { } m["kbb"] = { otherNames = {"Kachuyana", "Kaxuiana", "Kaxuiâna", "Kashuyana"}, } m["kbc"] = { otherNames = {"Caduveo", "Ediu-Adig", "Guaicurú", "Kadiweu", "Mbayá", "Mbayá-Guaycuru", "Waikurú"}, } m["kbd"] = { aliases = {"East Circassian"}, } m["kbe"] = { otherNames = {"Kaanytju", "Kandju", "Kaantyu", "Gandju", "Gandanju", "Kamdhue", "Kandyu", "Kanyu"}, } m["kbh"] = { } m["kbi"] = { } m["kbj"] = { otherNames = {"Kare", "Kare (Central African Republic)", "Bantoid Kare"}, } m["kbk"] = { otherNames = {"Koiari"}, } m["kbm"] = { } m["kbn"] = { otherNames = {"Kare (Central African Republic)", "Mbum Kare"}, } m["kbo"] = { } m["kbp"] = { otherNames = {"Kabiye", "Kabye"}, } m["kbq"] = { } m["kbr"] = { } m["kbs"] = { } m["kbt"] = { } m["kbu"] = { } m["kbv"] = { otherNames = {"Dera", "Dera (New Guinea)"}, } m["kbw"] = { } m["kbx"] = { } m["kbz"] = { } m["kcb"] = { } m["kcc"] = { } m["kcd"] = { } m["kce"] = { } m["kcf"] = { } m["kcg"] = { } m["kch"] = { } m["kci"] = { } m["kcj"] = { } m["kck"] = { } m["kcl"] = { otherNames = {"Kela", "Gela"}, } m["kcm"] = { } m["kcn"] = { otherNames = {"Ki-Nubi"}, } m["kco"] = { } m["kcp"] = { } m["kcq"] = { } m["kcr"] = { } m["kcs"] = { } m["kct"] = { } m["kcu"] = { otherNames = {"Kami"}, } m["kcv"] = { } m["kcw"] = { } m["kcx"] = { } m["kcy"] = { } m["kcz"] = { } m["kda"] = { otherNames = {"Gadang", "Gadhang", "Gadjang", "Kattang", "Kutthung"}, } m["kdc"] = { } m["kdd"] = { } m["kde"] = { } m["kdf"] = { } m["kdg"] = { } m["kdh"] = { } m["kdi"] = { otherNames = {"Kuman"}, } m["kdj"] = { } m["kdk"] = { } m["kdl"] = { } m["kdm"] = { } m["kdn"] = { } m["kdp"] = { } m["kdq"] = { } m["kdr"] = { } m["kdt"] = { } m["kdu"] = { otherNames = {"Kedaru", "Debri"}, -- Debri is subsumed for now as it lacks an ISO code, may need to be split } m["kdv"] = { otherNames = {"Kadu"}, } m["kdw"] = { } m["kdx"] = { } m["kdy"] = { } m["kdz"] = { otherNames = {"Ndaktup", "Ncha", "Bitwi"}, } m["kea"] = { otherNames = {"Cape Verdean Creole", "Kriolu", "Creole", "Barlavento", "Sotavento"}, } m["keb"] = { } m["kec"] = { } m["ked"] = { } m["kee"] = { } m["kef"] = { } m["keg"] = { } m["keh"] = { } m["kei"] = { } m["kej"] = { } m["kek"] = { aliases = {"Qʼeqchi"} } m["kel"] = { otherNames = {"Kela", "Yela"}, } m["kem"] = { } m["ken"] = { } m["keo"] = { } m["kep"] = { } m["keq"] = { } m["ker"] = { } m["kes"] = { } m["ket"] = { } m["keu"] = { } m["kev"] = { } m["kew"] = { otherNames = {"West Kewa", "East Kewa", "South Kewa", "Erave", "Pasuma"}, } m["kex"] = { } m["key"] = { } m["kez"] = { } m["kfa"] = { } m["kfb"] = { } m["kfc"] = { } m["kfd"] = { } m["kfe"] = { otherNames = {"Kota"}, } m["kff"] = { } m["kfg"] = { } m["kfh"] = { } m["kfi"] = { } m["kfj"] = { } m["kfk"] = { } m["kfl"] = { } m["kfn"] = { } m["kfo"] = { otherNames = {"Koro", "Koro Jula"}, -- the last name is misleading, as Jula is a diff. language } m["kfp"] = { } m["kfq"] = { } m["kfr"] = { aliases = {"Kutchi", "Cutchi", "Kachchhi", "Kutchhi"}, } m["kfs"] = { } m["kft"] = { } m["kfu"] = { } m["kfv"] = { } m["kfw"] = { otherNames = {"Kharam"}, } m["kfx"] = { otherNames = {"Kullu"}, } m["kfy"] = { } m["kfz"] = { } m["kga"] = { } m["kgb"] = { } m["kgd"] = { } m["kge"] = { } m["kgf"] = { } m["kgg"] = { } m["kgi"] = { } m["kgj"] = { } m["kgk"] = { } m["kgl"] = { } m["kgn"] = { otherNames = {"Keringani"}, } m["kgo"] = { } m["kgp"] = { } m["kgq"] = { } m["kgr"] = { } m["kgs"] = { } m["kgt"] = { } m["kgu"] = { } m["kgv"] = { } m["kgw"] = { } m["kgx"] = { } m["kgy"] = { } m["kha"] = { } m["khb"] = { aliases = {"Lue", "Tai Lü", "Tai Lue", "Dai Lue"}, } m["khc"] = { } m["khd"] = { } m["khe"] = { } m["khf"] = { } m["khh"] = { } m["khj"] = { } m["khl"] = { } m["kho"] = { } m["khp"] = { } m["khq"] = { otherNames = {"Western Songhay", "Koyra Chiini Songhay"}, } m["khr"] = { } m["khs"] = { } m["kht"] = { aliases = {"Tai Khamti"}, } m["khu"] = { } m["khv"] = { otherNames = {"Khwarshi", "Xvarshi", "Inkhokvari"}, } m["khw"] = { } m["khx"] = { } m["khy"] = { otherNames = {"Kele", "Kele (Congo)", "Kele (Democratic Republic of the Congo)", "Lokele"}, } m["khz"] = { } m["kia"] = { } m["kib"] = { } m["kic"] = { } m["kid"] = { } m["kie"] = { } m["kif"] = { } m["kig"] = { } m["kih"] = { } m["kii"] = { otherNames = {"Kichai"}, } m["kij"] = { } m["kil"] = { } m["kim"] = { otherNames = {"Tofalar", "Karagas"}, } m["kio"] = { } m["kip"] = { } m["kiq"] = { } m["kis"] = { } m["kit"] = { } m["kiv"] = { } m["kiw"] = { } m["kix"] = { } m["kiy"] = { otherNames = {"Faia"}, } m["kiz"] = { } m["kja"] = { } m["kjb"] = { } m["kjc"] = { } m["kjd"] = { } m["kje"] = { } m["kjg"] = { } m["kjh"] = { } m["kji"] = { } m["kjj"] = { otherNames = {"Khinalig", "Xinalug", "Xinalugh", "Khinalugh"}, } m["kjk"] = { } m["kjl"] = { } m["kjm"] = { } m["kjn"] = { otherNames = {"Uw Oykangand", "Uw Olkola", "Olkol", "Olgolo", "Uw-Oykangand", "Uw-Olgol", "Koko Wanggara", "Ogh-Undjan", "Undjan", "Kawarrangg", "Athima", "Uw", "Kunjen-Undjan-Athima"}, } m["kjo"] = { } m["kjp"] = { aliases = {"Phlou", "Eastern Pwo Karen"}, } m["kjq"] = { } m["kjr"] = { } m["kjs"] = { } m["kjt"] = { aliases = {"Phrae Pwo Karen", "Northeastern Pwo", "Northeastern Pwo Karen"}, } m["kju"] = { } m["kjx"] = { otherNames = {"Keriaka"}, } m["kjy"] = { } m["kjz"] = { } m["kka"] = { } m["kkb"] = { } m["kkc"] = { } m["kkd"] = { } m["kke"] = { } m["kkf"] = { } m["kkg"] = { } m["kkh"] = { aliases = {"Tai Khün", "Dai Kun"}, } m["kki"] = { otherNames = {"Kaguru"}, } m["kkj"] = { } m["kkk"] = { } m["kkl"] = { } m["kkm"] = { } m["kkn"] = { } m["kko"] = { otherNames = {"Kithonirishe"}, } m["kkp"] = { otherNames = {"Kok-Kaper", "Gugubera", "Koko-Pera"}, } m["kkq"] = { } m["kkr"] = { otherNames = {"Kir"}, } m["kks"] = { otherNames = {"Giiwo"}, } m["kkt"] = { } m["kku"] = { } m["kkv"] = { } m["kkw"] = { } m["kkx"] = { } m["kky"] = { } m["kkz"] = { } m["kla"] = { otherNames = {"Klamath"}, } m["klb"] = { } m["klc"] = { } m["kld"] = { otherNames = {"Kamilaroi", "Kamilarai", "Kamalarai", "Gamilaroi"}, } m["kle"] = { } m["klf"] = { } m["klg"] = { } m["klh"] = { } m["kli"] = { } m["klj"] = { otherNames = {"Turkic Khalaj", "Arghu"}, } m["klk"] = { otherNames = {"Kono"}, } m["kll"] = { } m["klm"] = { otherNames = {"Migum"}, } m["kln"] = { } m["klo"] = { } m["klp"] = { } m["klq"] = { } m["klr"] = { } m["kls"] = { } m["klt"] = { } m["klu"] = { } m["klv"] = { } m["klw"] = { otherNames = {"Tado"}, } m["klx"] = { } m["kly"] = { } m["klz"] = { } m["kma"] = { } m["kmb"] = { otherNames = {"North Mbundu"}, } m["kmc"] = { aliases = {"Southern Gam", "Southern Dong"}, } m["kmd"] = { } m["kme"] = { } m["kmf"] = { otherNames = {"Kare", "Kare (Papua New Guinea)"}, } m["kmg"] = { } m["kmh"] = { } m["kmi"] = { } m["kmj"] = { otherNames = {"Kumarbhag", "Kumarbhag Pahariya", "Kumar Paharia", "Malto"}, } m["kmk"] = { } m["kml"] = { otherNames = {"Lower Tanudan Kalinga", "Upper Tanudan Kalinga"}, } m["kmm"] = { otherNames = {"Kom"}, } m["kmn"] = { } m["kmo"] = { } m["kmp"] = { } m["kmq"] = { } m["kmr"] = { aliases = {"Kurmanji"}, } m["kms"] = { } m["kmt"] = { } m["kmu"] = { } m["kmv"] = { otherNames = {"Karipúna French Creole", "Amapá French Creole"}, } m["kmw"] = { otherNames = {"Kikomo", "Komo (Democratic Republic of the Congo)", "Komo", "Kikumu"}, } m["kmx"] = { } m["kmy"] = { } m["kmz"] = { otherNames = {"Khorasani Turkic"}, } m["kna"] = { otherNames = {"Dera", "Dera (Nigeria)"}, } m["knb"] = { } m["knd"] = { } m["kne"] = { aliases = {"Kankana-ey"}, } m["knf"] = { } m["kni"] = { } m["knj"] = { otherNames = {"Acateco", "Western Kanjobal"}, } m["knk"] = { } m["knl"] = { } m["knm"] = { -- two unrelated lects have this name; this is the Katukinian one otherNames = {"Kanamarí", "Katukina-Kanamari", "Kanamare", "Katukína", "Katukina"}, } m["kno"] = { otherNames = {"Kono", "Konnoh"}, } m["knp"] = { } m["knq"] = { } m["knr"] = { } m["kns"] = { } m["knt"] = { otherNames = {"Panoan Katukína", "Katukína", "Catuquina", "Waninawa", "Waninnawa", "Kamanawa", "Kamannaua", "Katukina do Jurua", "Katukina of Olinda", "Katukina of Sete Estreles", "Kanamari"}, } m["knu"] = { -- a dialect of 'kpe' otherNames = {"Kono"}, } m["knv"] = { } m["knx"] = { otherNames = {"Salako", "Selako", "Ahe"}, } m["kny"] = { } m["knz"] = { } m["koa"] = { } m["koc"] = { } m["kod"] = { } m["koe"] = { } m["kof"] = { } m["kog"] = { otherNames = {"Kogi", "Cogi", "Kagaba", "Cagaba", "Cágaba"}, } m["koh"] = { } m["koi"] = { } m["kok"] = { } m["kol"] = { otherNames = {"Kol", "Kol (Papua New Guina)"}, } m["koo"] = { } m["kop"] = { aliases = {"Waupe", "Kwato"}, } m["koq"] = { aliases = {"iKota", "Ikota", "Kota"}, } m["kos"] = { } m["kot"] = { aliases = {"Logone"}, } m["kou"] = { } m["kov"] = { } m["kow"] = { } m["koy"] = { otherNames = {"Denaakk'e"}, } m["koz"] = { } m["kpa"] = { } m["kpb"] = { } m["kpc"] = { otherNames = {"Kurripako"}, } m["kpd"] = { } m["kpe"] = { } m["kpf"] = { } m["kpg"] = { } m["kph"] = { } m["kpi"] = { } m["kpj"] = { } m["kpk"] = { } m["kpl"] = { } m["kpm"] = { aliases = {"K'Ho"}, } m["kpn"] = { } m["kpo"] = { } m["kpq"] = { } m["kpr"] = { } m["kps"] = { } m["kpt"] = { } m["kpu"] = { } m["kpv"] = { otherNames = {"Komi"}, } m["kpw"] = { } m["kpx"] = { otherNames = {"Mountain Koiali"}, } m["kpy"] = { } m["kpz"] = { } m["kqa"] = { } m["kqb"] = { } m["kqc"] = { } m["kqd"] = { } m["kqe"] = { } m["kqf"] = { } m["kqg"] = { } m["kqh"] = { } m["kqi"] = { } m["kqj"] = { } m["kqk"] = { } m["kql"] = { } m["kqm"] = { } m["kqn"] = { otherNames = {"Chikaonde", "Kawonde"}, } m["kqo"] = { } m["kqp"] = { } m["kqq"] = { } m["kqr"] = { } m["kqs"] = { } m["kqt"] = { } m["kqu"] = { } m["kqv"] = { } m["kqw"] = { } m["kqx"] = { } m["kqy"] = { } m["kqz"] = { } m["kra"] = { } m["krb"] = { } m["krc"] = { } m["krd"] = { } m["kre"] = { } m["krf"] = { otherNames = {"Koro"}, } m["krh"] = { } m["kri"] = { otherNames = {"Sierra Leonean Creole"}, } m["krj"] = { } m["krk"] = { } m["krl"] = { varieties = { { "North Karelian", "Northern Karelian" }, { "South Karelian", "Southern Karelian" }, { "Tver Karelian" } } } m["krm"] = { } m["krn"] = { } m["krp"] = { } m["krr"] = { otherNames = {"Krung", "Kreung", "Krüng"}, } m["krs"] = { otherNames = {"Gbaya"}, } m["kru"] = { aliases = {"Kurux"}, } m["krv"] = { otherNames = {"Kravet"}, } m["krw"] = { } m["krx"] = { } m["kry"] = { otherNames = {"Kryc", "Kryz"}, varieties = {"Jek", "Dzhek", "Cek", "Khaput", "Yergyudzh", "Alyk"}, } m["krz"] = { } m["ksa"] = { } m["ksb"] = { otherNames = {"Shambaa"}, } m["ksc"] = { } m["ksd"] = { otherNames = {"Kuanua"}, } m["kse"] = { } m["ksf"] = { } m["ksg"] = { } m["ksi"] = { } m["ksj"] = { } m["ksk"] = { } m["ksl"] = { } m["ksm"] = { } m["ksn"] = { } m["kso"] = { } m["ksp"] = { } m["ksq"] = { } m["ksr"] = { } m["kss"] = { } m["kst"] = { } m["ksu"] = { } m["ksv"] = { } m["ksw"] = { aliases = {"S'gaw Kayin", "S'gaw", "Sgaw", "White Karen"}, } m["ksx"] = { } m["ksy"] = { } m["ksz"] = { } m["kta"] = { } m["ktb"] = { } m["ktc"] = { } m["ktd"] = { } m["ktf"] = { } m["ktg"] = { otherNames = {"Kalkutungu", "Galgadungu", "Kalkutung", "Kalkadoon", "Galgaduun"}, } m["kth"] = { } m["kti"] = { otherNames = {"Kati"}, } m["ktj"] = { } m["ktk"] = { } m["ktl"] = { } m["ktm"] = { } m["ktn"] = { otherNames = {"Caritiana"}, } m["kto"] = { } m["ktp"] = { otherNames = {"Khatu"}, } m["ktq"] = { } m["kts"] = { } m["ktt"] = { } m["ktu"] = { otherNames = {"Munukutuba", "Kikongo-Kituba", "Kikongo", "Kikongo ya leta", "Kibulamatadi", "Kikwango", "Ikeleve", "Kizabave"}, } m["ktv"] = { } m["ktw"] = { otherNames = {"Cahto"}, } m["ktx"] = { } m["kty"] = { otherNames = {"Kango (Bas-Uélé District)"}, -- distinct in name, but not necessarily in identity, from 'kzy' } m["ktz"] = { otherNames = {"Zhuǀ'hoan", "ǂKxʼauǁʼein", "ǁAuǁei", "ǁAuǁen", "Auen", "Kaukau", "Koko", "Kung-Gobabis", "‡Kx'auǁ'ei", "ǂKx'auǁ'ein", "ǁX'auǁ'e", "Juǀ'hoansi", "Juǀʼhoan"}, } m["kub"] = { } m["kuc"] = { } m["kud"] = { otherNames = {"'Auhelawa"}, } m["kue"] = { otherNames = {"Simbu", "Chimbu"}, } m["kuf"] = { } m["kug"] = { } m["kuh"] = { } m["kui"] = { otherNames = {"Kuikúro-Kalapálo", "Kuikuro", "Apalakiri"}, } m["kuj"] = { } m["kuk"] = { } m["kul"] = { otherNames = {"Tof", "Korom Boye", "Akandi", "Akande", "Kande", "Richa"}, } m["kum"] = { } m["kun"] = { } m["kuo"] = { } m["kup"] = { } m["kuq"] = { } m["kus"] = { } m["kut"] = { } m["kuu"] = { } m["kuv"] = { } m["kuw"] = { } m["kux"] = { } m["kuy"] = { } m["kuz"] = { } m["kva"] = { } m["kvb"] = { } m["kvc"] = { } m["kvd"] = { otherNames = {"Kui"}, } m["kve"] = { } m["kvf"] = { } m["kvg"] = { } m["kvh"] = { } m["kvi"] = { } m["kvj"] = { } m["kvk"] = { } m["kvl"] = { } m["kvm"] = { } m["kvn"] = { } m["kvo"] = { } m["kvp"] = { } m["kvq"] = { } m["kvr"] = { } m["kvt"] = { } m["kvu"] = { } m["kvv"] = { } m["kvw"] = { } m["kvx"] = { } m["kvy"] = { } m["kvz"] = { } m["kwa"] = { } m["kwb"] = { otherNames = {"Kwa"}, } m["kwc"] = { } m["kwd"] = { } m["kwe"] = { } m["kwf"] = { } m["kwg"] = { } m["kwh"] = { } m["kwi"] = { otherNames = {"Awa", "Cuaiquer", "Awa Pit", "Awapit", "Kwaiker", "Coaiquer", "Quaiquer"}, } m["kwj"] = { } m["kwk"] = { aliases = {"Kwakʼwala"} } m["kwl"] = { } m["kwm"] = { } m["kwn"] = { } m["kwo"] = { } m["kwp"] = { } m["kwq"] = { } m["kwr"] = { } m["kws"] = { } m["kwt"] = { } m["kwu"] = { } m["kwv"] = { otherNames = {"Sara Dunjo"}, } m["kww"] = { } m["kwx"] = { } m["kwz"] = { } m["kxa"] = { } m["kxb"] = { } m["kxc"] = { } m["kxd"] = { aliases = {"Brunei"}, } m["kxe"] = { } m["kxf"] = { } m["kxh"] = { } m["kxi"] = { otherNames = {"Nabay", "Nabaay"}, } m["kxj"] = { } m["kxk"] = { } m["kxm"] = { aliases = {"Thai Khmer", "Surin Khmer"}, } m["kxn"] = { otherNames = {"Tanjong", "Kanowit-Tanjong Melanau"}, } m["kxo"] = { } m["kxp"] = { } m["kxq"] = { } m["kxr"] = { otherNames = {"Koro (Papua New Guinea)", "Koro"}, } m["kxs"] = { } m["kxt"] = { } m["kxu"] = { otherNames = {"Kui", "Kuy"}, } m["kxv"] = { } m["kxw"] = { } m["kxx"] = { } m["kxy"] = { } m["kxz"] = { } m["kya"] = { } m["kyb"] = { } m["kyc"] = { } m["kyd"] = { } m["kye"] = { } m["kyf"] = { } m["kyg"] = { } m["kyh"] = { otherNames = {"Karuk"}, } m["kyi"] = { } m["kyj"] = { } m["kyk"] = { } m["kyl"] = { } m["kym"] = { } m["kyn"] = { } m["kyo"] = { } m["kyp"] = { } m["kyq"] = { } m["kyr"] = { otherNames = {"Caravare", "Curuaia", "Kuruaia"}, } m["kys"] = { } m["kyt"] = { } m["kyu"] = { } m["kyv"] = { } m["kyw"] = { otherNames = {"Kurmali"}, } m["kyx"] = { otherNames = {"Konua"}, } m["kyy"] = { } m["kyz"] = { } m["kza"] = { } m["kzb"] = { } m["kzc"] = { } m["kzd"] = { } m["kzf"] = { otherNames = {"Tado", "Inde", "Pekava", "West Kaili"}, } m["kzg"] = { } m["kzh"] = { otherNames = {"Kenuzi-Dongola", "Andaandi", "Kenzi", "Mattoki"}, } m["kzi"] = { } m["kzj"] = { } m["kzk"] = { otherNames = {"Dororo", "Guliguli"}, } m["kzl"] = { } m["kzm"] = { } m["kzn"] = { } m["kzo"] = { } m["kzp"] = { } m["kzq"] = { } m["kzr"] = { aliases = {"Mbum East", "Lakka"}, } m["kzs"] = { } m["kzt"] = { } m["kzu"] = { } m["kzv"] = { } m["kzw"] = { -- contrast xoo, sai-kat, sai-xoc, the last of which the ISO conflated into this code otherNames = {"Kipeá", "Quipea", "Kamurú", "Camuru", "Dzubukuá", "Dzubucua", "Karirí", "Sabujá", "Sapoyá", "Pedra Branca"}, } m["kzx"] = { } m["kzy"] = { otherNames = {"Kango", "Kango (Tshopo District)"}, -- distinct in name, but not necessarily in identity, from 'kty' } m["kzz"] = { } return m 3noq5crf33bzu48k0iw2jk0rbdgjeup Modul:languages/data/exceptional/extra 828 33778 375380 374164 2026-09-22T05:36:21Z Hakimi97 2668 [[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]] 375380 Scribunto text/plain local m = {} m["aav-khs-pro"] = { aliases = {"Proto-Khasic"}, } m["aav-nic-pro"] = { } m["aav-pkl-pro"] = { } m["aav-pro"] = { -- mkh-pro will merge into this. } m["afa-pro"] = { aliases = {"Proto-Afro-Asiatic", "Hamito-Semitic"}, } m["alg-aga"] = { aliases = {"Agwam", "Agaam"}, } m["alg-pro"] = { } m["alv-ama"] = { } m["alv-bgu"] = { otherNames = {"Gubëeher", "Nyun Gubëeher", "Nun Gubëeher"}, } m["alv-bua-pro"] = { } m["alv-cng-pro"] = { } m["alv-edk-pro"] = { } m["alv-edo-pro"] = { } m["alv-fli-pro"] = { } m["alv-gbe-pro"] = { } m["alv-gng-pro"] = { } m["alv-gtm-pro"] = { aliases = {"Proto-Ghana-Togo Mountain"}, } m["alv-gwa"] = { } m["alv-hei-pro"] = { } m["alv-ido-pro"] = { } m["alv-igb-pro"] = { } m["alv-kwa-pro"] = { } m["alv-mum-pro"] = { } m["alv-nup-pro"] = { } m["alv-pro"] = { } m["alv-von-pro"] = { } m["alv-yor-pro"] = { } m["alv-yrd-pro"] = { } m["apa-pro"] = { aliases = {"Proto-Apache", "Proto-Southern Athabaskan"}, } m["aql-pro"] = { } m["art-adu"] = { aliases = {"Westron"}, } m["art-bel"] = { } m["art-blk"] = { } m["art-bsp"] = { } m["art-com"] = { } m["art-dtk"] = { } m["art-elo"] = { } m["art-gld"] = { } m["art-lap"] = { } m["art-man"] = { } m["art-mun"] = { } m["art-nav"] = { } m["art-vlh"] = { } m["ath-nic"] = { } m["ath-pro"] = { } m["auf-pro"] = { aliases = {"Proto-Arawan", "Proto-Arauan"}, } m["aus-alu"] = { otherNames = {"Ogh-Alungul", "Alngula"}, } m["aus-and"] = { aliases = {"Adithinngithigh"}, } m["aus-ang"] = { otherNames = {"Ogh-Anggula", "Anggula", "Ogh-Anggul", "Anggul"}, } m["aus-arn-pro"] = { } m["aus-bra"] = { aliases = {"Barranbinja", "Baranbinya", "Burranbinya", "Burrumbiniya", "Burrunbinya", "Barrumbinya", "Barren-binya", "Parran-binye"}, } m["aus-brm"] = { } m["aus-cww-pro"] = { } m["aus-dal-pro"] = { } m["aus-guw"] = { otherNames = {"Gowar", "Goowar", "Gooar", "Guar", "Gowr-burra", "Ngugi", "Mugee", "Wogee", "Gnoogee", "Chunchiburri", "Booroo-geen-merrie"}, } m["aus-lsw"] = { aliases = {"Little Swanport Tasmanian"}, } m["aus-mbi"] = { otherNames = {"Mbeiwum"}, } m["aus-ngk"] = { otherNames = {"Ngkot", "Nggoth"}, } m["aus-nyu-pro"] = { } m["aus-pam-pro"] = { } m["aus-tul"] = { otherNames = {"Dappil", "Dapil", "Toolooa", "Dulua", "Narung", "Dandan"}, } m["aus-uwi"] = { otherNames = {"Uwinjmil"}, } m["aus-wdj-pro"] = { } m["aus-won"] = { } m["aus-wul"] = { otherNames = {"Manbara", "Wulgurugaba", "Wulgurukaba", "Nhawalgaba"}, } m["aus-ynk"] = { -- contrast nny } m["awd-amc-pro"] = { otherNames = {"Western Maipuran"}, } m["awd-kmp-pro"] = { otherNames = {"Campa", "Kampan", "Campan", "Pre-Andine Maipurean"}, } m["awd-prw-pro"] = { otherNames = {"Paresí-Waurá", "Parecí–Xingú", "Paresí–Xingu", "Central Arawak", "Central Maipurean"}, } m["awd-ama"] = { } m["awd-ana"] = { aliases = {"Anauya"}, } m["awd-apo"] = { otherNames = {"Lapachu"}, } m["awd-cab"] = { aliases = {"Cabere", "Cávere", "Cavere"}, } m["awd-gnu"] = { otherNames = {"Guinao", "Inao", "Guniare", "Quinhau", "Guiano"}, } m["awd-kar"] = { aliases = {"Kariaí", "Kariai", "Cariyai", "Carihiahy"}, } m["awd-kaw"] = { aliases = {"Cawishana", "Cayuishana", "Kaishana", "Cauixana"}, } m["awd-kus"] = { aliases = {"Kustenaú", "Custenau", "Kutenabu"}, } m["awd-man"] = { } m["awd-mar"] = { aliases = {"Marawán"}, } m["awd-mpr"] = { aliases = {"Maypure", "Mejepure"}, } m["awd-mrt"] = { aliases = {"Mariate"}, } m["awd-nwk-pro"] = { aliases = {"Proto-Newiki"}, } m["awd-pai"] = { aliases = {"Paiconeca", "Paikone", "Paicone"}, } m["awd-pas"] = { aliases = {"Passé", "Pazé"}, } m["awd-pro"] = { otherNames = {"Proto-Arawakan", "Proto-Maipurean", "Proto-Maipuran"}, } m["awd-she"] = { aliases = {"Shebaya", "Shebaye"}, } m["awd-taa-pro"] = { otherNames = {"Proto-Ta-Arawakan", "Proto-Caribbean Northern Arawak"}, } m["awd-wai"] = { otherNames = {"Wainuma", "Wai", "Waima", "Wainumi", "Wainambí", "Waiwana", "Waipi", "Yanuma"}, } m["awd-war"] = { } m["awd-yum"] = { aliases = {"Jumana"}, } m["azc-caz"] = { aliases = {"Caxcan", "Kaskán"}, } m["azc-cup-pro"] = { } m["azc-ktn"] = { aliases = {"Gitanemuk"}, } m["azc-nah-pro"] = { } m["azc-nic"] = { } m["azc-num-pro"] = { } m["azc-pro"] = { } m["azc-tak-pro"] = { } m["azc-tat"] = { } m["ber-fog"] = { otherNames = {"El-Fogaha", "El-Foqaha", "Foqaha", "Fuqaha"}, } m["ber-pro"] = { } m["ber-zuw"] = { } m["bnt-bal"] = { } m["bnt-bon"] = { } m["bnt-boy"] = { } m["bnt-bwa"] = { } m["bnt-cmw"] = { otherNames = {"Bravanese", "Mwiini", "Mwini", "Chimwini", "Chimini", "Brava"}, } m["bnt-ind"] = { otherNames = {"Kɔlɔmɔnyi", "Kɔlɛ", "Kasaï Oriental"}, } m["bnt-lal"] = { } m["bnt-mpi"] = { } m["bnt-mpu"] = { } m["bnt-ngu-pro"] = { } m["bnt-phu"] = { aliases = {"Siphuthi"}, } m["bnt-pro"] = { } m["bnt-sab-pro"] = { } m["bnt-sbo"] = { } m["bnt-sts-pro"] = { } m["btk-pro"] = { } m["cau-abz-pro"] = { otherNames = {"Proto-Abazgi", "Proto-Abkhaz-Tapanta"}, } m["cau-and-pro"] = { aliases = {"Proto-Andi", "Proto-Andic"}, } m["cau-ava-pro"] = { aliases = {"Proto-Avar-Andian", "Proto-Avar-Andi", "Proto-Avar-Andic"}, } m["cau-cir-pro"] = { otherNames = {"Proto-Adyghe-Kabardian", "Proto-Adyghe-Circassian"}, } m["cau-drg-pro"] = { otherNames = {"Proto-Dargin"}, } m["cau-lzg-pro"] = { aliases = {"Proto-Lezgi", "Proto-Lezgian", "Proto-Lezgic"}, } m["cau-nec-pro"] = { } m["cau-nkh-pro"] = { } m["cau-nwc-pro"] = { } m["cau-tsz-pro"] = { otherNames = {"Proto-Tsezic", "Proto-Didoic"}, } m["cba-ata"] = { otherNames = {"Atanque", "Cancuamo", "Kankuamo", "Kankwe", "Kankuí", "Atanke"}, } m["cba-cat"] = { otherNames = {"Catio Chibcha", "Old Catio"}, } m["cba-dor"] = { otherNames = {"Chumulu", "Changuena", "Changuina", "Chánguena", "Gualaca"}, } m["cba-dui"] = { } m["cba-hue"] = { otherNames = {"Güetar", "Guetar", "Brusela"}, } m["cba-nut"] = { otherNames = {"Nutabane"}, } m["cba-pro"] = { } m["ccs-pro"] = { } m["ccs-gzn-pro"] = { aliases = {"Proto-Karto-Zan"}, } m["cdc-cbm-pro"] = { otherNames = {"Proto-Central-Chadic", "Proto-Biu-Mandara"}, } m["cdc-mas-pro"] = { } m["cdc-pro"] = { } m["cdd-pro"] = { } m["cel-bry-pro"] = { aliases = {"Proto-Brittonic", "Common Brythonic", "Common Brittonic"}, } m["cel-gal"] = { } m["cel-gau"] = { } m["cel-pro"] = { } m["chi-pro"] = { } m["chm-pro"] = { } m["cmc-pro"] = { } m["crp-bip"] = { } m["crp-cpr"] = { } m["crp-gep"] = { aliases = {"Greenlandic Pidgin", "Greenlandic Eskimo Pidgin"}, } m["crp-kia"] = { } m["crp-mar"] = { otherNames = {"Jamaican Maroon Spirit Possession Language"}, } m["crp-mpp"] = { aliases = {"Macao Pidgin Portuguese"}, } m["crp-rsn"] = { } m["crp-slb"] = { otherNames = {"Solombala-English", "Solombala English-Russian Pidgin"}, } m["crp-spp"] = { } m["crp-tpr"] = { } m["csu-bba-pro"] = { } m["csu-maa-pro"] = { } m["csu-pro"] = { } m["csu-sar-pro"] = { } m["cus-ash"] = { otherNames = {"Ashraf", "Af-Ashraaf"}, varieties = { {"Marka, Lower Shabelle"}, "Shingani"}, } m["cus-hec-pro"] = { } m["cus-som-pro"] = { otherNames = {"Proto-Sam", "Proto-Macro-Somali"}, } m["cus-sou-pro"] = { otherNames = {"Proto-Rift"}, } m["cus-pro"] = { } m["dmn-dam"] = { } m["dra-bry"] = { aliases = {"Byari"}, } m["dra-cen-pro"] = { } m["dra-mkn"] = { aliases = {"Nadugannada"}, } m["dra-nor-pro"] = { } m["dra-okn"] = { aliases = {"Halegannada"}, } m["dra-ote"] = { } m["dra-pro"] = { } m["dra-sdo-pro"] = { aliases = {"Proto-South Dravidian"}, } m["dra-sdt-pro"] = { aliases = {"Proto-South-Central Dravidian"}, } m["dra-sou-pro"] = { aliases = {"Proto-Southern Dravidian"}, } m["egx-dem"] = { aliases = {"Demotic Egyptian", "Enchorial"}, } m["dmn-pro"] = { } m["dmn-mdw-pro"] = { } m["dru-pro"] = { } m["ero-gsz"] = { } m["ero-nya"] = { } m["ero-tau"] = { } m["esx-esk-pro"] = { } m["esx-ink"] = { } m["esx-inq"] = { } m["esx-inu-pro"] = { } m["esx-pro"] = { } m["esx-tut"] = { } m["euq-pro"] = { aliases = {"Proto-Vasconic"}, } m["gba-pro"] = { } m["gem-pro"] = { aliases = {"Common Germanic"}, } m["gme-bur"] = { aliases = {"Burgundish", "Burgundic"}, } m["gme-cgo"] = { } m["gmq-gut"] = { } m["gmq-jmk"] = { aliases = {"Jamtlandic"}, } m["gmq-mno"] = { } m["gmq-oda"] = { } m["gmq-ogt"] = { aliases = {"Old Gotlandic"}, } m["gmq-osw"] = { } m["gmq-pro"] = { aliases = {"Proto-Scandinavian", "Primitive Norse", "Proto-Nordic", "Ancient Nordic", "Ancient Scandinavian", "Old Nordic", "Old Scandinavian", "Proto-North Germanic", "North Proto-Germanic", "Common Scandinavian"}, } m["gmq-scy"] = { } m["gmw-bgh"] = { } m["gmw-cfr"] = { varieties = {"Mittelfränkisch", "Ripuarian", "Moselle Franconian", "Colognian", "Kölsch"}, } m["gmw-ecg"] = { varieties = {"Thuringian", "Thüringisch", "Upper Saxon", "Upper Saxon German", "Obersächsisch", "Lusatian", "Erzgebirgisch", "Silesian", "Silesian German", "High Prussian"}, } m["gmw-fin"] = { aliases = {"Fingal"}, } m["gmw-gts"] = { aliases = {"Gottscheerisch"}, } m["gmw-jdt"] = { } m["gmw-msc"] = { } m["gmw-pro"] = { } m["gmw-rfr"] = { aliases = {"Rheinfränkisch", "Rhenish Franconian"}, varieties = {"Hessian", "Lorraine Franconian", "Lorrainian", "Lothringisch", "Palatine German", "Pfälzisch", "Pälzisch", "Palatinate German"}, } m["gmw-stm"] = { aliases = {"Satu Mare Swabian", "Sathmarschwäbisch", "Sathmarisch"}, } m["gmw-tsx"] = { aliases = {"Siebenbürger Saxon"}, } m["gmw-vog"] = { } m["gmw-zps"] = { aliases = {"Zipser", "Zipserisch", "Outzäpsersch"}, } m["gn-cls"] = { } m["grk-cal"] = { aliases = {"Italian Greek", "Bova"}, } m["grk-ita"] = { aliases = {"Griko", "Grico", "Grecanic"}, } m["grk-mar"] = { aliases = {"Mariupolitan Greek", "Rumeíka", "Rumeika"}, } m["grk-pro"] = { aliases = {"Proto-Greek"}, } m["hmn-pro"] = { } m["hmx-mie-pro"] = { } m["hmx-pro"] = { } m["hyx-pro"] = { } m["iir-nur-pro"] = { } m["iir-pro"] = { } m["ijo-pro"] = { aliases = {"Proto-Ijaw"}, } m["inc-apa"] = { aliases = {"Apabhraṃśa"}, } m["inc-ash"] = { aliases = {"Asokan Prakrit", "Aśokan Prakrit"}, } m["inc-dng-pro"] = { } m["inc-kam"] = { } m["inc-kho"] = { } m["inc-khr"] = { } m["inc-krd-pro"] = { } m["inc-mas"] = { } m["inc-mbn"] = { } m["inc-mgu"] = { } m["inc-mor"] = { aliases = {"Middle Oriya"}, } m["inc-oas"] = { } m["inc-oaw"] = { aliases = {"Early Awadhi"}, } m["inc-obn"] = { } m["inc-ogu"] = { otherNames = {"Old Western Rajasthani"}, } m["inc-ohi"] = { aliases = {"Dehlavi"}, } m["inc-oor"] = { aliases = {"Old Oriya"}, } m["inc-opa"] = { } m["inc-pro"] = { } m["inc-sar"] = { } m["ine-ana-pro"] = { } m["ine-bsl-pro"] = { } m["ine-kal"] = { aliases = {"Kalašmaic", "Kalasmaic"}, } m["ine-pae"] = { } m["ine-pro"] = { } m["ine-toc-pro"] = { } m["itc-psa"] = { } m["mis-idn"] = { } m["mis-tdl"] = { } m["mis-tdt"] = { } m["mis-xnu"] = { } m["ngf-bin-pro"] = { } m["njo-jgl"] = { } m["njo-mng"] = { } m["paa-kmn"] = { } m["paa-lei"] = { } m["poz-nes"] = { } m["poz-pcc-pro"] = { } m["roa-can"] = { } m["roa-ona"] = { } m["sai-gua"] = { } m["sai-peb"] = { } m["sem-sam"] = { } m["sit-aao-pro"] = { } m["sit-ban"] = { } m["sit-bdi-pro"] = { } m["sit-ers-pro"] = { } m["sit-khb-pro"] = { } m["sit-khp-pro"] = { } m["sit-khw-pro"] = { } m["sit-kon-pro"] = { } m["sit-nas-pro"] = { } m["sit-tng-pro"] = { } m["tbq-brm-pro"] = { } m["tup-kaw"] = { } m["xme-old"] = { } m["xme-mid"] = { aliases = {"Atropatenian"}, } m["xme-ker"] = { otherNames = {"Kermanian", "Central Iranian Dialects", "Central Plateau Dialects", "Central Iranian", "South Median", "Gazi", "Soi", "Sohi", "Abuzeydabadi", "Abyanehi", "Farizandi", "Jowshaqani", "Nashalji", "Qohrudi", "Yarandi", "Tari", "Sedehi", "Ardestani", "Zefrehi", "Isfahani", "Kafroni", "Varzenehi", "Khuri", "Nayini", "Anaraki", "Zoroastrian Dari", "Behdināni", "Behdinani", "Gabri", "Gavrŭni", "Gavruni", "Gabrōni", "Gabroni", "Kermani", "Yazdi", "Bidhandi", "Bijagani", "Chimehi", "Hanjani", "Komjani", "Naraqi", "Qalhari", "Varani", "Zori"}, } m["xme-taf"] = { } m["xme-ttc-pro"] = { } m["xme-kls"] = { aliases = {"Kalāsuri", "Kalasur", "Kalāsur"}, } m["xme-klt"] = { } m["xme-ott"] = { otherNames = {"Old Tatic", "Old Azeri", "Azari", "Azeri", "Āḏarī", "Adari", "Adhari"}, } m["ira-kms-pro"] = { } m["ira-mpr-pro"] = { } m["ira-pat-pro"] = { } m["ira-pro"] = { } m["ira-zgr-pro"] = { } m["xsc-pro"] = { } m["xsc-sar-pro"] = { } m["xsc-skw-pro"] = { } m["xsc-sak-pro"] = { aliases = {"Proto-Sakan"}, } m["ira-sym-pro"] = { } m["ira-sgi-pro"] = { } m["ira-mny-pro"] = { } m["ira-shy-pro"] = { } m["ira-shr-pro"] = { } m["ira-sgc-pro"] = { aliases = {"Proto-Sogdian"}, } m["ira-wnj"] = { aliases = {"Old Vanji", "Vanchi", "Vanži", "Wanji"}, } m["iro-ere"] = { } m["iro-min"] = { } m["iro-nor-pro"] = { } m["iro-pro"] = { } m["itc-pro"] = { } m["jpx-hcj"] = { aliases = {"Hachijo"}, } m["jpx-pro"] = { } m["jpx-ryu-pro"] = { } m["kar-pro"] = { } m["kca-eas"] = { } m["kca-nor"] = { } m["kca-pro"] = { } m["kca-sou"] = { } m["khi-kho-pro"] = { } m["khi-kun"] = { otherNames = {"ǃOǃKung", "ǃ'OǃKung", "Kung", "Ekoka ǃKung", "Ekoka Kung", "Sekele"}, } m["ko-ear"] = { } m["kro-pro"] = { } m["ku-pro"] = { } m["map-ata-pro"] = { } m["map-bms"] = { } m["map-pro"] = { } m["mis-hkl"] = { aliases = {"Kelantan Peranakan Chinese", "Kelantan Peranakan Hokkien", "Hokkien Kelantan", "Kelantan Local Hokkien"} } m["mis-isa"] = { } m["mis-jie"] = { aliases = {"Chieh", "Kjet"}, } m["mis-jzh"] = { aliases = {"Haihua"}, } m["mis-kas"] = { aliases = {"Cassite", "Kassitic", "Kaššite"}, } m["mis-mmd"] = { otherNames = {"Mimi of Gaudefroy-Demombynes", "Mimi-D"}, } m["mis-mmn"] = { otherNames = {"Mimi-N"}, } m["mis-phi"] = { aliases = {"Philistian", "Philistinian"}, } m["mis-rou"] = { aliases = {"Ruanruan", "Ruan-ruan", "Juan-juan"}, } m["mis-tnw"] = { aliases = {"Tangwanghua"}, } m["mis-tuh"] = { aliases = {"'Azha"}, } m["mis-tuo"] = { aliases = {"Tabghach", "Taghbach"}, } m["mis-wuh"] = { aliases = {"Wuwan", "Awar"}, } m["mis-xbi"] = { aliases = {"Serbi", "Shirwi"}, } m["mjg-mgl"] = { aliases = {"Huzhu", "Huzhu Monguor"}, } m["mjg-mgr"] = { aliases = {"Minhe", "Minhe Monguor"}, } m["mkh-asl-pro"] = { } m["mkh-ban-pro"] = { } m["mkh-kat-pro"] = { } m["mkh-khm-pro"] = { } m["mkh-kmr-pro"] = { } m["mkh-mmn"] = { } m["mkh-mnc-pro"] = { } m["mkh-mvi"] = { } m["mkh-pal-pro"] = { } m["mkh-pea-pro"] = { } m["mkh-pkn-pro"] = { } m["mkh-pro"] = { --This will be merged into 2015 aav-pro. } m["mnw-tha"] = { aliases = {"Raman", "Thai Raman", "Siamese Mon"}, } m["mkh-vie-pro"] = { } m["mns-cen"] = { } m["mns-nor"] = { } m["mns-pro"] = { } m["mns-sou"] = { } m["mun-pro"] = { aliases = {"Proto-Mundan"}, } m["myn-chl"] = { -- the stage after ''emy'' otherNames = {"Cholti", "Colonial Ch'olti'", "Colonial Cholti"}, } m["myn-pro"] = { aliases = {"Proto-Maya"}, } m["nai-ala"] = { otherNames = {"Alasapa", "Pinto"}, } m["nai-bay"] = { otherNames = {"Bayougoula", "Bayou Goula", "Ischenoca"}, -- tribe merged with "Mougulasha", "Mongoulacha", "Mugulasha", "Mougulasha", "Muglahsa", "Muglasha", "Muguasha", "Imongolosha", "Houma", "Acolapissa" } m["nai-cal"] = { } m["nai-chi"] = { } m["nai-chu-pro"] = { aliases = {"Proto-Chumashan"}, } m["nai-cig"] = { } m["nai-ckn-pro"] = { aliases = {"Proto-Chinook"}, } m["nai-guz"] = { aliases = {"Guazacapan"}, } m["nai-hit"] = { otherNames = {"Atcik-hata", "At-pasha-shliha"}, } m["nai-ipa"] = { otherNames = {"'Iipay 'aa", "Northern Diegueño", "Diegueño"}, } m["nai-jtp"] = { otherNames = {"Xutiapa", "Jalapa", "Xalapa"}, } m["nai-jum"] = { aliases = {"Jumaitepeque", "Jumaytepec"}, } m["nai-kat"] = { otherNames = {"Kathlamet Chinook"}, } m["nai-klp-pro"] = { } m["nai-knm"] = { } m["nai-kum"] = { otherNames = {"Kumiai", "Central Diegueño", "Diegueño"}, } m["nai-mac"] = { aliases = {"Macorís", "Macorix", "Mazorij", "Mazorig", "Mazoriges"}, } m["nai-mdu-pro"] = { aliases = {"Proto-Maiduan"}, } m["nai-miz-pro"] = { aliases = {"Proto-Mixe-Zoquean"}, } m["nai-mus-pro"] = { aliases = {"Proto-Muskhogean", "Proto-Muskogee"}, } m["nai-nao"] = { } m["nai-nrs"] = { } m["nai-okw"] = { } m["nai-per"] = { } m["nai-pic"] = { } m["nai-plp-pro"] = { } m["nai-pom-pro"] = { aliases = {"Proto-Pomoan"}, } m["nai-qng"] = { } m["nai-sca-pro"] = { -- NB 'sio-pro' "Proto-Siouan" which is Proto-Western Siouan } m["nai-sin"] = { aliases = {"Sinacantan", "Zinacantán", "Zinacantan"}, } m["nai-sln"] = { } m["nai-spt"] = { aliases = {"Shahaptin"}, } m["nai-tap"] = { otherNames = {"Tapachulteca", "Tapachulteco", "Tapachula"}, } m["nai-taw"] = { } m["nai-teq"] = { otherNames = {"Tequistlateco", "Tequistlateca", "Chontal", "Chontol of Oaxaca", "Oaxaca Chontal", "Oaxacan Chontal"}, } m["nai-tip"] = { otherNames = {"Tipay", "Tiipai", "Tiipay", "Jamul Tiipay", "Southern Digueño", "Diegueño"}, } m["nai-tot-pro"] = { } m["nai-tsi-pro"] = { } m["nai-utn-pro"] = { otherNames = {"Proto-Miwok-Costanoan"}, } m["nai-wai"] = { aliases = {"Guaycura", "Waicura"}, } m["nai-wji"] = { otherNames = {"Jicaque of El Palmar", "Sula"}, } m["nai-yup"] = { aliases = {"Jupiltepeque", "Yupiltepec", "Jupiltepec", "Xupiltepec"}, } m["nan-dat"] = { aliases = {"Datian"}, } m["nan-hbl"] = { aliases = {"Hokkienese", "Quanzhang", "Fukien", "Banlam", "Banlamese", "Ban-lam"}, } m["nan-hlh"] = { aliases = {"Hailufeng", "Hoklo Min", "Hai Lok Hong"}, } m["nan-lnx"] = { aliases = {"Longyan", "Liongna"}, } m["nan-tws"] = { aliases = {"Teochew Min", "Chiuchow", "Teo-Swa", "Teo-Swa Min", "Tio-Sua"}, } m["nan-zhe"] = { aliases = {"Zhenan"}, } m["nan-zsh"] = { aliases = {"Sanxiang", "Samheung", "Sahiu"}, } m["ngf-pro"] = { } m["nic-bco-pro"] = { } m["nic-bod-pro"] = { } m["nic-eov-pro"] = { } m["nic-gns-pro"] = { } m["nic-grf-pro"] = { } m["nic-gur-pro"] = { } m["nic-jkn-pro"] = { } m["nic-lcr-pro"] = { } m["nic-ogo-pro"] = { } m["nic-ovo-pro"] = { } m["nic-plt-pro"] = { } m["nic-pro"] = { } m["nic-ubg-pro"] = { } m["nic-ucr-pro"] = { } m["nic-vco-pro"] = { } m["nub-har"] = { aliases = {"Ḥarāza"}, } m["nub-pro"] = { } m["omq-cha-pro"] = { } m["omq-maz-pro"] = { aliases = {"Proto-Mazatecan"}, } m["omq-mix-pro"] = { } m["omq-mxt-pro"] = { } m["omq-otp-pro"] = { } m["omq-pro"] = { aliases = {"Proto-Otomanguean", "Proto-Oto-Mangue"}, } m["omq-sjq"] = { aliases = {"Chatino Sign Language", "San Juan Quiahije Chatino Sign Language"}, } m["omq-tel"] = { } m["omq-teo"] = { } m["omq-tri-pro"] = { } m["omq-zap-pro"] = { } m["omq-zpc-pro"] = { } m["omv-aro-pro"] = { } m["omv-diz-pro"] = { aliases = {"Proto-Maji"}, } m["omv-pro"] = { } m["oto-otm-pro"] = { } m["oto-pro"] = { } m["paa-kwn"] = { } m["paa-nha-pro"] = { } m["paa-nun"] = { } m["phi-din"] = { } m["phi-kal-pro"] = { aliases = {"Proto-Calamian"}, } m["phi-nag"] = { } m["phi-pro"] = { } m["poz-abi"] = { otherNames = {"Sembuak", "Tubu"}, } m["poz-bal"] = { } m["poz-btk-pro"] = { } m["poz-cet-pro"] = { } m["poz-hce-pro"] = { otherNames = {"Proto-South Halmahera - West New Guinea"}, } m["poz-lgx-pro"] = { } m["poz-mcm-pro"] = { } m["poz-mic-pro"] = { } m["poz-mly-pro"] = { } m["poz-msa-pro"] = { } m["poz-oce-pro"] = { } m["poz-pep-pro"] = { aliases = {"Proto-Eastern-Polynesian", "Proto-East Polynesian", "Proto-East-Polynesian"}, } m["poz-pnp-pro"] = { } m["poz-pol-pro"] = { } m["poz-pro"] = { otherNames = {"Proto-Western Malayo-Polynesian"}, -- Western is subsumed into general Proto-MP } m["poz-sml"] = { aliases = {"Sarawak"}, } m["poz-ssw-pro"] = { } m["poz-swa-pro"] = { } m["poz-ter"] = { aliases = {"Terengganu"}, } m["pqe-pro"] = { } m["pra-niy"] = { } m["qfa-adm-pro"] = { } m["qfa-bet-pro"] = { aliases = {"Proto-Tai-Be"}, } m["qfa-cka-pro"] = { } m["qfa-hur-pro"] = { } m["qfa-kad-pro"] = { } m["qfa-kms-pro"] = { } m["qfa-kor-pro"] = { } m["qfa-kra-pro"] = { } m["qfa-lic-pro"] = { } m["qfa-onb-pro"] = { aliases = {"Proto-Ong-Be", "Proto-Bê"}, } m["qfa-ong-pro"] = { } m["qfa-tak-pro"] = { aliases = {"Proto-Tai-Kadai"}, } m["qfa-yen-pro"] = { } m["qfa-yuk-pro"] = { } m["qwe-kch"] = { otherNames = {"Kichwa shimi", "Runashimi", "Runa", "Quichua", "Quecha", "Inga", "Chimborazo", "Imbabura Highland Kichwa", "Cañar Highland Quecha", "Quechua"}, } m["qwe-pro"] = { } m["roa-ang"] = { otherNames = {"Craonnais", "Baugeois", "Saumurois"}, } m["roa-bbn"] = { otherNames = {"Bourbonnais", "Berrichon", "Moulins", "Allier", "Nivernais", "Haut-Berrichon", "Bas-Berrichon"}, } m["roa-brg"] = { otherNames = {"Burgundian", "Bregognon", "Dijonnais", "Morvandiau", "Morvandeau", "Morvan", "Bourguignon-Morvandiau", "Mâconnais", "Brionnais", "Brionnais-Charolais", "Auxerrois", "Beaunois", "Langrois", "Valsaônois", "Verduno-Chalonnais", "Sédelocien"}, } m["roa-cha"] = { otherNames = {"Bassignot", "Langrois", "Sennonais", "Vallage", "Troyen", "Briard", "Der", "Perthois", "Rémois", "Argonnais", "Porcien", "Ardennais", "Sugny"}, } m["roa-fcm"] = { otherNames = {"Frainc-Comtou", "Comtois", "Jurassien", "Ajoulot", "Vâdais", "Taignon", "Bisontin", "Bousbot"}, } m["roa-gal"] = { } m["roa-gib"] = { } m["roa-gis"] = { } m["roa-leo"] = { } m["roa-lor"] = { otherNames = {"Gaumais", "Vosgien", "Welche", "Argonnais", "Longovicien", "Messin", "Nancéien", "Spinalien", "Déodatien"}, } m["roa-oca"] = { } m["roa-ole"] = { } m["roa-opt"] = { aliases = {"Galician-Portuguese", "Galician Portuguese", "Medieval Galician", "Medieval Portuguese", "Old Galician", "Old Portuguese"}, } m["roa-orl"] = { otherNames = {"Beauceron", "Solognot", "Gâtinais", "Blaisois", "Vendômois"}, } m["roa-poi"] = { otherNames = {"Poitevin", "Saintongeais", "Maraîchin"}, } m["roa-tar"] = { } m["sai-all"] = { otherNames = {"Alyentiyak", "Huarpe", "Warpe"}, } m["sai-and"] = { -- not to be confused with 'cbc' or 'ano' otherNames = {"Miranya", "Miranha", "Miranha Carapana-Tapuya", "Miraña-Carapana-Tapuyo", "Andokero", "Miranya-Karapana-Tapuyo", "Miraña", "Carapana"}, } m["sai-ayo"] = { aliases = {"Ayoman", "Ayamán", "Ayaman"}, } m["sai-bae"] = { aliases = {"Baenã", "Baenán", "Baena"}, } m["sai-bag"] = { otherNames = {"Patagón de Bagua"}, } m["sai-bet"] = { otherNames = {"Betoy", "Betoya", "Betoye", "Betoi-Jirara", "Jirara"}, } m["sai-bor-pro"] = { otherNames = {"Proto-Bora-Muinane", "Proto-Bora-Muiname"}, } m["sai-cac"] = { otherNames = {"Kakán", "Diaguita", "Cacan", "Kakan", "Calchaquí", "Chaka", "Kaka", "Kaká", "Caca", "Caca-Diaguita", "Catamarcano", "Capayán", "Capayana", "Yacampis"}, } m["sai-caq"] = { otherNames = {"Cara", "Kara"}, } m["sai-car-pro"] = { } m["sai-cat"] = { } m["sai-cer-pro"] = { otherNames = {"Proto-Amazonian Jê"}, } m["sai-chi"] = { } m["sai-chn"] = { aliases = {"Chana"}, } m["sai-chp"] = { aliases = {"Txapacura", "Xapacura", "Guapore", "Šapakura", "Txapakura", "Txapakúra", "Xapakúra"}, } m["sai-chr"] = { aliases = {"Charrúa", "Charruá"}, } m["sai-chu"] = { aliases = {"Churoya"}, } m["sai-cje-pro"] = { otherNames = {"Proto-Akuwẽ"}, } m["sai-cmg"] = { aliases = {"Comechingón", "Comechingona", "Comechingone"}, } m["sai-cno"] = { otherNames = {"Chonos", "Caucau"}, } m["sai-cnr"] = { aliases = {"Cañar"}, } m["sai-coe"] = { aliases = {"Koeruna"}, } m["sai-col"] = { aliases = {"Colan"}, } m["sai-cop"] = { } m["sai-crd"] = { otherNames = {"Coroado"}, } m["sai-ctq"] = { aliases = {"Catuquinarú", "Katukinaru"}, } m["sai-cul"] = { otherNames = {"Culle", "Kulyi", "Ilinga", "Linga"}, } m["sai-cva"] = { } m["sai-esm"] = { otherNames = {"Esmeraldeño", "Atacame", "Takame"}, } m["sai-ewa"] = { } m["sai-gam"] = { aliases = {"Gamella", "Acobu", "Curinsi", "Barbados"}, } m["sai-gay"] = { aliases = {"Gayon"}, } m["sai-gmo"] = { otherNames = {"Wamo", "Santa Rosa", "San Jose", "Barinas", "Guamotey", "Guama"}, } m["sai-gue"] = { aliases = {"Guenoa"}, } m["sai-hau"] = { otherNames = {"Manek'enk"}, } m["sai-jee-pro"] = { otherNames = {"Proto-Gê", "Proto-Jean", "Proto-Gean", "Proto-Jê-Kaingang", "Proto-Ye"}, } m["sai-jko"] = { aliases = {"Geicó", "Jeicó", "Jaikó", "Geikó", "Yeikó", "Jeiko", "Geico", "Jeico", "Jaiko", "Geiko", "Yeiko", "Eyco"}, } m["sai-jrj"] = { } m["sai-kat"] = { -- contrast xoo, kzw, sai-xoc otherNames = {"Catrimbi", "Catembri", "Kariri de Mirandela", "Mirandela", "Kariri", "Kiriri"}, } m["sai-mal"] = { aliases = {"Malali"}, } m["sai-mar"] = { } m["sai-mat"] = { otherNames = {"Matanauí", "Matanaui", "Matanawü", "Mitandua", "Moutoniway"}, } m["sai-mcn"] = { aliases = {"Mokana"}, } m["sai-men"] = { aliases = {"Menién"}, } m["sai-mil"] = { otherNames = {"Milykayak", "Huarpe", "Warpe"}, } m["sai-mlb"] = { aliases = {"Malibú", "Malebú"}, } m["sai-msk"] = { aliases = {"Masakara", "Masacará", "Masacara"}, } m["sai-muc"] = { otherNames = {"Mucuchi", "Mokochi", "Mocochí", "Mirripú", "Maripú", "Mucuchí-Maripú"}, } m["sai-mue"] = { aliases = {"Muellamués"}, } m["sai-muz"] = { } m["sai-mys"] = { otherNames = {"Mayna", "Maina", "Rimachu"}, } m["sai-nat"] = { otherNames = {"Natu", "Peagaxinan"}, } m["sai-nje-pro"] = { otherNames = {"Proto-Core Jê"}, } m["sai-opo"] = { otherNames = {"Opon", "Opón-Karare", "Opón-Carare", "Carare", "Carare-Opón"}, } m["sai-oto"] = { aliases = {"Otomako", "Otomacan", "Otomac", "Otomak"}, } m["sai-pal"] = { } m["sai-pam"] = { aliases = {"Pamiwa"}, } m["sai-par"] = { aliases = {"Paratio", "Prarto"}, } m["sai-pnz"] = { aliases = {"Pansaleo"}, } m["sai-prh"] = { } m["sai-ptg"] = { otherNames = {"Patagón de Perico"}, } m["sai-pur"] = { aliases = {"Purukoto", "Purucotó", "Purucoto"}, } m["sai-pyg"] = { aliases = {"Payawá", "Payagua"}, } m["sai-pyk"] = { aliases = {"Gavião-Pykobjê", "Pykobjê-Gavião", "Gavião", "Pyhcopji", "Gavião-Pyhcopji"}, } m["sai-qmb"] = { otherNames = {"Kimbaya", "Quindío", "Quindio", "Quindo"}, } m["sai-qtm"] = { aliases = {"Quitemoca"}, } m["sai-rab"] = { } m["sai-ram"] = { } m["sai-sac"] = { otherNames = {"Sacata", "Zácata", "Chillao"}, } m["sai-san"] = { aliases = {"Sanavirón", "Sanabirón", "Sanabiron", "Sanavirona", "Zanavirona"}, } m["sai-sap"] = { aliases = {"Zapará", "Zapara"}, } m["sai-sec"] = { otherNames = {"Sek", "Sec"}, } m["sai-sin"] = { otherNames = {"Cenúfana", "Zenúfana", "Cinifaná", "Sinufana", "Sinú", "Cenú", "Zenú", "Finzenú", "Fincenú", "Pancenú", "Sutagao"}, } m["sai-sje-pro"] = { } m["sai-tab"] = { otherNames = {"Aconipa"}, } m["sai-tal"] = { otherNames = {"Atalán", "Tallan", "Tallanca", "Atalan", "Sek"}, } m["sai-tap"] = { otherNames = {"Tapayúna", "Kajkwakhrattxi"}, } m["sai-tar-pro"] = { } m["sai-teu"] = { aliases = {"Tehues", "Teuéx"}, } m["sai-tim"] = { otherNames = {"Cuica", "Timote-Cuica"}, } m["sai-tpr"] = { aliases = {"Taparito"}, } m["sai-trr"] = { otherNames = {"Caratiú"}, } m["sai-wai"] = { aliases = {"Waitaka", "Waitacá", "Waitaca", "Goytacá", "Goitacá", "Guaitacá", "Guiatacá", "Guiatacás", "Goiatacá", "Goiatacás", "Guaiatacá", "Goytacaz", "Goitacaz", "Goyataca", "Aitacaz", "Uetacaz", "Uetacá", "Outacá", "Ouetacá", "Eutacá", "Itacaz", "Vaitacá"}, } m["sai-way"] = { aliases = {"Wajumará", "Wajumara", "Wayumará", "Azumara", "Guimara"}, } m["sai-wit-pro"] = { otherNames = {"Proto-Huitotoan", "Proto-Uitotoan"}, } m["sai-wnm"] = { otherNames = {"Wañam", "Wanyam", "Huanyam", "Uanham", "Abitana"}, } m["sai-xoc"] = { -- contrast xoo, kzw, sai-kat otherNames = {"Xoco", "Chocó", "Shokó", "Shoko", "Shocó", "Shoco", "Choco", "Chocaz", "Kariri-Xocó", "Kariri-Xoco", "Kariri-Shoko", "Cariri-Chocó", "Xukuru-Kariri", "Xucuru-Kariri", "Xucuru-Cariri", "Xukurú-Kirirí"}, } m["sai-yao"] = { aliases = {"Yao", "Jaoi", "Yaoi", "Yaio", "Anacaioury"}, } m["sai-yar"] = { -- not the same family as 'suy' aliases = {"Yaruma"}, } m["sai-yri"] = { aliases = {"Jurí"}, } m["sai-yup"] = { otherNames = {"Yupuá", "Yupúa", "Jupua", "Jupuá", "Jupúa", "Hiupiá", "Yupuá-Duriña", "Duriña"}, } m["sai-yur"] = { aliases = {"Yurumangui", "Yurimangí", "Yurimangi", "Yurimanguí", "Yurimangui"}, } m["sal-pro"] = { aliases = {"Proto-Salishan"}, } m["sdv-daj-pro"] = { } m["sdv-eje-pro"] = { } m["sdv-nil-pro"] = { } m["sdv-nyi-pro"] = { } m["sdv-tmn-pro"] = { } m["sel-nor"] = { aliases = {"Taz Selkup"}, } m["sel-pro"] = { } m["sel-sou"] = { } m["sem-amm"] = { } m["sem-amo"] = { aliases = {"Amoritic"}, } m["sem-cha"] = { aliases = {"Cheha", "Čäha", "Čäxa"}, } m["sem-dad"] = { otherNames = {"Dadanite", "Lihyanite", "Lihyanitic"}, } m["sem-dum"] = { } m["sem-has"] = { } m["sem-his"] = { otherNames = {"Thamudic E"}, } m["sem-mhr"] = { otherNames = {"Muher Gurage", "Muxar", "Muxər", "Muhər", "Muḫər"}, } m["sem-pro"] = { } m["sem-saf"] = { } m["sem-srb"] = { } m["sem-tay"] = { otherNames = {"Taymanite", "Thamudic A"}, } m["sem-tha"] = { } m["sem-wes-pro"] = { } m["sio-pro"] = { -- NB this is not Proto-Siouan-Catawban 'nai-sca-pro' } m["sit-bok"] = { otherNames = {"Ramo", "Pailibo"}, } m["sit-bai-pro"] = { } m["sit-cai"] = { } m["sit-cha"] = { } m["sit-hrs-pro"] = { } m["sit-jap"] = { otherNames = {"Chabao", "Kuru"}, } m["sit-kha-pro"] = { } m["sit-liz"] = { } m["sit-lnj"] = { } m["sit-lrn"] = { } m["sit-luu-pro"] = { } m["sit-prn"] = { } m["sit-pro"] = { } m["sit-sit"] = { otherNames = {"Eastern rGyalrong", "rGyalrong", "Rgyalrong", "rGyalrongic", "Gyalrong", "Gyarong", "rGyarong", "Gyarung", "Jiarong", "Jiarongyu", "Jyarong", "Jyarung", "Yelong", "Kuru"}, } m["sit-tam-pro"] = { aliases = {"Proto-Tamang"}, } m["sit-tan-pro"] = { } m["sit-tgm"] = { } m["sit-tos"] = { } m["sit-tsh"] = { otherNames = {"Caodeng", "Sidaba", "rGyalrong", "Rgyalrong", "Jiarong", "Gyarung", "Kuru"}, } m["sit-zbu"] = { otherNames = {"Ribu", "Rdzong'bur", "Rdzongmbur", "Showu", "rGyalrong", "Rgyalrong", "Jiarong", "Gyarung", "Kuru"}, } m["sla-pro"] = { aliases = {"Common Slavic"}, } m["smi-pro"] = { aliases = {"Proto-Sami"}, } m["son-pro"] = { aliases = {"Proto-Songhai"}, } m["sqj-pro"] = { } m["ssa-klk-pro"] = { aliases = {"Proto-Rub"}, } m["ssa-kom-pro"] = { } m["ssa-pro"] = { } m["syd-pro"] = { } m["tai-pro"] = { } m["tai-swe-pro"] = { } m["tbq-bdg-pro"] = { } m["tbq-blg"] = { aliases = {"Pai-lang", "Pailang"}, } m["tbq-gkh"] = { aliases = {"Gɔkhý", "Gɔkhy", "Gouke"}, } m["tbq-kuk-pro"] = { otherNames = {"Proto-Kukish"}, } m["tbq-lal-pro"] = { } m["tbq-laz"] = { otherNames = {"Lare", "Shuitianhua"}, } m["tbq-lob-pro"] = { } m["tbq-lol-pro"] = { otherNames = {"Proto-Yi", "Proto-Ngwi", "Proto-Nisoic"}, } m["tbq-mil"] = { } m["tbq-mor"] = { aliases = {"Morān"}, } m["tbq-ngo"] = { otherNames = {"Ngachang", "Achang"}, } -- tbq-pro is now etymology-only m["trk-dkh"] = { aliases = {"Dukha"}, } m["trk-eog"] = { } m["trk-oat"] = { } m["trk-pro"] = { } m["tup-gua-pro"] = { } m["tup-kab"] = { aliases = {"Kabixiana", "Cabixiana", "Cabishiana", "Kapishana", "Capishana", "Kapišana", "Cabichiana", "Capichana", "Capixana"}, } m["tuw-alk"] = { aliases = {"Alechuka"}, } m["tuw-bal"] = { } m["tuw-kkl"] = { aliases = {"Chinese Kyakala"}, } m["tuw-kli"] = { aliases = {"Kilen", "Kirin", "Kila", "Hezhe", "Qile'en"}, } m["tup-pro"] = { } m["tuw-pro"] = { } m["tuw-sol"] = { } m["urj-fin-pro"] = { } m["urj-koo"] = { aliases = {"Old Permian"}, } m["urj-kuk"] = { aliases = {"Kukkuzi Votic", "Kukkuzi Ingrian", "Kukkusi"}, } m["urj-kya"] = { } m["urj-mdv-pro"] = { } m["urj-prm-pro"] = { } m["urj-pro"] = { otherNames = {"Proto-Finno-Ugric", "Proto-Finno-Permic"}, -- PFU and PFP are subsumed into PU per [[Wiktionary:Beer parlour/2015/January#Merging Finno-Volgaic, Finno-Samic, Finno-Permic and Finno-Ugric into Uralic]] } m["urj-ugr-pro"] = { } m["xgn-pro"] = { } m["xnd-pro"] = { otherNames = {"Proto-Na-Dené", "Proto-Athabaskan-Eyak-Tlingit"}, } m["yok-bvy"] = { otherNames = {"Tulamni-Hometwoli", "Tulamni", "Tulamne", "Tuolumne", "Tawitchi", "Hometwoli", "Taneshach"}, } m["yok-dly"] = { otherNames = {"Far Northern Valley Yokuts", "Yachikumne", "Yachikumni", "Chulamni", "Lower San Joaquin", "Lakisamni", "Tawalimni"}, } m["yok-gsy"] = { } m["yok-kry"] = { otherNames = {"Choinimni", "Choynimni", "Ayticha", "Kocheyali", "Ayitcha", "Michahay", "Chukaymina", "Chukaimina"}, } m["yok-nvy"] = { otherNames = {"Chukchansi", "Kechayi", "Dumna", "Chawchila", "Noptinte", "Nopṭinṭe", "Nopthrinthre", "Nopchinchi", "Takin"}, } m["yok-ply"] = { otherNames = {"Paleuyami", "Altinin", "Poso Creek", "Poso Creek Yokuts"}, } m["yok-svy"] = { otherNames = {"Yawelmani", "Tachi", "Koyeti", "Nutunutu", "Chunut", "Wo'lasi", "Choynok", "Choinok", "Wechihit"}, } m["yok-tky"] = { otherNames = {"Wikchamni", "Wukchamni", "Wukchumni", "Yawdanchi"}, } m["ypk-pro"] = { } m["yrk-for"] = { } m["yrk-tun"] = { } m["zhx-min-pro"] = { } m["zhx-sht"] = { otherNames = {"Xiangnan Tuhua", "Yuebei Tuhua", "Shipo", "Shina"}, } m["zhx-sic"] = { otherNames = {"Sichuanese Mandarin"}, } m["zhx-tai"] = { aliases = {"Toishanese"}, } m["zle-ono"] = { } m["zle-ort"] = { } m["zlm-coa"] = { } m["zlm-pah"] = { } m["zls-chs"] = { } m["zlw-ocs"] = { } m["zlw-opl"] = { } m["zlw-osk"] = { } m["zlw-slv"] = { } return m k4oi3h042u52duml7dvr2ovzu55yu2m Modul:languages/data/3/t/extra 828 33780 375379 373750 2026-09-22T05:36:14Z Hakimi97 2668 [[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]] 375379 Scribunto text/plain local m = {} m["taa"] = { otherNames = {"Tanana", "Middle Tanana"}, } m["tab"] = { aliases = {"Tabassaran"}, } m["tac"] = { } m["tad"] = { } m["tae"] = { } m["taf"] = { } m["tag"] = { } m["taj"] = { } m["tak"] = { } m["tal"] = { } m["tan"] = { } m["tao"] = { otherNames = {"Tao"}, } m["tap"] = { } m["tar"] = { } m["tas"] = { otherNames = {"Tay Boi Pidgin French", "Vietnamese Pidgin French"}, } m["tau"] = { otherNames = {"Tabesna", "Nabesna"}, } m["tav"] = { } m["taw"] = { } m["tax"] = { } m["tay"] = { } m["taz"] = { } m["tba"] = { } m["tbc"] = { } m["tbd"] = { } m["tbe"] = { } m["tbf"] = { } m["tbg"] = { } m["tbh"] = { } m["tbi"] = { otherNames = {"Ingessana", "Gaahmg"}, } m["tbj"] = { } m["tbk"] = { } m["tbl"] = { aliases = {"Tagabili"}, } m["tbm"] = { } m["tbn"] = { } m["tbo"] = { } m["tbp"] = { otherNames = {"Diebroud", "Dabra"}, } m["tbr"] = { } m["tbs"] = { } m["tbt"] = { otherNames = {"Tembo"}, } m["tbu"] = { otherNames = {"Tubare"}, } m["tbv"] = { } m["tbw"] = { } m["tbx"] = { otherNames = {"Middle Watut"}, } m["tby"] = { } m["tbz"] = { } m["tca"] = { otherNames = {"Tikuna"}, } m["tcb"] = { } m["tcc"] = { } m["tcd"] = { } m["tce"] = { } m["tcf"] = { } m["tcg"] = { } m["tch"] = { } m["tci"] = { } m["tck"] = { } m["tcl"] = { otherNames = {"Taman", "Taman (Burma)"}, } m["tcm"] = { } m["tco"] = { } m["tcp"] = { otherNames = {"Tawr"}, } m["tcq"] = { } m["tcs"] = { otherNames = {"Big Thap", "Blaikman", "Brokan", "Broken", "Broken English", "Cape York Creole", "Lockhart Creole", "Papuan Pidgin English", "Torres Strait Brokan", "Torres Strait Broken", "Torres Strait Pidgin", "Yumplatok"}, } m["tct"] = { } m["tcu"] = { } m["tcw"] = { } m["tcx"] = { } m["tcy"] = { } m["tcz"] = { otherNames = {"Thado"}, } m["tda"] = { } m["tdb"] = { } m["tdc"] = { } m["tdd"] = { aliases = {"Tai Nuea", "Dehong Dai", "Tai Dehong", "Tai Le", "Chinese Shan", "Chinese Tai"}, } m["tde"] = { } m["tdf"] = { otherNames = {"Taliang", "Tariang", "Kasseng"}, } m["tdg"] = { } m["tdh"] = { } m["tdi"] = { } m["tdj"] = { } m["tdk"] = { } m["tdl"] = { otherNames = {"Tapshin"}, } m["tdm"] = { otherNames = {"Taruamá"}, } m["tdn"] = { } m["tdo"] = { } m["tdq"] = { } m["tdr"] = { } m["tds"] = { otherNames = {"Taori"}, } m["tdt"] = { otherNames = {"Tetum Dili", "Tetun Prasa", "Tétum Praça", "Tetun-Dili", "Tetun-Prasa"}, } m["tdv"] = { } m["tdy"] = { } m["tea"] = { } m["teb"] = { } m["tec"] = { } m["ted"] = { } m["tee"] = { } m["tef"] = { } m["teg"] = { } m["teh"] = { otherNames = {"Patagón", "Chon", "Chon Patagón", "Chon Patagon", "Aoniken", "Aonikenk", "Inaquen", "Aonek'o 'ajen"}, } m["tei"] = { } m["tek"] = { } m["tem"] = { otherNames = {"Timne", "Themne", "KaThemne"}, } m["ten"] = { otherNames = {"Tama"}, } m["teo"] = { } m["tep"] = { } m["teq"] = { } m["ter"] = { } m["tes"] = { } m["tet"] = { otherNames = {"Tetun"}, } m["teu"] = { } m["tev"] = { } m["tew"] = { otherNames = {"Tano", "Santa Clara Tewa", "San Ildefonso Tewa", "Tesuque Tewa", "Nambe Tewa", "Ohkay Owingeh", "Pojoaque"}, } m["tex"] = { } m["tey"] = { } m["tez"] = { otherNames = {"Tin Sert"}, } m["tfi"] = { } m["tfn"] = { otherNames = {"Tanaina"}, } m["tfo"] = { } m["tfr"] = { } m["tft"] = { } m["tga"] = { } m["tgb"] = { } m["tgc"] = { } m["tgd"] = { } m["tge"] = { } m["tgf"] = { otherNames = {"Chalikha", "Chalipkha", "Tshali", "Tshalingpa"}, } m["tgh"] = { } m["tgi"] = { } m["tgn"] = { } m["tgo"] = { } m["tgp"] = { } m["tgq"] = { } m["tgr"] = { } m["tgs"] = { } m["tgt"] = { } m["tgu"] = { } m["tgv"] = { } m["tgw"] = { } m["tgx"] = { } m["tgy"] = { } m["thc"] = { } m["thd"] = { otherNames = {"Thaayorre", "Thayore"}, } m["the"] = { } m["thf"] = { } m["thh"] = { } m["thi"] = { } m["thk"] = { } m["thl"] = { } m["thm"] = { aliases = {"Aheu", "So (Thavung)"}, } m["thn"] = { } m["thp"] = { } m["thq"] = { } m["thr"] = { } m["ths"] = { } m["tht"] = { } m["thu"] = { } m["thy"] = { } m["tic"] = { } m["tif"] = { } m["tig"] = { } m["tih"] = { } m["tii"] = { } m["tij"] = { } m["tik"] = { } m["til"] = { } m["tim"] = { } m["tin"] = { } m["tio"] = { } m["tip"] = { } m["tiq"] = { } m["tis"] = { } m["tit"] = { } m["tiu"] = { } m["tiv"] = { otherNames = {"Tivi"}, } m["tiw"] = { } m["tix"] = { otherNames = {"Isleta", "Isleta Tiwa", "Isleta Pueblo", "Sandia", "Sandia Tiwa", "Sandia Pueblo"}, } m["tiy"] = { aliases = {"Teduray"}, } m["tiz"] = { } m["tja"] = { } m["tjg"] = { } m["tji"] = { } m["tjl"] = { aliases = {"Red Tai (Myanmar)", "Red Shan", "Shan Bamar", "Shan Kalee", "Shan Ni", "Tai Laeng", "Tai Lai", "Tai Leng", "Tai Nai", "Tai Naing"}, } m["tjm"] = { } m["tjn"] = { } m["tjs"] = { } m["tju"] = { } m["tjw"] = { otherNames = {"Djabwurrung", "Djab Wurrung", "Tjapwurrung"}, } m["tka"] = { } m["tkb"] = { } m["tkd"] = { } m["tke"] = { } m["tkf"] = { } m["tkl"] = { } m["tkm"] = { } m["tkn"] = { aliases = {"Tokunoshima", "Toku-no-Shima"}, } m["tkp"] = { } m["tkq"] = { } m["tkr"] = { otherNames = {"Caxur", "Tsaxur"}, } m["tks"] = { otherNames = {"Takestani"}, } m["tkt"] = { } m["tku"] = { } m["tkv"] = { otherNames = {"Pano"}, } m["tkw"] = { } m["tkx"] = { } m["tkz"] = { } m["tla"] = { } m["tlb"] = { } m["tlc"] = { otherNames = {"Yecuatla Totonac"}, } m["tld"] = { } m["tlf"] = { } m["tlg"] = { } m["tlh"] = { } m["tli"] = { } m["tlj"] = { } m["tlk"] = { } m["tll"] = { varieties = {"Indanga"}, } m["tlm"] = { } m["tln"] = { } m["tlo"] = { } m["tlp"] = { } m["tlq"] = { } m["tlr"] = { } m["tls"] = { } m["tlt"] = { otherNames = {"Sou Nama"}, } m["tlu"] = { } m["tlv"] = { otherNames = {"Soboyo"}, } m["tlx"] = { } m["tly"] = { otherNames = {"Talyshi", "Talishi", "Taleshi", "Tolashi", "Asalemi", "Anbarani"}, } m["tma"] = { otherNames = {"Tama"}, } m["tmb"] = { } m["tmc"] = { } m["tmd"] = { } m["tme"] = { } m["tmf"] = { } m["tmg"] = { } m["tmh"] = { otherNames = {"Tamashek", "Tamahaq", "Tamajaq", "Tamasheq"}, } m["tmi"] = { } m["tmj"] = { } m["tml"] = { } m["tmm"] = { } m["tmn"] = { otherNames = {"Taman"}, } m["tmo"] = { } m["tmq"] = { } m["tms"] = { } m["tmt"] = { } m["tmu"] = { otherNames = {"Turu"}, } m["tmv"] = { otherNames = {"Tembo"}, } m["tmw"] = { } m["tmy"] = { } m["tmz"] = { } m["tna"] = { } m["tnb"] = { } m["tnc"] = { } m["tnd"] = { } m["tne"] = { } m["tng"] = { } m["tnh"] = { } m["tni"] = { } m["tnk"] = { } m["tnl"] = { } m["tnm"] = { } m["tnn"] = { } m["tno"] = { } m["tnp"] = { } m["tnq"] = { aliases = {"Taino"}, } m["tnr"] = { } m["tns"] = { } m["tnt"] = { } m["tnu"] = { } m["tnv"] = { aliases = {"Tangchangya", "Tonchongya", "Tongchongya"}, } m["tnw"] = { } m["tnx"] = { } m["tny"] = { } m["tnz"] = { otherNames = {"Tonga"}, } m["tob"] = { otherNames = {"Chaco Sur", "Namqom", "Qom", "Toba Qom"}, } m["toc"] = { } m["tod"] = { } m["tof"] = { } m["tog"] = { otherNames = {"Kitonga", "Chitonga", "Siska", "Sisya", "Tonga", "Western Nyasa"}, } m["toh"] = { otherNames = {"Gitonga", "Tonga"}, } m["toi"] = { otherNames = {"Tonga", "Chitonga", "Plateau Tonga", "Zambezi"}, } m["toj"] = { } m["tok"] = { } m["tol"] = { otherNames = {"Smith River", "Smith River Tolowa"}, } m["tom"] = { } m["too"] = { } m["top"] = { } m["toq"] = { } m["tor"] = { } m["tos"] = { } m["tou"] = { } m["tov"] = { } m["tow"] = { otherNames = {"Towa"}, } m["tox"] = { } m["toy"] = { } m["toz"] = { } m["tpa"] = { } m["tpc"] = { } m["tpe"] = { } m["tpf"] = { } m["tpg"] = { } m["tpi"] = { otherNames = {"Melanesian Pidgin English", "Neo-Melanesian", "New Guinea Pidgin"}, } m["tpj"] = { } m["tpk"] = { otherNames = {"Coastal Tupi", "Tupiniquim"}, } m["tpl"] = { } m["tpm"] = { } m["tpn"] = { } m["tpo"] = { } m["tpp"] = { } m["tpq"] = { } m["tpr"] = { } m["tpt"] = { } m["tpu"] = { } m["tpv"] = { } m["tpw"] = { aliases = {"Classical Tupi"}, } m["tpx"] = { } m["tpy"] = { } m["tpz"] = { } m["tqb"] = { } m["tql"] = { } m["tqm"] = { } m["tqn"] = { } m["tqo"] = { } m["tqp"] = { } m["tqq"] = { } m["tqr"] = { } m["tqt"] = { } m["tqu"] = { } m["tqw"] = { } m["tra"] = { } m["trb"] = { } m["trc"] = { } m["trd"] = { } m["tre"] = { } m["trf"] = { } m["trg"] = { } m["trh"] = { } m["tri"] = { otherNames = {"Trio", "Tiriyó", "Tarano"}, } m["trj"] = { } m["trl"] = { } m["trm"] = { } m["trn"] = { otherNames = {"Trinitario Moxos", "Moxo", "Moxos", "Mojo", "Moxa"}, } m["tro"] = { otherNames = {"Tarao Naga", "Taraotrong", "Tarau"}, } m["trp"] = { } m["trq"] = { } m["trr"] = { } m["trs"] = { } m["trt"] = { } m["tru"] = { } m["trv"] = { otherNames = {"Seediq"}, } m["trw"] = { } m["trx"] = { otherNames = {"Tringus", "Tringgus-Sembaan Bidayuh"}, } m["try"] = { aliases = {"Tai Turung"}, } m["trz"] = { } m["tsa"] = { } m["tsb"] = { } m["tsc"] = { } m["tsd"] = { } m["tse"] = { } m["tsg"] = { aliases = {"Sūg"}, } m["tsh"] = { } m["tsi"] = { } m["tsj"] = { otherNames = {"Sharchop"}, } m["tsl"] = { } m["tsm"] = { } m["tsp"] = { } m["tsq"] = { } m["tsr"] = { } m["tss"] = { } m["tsu"] = { } m["tsv"] = { } m["tsw"] = { } m["tsx"] = { } m["tsy"] = { } m["tta"] = { } m["ttb"] = { } m["ttc"] = { } m["ttd"] = { } m["tte"] = { otherNames = {"Tubetube"}, } m["ttf"] = { } m["ttg"] = { } m["tth"] = { } m["tti"] = { } m["ttj"] = { aliases = {"Rutooro"}, } m["ttk"] = { aliases = {"Totoró"}, } m["ttl"] = { } m["ttm"] = { } m["ttn"] = { } m["tto"] = { } m["ttp"] = { } m["ttr"] = { } m["tts"] = { aliases = {"Isanese", "Isaan", "Issan", "Northeastern Thai"}, } m["ttt"] = { otherNames = {"Caucasian Tat", "Muslim Tat", "Armeno-Tat"}, } m["ttu"] = { } m["ttv"] = { } m["ttw"] = { otherNames = {"Tutoh"}, } m["tty"] = { } m["ttz"] = { } m["tua"] = { } m["tub"] = { } m["tuc"] = { } m["tud"] = { } m["tue"] = { } m["tuf"] = { } m["tug"] = { } m["tuh"] = { } m["tui"] = { } m["tuj"] = { } m["tul"] = { } m["tum"] = { } m["tun"] = { } m["tuo"] = { } m["tuq"] = { otherNames = {"Teda"}, } m["tus"] = { } m["tuu"] = { } m["tuv"] = { } m["tux"] = { } m["tuy"] = { } m["tuz"] = { } m["tva"] = { } m["tvd"] = { } m["tve"] = { } m["tvk"] = { } m["tvl"] = { } m["tvm"] = { } m["tvn"] = { } m["tvo"] = { } m["tvs"] = { } m["tvt"] = { } m["tvu"] = { otherNames = {"Tunen-Aling'a"}, } m["tvw"] = { } m["tvx"] = { } m["tvy"] = { otherNames = {"Bidau Creole Portuguese"}, } m["twa"] = { } m["twb"] = { } m["twc"] = { } m["twe"] = { otherNames = {"Tewa"}, } m["twf"] = { aliases = {"Northern Tiwa"}, } m["twg"] = { } m["twh"] = { aliases = {"Tai Khao", "White Tai"}, } m["twm"] = { } m["twn"] = { } m["two"] = { } m["twp"] = { } m["twq"] = { } m["twr"] = { } m["twt"] = { } m["twu"] = { } m["tww"] = { } m["twy"] = { otherNames = {"Taboyan"}, } m["txa"] = { } m["txb"] = { otherNames = {"West Tocharian", "Kuchean"}, } m["txc"] = { } m["txe"] = { } m["txg"] = { } m["txj"] = { } m["txh"] = { } m["txi"] = { } m["txm"] = { } m["txn"] = { } m["txo"] = { } m["txq"] = { } m["txr"] = { } m["txs"] = { } m["txt"] = { } m["txu"] = { } m["txx"] = { } m["tya"] = { } m["tye"] = { } m["tyh"] = { } m["tyi"] = { } m["tyj"] = { aliases = {"Tai Yo", "Tai Mène", "Tai Maen"}, } m["tyl"] = { } m["tyn"] = { } m["typ"] = { otherNames = {"Gugu Thaypan", "Thaypan", "Kuku Thaypan", "Agu Alaya", "Awu Alaya", "Alaya", "Gugu-Rarmul", "Koko-Rarmul", "Rarmul"}, } m["tyr"] = { aliases = {"Red Tai (Vietnam)"}, } m["tys"] = { aliases = {"Sa Pa", "Tày Sa Pa", "Tai Sapa"}, } m["tyt"] = { } m["tyu"] = { } m["tyv"] = { aliases = {"Tyvan"}, } m["tyx"] = { } m["tyz"] = { aliases = {"Tay", "Tho", "Bao Yen", "Cao Bang"}, -- Both Bao Lac and Trung Khanh are located in Cao Bang. varieties = {"Central Tày", "Eastern Tày", "Northern Tày", "Southern Tày", "Tày Bao Lac", "Tày Trung Khanh"}, } m["tza"] = { } m["tzh"] = { } m["tzj"] = { aliases = {"Tzutujil"}, } m["tzl"] = { } m["tzm"] = { } m["tzn"] = { } m["tzo"] = { } m["tzx"] = { otherNames = {"Karawari"}, } return m t8jijj9iofqeq1rud9560uhc9d65jn2 Modul:gender and number/data 828 33823 375354 231400 2026-09-22T03:13:31Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/89447487|89447487]]) 375354 Scribunto text/plain local data = {} local insert = table.insert -- A list of all possible "parts" that a specification can be made out of. For each part, we list the class it's in -- (gender, animacy, etc.), the associated category (if any) and the display form. In a given gender/number spec, only -- one part of each class is allowed. `display` is how the code is diplayed to the user and should normally be wrapped -- in <abbr title="tooltip">...</abbr> with an explanatory tooltip. If not, it will automatically be wrapped in this -- fashion. If `req` is true, a category "Requests for TYPE in LANG entries" will be generated, except for the code "?", -- which is special-cased; TYPE is "gender" unless the POS is "verb", in which case it is "aspect". data.codes = { ["?"] = {type = "other", req = true, display = '<abbr title="gender incomplete">?</abbr>'}, -- FIXME: The following should be either eliminated in favor of g! or converted to a general "gender/number unattested". ["?!"] = {type = "other", display = "genus tidak disahkan"}, -- Genders ["m"] = {type = "gender", cat = "POS maskulin", display = '<abbr title="masculine gender">m</abbr>'}, ["f"] = {type = "gender", cat = "POS feminin", display = '<abbr title="feminine gender">f</abbr>'}, ["n"] = {type = "gender", cat = "POS neuter", display = '<abbr title="neuter gender">n</abbr>'}, ["c"] = {type = "gender", cat = "POS genus umum", display = '<abbr title="common gender">c</abbr>'}, ["gneut"] = {type = "gender", cat = "POS bebas genus", display = "gender-neutral"}, ["g!"] = {type = "gender", display = "genus tidak disahkan"}, ["g?"] = {type = "gender", req = true, display = "genus tidak khusus"}, -- Animacy -- Animate = either animal or personal (for Russian, etc.) ["an"] = {type = "animacy", cat = "POS bernyawa", display = '<abbr title="animate">anim</abbr>'}, ["in"] = {type = "animacy", cat = "POS tidak bernyawa", display = '<abbr title="inanimate">inan</abbr>'}, -- Animal (for Ukrainian, Belarusian, Polish, etc.) ["anml"] = {type = "animacy", cat = "POS haiwan", display = "animal"}, -- Personal (for Ukrainian, Belarusian, Polish, etc.) ["pr"] = {type = "animacy", cat = "POS peribadi", display = '<abbr title="personal">pers</abbr>'}, ["np"] = {type = "animacy", cat = "POS bukan peribadi", display = '<abbr title="nonpersonal">npers</abbr>'}, ["an!"] = {type = "animacy", display = "kebernyawaan tidak disahkan"}, ["an?"] = {type = "animacy", req = true, display = "kebernyawaan tidak khusus"}, -- Definiteness ["def"] = {type = "definiteness", cat = "POS muktamad", display = '<abbr title="definite">def</abbr>'}, ["indef"] = {type = "definiteness", cat = "POS tak muktamad", display = '<abbr title="indefinite">indef</abbr>'}, -- Virility (for Polish) ["vr"] = {type = "virility", cat = "POS jantan", display = '<abbr title="virile (= masculine personal)">vir</abbr>'}, ["nv"] = {type = "virility", cat = "POS bukan jantan", display = '<abbr title="nonvirile (= other than masculine personal)">nvir</abbr>'}, -- Numbers ["s"] = {type = "number", display = '<abbr title="singular number">sg</abbr>'}, ["d"] = {type = "number", cat = "dualia tantum", display = '<abbr title="dual number">du</abbr>'}, ["p"] = {type = "number", cat = "pluralia tantum", display = '<abbr title="plural number">pl</abbr>'}, ["num!"] = {type = "number", display = "bilangan tidak disahkan"}, ["num?"] = {type = "number", req = true, display = "bilangan tidak khusus"}, -- Verb qualifiers ["impf"] = {type = "aspect", cat = "POS tak sempurna", display = '<abbr title="imperfective aspect">impf</abbr>'}, ["pf"] = {type = "aspect", cat = "POS sempurna", display = '<abbr title="perfective aspect">pf</abbr>'}, ["asp!"] = {type = "aspect", display = "aspek tidak disahkan"}, ["asp?"] = {type = "aspect", req = true, display = "aspek tidak khusus"}, } -- Combined codes that are equivalent to giving multiple specs. `mf` is the same as specifying two separate specs, -- one with `m` in it and the other with `f`. `mfbysense` is similar but is used for nouns that can be either masculine -- or feminine according as to whether they refer to masculine or feminine beings. local combinations = { ["biasp"] = {codes = {"impf", "pf"}}, ["anin"] = {codes = {"an", "in"}}, -- "bianimate" doesn't exist as a linguistic term } for _, comb in ipairs{"mf", "mn", "fm", "fn", "cn", "nm", "nf", "nc", "mfn", "mnf", "fmn", "fnm", "nmf", "nfm"} do local codes = {} for ch in comb:gmatch(".") do insert(codes, ch) end combinations[comb] = {codes = codes} combinations[comb .. "equiv"] = {codes = codes, display = '<abbr title="different genders do not affect the meaning">same meaning</abbr>'} if comb == "mf" or comb == "fm" then combinations[comb .. "bysense"] = {codes = codes, cat = "masculine and feminine POS by sense", display = '<abbr title="according to the gender of the referent">by sense</abbr>'} end end data.combinations = combinations -- Categories when multiple gender/number codes of a given type occur in different specs (two or more of the same type -- cannot occur in a single spec). data.multicode_cats = { ["gender"] = "POS dengan berbilang genus", ["animacy"] = "POS dengan berbilang kebernyawaan", ["aspect"] = "POS dwiaspek", } return data 1znket0vk37umsxisddunv6ol7cgsjt Modul:module categorization 828 35046 375363 255151 2026-09-22T04:27:35Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92367848|92367848]]) 375363 Scribunto text/plain local export = {} local put_module = "Modul:parse utilities" local rsplit = mw.text.split local rfind = mw.ustring.find local unpack = unpack or table.unpack -- Lua 5.2 compatibility local insert = table.insert local keyword_to_module_type = { common = "Language-specific utility", utilities = "Language-specific utility", headword = "Headword-line", translit = "transliterasi", infl = "fleksi", inflection = "fleksi", decl = "fleksi", declension = "fleksi", adecl = "fleksi", conj = "fleksi", conjugation = "fleksi", noun = "fleksi", nouns = "fleksi", pronoun = "fleksi", pronouns = "fleksi", verb = "fleksi", verbs = "fleksi", adjective = "fleksi", adjectives = "fleksi", adj = "fleksi", nominal = "fleksi", nominals = "fleksi", pron = "sebutan", pronun = "sebutan", pronunc = "sebutan", pronunciation = "sebutan", IPA = "sebutan", stripdiacritics = "penyiatan diakritik", sortkey = "penjana kunci isih", } -- If a module type is here, we will generate a lang-specific module-type category such as -- [[:Category:Pali inflection modules]]. local module_type_generates_lang_specific_cat = { ["fleksi"] = true, ["data"] = true, ["kes ujian"] = true, } -- If a module type is here, we will fetch extra languages to generate categories for. The value is a module that -- returns a function that fetches all the languages that use a given module for -- transliteration/diacritic-stripping/sortkey generation. local languages_from_module_type = { ["transliterasi"] = "Modul:languages/byTranslitModule", ["kes ujian transliterasi"] = "Modul:languages/byTranslitModule", ["penyiatan diakritik"] = "Modul:languages/byStripDiacriticsModule", ["penjana kunci isih"] = "Modul:languages/bySortkeyModule", } local function get_transliteration_family_cats(objs) local families_to_check = {"aav", "afa", "dra", "inc", "ira", "map", "sit", "sla", "trk", "urj"} local result = {} for _, famcode in ipairs(families_to_check) do for _, obj in ipairs(objs) do if obj:hasType("language") and obj:inFamily(famcode) then local fam = require("Modul:families").getByCode(famcode) insert(result, "Modul transliterasi bahasa-bahasa " .. fam:getCanonicalName()) break end end end return result end -- If a module type is here, we will fetch extra categories given the generated languages/scripts/families. local extra_cats_from_module_type = { ["transliterasi"] = get_transliteration_family_cats, } local module_type_patterns = { {"/data%f[-/%z]", "data"}, {"/testcases%f[-/%z]", function(typ) if typ == "sebutan" then return "kes ujian sebutan" elseif typ == "transliterasi" then return "kes ujian transliterasi" else return "kes ujian" end end}, } -- Split an argument on comma, but not comma followed by whitespace. local function split_on_comma(val) if val:find(",%s") then return require(put_module).split_on_comma(val) else return rsplit(val, ",") end end local function get_lang_or_script(code) return code == "-" and code or require("Modul:languages").getByCode(code, nil, "allow etym") or require("Modul:languages").getByCode(code .. "-pro", nil, "allow etym") or require("Modul:scripts").getByCode(code) end local function obj_code(obj) if obj == "-" then return obj end return obj:getCode() end local function infer_lang_or_script_code(name) local hyphen_parts = rsplit(name, "%-") for i = #hyphen_parts - 1, 1, -1 do local code = table.concat(hyphen_parts, "-", 1, i) local obj = get_lang_or_script(code) if obj then local rest = table.concat(hyphen_parts, "-", i + 1) return obj, rest end end return nil, nil end local function infer_lang_and_script_codes(name) local objs = {} while true do local obj, rest = infer_lang_or_script_code(name) if not obj then return objs, name end if #objs > 0 and obj:getCode() == "to" then -- skip 'to' in e.g. [[Modul:ks-Arab-to-Deva-translit]]; it's not Tongan else insert(objs, obj) end name = rest end end --[==[ Main entry point called from another module. ]==] function export.categorize_module(data) local pagename, return_raw, noerror = data.pagename, data.return_raw, data.noerror local langlist, module_type, return_cats = data.langlist, data.module_type, data.return_cats local title if pagename then title = mw.title.new(pagename, 'Modul') else title = mw.title.getCurrentTitle() -- Fuckme, sometimes this function is called with a faked frame and a title with the namespace already chopped out, -- so this test cannot be done in that case. if title.nsText ~= "Modul" then error(("This template should only be used in the Module namespace, not on page '%s'."):format(title.fullText)) end pagename = title.fullText end local subpage = title.subpageText local null_return_value = return_raw and {} or "" -- To ensure no categories are added on documentation pages. if subpage == "doc" then return null_return_value end local categories = {} local function insert_cat(cat, sortkey) for _, existing_cat in ipairs(categories) do if existing_cat.name == cat then return end end insert(categories, {name = cat, sort = sortkey}) end local root_pagename if subpage ~= pagename then root_pagename = title.rootText else root_pagename = pagename end root_pagename = root_pagename:gsub("^Modul:", "") -- Take the module type(s) from type= if given, or infer from the pagename. local module_types if module_type then module_types = {} local module_type_specs = split_on_comma(module_type) for _, spec in ipairs(module_type_specs) do local modtype, sortkey = spec:match("^(.-):(.*)$") modtype = modtype or spec sortkey = sortkey and sortkey:gsub("_", " ") or nil insert(module_types, {type = modtype, sort = sortkey}) end else local module_type_keyword = root_pagename:match("[-%a]+[- ]([^/]+)%f[/%z]") if not module_type_keyword then if noerror then return null_return_value else error(("Could not extract module type from root pagename '%s'"):format(root_pagename)) end end module_type = keyword_to_module_type[module_type_keyword] if not module_type then if noerror then return null_return_value else error(("Did not recognize inferred module-type keyword '%s' from root pagename '%s'"):format( module_type_keyword, root_pagename)) end end module_types = {{type = module_type}} end -- Look for additional module type(s) inferred by pattern. for _, pattern_spec in ipairs(module_type_patterns) do local pattern, inferred_type = unpack(pattern_spec) if rfind(pagename, pattern) then local function insert_module_type(typ) require("Modul:table").insertIfNot(module_types, typ, {key = function(obj) return obj.type end}) end if type(inferred_type) == "string" then insert_module_type({type = inferred_type}) else local addl_types = {} for _, typ in ipairs(module_types) do insert(addl_types, {type = inferred_type(typ.type), sort = typ.sort}) end for _, typ in ipairs(addl_types) do insert_module_type(typ) end end end end -- If 1= specified, take the languages/scripts directly from there. Otherwise, (a) try to extract one or more -- languages/scripts from the pagename (e.g. [[Modul:uk-be-headword]] -> Ukrainian and Belarusian (languages); -- [[Modul:bho-Kthi-translit]] -> Bhojpuri (language) and Kaithi (script); [[Modul:Deva-Kthi-translit]] -> -- Devanagari and Kaithi (scripts)); and (b) if the specified or inferred module type(s) contain a type listed in -- languages_from_module_type[], use the function referenced there to extract additional languages (i.e. all the -- languages that use the module we are processing). local inferred_objs if langlist then inferred_objs = {} for _, code in ipairs(rsplit(langlist, ",")) do -- We need to have an indicator of families because we allow bare family codes to stand for proto-languages. if code:find("^fam:") then code = code:gsub("^fam:", "") local family = require("Modul:families").getByCode(code) or error(("Unrecognized family code '%s' in [[Modul:module categorization]]"):format(code)) local descendants = family:getDescendantCodes() for _, desc in ipairs(descendants) do local obj = get_lang_or_script(desc) if obj then -- make sure we skip families without proto-languages insert(inferred_objs, obj) end end else local obj = get_lang_or_script(code) if not obj then error(("Unrecognized language or script code '%s'"):format(code)) end insert(inferred_objs, obj) end end else inferred_objs = infer_lang_and_script_codes(root_pagename) for _, modtype in ipairs(module_types) do local languages_extractor = languages_from_module_type[modtype.type] if languages_extractor then local langs = require(languages_extractor)(root_pagename) if langs then for _, obj in ipairs(langs) do require("Modul:table").insertIfNot(inferred_objs, obj, {key = obj_code}) end end end end if #inferred_objs == 0 then if noerror then return null_return_value else error(("Could not infer any languages or scripts from root pagename '%s'"):format(root_pagename)) end end end if pagename:find("^Modul:Pengguna:") then insert_cat("Modul kotak pasir pengguna") elseif pagename:find("/sandbox") then insert_cat("Modul kotak pasir") else for _, modtype in ipairs(module_types) do for _, obj in ipairs(inferred_objs) do local function insert_overall_module_type_cat(sortkey) if modtype.type ~= "-" then insert_cat("Modul " .. modtype.type, modtype.sort or sortkey) end end if obj == "-" then insert_overall_module_type_cat() else if obj:hasType("script") and modtype.type ~= "-" then insert_cat("Modul " .. modtype.type .. " mengikut tulisan", obj:getCanonicalName()) end local function construct_lang_or_sc_cat(obj, suffix) local prefix if obj:hasType("language") then prefix = obj:getFullName() else prefix = obj:getCategoryName() end return suffix .. " bahasa " .. prefix end insert_cat(construct_lang_or_sc_cat(obj, "Modul"), modtype.type) insert_overall_module_type_cat(obj:getCanonicalName()) if module_type_generates_lang_specific_cat[modtype.type] then insert_cat(construct_lang_or_sc_cat(obj, "Modul bahasa " .. mw.getContentLanguage():lcfirst(modtype.type))) end end end if extra_cats_from_module_type[modtype.type] then local extra_cats = extra_cats_from_module_type[modtype.type](inferred_objs) for _, cat in ipairs(extra_cats) do insert_cat(cat) end end end end for i, catspec in ipairs(categories) do if catspec.sort then categories[i] = ("%s|%s"):format(catspec.name, catspec.sort) else categories[i] = catspec.name end end if return_cats then return table.concat(categories, ",") elseif return_raw then return categories else for i, cat in ipairs(categories) do categories[i] = "[[Kategori:" .. cat .. "]]" end return table.concat(categories) end end --[==[ Main entry point called from a template. ]==] function export.categorize(frame) local params = { [1] = true, -- comma-separated list of languages; by default, inferred from module name type = true, [2] = {alias_of = "type"}, pagename = true, -- for testing return_cats = {type = "boolean"}, -- for testing } local parent_args = frame:getParent().args local args = require("Modul:parameters").process(parent_args, params) return export.categorize_module { pagename = args.pagename, langlist = args[1], module_type = args.type, return_cats = args.return_cats, } end --[==[Table used in the documentation to {{tl|module cat}}.]==] function export.keyword_to_module_type_table() local parts = {} local function ins(text) insert(parts, text) end ins('{|class="wikitable"') ins("! Kata kunci !! Jenis modul disimpulkan") local keywords = {} for k, v in pairs(keyword_to_module_type) do insert(keywords, k) end table.sort(keywords) for _, keyword in ipairs(keywords) do ins("|-") ins(("| <code>%s</code> || <code>%s</code>"):format(keyword, keyword_to_module_type[keyword])) end ins("|}") return table.concat(parts, "\n") end return export nsy77mhqd1amfc11l0mjbxjw2vhai2t Modul:audio 828 48711 375391 226683 2026-09-22T07:36:06Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/88355452|88355452]]) 375391 Scribunto text/plain local export = {} local headword_data_module = "Module:headword/data" local IPA_module = "Module:IPA" local labels_module = "Module:labels" local links_module = "Module:links" local parameters_module = "Module:parameters" local qualifier_module = "Module:qualifier" local references_module = "Module:references" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local template_styles_module = "Module:TemplateStyles" local utilities_module = "Module:utilities" local audio_styles_css = "audio/styles.css" local function track(page) require("Module:debug/track")("audio/" .. page) return true end local function wrap_qualifier_css(text, suffix) return require(qualifier_module).wrap_qualifier_css(text, suffix) end --[==[ Display a box that can be used to play an audio file. `data` is a table containing the following fields: * `lang` ('''required'''): language object for the audio files; * `file` ('''required'''): file containing the audio; * `caption`: Caption to display before the audio box; normally {"Audio"}, and does not usually need to be changed; * `nocaption`: If specified, don't display the caption; * `q`: {nil} or a list of left regular qualifier strings, formatted using {format_qualifier()} in [[Module:qualifier]] and displayed before the audio box and after the caption (and any accent qualifiers); * `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the audio box (and after any accent qualifiers); * `a`: {nil} or a list of left accent qualifier strings, formatted using {format_qualifiers()} in [[Module:accent qualifier]] and displayed before the audio box and after the caption; * `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the homophone in question; * `refs`: {nil} or a list of references or reference specs to add directly after the audio box; the value of a list item is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}}) and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the actual reference, located in the {{cd|<nowiki><references /></nowiki>}} section; * `text`: Text of the audio snippet; if specified, should be an object of the form passed to {full_link()} in [[Module:links]], including a `lang` field containing the language of the text (usually the same as `data.lang`); displayed before the audio box, after any regular and accent qualifiers; * `IPA`: IPA of the audio snippet, or a list of IPA specs; if specified, should be surrounded by slashes or brackets, and will be processed using {format_IPA_multiple()} in [[Module:IPA]] and displayed before the audio box, after any regular and accent qualifiers and after the text of the audio snippet, if given; * `nocat`: If true, suppress categorization; * `sort`: Sort key for categorization. ]==] function export.format_audio(data) local langname = data.lang:getFullName() local cats = { "Perkataan bahasa " .. langname .. " dengan sebutan audio" } local function format_a(a) if a and a[1] then return require(labels_module).show_labels { lang = data.lang, labels = a, mode = "accent", nocat = true, open = false, close = false, no_track_already_seen = true, } end return nil end local function format_q(q) if q and q[1] then return require(qualifier_module).format_qualifier(q, false, false) end return nil end local function make_td_if(text) if text == "" then return text end return "<td>" .. text .. "</td>" end -- Generate the full text preceding the audio box. local pretext_parts = {} local function ins(text) table.insert(pretext_parts, text) end local formatted_accent_labels, formatted_qualifiers, formatted_text, formatted_ipa formatted_accent_labels = format_a(data.a) formatted_qualifiers = format_q(data.q) if data.text then formatted_text = require(links_module).full_link(data.text, "term", true) end if data.IPA then local ipa_cats local ipa = data.IPA if type(ipa) == "string" then ipa = {ipa} end local ipa_items = {} for _, ipa_item in ipairs(ipa) do table.insert(ipa_items, {pron = ipa_item}) end formatted_ipa, ipa_cats = require(IPA_module).format_IPA_multiple(data.lang, ipa_items, nil, "no count", "raw") if ipa_cats[1] then require(table_module).extend(cats, ipa_cats) end end local has_qual = formatted_accent_labels or formatted_qualifiers if not data.nocaption then -- Track uses of caption (3=). Over time as we eliminate most of them, we can use this to find and -- eliminate the remainder. if data.caption then track("caption") end ins(data.caption or "Audio") if has_qual then ins(" " .. wrap_qualifier_css("(", "brac")) end end if formatted_accent_labels then ins(formatted_accent_labels) if formatted_qualifiers then ins(wrap_qualifier_css(",", "comma") .. " ") end end if formatted_qualifiers then ins(formatted_qualifiers) end if has_qual then if not data.nocaption then ins(wrap_qualifier_css(")", "brac")) end end if (formatted_text or formatted_ipa) and (has_qual or not data.nocaption) then ins(wrap_qualifier_css(";", "semicolon") .. " ") end if formatted_text then ins(formatted_text) if formatted_ipa then ins(" ") end end ins(formatted_ipa) if not data.nocaption then ins(wrap_qualifier_css(":", "colon")) end local pretext = make_td_if(table.concat(pretext_parts)) -- Generate the full text following the audio box. local posttext_parts = {} local function ins(text) table.insert(posttext_parts, text) end local formatted_post_accent_labels = format_a(data.aa) local formatted_post_qualifiers = format_q(data.qq) local formatted_references = data.refs and require(references_module).format_references(data.refs) or nil if formatted_references then ins(formatted_references) end if formatted_post_accent_labels or formatted_post_qualifiers then if formatted_references then ins(" ") end ins(wrap_qualifier_css("(", "brac")) if formatted_post_accent_labels then ins(formatted_post_accent_labels) if formatted_post_qualifiers then ins(wrap_qualifier_css(",", "comma") .. " ") end end if formatted_post_qualifiers then ins(formatted_post_qualifiers) end ins(wrap_qualifier_css(")", "brac")) end if data.bad then table.insert(cats, langname .. " terms with nonstandard or incorrect audio pronunciations") ins(" " .. require(qualifier_module).wrap_css("Note: this pronunciation may be nonstandard or incorrect: " .. data.bad, "bad-audio-note")) end local posttext = make_td_if(table.concat(posttext_parts)) local template = [=[ <tr>%s<td class="audiofile">[[File:%s|noicon|175px]]</td><td class="audiometa" style="font-size: 80%%;">([[:File:%s|file]])</td>%s</tr>]=] local text = template:format(pretext, data.file, data.file, posttext) text = '<table class="audiotable" style="vertical-align: middle; display: inline-block; list-style: none; line-height: 1em; border-collapse: collapse; margin: 0;">' .. text .. "</table>" local stylesheet = require(template_styles_module)(audio_styles_css) local categories = data.nocat and "" or cats[1] and require(utilities_module).format_categories(cats, data.lang, data.sort) or "" return stylesheet .. text .. categories end --[==[ FIXME: Old entry point for formatting multiple audios in a single table. Not used anywhere and needs rewriting to the standard of format_audio(). Meant to be called from a module. `data` is a table containing the following fields: <pre> { lang = LANGUAGE_OBJECT, audios = {{file = "FILENAME", qualifiers = nil or {"QUALIFIER", "QUALIFIER", ...}}, ...}, caption = nil or "CAPTION" } </pre> Here: * `lang` is a language object. * `audios` is the list of audio files to display. FILENAME is the name of the audio file without a namespace. QUALIFIER is a qualifier string to display after the specific audio file in question, formatted using {format_qualifier()} in [[Module:qualifier]]. * `caption`, if specified, adds a caption before the audio file. ]==] function export.format_multiple_audios(data) local audiocats = { "Perkataan bahasa " .. data.lang:getFullName() .. " dengan sebutan audio" } local rows = { } local caption = data.caption for _, audio in ipairs(data.audios) do local qualifiers = audio.qualifiers local function repl(key) if key == "file" then return audio.file elseif key == "caption" then if not caption then return "" end return "<td rowspan=" .. #data.audios .. ">" .. caption .. ":</td>" elseif key == "qualifiers" then if not qualifiers or not qualifiers[1] then return "" end return "<td>" .. require(qualifier_module).format_qualifier(qualifiers) .. "</td>" end end local template = [=[ <tr>{{{caption}}} <td class="audiofile">[[File:{{{file}}}|noicon|175px]]</td> <td class="audiometa" style="font-size: 80%;">([[:File:{{{file}}}|file]])</td> {{{qualifiers}}}</tr>]=] local text = (mw.ustring.gsub(template, "{{{([a-z0-9_:]+)}}}", repl)) table.insert(rows, text) caption = nil end local function repl(key) if key == "rows" then return table.concat(rows, "\n") end end local template = [=[ <table class="audiotable" style="vertical-align: middle; display: inline-block; list-style: none; line-height: 1em; border-collapse: collapse;"> {{{rows}}} </table> ]=] local stylesheet = require(template_styles_module)(audio_styles_css) local text = mw.ustring.gsub(template, "{{{([a-z0-9_:]+)}}}", repl) local categories = data.nocat and "" or #audiocats > 0 and require(utilities_module).format_categories(audiocats, data.lang, data.sort) or "" -- remove newlines due to HTML generator bug in MediaWiki(?) - newlines in tables cause list items to not end correctly text = mw.ustring.gsub(text, "\n", "") return stylesheet .. text .. categories end --[==[ Construct the `text` object passed into {format_audio()}, from raw-ish arguments (essentially, the output of {process()} in [[Module:parameters]]). On entry, `args` contains the following fields: * `lang` ('''required'''): Language object. * `text`: Text. If this isn't defined and neither are any of `gloss`, `tr`, `ts`, `pos`, `lit` or `genders`, the function returns {nil}. * `gloss`: Gloss of text. * `tr`: Manual transliteration of text. * `ts`: Transcription of text. * `pos`: Part of speech of text. * `lit`: Literal meaning of text. * `genders`: List of gender/number spec(s) of text. * `sc`: Optional script object of text (rarely needs to be set). * `pagename`: Pagename; used in place of `text` when `text` is unset but other text-related parameters are set. If not specified, taken from the actual pagename. ]==] function export.construct_audio_textobj(args) local textobj if args.text or args.gloss or args.tr or args.ts or args.pos or args.lit or args.genders and args.genders[1] then local text = args.text or args.pagename or mw.loadData("Module:headword/data").pagename textobj = { lang = args.lang, alt = wrap_qualifier_css("“", "quote") .. text .. wrap_qualifier_css("”", "quote"), gloss = args.gloss, tr = args.tr, ts = args.ts, pos = args.pos, lit = args.lit, genders = args.genders, sc = args.sc, } end return textobj end --[==[ Entry point for {{tl|audio}} template. ]==] function export.show(frame) local parent_args = frame:getParent().args local compat = parent_args.lang local offset = compat and 0 or 1 local params = { [compat and "lang" or 1] = {required = true, type = "language", default = "en"}, [1 + offset] = {required = true, default = "Example.ogg"}, [2 + offset] = {}, ["q"] = {type = "qualifier"}, ["qq"] = {type = "qualifier"}, ["a"] = {type = "labels"}, ["aa"] = {type = "labels"}, ["ref"] = {type = "references"}, ["IPA"] = {sublist = true}, ["text"] = {}, ["t"] = {}, ["gloss"] = {alias_of = "t"}, ["tr"] = {}, ["ts"] = {}, ["pos"] = {}, ["lit"] = {}, ["g"] = {sublist = true}, ["sc"] = {type = "script"}, ["bad"] = {}, ["nocat"] = {type = "boolean"}, ["sort"] = {}, ["pagename"] = {}, } local args = require(parameters_module).process(parent_args, params) local lang = args[compat and "lang" or 1] -- Needed in construct_audio_textobj(). args.lang = lang local textobj = export.construct_audio_textobj(args) local caption = args[2 + offset] local nocaption if caption == "-" then caption = nil nocaption = true end if caption then -- Remove final colon if given, to avoid two colons. caption = caption:gsub(":$", "") end local data = { lang = lang, file = args[1 + offset], caption = caption, nocaption = nocaption, q = args.q, qq = args.qq, a = args.a, aa = args.aa, refs = args.ref, text = textobj, IPA = args.IPA, bad = args.bad, nocat = args.nocat, sort = args.sort, } return export.format_audio(data) end return export anshp52gfmccr0u5s0khcpl3vgtaamp Modul:headword utilities 828 54854 375352 223188 2026-09-22T03:13:11Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[:en:Special:Diff/92717011|92717011]]) 375352 Scribunto text/plain local export = {} local require_when_needed = require("Module:utilities/require when needed") local affix_module = "Module:affix" local debug_track_module = "Module:debug/track" local decorations_module = "Module:decorations" local en_utilities_module = "Module:en-utilities" local fun_is_callable_module = "Module:fun/isCallable" local headword_module = "Module:headword" local headword_data_module = "Module:headword/data" local languages_module = "Module:languages" local links_module = "Module:links" local parameters_module = "Module:parameters" local parse_interface_module = "Module:parse interface" local parse_utilities_module = "Module:parse utilities" local string_pattern_escape_module = "Module:string/patternEscape" local string_replacement_escape_module = "Module:string/replacementEscape" local string_utilities_module = "Module:string utilities" local table_module = "Module:table" local yesno_module = "Module:yesno" local dump = mw.dumpObject local unpack = unpack or table.unpack -- Lua 5.2 compatibility local insert = table.insert local concat = table.concat local remove = table.remove local sort = table.sort local deep_equals = require_when_needed(table_module, "deepEquals") local extend = require_when_needed(table_module, "extend") local insert_if_not = require_when_needed(table_module, "insertIfNot") local list_to_set = require_when_needed(table_module, "listToSet") local serial_comma_join = require_when_needed(table_module, "serialCommaJoin") local shallow_copy = require_when_needed(table_module, "shallowCopy") local split = require_when_needed(string_utilities_module, "split") local ugsub = require_when_needed(string_utilities_module, "gsub") local umatch = require_when_needed(string_utilities_module, "match") local pattern_escape = require_when_needed(string_pattern_escape_module) local replacement_escape = require_when_needed(string_replacement_escape_module) local escape_wikicode = require_when_needed(parse_utilities_module, "escape_wikicode") local parse_inline_modifiers = require_when_needed(parse_utilities_module, "parse_inline_modifiers") local term_contains_top_level_html = require_when_needed(parse_utilities_module, "term_contains_top_level_html") local get_lang_by_code = require_when_needed(languages_module, "getByCode") local is_callable = require_when_needed(fun_is_callable_module) local format_decorations = require_when_needed(decorations_module, "format_decorations") local function split_on_comma(val) if val:find(",") then return require(parse_interface_module).split_on_comma(val) else return {val} end end local function ine(val) if val == "" then return nil else return val end end --[=[ Add decorations to a term. `termobj` is the object describing the term, which should optionally contain: * left qualifiers in `q`, an array of strings; * right qualifiers in `qq`, an array of strings; * left labels in `l`, an array of strings; * right labels in `ll`, an array of strings; * references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text` (formatted reference text) and optionally `name` and/or `group`; `text` is the text of the term itself, and `lang` is the language object. ]=] local function add_decorations(text, termobj, lang) local function field_non_empty(field) local list = termobj[field] if not list then return nil end if type(list) ~= "table" then error(("Internal error: Wrong type for `termobj.%s`=%s, should be \"table\""):format( field, mw.dumpObject(list))) end return list[1] end if field_non_empty("q") or field_non_empty("qq") or field_non_empty("l") or field_non_empty("ll") or field_non_empty("refs") then text = format_decorations { lang = lang, text = text, q = termobj.q, qq = termobj.qq, l = termobj.l, ll = termobj.ll, refs = termobj.refs, } end return text end local param_mods = { id = {}, -- disabled when `is_head = true` q = {type = "qualifier"}, qq = {type = "qualifier"}, l = {type = "labels"}, ll = {type = "labels"}, -- [[Module:headword]] expects part references in `.refs`. ref = {item_dest = "refs", type = "references", store = "insert-flattened"}, } local optional_param_mods = { g = {item_dest = "genders", type = "genders"}, alt = {}, lang = {type = "language"}, sc = {type = "script"}, t = {item_dest = "gloss"}, gloss = {}, pos = {}, lit = {}, tr = {}, ts = {}, face = {}, nolinkinfl = {type = "boolean"}, } local optional_headword_param_mods = { sc = {type = "script"}, tr = {}, ts = {}, } --[==[ Parse a single inflection or headword form or list of such forms. In either case, inline modifiers may be attached. `data` is an object with the following fields: * `val`: The raw value to parse. Required. * `paramname`: The name of the parameter from which the value was taken; used in error messages. Required. * `is_head`: We are parsing a headword parameter (a value which goes into the `heads` field of `data`). This changes the allowed modifiers, disabling the `id` modifier and only allowing a subset of optional modifiers. * `frob`: An optional function of one value to apply to the form after inline modifiers have been removed (i.e. to apply to the `.term` field of the returned object). * `include_mods`: List of extra inline modifiers to include, besides the default ones (see below). Each list item is either a string specifying a recognized extra inline modifier (see `optional_param_mods` in the code), or a two-item list of modifier name and modifier spec, where the spec should follow the syntax for modifier specs in `parse_inline_modifiers` in [[Module:parse utilities]]. * `exclude_mods`: List of default inline modifiers to not include. * `splitchar`: If specified, the value in `val` can be a list of forms to parse, separated by the value of `splitchar` (which is a Lua pattern, as in `parse_inline_modifiers` in [[Module:parse utilities]]). Most commonly, `splitchar` is a single comma and the values are comma-separated (in this case, splitting will not happen if a space follows the comma). * `parse_lang_prefix`: If specified, allow a language prefix to precede a form, and if found, store into the `.lang` field of the returned object. * `preserve_splitchar`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_inline_modifiers` in [[Module:parse utilities]]. Returns an object suitable for storing as one element of one of the lists in `headdata.inflections`, where `headdata` is the structure passed to [[Module:headword]]. If `splitchar` is specified, howeve, the return value is a list of such objects. The following default inline modifiers are currently recognized: * `q`: Left qualifier. * `qq`: Right qualifier. * `l`: Comma-separated list of left labels. No space should follow the comma. * `ll`: Comma-separated list of right labels. No space should follow the comma. * `ref`: Reference or references. See {{tl|IPA}} for the syntax. * `id`: Sense ID, in case there are multiple senses. See {{tl|l}}. The following are the recognized additional inline modifiers: * `g`: Comma-separated list of genders. * `alt`: Display text. * `lang`: Language code of language of the form, if different from the language of the headword. * `sc`: Script code of script of the form. Almost never needed. * `t`: Gloss for the form. * `gloss`: Gloss for the form (alias for `t`). * `pos`: Part of speech of the form. * `lit`: Literal meaning of the form. * `tr`: Manual transliteration of the form. * `ts`: Transcription of the form, for languages where the transliteration differs markedly from the pronunciation. * `face`: Face to display the form in, e.g. {"hypothetical"} for a hypothetical form (unlinkable and displayed in italics). * `nolinkinfl`: Make the form unlinkable. ]==] function export.parse_term_with_modifiers(data) local paramname, val, frob = data.paramname, data.val, data.frob local function generate_obj(term, parse_err) if frob then term = frob(term, parse_err) end if data.parse_lang_prefix and term:find(":") then return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix { term = term, paramname = paramname, parse_lang_prefix = true, parse_err = parse_err, } else return {term = term} end end -- Check for inline modifier, e.g. מרים<tr:Miryem>. But exclude top-level HTML entry with <span ...>, -- <sup> or similar in it. if (val:find("<", nil, true) or data.splitchar) and not term_contains_top_level_html(val) and -- don't parse inline modifiers if is_head and the value begins with a ~ (link modifier syntax) (not data.is_head or not val:find("^~")) then local param_mods = param_mods if data.is_head then param_mods = shallow_copy(param_mods) param_mods.id = nil end if data.include_mods or data.exclude_mods then if not data.is_head then -- already copied when data.is_head param_mods = shallow_copy(param_mods) end if data.include_mods then local optional_mods = data.is_head and optional_headword_param_mods or optional_param_mods for _, mod in ipairs(data.include_mods) do if type(mod) == "table" then if #mod ~= 2 then error(("Internal error: Modifier spec %s in `include_mods` should be of length 2"):format( dump(mod))) end local modkey, modvalue = unpack(mod) param_mods[modkey] = modvalue elseif not optional_mods[mod] then error(("Internal error: Unrecognized modifier spec %s in `include_mods`"):format( dump(mod))) else param_mods[mod] = optional_mods[mod] end end end if data.exclude_mods then for _, mod in ipairs(data.exclude_mods) do if not param_mods[mod] then error(("Internal error: Modifier spec %s in `exclude_mods` not found among existing modifiers" ):format(dump(mod))) else param_mods[mod] = nil end end end end return parse_inline_modifiers(val, { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, splitchar = data.splitchar, preserve_splitchar = data.preserve_splitchar, delimiter_key = data.delimiter_key, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, pre_normalize_modifiers = data.pre_normalize_modifiers, }) else local retval = generate_obj(val) if data.splitchar then retval = {retval} end return retval end end --[==[ Parse a list of inflection forms that may have inline modifiers attached. `data` is an object with the following fields: * `forms`: The list of raw values to parse. Required. * `paramname`: The name of the first parameter from which the value was taken; used in error messages. If this is a two-element list, the first element is the first parameter and the second element is the prefix of the remaining parameters. Parameter names that are numbers are handled correctly, as are those with \1 in it marking where the parameter index goes. Required. * `qualifiers`: If specified, a possibly gappy list of left qualifiers to add to the parsed terms (for compatibility purposes). * `splitchar`: As in `parse_term_with_modifiers()`. The resulting per-term lists will be flattened. * `frob`, `include_mods`, `exclude_mods`, `is_head`, `preserve_splitchar`, `parse_lang_prefix`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_with_modifiers()`. Returns a list of objects, suitable for storing as one of the lists in `headdata.inflections` (once a label is added), where `headdata` is the structure passed to [[Module:headword]]. ]==] function export.parse_term_list_with_modifiers(data) local paramname, forms = data.paramname, data.forms local qualifiers = data.qualifiers local first, restpref if type(paramname) == "table" then first = paramname[1] restpref = paramname[2] else first = paramname restpref = paramname end local terms = {} data = shallow_copy(data) for i, val in ipairs(forms) do data.paramname = i == 1 and first or type(restpref) == "number" and restpref + i - 1 or restpref:find("\1", nil, true) and restpref:gsub("\1", tostring(i)) or restpref .. i data.val = val local parsed = export.parse_term_with_modifiers(data) if qualifiers and qualifiers[i] then if data.splitchar then for _, term in ipairs(parsed) do term.q = {qualifiers[i]} end else parsed.q = {qualifiers[i]} end end if data.splitchar then extend(terms, parsed) else terms[i] = parsed end end return terms end --[==[ Construct a link to [[Appendix:Glossary]] for `entry`. If `text` is specified, it is the display text; otherwise, `entry` is used. ]==] function export.glossary_link(entry, text) text = text or entry return "[[Lampiran:Glosari#" .. entry .. "|" .. text .. "]]" end function export.replace_glossary_links_in_label(label) if label:find("<<", nil, true) then label = label:gsub("<<(.-)|(.-)>>", export.glossary_link):gsub("<<(.-)>>", export.glossary_link) end return label end --[==[ Insert a fixed inflection (a label not associated with any inflection values) into an `inflections` field. The `inflections` field will be initialized if needed. `data` is an object with the following fields: * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs to be set for nested inflections, which are specified for an inflection object rather than the headword data structure as a whole. * `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required. * `originating_term`: The term object from which this label is derived. If specified, decorations will be taken from this object. ]==] function export.insert_fixed_inflection(data) local headdata, origterm, label = data.headdata, data.originating_term, data.label local inflobj = data.inflobj or headdata inflobj.inflections = inflobj.inflections or {} if not origterm then insert(inflobj.inflections, { label = export.replace_glossary_links_in_label(label) }) else if origterm.id then error(("It doesn't make sense to pass in an ID '%s' for label '%s' in conjunction with a term value '%s'" ):format(origterm.id, label, origterm.term)) end origterm = shallow_copy(origterm) -- Preserve decorations origterm.term = nil origterm.label = export.replace_glossary_links_in_label(label) insert(inflobj.inflections, origterm) end end --[==[ Insert previously-parsed terms into an `inflections` field. The `inflections` field will be initialized if needed. `data` is an object with the following fields: * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs to be set for nested inflections, which are specified for an inflection object rather than the headword data structure as a whole. * `terms`: The list of parsed terms. If {nil} or omitted, nothing happens unless `request` is set. * `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required. * `no_label`: If the term is {"-"} and there are no other terms, insert a fixed label with this value. Defaults to {"no "} plus the label. * `usually_no_label`: If the term is {"-"} and there are other terms, insert a fixed label with this value. Defaults to {"usually no "} plus the label. * `cats`: List of categories to insert when terms are given that are not {"-"}. Each category is a string naming a full category to insert (including the appropriate language name prefixed). * `no_cats`: List of categories to insert when a term is given as {"-"}. * `usually_no_cats`: List of categories to insert when a term is given as {"-"} and additional terms are specified as well (representing, e.g. for the inflection {"plural"}, a term which usually has no plural but does under some circumstances). If omitted, both the categories in `cats` and `no_cats` are inserted. * `accel`: If specified, a full accelerator object to add to the inflections. * `request`: If specified and no terms are given, insert a label with a request for inflections to be given. * `enable_auto_translit`: If specified and terms are given, display automatic transliteration of the terms. The return value indicates whether the inflection exists and how many terms are in it. It is an object with the following fields: * `exists`: {"yes"} if one or more terms were specified; {"no"} if the value was given as {"-"}; {"usually no"} if the first value was given as {"-"} but additional terms were supplied; otherwise {nil}, indicating that the status is unspecified. * `numterms`: Number of terms in the inflection. Will be 0 unless `exists` has the value {"yes"} or {"usually no"}. * `request`: True if no terms were specified but a term request was inserted into the inflection (because `data.request` was specified). Otherwise {nil}. ]==] function export.insert_inflection(data) local headdata, terms, label = data.headdata, data.terms, data.label local inflobj = data.inflobj or headdata local retval = {} local accel = data.accel if data.accel_form then if accel then error("Internal error: can't specify both data.accel and data.accel_form") end if headdata.heads then local lemmas = {} local lemma_translits = {} for i, headobj in ipairs(headdata.heads) do lemmas[i] = headobj.term if lemmas[i] == "+" then error("Internal error: If you use data.accel_form, you should have resolved all occurrences of + in heads appropriately") end lemma_translits[i] = headobj.tr end accel = { lemma = lemmas, lemma_translit = lemma_translits, form = data.accel_form, } else accel = { form = data.accel_form, } end end local function insert_cats(cats) for _, cat in ipairs(cats) do insert(headdata.categories, cat) end end if terms and terms[1] then terms = shallow_copy(terms) if terms[1].term == "-" then if terms[2] then export.insert_fixed_inflection { headdata = headdata, inflobj = inflobj, originating_term = terms[1], label = data.usually_no_label or "biasanya tiada " .. label, } remove(terms, 1) retval.numterms = #terms retval.exists = "usually no" if data.usually_no_cats then insert_cats(data.usually_no_cats) else if data.no_cats then insert_cats(data.no_cats) end if data.cats then insert_cats(data.cats) end end else export.insert_fixed_inflection { headdata = headdata, inflobj = inflobj, originating_term = terms[1], label = data.no_label or "tiada " .. label, } retval.numterms = 0 retval.exists = "no" if data.no_cats then insert_cats(data.no_cats) end return retval end else retval.numterms = #terms retval.exists = "yes" if data.cats then insert_cats(data.cats) end end if data.check_missing then error("Internal error: check_missing support removed; use checkredlinks=true in [[Module:headword]]") end terms.label = export.replace_glossary_links_in_label(label) if accel then terms.accel = accel end terms.enable_auto_translit = data.enable_auto_translit inflobj.inflections = inflobj.inflections or {} insert(inflobj.inflections, terms) elseif data.request then inflobj.inflections = inflobj.inflections or {} insert(inflobj.inflections, { label = export.replace_glossary_links_in_label(label), request = true, }) retval.numterms = 0 -- retval.exists = nil retval.request = true else retval.numterms = 0 -- retval.exists = nil end return retval end --[==[ Parse raw arguments from `forms` for inline modifiers, and insert the resulting terms (which should not require significant additional processing) into `headdata.inflections`. `data` is an object with the following fields: * `forms`: The list of raw values to parse. If {nil} or omitted, nothing happens. * `headdata`: The headword structure passed to [[Module:headword]]. Required. * `paramname`: As in `parse_term_list_with_modifiers()`. Required. * `label`: As in `insert_inflection()`. Required. * `qualifiers`, `frob`, `include_mods`, `exclude_mods`, `is_head`, `splitchar`, `preserve_splitchar`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_list_with_modifiers()`. * `accel`: As in `insert_inflection()`. Return value is as in `insert_inflection()`. ]==] function export.parse_and_insert_inflection(data) local forms = data.forms if forms and forms[1] then data = shallow_copy(data) data.forms = forms data.terms = export.parse_term_list_with_modifiers(data) return export.insert_inflection(data) end return { numterms = 0 } end --[==[ Canonicalize a single term or term-like object or a list of either into a list of term-like objects. `abterms` is the term or list to canonicalize, and `field` is the name of the field holding the term (defaulting to {"term"}). This does the minimal work necessary, meaning that the return value may partly or completely share memory with the value passed in. As a special case, if `abterms` is {nil}, {nil} is returned. If `origin_val` is specified, add a field `origin` containing the value of `origin_val` to each resulting term-like object (in this case, the object will be copied a necessary to avoid side-effecting the passed-in objects). ]==] function export.canonicalize_termobj_list(abterms, field, origin_val) if abterms == nil then return nil end field = field or "term" if type(abterms) == "string" then return {{[field] = abterms, origin = origin_val}} elseif not abterms[1] then if origin_val ~= nil then abterms = shallow_copy(abterms) abterms.origin = origin_val end return {abterms} else -- Check if already in full list term and return directly if so (unless `origin_val` is given, in which case we -- need to shallow-copy both the list and each term in it). local must_convert = false for _, term in ipairs(abterms) do if type(term) == "string" then must_convert = true break end end if not must_convert then if origin_val ~= nil then abterms = shallow_copy(abterms) for i, abterm in ipairs(abterms) do abterms[i] = shallow_copy(abterm) abterms[i].origin = origin_val end end return abterms end end local retval = {} for _, term in ipairs(abterms) do if type(term) == "string" then insert(retval, {[field] = term, origin = origin_val}) else if origin_val ~= nil then term = shallow_copy(term) term.origin = origin_val end insert(retval, term) end end return retval end --[==[ Combine two sets of decorations. If either is {nil}, just return the other, and if both are {nil}, return {nil}. ]==] function export.combine_decorations(decs1, decs2) if not decs1 and not decs2 then return nil end if not decs1 then return decs2 end if not decs2 then return decs1 end local combined = shallow_copy(decs1) for _, dec in ipairs(decs2) do insert_if_not(combined, dec) end return combined end function export.combine_qualifiers_or_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use combine_decorations instead") end --[==[ Combine the decorations (qualifiers, labels, references) and ID's of two term objects. `destobj` is the "destination term object" into which the combined properties are written, and `srcobj` is the "source object" into which the properties are merged. `destobj` is side-effected (but the lists inside of `destobj` are not); if this is undesirable, make sure to shallow-copy `destobj` first. If both objects have values for a given decoration, the values of `destobj` come first. If both objects have a value for `id`, the values must match or an error is thrown; otherwise, the resulting value of `id` comes from whichever one is defined. '''NOTE:''' This may not be the correct behavior when deduplicating a list of term objects. See `insert_termobj_combining_duplicates` for a different approach. ]==] function export.combine_termobj_decorations(destobj, srcobj) destobj.q = export.combine_decorations(destobj.q, srcobj.q) destobj.qq = export.combine_decorations(destobj.qq, srcobj.qq) destobj.l = export.combine_decorations(destobj.l, srcobj.l) destobj.ll = export.combine_decorations(destobj.ll, srcobj.ll) destobj.refs = export.combine_decorations(destobj.refs, srcobj.refs) if destobj.id and srcobj.id and destobj.id ~= srcobj.id then -- FIXME: We probably want to pass in an error function error(("Can't specify two different ID's %s and %s when combining objects"):format(srcobj.id, destobj.id)) end destobj.id = destobj.id or srcobj.id return destobj end function export.combine_termobj_qualifiers_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use combine_termobj_decorations instead") end function export.termobj_has_decorations(obj) return obj.q and obj.q[1] or obj.qq and obj.qq[1] or obj.l and obj.l[1] or obj.ll and obj.ll[1] or obj.refs and obj.refs[1] end function export.termobj_has_qualifiers_or_labels(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use termobj_has_decorations instead") end local function one_decoration_equal(prop1, prop2) local prop1_is_nil = not prop1 or not prop1[1] local prop2_is_nil = not prop2 or not prop2[1] if prop1_is_nil and prop2_is_nil then return true end if prop1_is_nil or prop2_is_nil then return false end return deep_equals(prop1, prop2) end function export.termobj_decorations_equal(obj1, obj2) return one_decoration_equal(obj1.q, obj2.q) and one_decoration_equal(obj1.qq, obj2.qq) and one_decoration_equal(obj1.l, obj2.l) and one_decoration_equal(obj1.ll, obj2.ll) and one_decoration_equal(obj1.refs, obj2.refs) and obj1.id == obj2.id end function export.termobj_ancillary_properties_equal(...) -- FIXME: Added 2026-09-17. Remove after a month. error("Use termobj_decorations_equal instead") end function export.convert_termobj_to_formobj(termobj) local formobj = { form = termobj.term, translit = termobj.tr, } local footnotes local function mods_to_footnote(mod_prefix, mod_vals) if mod_vals and mod_vals[1] then footnotes = footnotes or {} for _, val in ipairs(mod_vals) do insert(footnotes, "[" .. mod_prefix .. ":" .. val .. "]") end end end mods_to_footnote("q", termobj.q) mods_to_footnote("qq", termobj.qq) mods_to_footnote("l", termobj.l) mods_to_footnote("ll", termobj.ll) mods_to_footnote("ref", termobj.refs) mods_to_footnote("id", termobj.id and {termobj.id} or nil) formobj.footnotes = footnotes return formobj end local recognized_multi_mods = { q = "q", qq = "qq", l = "l", ll = "ll", ref = "refs", } local recognized_single_mods = { id = "id", } function export.add_footnote_to_termobj(termobj, footnote) local stripped_footnote = footnote:match("^%[(.*)%]$") if not stripped_footnote then error("Internal error: Footnote should be surrounded by brackets at this stage: " .. footnote) end local prefix, rest = stripped_footnote:match("^([a-z]+):(.+)$") local field, is_multi if prefix then if recognized_multi_mods[prefix] then field = recognized_multi_mods[prefix] is_multi = true elseif recognized_single_mods[prefix] then field = recognized_single_mods[prefix] is_multi = false end end if not field then rest = stripped_footnote field = "l" is_multi = true end if is_multi then if not termobj[field] then termobj[field] = {} end insert(termobj[field], rest) else if termobj[field] and termobj[field] ~= rest then error(("Can't set two values for '%s': '%s' and '%s'"):format(field, termobj[field], rest)) end termobj[field] = rest end end function export.convert_formobj_to_termobj(formobj) local termobj = { term = formobj.form, tr = formobj.translit, } if formobj.footnotes then for _, footnote in ipairs(formobj.footnotes) do export.add_footnote_to_termobj(termobj, footnote) end end return termobj end local function extract_termobj_field_modifiers(fieldval) return fieldval:match("^([*+]?)(.*)$") end function export.remove_termobj_field_modifiers(termobj) local function remove_field_modifiers(field) if termobj[field] and termobj[field][1] then local any_field_modifiers = false for _, val in ipairs(termobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods ~= "" then any_field_modifiers = true break end end local new_field = {} if any_field_modifiers then for _, val in ipairs(termobj[field]) do local _, field_without_mods = extract_termobj_field_modifiers(val) insert_if_not(new_field, field_without_mods) end termobj[field] = new_field end end end remove_field_modifiers("q") remove_field_modifiers("qq") remove_field_modifiers("l") remove_field_modifiers("ll") remove_field_modifiers("refs") end function export.insert_termobj_combining_duplicates(destobjs, termobj) for _, destobj in ipairs(destobjs) do if destobj.term == termobj.term and destobj.tr == termobj.tr then -- Form already present; maybe combine footnotes. local function combine_field_values(field) if termobj[field] and termobj[field][1] then -- Check to see if there are existing values with *; if so, remove them. if destobj[field] and destobj[field][1] then local any_values_with_asterisk = false for _, val in ipairs(destobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods:find("%*") then any_values_with_asterisk = true break end end if any_values_with_asterisk then local filtered_values = {} for _, val in ipairs(destobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if not field_mods:find("%*") then insert(filtered_values, val) end end if filtered_values[1] then destobj[field] = filtered_values else destobj[field] = nil end end end local any_footnotes_with_plus = false for _, val in ipairs(termobj[field]) do local field_mods, _ = extract_termobj_field_modifiers(val) if field_mods:find("%+") then any_footnotes_with_plus = true break end end if any_footnotes_with_plus then if not destobj[field] then destobj[field] = {} else destobj[field] = shallow_copy(destobj[field]) end for _, val in ipairs(termobj[field]) do local already_seen = false local field_mods, field_without_mods = extract_termobj_field_modifiers(val) if field_mods:find("%+") then for _, existing_val in ipairs(destobj[field]) do local _, existing_field_without_mods = extract_termobj_field_modifiers(existing_val) if existing_field_without_mods == field_without_mods then already_seen = true break end end if not already_seen then insert(destobj[field], val) end end end end end end combine_field_values("q") combine_field_values("qq") combine_field_values("l") combine_field_values("ll") combine_field_values("refs") if destobj.id and termobj.id and destobj.id ~= termobj.id then -- FIXME: We probably want to pass in an error function error(("Can't specify two different ID's %s and %s when combining objects"):format(termobj.id, destobj.id)) end destobj.id = destobj.id or termobj.id return end end insert(destobjs, termobj) end export.allowed_special_indicators = { ["first"] = true, ["first-second"] = true, ["first-last"] = true, ["second"] = true, ["last"] = true, ["each"] = true, ["+"] = true, -- requests the default behavior with preposition handling } --[==[ Check for special indicators (values such as {"+first"} or {"+first-last"} that are used in a `pl`, `f`, etc. argument and indicate how to inflect a multiword term). If `form` is such an indicator, the return value is `form` minus the initial `+` sign; otherwise, if form begins with a `+` sign, an error is thrown; otherwise the return value is nil. ]==] function export.get_special_indicator(form, noerror) if form:find("^%+") then form = form:gsub("^%+", "") if not export.allowed_special_indicators[form] then if noerror then return nil end local indicators = {} for indic, _ in pairs(export.allowed_special_indicators) do insert(indicators, "+" .. indic) end sort(indicators) error("Special inflection indicator beginning with '+' can only be " .. mw.text.listToText(indicators) .. ": +" .. form) end return form end return nil end local function add_endings(bases, endings) local retval = {} if type(bases) ~= "table" then bases = {bases} end if type(endings) ~= "table" then endings = {endings} end for _, base in ipairs(bases) do for _, ending in ipairs(endings) do insert(retval, base .. ending) end end return retval end --[==[ Inflect a possibly multiword or hyphenated term `form` using the function `inflect`, which is a function of one argument that is called on a single word to inflect and should return either the inflected word or a list of inflected words. `special` indicates how to inflect the multiword term and should be e.g. {"first"} to inflect only the first word, {"first-last"} to inflect the first and last words, {"each"} to inflect each word, etc. See `allowed_special_indicators` above for the possibilities. If `special` is `+`, or is omitted and the term is multiword (i.e. containing a space character), and `prepositions` is supplied, the function checks for multiword or hyphenated terms containing the prepositions in `prepositions`, e.g. Italian [[senso di marcia]] or [[medaglia d'oro]] or Portuguese [[tartaruga-do-mar]]. If such a term is found, only the first word is inflected. Otherwise, the default is {"first-last"}. `prepositions` is a list of Lua patterns matching prepositions. The patterns will automatically have the separator character (space or hyphen) added to the left side but not the right side, so they should contain a space character (which will automatically be converted to the appropriate separator) on the right side unless the preposition is joined on the right side with an apostrophe. Examples of preposition patterns for Italian are {"di "}, {"sull'"} and {"d?all[oae] "} (which matches {"dallo "}, {"dalle "}, {"alla "}, etc.). The return value is always either a list of inflected multiword or hyphenated terms, or nil if `special` is omitted and `form` is not multiword. (If `special` is specified and `form` is not multiword or hyphenated, an error results.) ]==] function export.handle_multiword(form, special, inflect, prepositions, sep) sep = sep or form:find(" ") and " " or "%-" local raw_sep = sep == " " and " " or "-" -- Used to add regex version of separator in the replacement portion of ugsub() or :gsub() local sep_replacement = sep == " " and " " or "%%-" -- Given a Lua pattern, replace space with the appropriate separator. local function hack_re(re) if sep == " " then return re end return (re:gsub(" ", sep_replacement)) end if special == "first" then local first, rest = form:match(hack_re("^(.-)( .*)$")) if not first then error("Special indicator 'first' can only be used with a multiword term: " .. form) end return add_endings(inflect(first), rest) elseif special == "second" then local first, second, rest = form:match(hack_re("^([^ ]+ )([^ ]+)( .*)$")) if not first then error("Special indicator 'second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings({first}, inflect(second)), rest) elseif special == "first-second" then local first, space, second, rest = form:match(hack_re("^([^ ]+)( )([^ ]+)( .*)$")) if not first then error("Special indicator 'first-second' can only be used with a term with three or more words: " .. form) end return add_endings(add_endings(add_endings(inflect(first), space), inflect(second)), rest) elseif special == "each" then local terms = split(form, sep) if #terms < 2 then error("Special indicator 'each' can only be used with a multiword term: " .. form) end for i, term in ipairs(terms) do terms[i] = inflect(term) if i > 1 then terms[i] = add_endings(raw_sep, terms[i]) end end local result = "" for _, term in ipairs(terms) do result = add_endings(result, term) end return result elseif special == "first-last" then local first, middle, last = form:match(hack_re("^(.-)( .* )(.-)$")) if not first then first, middle, last = form:match(hack_re("^(.-)( )(.*)$")) end if not first then error("Special indicator 'first-last' can only be used with a multiword term: " .. form) end return add_endings(add_endings(inflect(first), middle), inflect(last)) elseif special == "last" then local rest, last = form:match(hack_re("^(.* )(.-)$")) if not rest then error("Special indicator 'last' can only be used with a multiword term: " .. form) end return add_endings(rest, inflect(last)) elseif special and special ~= "+" then error("Unrecognized special=" .. special) end -- Only do default behavior if special indicator '+' explicitly given or separator is space; otherwise we will -- break existing behavior with hyphenated words. if (special == "+" or sep == " ") and form:find(sep) then if prepositions then -- check for prepositions in the middle of the word; do it this way so we can handle -- more than one word before the preposition (and usually inflect each word) for _, prep in ipairs(prepositions) do local first, space_prep_rest = umatch(form, hack_re("^(.-)( " .. prep .. ".*)$")) if first then return add_endings(inflect(first), space_prep_rest) end end end -- multiword or hyphenated expressions default to first-last; we need to pass in the separator to avoid -- problems with multiword terms containing hyphens in the individual words return export.handle_multiword(form, "first-last", inflect, prepositions, sep) end return nil end local function link_hyphen_split_component(word, data) if data.link_hyphen_split_component then return data.link_hyphen_split_component(word) else return "[[" .. word .. "]]" end end -- Default function to split a word on apostrophes. Don't split apostrophes at the beginning or end of a word (e.g. -- [['ndrangheta]] or [[po']]). Handle multiple apostrophes correctly, e.g. [[l'altr'ieri]] -> [[l']][altr']][[ieri]]. function export.default_split_apostrophe(word, data) local apostrophe_parts = split(word, "'", true, true) local linked_apostrophe_parts = {} local apostrophes_at_beginning = "" local i = 1 -- Apostrophes at beginning get attached to the first word after (which will always exist but may -- be blank if the word consists only of apostrophes). while i < #apostrophe_parts do -- <, not <=, in case the word consists only of apostrophes local apostrophe_part = apostrophe_parts[i] i = i + 1 if apostrophe_part == "" then apostrophes_at_beginning = apostrophes_at_beginning .. "'" else break end end apostrophe_parts[i] = apostrophes_at_beginning .. apostrophe_parts[i] -- Now, do the remaining parts. A blank part indicates more than one apostrophe in a row; we join -- all of them to the preceding word. while i <= #apostrophe_parts do local apostrophe_part = apostrophe_parts[i] if apostrophe_part == "" then linked_apostrophe_parts[#linked_apostrophe_parts] = linked_apostrophe_parts[#linked_apostrophe_parts] .. "'" elseif i == #apostrophe_parts then insert(linked_apostrophe_parts, apostrophe_part) else insert(linked_apostrophe_parts, apostrophe_part .. "'") end i = i + 1 end for j, tolink in ipairs(linked_apostrophe_parts) do linked_apostrophe_parts[j] = link_hyphen_split_component(tolink, data) end return concat(linked_apostrophe_parts) end --[=[ Auto-add links to a word that should not have spaces but may have hyphens and/or apostrophes. We split off final punctuation, then split on hyphens if `data.split_hyphen` is given, and also split on apostrophes if `data.split_apostrophe` is given. We only split on hyphens if they are in the middle of the word, not at the beginning or end (hyphens at the beginning or end indicate suffixes or prefixes, respectively). `include_hyphen_prefixes`, if given, is a set of prefixes (not including the final hyphen) where we should include the final hyphen in the prefix. Hence, e.g. if "anti" is in the set, a Portuguese word like [[anti-herói]] "anti-hero" will be split [[anti-]][[herói]] (whereas a word like [[código-fonte]] "source code" will be split as [[código]]-[[fonte]]). If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French [[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word to split, and the return value should be the split and linked word. ]=] local function add_single_word_links(space_word, data, term_has_spaces) local space_word_no_punct, punct local punct_pattern = data.punctuation if punct_pattern and is_callable(punct_pattern) then space_word_no_punct, punct = punct_pattern(space_word) else if punct_pattern == nil then punct_pattern = "[,;:?!]" end space_word_no_punct, punct = umatch(space_word, "^(.*)(" .. punct_pattern .. ")$") end space_word_no_punct = space_word_no_punct or space_word punct = punct or "" local words if space_word_no_punct:sub(1, 1) == "-" or space_word_no_punct:sub(-1) == "-" then -- don't split prefixes and suffixes words = {space_word_no_punct} else local splitter if term_has_spaces then splitter = data.split_hyphen_when_space else splitter = data.split_hyphen_when_no_space end if is_callable(splitter) then words = splitter(space_word_no_punct) if type(words) == "string" then return words .. punct end end end if not words then local split_hyphen if term_has_spaces then split_hyphen = data.split_hyphen_when_space else split_hyphen = data.split_hyphen_when_no_space if split_hyphen == nil then -- default to true; use `false` to avoid this split_hyphen = true end end if split_hyphen then words = split(space_word_no_punct, "-", true, true) else words = {space_word_no_punct} end end local linked_words = {} for j, word in ipairs(words) do if j < #words and data.include_hyphen_prefixes and data.include_hyphen_prefixes[word] then word = "[[" .. word .. "-]]" elseif j > 1 and data.include_hyphen_suffixes and data.include_hyphen_suffixes[word] then word = "[[-" .. word .. "]]" else -- Don't split on apostrophes if the word is in `no_split_apostrophe_words`. if (not data.no_split_apostrophe_words or not data.no_split_apostrophe_words[word]) and data.split_apostrophe and word:find("'", nil, true) then if data.split_apostrophe == true then word = export.default_split_apostrophe(word, data) else -- custom apostrophe splitter/linker word = data.split_apostrophe(word) end elseif word ~= "" then -- avoid -[[]]- (e.g. f--k) word = link_hyphen_split_component(word, data) end if j < #words then word = word .. "-" end end insert(linked_words, word) end return concat(linked_words) .. punct end --[=[ Auto-add links to a multiword term. `data` contains fields customizing how to do this. By default we proceed as follows: (1) If the term already has embedded links in it, they are left unchanged. (2) Otherwise, if there are spaces present, we split on spaces and link each word separately. (3) If a given space-separated component ends in punctuation (defaulting to [,;:?!]), it is separated off, the remainder of the algorithm run, and the punctuation pasted back on. (4) If there are hyphens in a given space-separated component, we may link each hyphenated term separately depending on the settings in `data`. Normally the hyphens are not included in the linked terms, but this can be overridden for specific prefixes and/or suffixes. By default, if there are spaces in the multiword term, we do not link hyphenated components (because of cases like "boire du petit-lait" where "petit-lait" should be linked as a whole), but do so otherwise (e.g. for "avant-avant-hier"); this can overridden for cases like "croyez-le ou non". Cases where only some of the hyphens should be split can always be handled by explicitly specifying the head (e.g. "Nord-Pas-de-Calais" given as head=[[Nord]]-[[Pas-de-Calais]]). (5) If there are apostrophes in a given component, we may link each apostrophe-separated term separately depending on the settings in `data`, including the apostrophe in the link to its left (so we split "de l'eau" as "[[de]] [[l']][[eau]]"). The settings in `data` are as follows: `split_hyphen_when_no_space`: Whether to split on hyphens when the term has no spaces. Defaults to true if set to `nil`. This can be a function of one argument, to implement a custom splitting algorithm for hyphen-separated terms. If this returns [FIXME: FINISH ME ...] If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French [[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word to split, and the return value should be the split and linked word. We don't always split on hyphens because of cases like "boire du petit-lait" where "petit-lait" should be linked as a whole, but provide the option to do it for cases like "croyez-le ou non". If there's no space, however, then it makes sense to split on hyphens by `no_split_apostrophe_words` and `include_hyphen_prefixes` allow for special-case handling of particular words and are as described in the comment above add_single_word_links(). ]=] function export.add_links_to_multiword_term(term, data) if term:match("[%[%]]") then return term end local words = split(term, " ", true, true) local term_has_spaces = #words > 1 local linked_words = {} for _, word in ipairs(words) do insert(linked_words, add_single_word_links(word, data, term_has_spaces)) end local retval = concat(linked_words, " ") -- If we ended up with a single link consisting of the entire term, -- remove the link. return retval:match("^%[%[([^%[%]]*)%]%]$") or retval end local function canonicalize_begin_end_spec(spec) local from, to = spec:match("^(.-):(.*)$") if not from then from = spec to = "" end return from, to end --[==[ Given a `linked_term` that is the output of add_links_to_multiword_term(), apply modifications as given in `modifier_spec` to change the link destination of subterms (normally single-word non-lemma forms; sometimes collections of adjacent words). This is usually used to link non-lemma forms to their corresponding lemma, but can also be used to replace a span of adjacent separately-linked words to a single multiword lemma. The format of `modifier_spec` is one or more semicolon-separated subterm specs, where each such spec is of the form SUBTERM:DEST, where SUBTERM is one or more words in the `linked_term` but without brackets in them, and DEST is the corresponding link destination to link the subterm to. Any occurrence of ~ in DEST is replaced with SUBTERM. Alternatively, a single modifier spec can be of the form BEGIN[FROM:TO], which is equivalent to writing BEGINFROM:BEGINTO (see example below). For example, given the source phrase [[il bue che dice cornuto all'asino]] "the pot calling the kettle black" (literally "the ox that calls the donkey horned/cuckolded"), the result of calling add_links_to_multiword_term() is [[il]] [[bue]] [[che]] [[dice]] [[cornuto]] [[all']][[asino]]. With a modifier_spec of 'dice:dire', the result is [[il]] [[bue]] [[che]] [[dire|dice]] [[cornuto]] [[all']][[asino]]. Here, based on the modifier spec, the non-lemma form [[dice]] is replaced with the two-part link [[dire|dice]]. Another example: given the source phrase [[chi semina vento raccoglie tempesta]] "sow the wind, reap the whirlwind" (literally (he) who sows wind gathers [the] tempest"). The result of calling add_links_to_multiword_term() is [[chi]] [[semina]] [[vento]] [[raccoglie]] [[tempesta]], and with a modifier_spec of 'semina:~re; raccoglie:~re', the result is [[chi]] [[seminare|semina]] [[vento]] [[raccogliere|raccoglie]] [[tempesta]]. Here we use the ~ notation to stand for the non-lemma form in the destination link. A more complex example is [[se non hai altri moccoli puoi andare a letto al buio]], which becomes [[se]] [[non]] [[hai]] [[altri]] [[moccoli]] [[puoi]] [[andare]] [[a]] [[letto]] [[al]] [[buio]] after calling add_links_to_multiword_term(). With the following modifier_spec: 'hai:avere; altr[i:o]; moccol[i:o]; puoi: potere; andare a letto:~; al buio:~', the result of applying the spec is [[se]] [[non]] [[avere|hai]] [[altro|altri]] [[moccolo|moccoli]] [[potere|puoi]] [[andare a letto]] [[al buio]]. Here, we rely on the alternative notation mentioned above for e.g. 'altr[i:o]', which is equivalent to 'altri:altro', and link multiword subterms using e.g. 'andare a letto:~'. (The code knows how to handle multiword subexpressions properly, and if the link text and destination are the same, only a single-part link is formed.) ]==] function export.apply_link_modifiers(linked_term, modifier_spec, lang) local split_modspecs = split(modifier_spec, "%s*;%s*") for j, modspec in ipairs(split_modspecs) do local id if modspec:find("<") then local rest rest, id = modspec:match("^(.*)<id:(.-)>$") if rest then modspec = rest end end local subterm, dest, otherlang local begin_spec, rest, end_spec = modspec:match("^%[(.-)%]([^:]*)%[(.-)%]$") if begin_spec then local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec) local end_from, end_to = canonicalize_begin_end_spec(end_spec) subterm = begin_from .. rest .. end_from dest = begin_to .. rest .. end_to end if not subterm then rest, end_spec = modspec:match("^([^:]*)%[(.-)%]$") if rest then local end_from, end_to = canonicalize_begin_end_spec(end_spec) subterm = rest .. end_from dest = rest .. end_to end end if not subterm then begin_spec, rest = modspec:match("^%[(.-)%]([^:]*)$") if begin_spec then local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec) subterm = begin_from .. rest dest = begin_to .. rest end end if not subterm then subterm, dest = modspec:match("^(.-)%s*:%s*(.*)$") if subterm and subterm ~= "^" and subterm ~= "$" then local langdest -- Parse off an initial language code (e.g. 'en:Higgs', 'la:minūtia' or 'grc:σκατός'). Also handle -- Wikipedia prefixes ('w:Abatemarco' or 'w:it:Colle Val d'Elsa'). otherlang, langdest = dest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if otherlang == "w" then local foreign_wikipedia, foreign_term = langdest:match("^([A-Za-z0-9._-]+):([^ ].*)$") if foreign_wikipedia then otherlang = otherlang .. ":" .. foreign_wikipedia langdest = foreign_term end dest = ("%s:%s"):format(otherlang, langdest) otherlang = nil elseif otherlang then otherlang = get_lang_by_code(otherlang, true, "allow etym") dest = langdest end end end if not subterm then if modspec == "?" or modspec == "!" then subterm = "$" dest = modspec elseif modspec == "..." or modspec == "...?" then subterm = "$" dest = " " .. modspec elseif modspec:find("^[A-Z]$") then -- X, Y, etc. by themselves are unlinked, to help with snowclones subterm = modspec dest = "_" else subterm = modspec dest = "~" end end if subterm == "^" then linked_term = dest:gsub("_", " ") .. linked_term elseif subterm == "$" then linked_term = linked_term .. dest:gsub("_", " ") else if subterm:find("[", nil, true) then error(("Subterm '%s' in modifier spec '%s' cannot have brackets in it"):format( escape_wikicode(subterm), escape_wikicode(modspec))) end local escaped_subterm = pattern_escape(subterm) local subterm_re = "%[%[" .. escaped_subterm:gsub("(%%?[ ',%-])", "%%]*%1%%[*") .. "%]%]" local expanded_dest if dest:find("~", nil, true) then expanded_dest = dest:gsub("~", replacement_escape(subterm)) else expanded_dest = dest end if otherlang then expanded_dest = expanded_dest .. "#" .. otherlang:getCanonicalName() end local subterm_replacement if expanded_dest == "_" then subterm_replacement = subterm if id then error("Can't supply <id:...> with an unlinked subterm") end if otherlang then error("Can't supply prefixed language with an unlinked subterm") end elseif id or otherlang then if id and expanded_dest:find("[", nil, true) then error("Can't supply <id:...> with destination with embedded brackets") end subterm_replacement = require(links_module).language_link { lang = otherlang or lang, term = expanded_dest, alt = subterm, id = id, } elseif expanded_dest:find("[", nil, true) then -- Use the destination directly if it has brackets in it (e.g. to put brackets around parts of a word). subterm_replacement = expanded_dest elseif expanded_dest == subterm then subterm_replacement = "[[" .. subterm .. "]]" else subterm_replacement = "[[" .. expanded_dest .. "|" .. subterm .. "]]" end local escaped_subterm_replacement = replacement_escape(subterm_replacement) local replaced_linked_term = ugsub(linked_term, subterm_re, escaped_subterm_replacement) if replaced_linked_term == linked_term then mw.log(("Attempted to replace %s with %s in %s"):format(subterm_re, escaped_subterm_replacement, linked_term)) error(("Subterm '%s' could not be located in %slinked expression %s, or replacement same as subterm"):format( subterm, j > 1 and "intermediate " or "", escape_wikicode(linked_term))) else linked_term = replaced_linked_term end end end return linked_term end local inflection_to_cats = { plural = { filter_plpos = function(plpos) -- plurals also occur with determiners, adjectives etc. and we don't want to generate categories like -- 'countable determiners', 'countable adjectives', etc. Note that the passed-in `plpos` has `proper nouns` -- converted to `nouns`. return plpos == "Kata nama" end, cats = {"{plpos} terhitung"}, no_cats = {"{plpos} tak terhitung"}, }, comparative = { cats = {"{plpos} bandingan"}, no_cats = {"{plpos} bukan bandingan"}, }, ["female equivalent"] = { cats = {"{plpos} dengan padanan genus lain"}, }, ["male equivalent"] = { cats = {"{plpos} dengan padanan genus lain"}, }, } --[=[ Validate the items in `items` against the list or set of valid items in `valid_items`. If `field` is given, fetch the item to check from that-named field of each object in `items`; otherwise use the items in `items` directly. If an error occurs, `item_type` specifies the type of item to mention in the error message, which will also list the allowed items (either taken directly from `valid_items` if a list, or from the sorted keys if a set). ]=] local function validate_items(data) local items, field, valid_items, item_type = data.items, data.field, data.valid_items, data.item_type local valid_set if valid_items[1] then valid_set = list_to_set(valid_items) else valid_set = valid_items end for _, item in ipairs(items) do if field then item = item[field] end if not valid_set[item] then local valid_list if valid_items[1] then valid_list = valid_items else valid_list = {} for valid_item, _ in pairs(valid_items) do insert(valid_list, valid_item) end table.sort(valid_list) end error(("Invalid %s: %s; expected one of %s"):format(item_type, item, mw.text.listToText(valid_list))) end end end local Headdata = {} function Headdata:get_canonicalized_plpos() return (self.pos_category:gsub("Kata nama khas", "Kata nama")) end --[==[ Canonicalize a category. The category string will have the full language name (i.e. the name of the L2 language under which an entry is inserted, which may a parent language if the language in question is an etymology-only language) prepended to it, and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech (with some canonicalization; specifically, `proper nouns` is converted to `nouns` when replacing `{plpos}`). To specify a full category and not have the language name prepended to it, precede it with {"Category:"}, which will be removed. ]==] function Headdata:canonicalize_category(category) if category:find("{plpos}") then local plpos = self:get_canonicalized_plpos() category = category:gsub("{plpos}", plpos) end if category:find("^Kategori:") then return (category:gsub("^Kategori:", "")) else return category .. " bahasa " .. self.langfullname end end --[==[ Canonicalize a list of categories according to the process described in `Headdata:canonicalize_category`. This simply loops over each category in `categories` and calls `Headdata:canonicalize_category` on each one. ]==] function Headdata:canonicalize_categories(categories) if not categories then return categories end local canon_cats = {} for _, cat in ipairs(categories) do insert(canon_cats, self:canonicalize_category(cat)) end return canon_cats end --[==[ Insert a category into the `categories` list in the headword `data` structure. `category` is normally a string naming the category, which will have the full language name prepended to it and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech (with some canonicalization; specifically, `proper nouns` is converted to `nouns` when replacing `{plpos}`). To specify a full category and not have the language name prepended to it, precede it with {"Category:"}. ]==] function Headdata:insert_category(category) insert(self.categories, self:canonicalize_category(category)) end --[==[ Validate the genders in `genders` (a list of gender spec objects, as produced by {type = "genders"} in [[Module:parameters]] and accepted by [[Module:gender and number]]), checking that all specified genders are in the list given in `valid_genders`. Optional `props` controls how the validation happens. In particular, unless `props.no_augment` is given, then for any gender beginning with `m`, if a corresponding gender beginning with `f` occurs, analogous genders beginning with `mf`, `mfbysense` and `mfequiv` are also allowed. For example, if `m-d` (masculine dual) and `f-d` (feminine dual) both occur, genders `mf-d`, `mfbysense-d` and `mfequiv-d` are also allowed. If a disallowed gender is given, an error occurs, giving the disallowed gender along with the list of all allowed genders. ]==] function Headdata:validate_genders(genders, valid_genders, props) if not genders then return end props = props or {} local gender_type, no_augment = props.gender_type, props.no_augment gender_type = gender_type or "headword" local valid_gender_set = list_to_set(valid_genders) local augmented_gender_set if no_augment then augmented_gender_set = valid_gender_set else augmented_gender_set = {} for g, _ in pairs(valid_gender_set) do augmented_gender_set[g] = true if g:find("^m") and not g:find("^mf") and valid_gender_set[g:gsub("^m", "f")] then augmented_gender_set[g:gsub("^m", "mf")] = true augmented_gender_set[g:gsub("^m", "mfbysense")] = true augmented_gender_set[g:gsub("^m", "mfequiv")] = true end end end validate_items { items = genders, field = "spec", valid_items = augmented_gender_set, item_type = ("%s gender"):format(gender_type), } end --[==[ Parse an inflection specified in `field`, the name of a parameter holding an inflection. If the parameter is numeric, the field should be given as a number (as with the `params` structure passed to [[Module:parameters]]), not a string containing the representation of a number. The field can specify multiple comma-separated terms, and each term can have associated inline modifiers that will be parsed (unless there is top-level HTML in the parameter, i.e. HTML not contained inside an inline modifier, e.g. as may be generated by using {{tl|l}} or similar template inside a parameter). This is a wrapper around the top-level `parse_term_with_modifiers()` function. `props` is an optional structure containing additional properties, including all additional properties documented for the top-level `parse_term_with_modifiers()` function. If the parameter in `field` is unspecified, the return value of this function will be an empty list, not {nil}, so it is always safe to iterate over the return value. By default, the allowed modifiers are the same as for `parse_term_with_modifiers()`, except that (normally) the `<tr:...>` modifier will be allowed if `include_tr` was specified in the original call to `process_headword()`; likewise for the `<ts:...>` modifier if `include_ts` was specified and the `<sc:...>` modifier if `include_sc` was specified. If If you pass in your own `include_mods` list of additional allowed modifiers, it will (normally) automatically be augmented with {"tr"}, {"ts"} and/or {"sc"} if `include_tr`, `include_ts` and/or `include_sc` was specified when calling `process_headword()`. To disable automatic augmentation of these modifiers (whether or not you specify an `include_mods` property), specify {no_augment_include_mods = true} in `props`. ]==] function Headdata:parse_inflection(field, props) local val = self.process_props.args[field] if not val then return {} end props = props and shallow_copy(props) or {} local include_mods = props.include_mods local data = self.process_props.data if not props.no_augment_include_mods and (data.include_tr or data.include_ts or data.include_sc) then include_mods = include_mods and shallow_copy(include_mods) or {} if data.include_tr then insert_if_not(include_mods, "tr") end if data.include_ts then insert_if_not(include_mods, "ts") end if data.include_sc then insert_if_not(include_mods, "sc") end end props.val = val props.paramname = field props.splitchar = props.splitchar or "," props.include_mods = include_mods return export.parse_term_with_modifiers(props) or {} end --[==[ Insert previously-parsed terms into the `inflections` of the headword `data` structure. This is a wrapper around the top-level `insert_inflection()` function. `terms` is the list of parsed terms. (If {nil}, nothing happens unless `request` is set in `props`.) `label` is the the label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) `props` is an optional structure containing additional properties, including all additional properties documented for the top-level `insert_inflection()` function. Unless `no_auto_cats` is given in `props`, certain labels automatically trigger the insertion of additional categories in specific circumstances. This is controlled by the `inflection_to_cats` structure in [[Module:headword utilities]]. For example, if the part of speech is {"nouns"} or {"proper nouns"} and the label (after removing any links and `<<...>>` glossary specs) is {"plural"}, an additional category <code><var>lang</var> countable nouns</code> will be added if a plural value is given (i.e. the value is not {"-"}). If the value is {"-"} (which indicates that there is no plural and triggers the insertion of the fixed inflection label {"no plural"}), <code><var>lang</var> uncountable nouns</code> will be inserted instead, and if both {"-"} and a value are given (which triggers the insertion of the {"usually no plural"} fixed inflection label), both categories are added. Similar categories are inserted when a comparative is given (with a label {"comparative"}), and if the label is {"female equivalent"} or {"male equivalent"} and the value is not {"-"}, a category such as <code><var>lang</var> nouns with other-gender equivalents</code> is inserted. ]==] function Headdata:insert_inflection(terms, label, props) props = props and shallow_copy(props) or {} if not props.no_auto_cats then local bare_label = label if bare_label:find("[[", nil, true) then bare_label = require(links_module).remove_links(bare_label) end if bare_label:find("<<", nil, true) then bare_label = bare_label:gsub("<<.-|(.-)>>", "%1"):gsub("<<(.-)>>", "%1") end local cats = inflection_to_cats[bare_label] if cats then if not cats.filter_plpos or cats.filter_plpos(self:get_canonicalized_plpos()) then if props.cats == nil then props.cats = self:canonicalize_categories(cats.cats) end if props.usually_no_cats == nil then props.usually_no_cats = self:canonicalize_categories(cats.usually_no_cats) end if props.no_cats == nil then props.no_cats = self:canonicalize_categories(cats.no_cats) end end end end props.headdata = self props.terms = terms props.label = label return export.insert_inflection(props) end --[==[ Insert a "fixed" inflection (a label without associated values) into the `inflections` table of the headword `data` structure, labeled according to `label` (which can have glossary links in it specified using `<<...>>`, exactly as for `:insert_inflection()`). An example label (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"}, specifying that the noun belongs to the hidden-''g'' declension. This is a direct wrapper around the top-level function `insert_fixed_inflection()`; see that function for more details on optional `props`. ]==] function Headdata:insert_fixed_inflection(label, props) props = props and shallow_copy(props) or {} props.headdata = self props.label = label export.insert_fixed_inflection(props) end --[==[ Parse the inflection(s) specified in `field` and insert them into the `inflections` table of the headword `data` structure, labeled according to `label`. This is equivalent to calling {terms = data:parse_inflection(field, props)} followed by {return data:insert_inflection(terms, label, props)} and behaves the same as the combination of those two functions. See their documentation for more details. ]==] function Headdata:parse_and_insert_inflection(field, label, props) local terms = self:parse_inflection(field, props) return self:insert_inflection(terms, label, props) end --[==[ Generate an inflection that may be specified explicitly or defaulted (which involves looping over the specified or defaulted heads and determining the script of each one, since the formation of the default depends on the script). `data` is the data object passed into the POS handler. `terms` is the list of terms to process. Those where the term itself is not `+` will be returned unchanged, while those where the term is `+` will be handled by generating the appropriate inflections from the headwords using `make_inflection` (which is passed three arguments, `head`, `tr` and `sccode`, i.e. the script code of `head`) and should return two values, term and translit, either of which can be nil. A nil head will be ignored, and otherwise the decorations specified on the `+` term will be combined with the decorations specified on the head. The return value is a list of inflections where no requests for the default inflection remain. ]==] function Headdata:resolve_special(terms, handle_special, props) props = props or {} local infls = {} local is_special = props.is_special or function(_data, infl) return infl.term == "+" end for _, termobj in ipairs(terms) do if not is_special(self, termobj) then insert(infls, termobj) else for _, headobj in ipairs(self.heads) do local head = headobj.term or self.pagename local head_no_links if props.with_links then head = head:find("%[") and head or require(headword_module).add_multiword_links(head, not headobj.term) head_no_links = require(links_module).remove_links(head) else head = require(links_module).remove_links(head) head_no_links = head end local newterms = handle_special { head = head, tr = headobj.tr, infl = termobj, sc = self.lang:findBestScript(head_no_links), } if newterms then newterms = export.canonicalize_termobj_list(newterms, "term", "resolve_special") for _, newterm in ipairs(newterms) do if not props.no_combine_handle_special_retval_with_origin then export.combine_termobj_decorations(newterm, termobj) end if not props.no_combine_handle_special_retval_with_head then export.combine_termobj_decorations(newterm, headobj) end insert(infls, newterm) end end end end end return infls end --[==[ Add the current page to a tracking page named `Wiktionary:Tracking/``lang``-headword/``page```, where ``lang`` is the language code of the current language. For example, if the current language is `mak` and `page` is {"redundant-lon"}, the current page will get added to the tracking page `Wiktionary:Tracking/mak-headword/redundant-lon`. All pages added to that tracking page can be seen by going to [[Special:WhatLinksHere/Wiktionary:Tracking/mak-headword/redundant-lon]]. This is typically used to track issues occurring in user-specified parameters that do not rise to the level of errors (e.g. redundant parameters, deprecated usages or other dispreferred values). ]==] function Headdata:track(page) return require(debug_track_module)(self.langcode .. "-headword/" .. page) end local boolean_param = {type = "boolean"} --[==[ Process an arbitrary headword in an arbitrary language, handling generic and language-specific arguments and calling `full_headword()` in [[Module:headword]]. This is intended for use in implementing headword modules (e.g. [[Module:uz-headword]] for Uzbek, [[Module:mn-headword]] for Uzbek, [[Module:gsw-headword]] for Alemannic German, etc.) and provides a general implementation of such modules. On input, `data` is an object with the following fields: * `lang`: The language object of the language being handled. '''Required.''' Use the special value {false} to indicate that the language is specified by the user in {{para|1}}. * `frame`: The frame object passed into the `show()` function of your module, which implements headword-handling for all parts of speech in the module, including a generic POS-handling template (e.g. {{tl|uz-head}} or {{tl|mn-head}}), which allows arbitrary parts of speech to be handled. '''Required.''' * `pos_functions`: A table listing, for each part of speech requiring special handling, the extra parameters (if any) that the part of speech accepts, along with how to handle them. See the examples below. '''Required.''' * `validate_lang`: If {lang = true} is specified, this is a function of one argument (a language object, based on the language specified in {{para|1}}) that should throw an error if the language object is disallowed. If omitted, all languages are allowed. * `numbered_head`: If true, explicit headwords are specified in a numbered param instead of in {{para|head}}. The param used is usually {{para|1}}, but is {{para|2}} for generic POS templates such as {{tl|mn-head}} or for POS-specific templates when {lang = true} is specified (e.g. {{tl|arb-noun}}), and is {{para|3}} for generic POS templates when {lang = true} is spcified (e.g. {{tl|arb-head}}). * `include_tr`: If true, allow explicit transliterations to be specified. The transliteration(s) for the headword(s) themselves is/are specified in {{para|tr}} or through the {{cd|<tr:...>}} inline modifier on headwords, and transliterations of inflections are specified through the {{cd|<tr:...>}} inline modifier. This should generally be given when a headword for the language may be in a script other than Latin. * `include_ts`: If true, allow explicit transcriptions to be specified. The transcription(s) for the headword(s) themselves is/are specified in {{para|ts}} or through the {{cd|<ts:...>}} inline modifier on headwords, and transcriptions of inflections are specified through the {{cd|<ts:...>}} inline modifier. This should generally only be given for certain languages where the spelling is radically different from the pronunciation (e.g. in cuneiform languages such as Hittite and Akkadian, and potentially in Tibetan), and represents a pronunciation-based rendering (usually not direct IPA). * `include_sc`: If true, allow an explicit script code to be specified. The overall script code for the headword(s) themselves can be specified using {{para|sc}}, and per-headword or per-inflection script codes are specified using the {{cd|<sc:...>}} inline modifier. This should generally be given when a language supports multiple scripts. * `infls`: An inflections structure specifying extra generic parameters that apply to all parts of speech and how to handle them. The format is the same as for the `infls` structure in `pos_functions`. * `augment_params`: A callback function to add extra generic parameters, or modify existing generic parameters in the `params` structure; but in general, extra generic parameters should be added through the `infls` structure instead. This callback should not be used to add part-of-speech-specific parameters; those are handled through the appropriate setting in `pos_functions`. It is passed two single arguments, the `headdata` object and the `params` table to be augmented. See below for extra fields stored in the `process_props` structure of the `headdata` object. This function is called after initializing the `params` table and processing the overall `infls` structure, but just before adding part-of-speech-specific parameters (from `pos_functions`) to `params`. Thus, it can override any generic parameters but may itself be overridden by a part-of-speech-specific parameter. * `augment_headdata`: A callback function to modify the `headdata` object passed to `full_headword()` in [[Module:headword]]. This should not be used to for part-of-speech-specific parameter handling; this is handled through the appropriate setting in `pos_functions`. It is passed a single argument, the `headdata` object, as for `augment_params`; but the `process_props` structure and other fields will be more filled out, as this callback is called later. This can be used, for example, to override the value of a generic setting in `headdata` (e.g. [[Module:uz-headword]] uses this to mark non-Latin terms as variant forms by setting `headdata.var`, being careful not to override a value already set by the user). This function is called after initializing the `headdata` table with all information taken from generic parameters and processing generic parameters specified in the overall `infls` structure, and just before processing the appropriate part-of-speech-specific `infls` structure in `pos_functions` (which in turn is followed by any handler function in `pos_functions`). Thus, it can override any value set during generic parameter processing but may itself be overridden by a part-of-speech-specific parameter or handler. * `force_cat`: If true, add the headword to the appropriate categories even on non-mainspace pages. This can be used for testing category handling in sample template calls on userspace test pages or template documentation pages. It should not be set in production code. * `enable_auto_translit`: If true, turn on automatic transliteration of inflections at a global level (i.e. applying to all inflections). This has no effect on headwords, which are automatically transliterated by default if in a non-Latin script and automatic transliteration is available for the language. You can also set this value for particular inflections in the `insert_inflection()` function. The `headdata` headword data structure has an extra field in it called `process_props` that is specific to the `process_headword()` function, containing various extra properies. As the operation of `process_headword()` proceeds, this object gets filled out with more fields. For example, once parameter parsing happens, the resulting values are available in the `args` field of `process_props`. The following fields are found in `process_props` (note that `poscat`, the canonicalized plural part of speech of the headword being processed, is *not* present here; it's directly on `headdata`): * `namespace`: The name of the current namespace; an empty string for the mainspace. This references the namespace of the actual page and isn't affected by the {{para|pagename}} parameter. * `indexing_poscat`: The canonicalized part of speech of the headword used to index into `pos_functions`. This is the same as `poscat` for specific part-of-speech templates such as {{tl|uz-noun}}, but has the value {"head"} for generic part-of-speech templates such as {{tl|uz-head}}. (Note that `poscat` is directly available on `headdata`.) * `generic_pos_template`: True if a generic POS templates like {{tl|uz-head}} or {{tl|mn-head}} was used. (This is signaled by omitting the invocation parameter {{para|1}} to `process_headword`.) * `lang_in_1`: True if the language code is to be fetched from {{para|1}}. * `pos_param`: The parameter holding the part of speech, if a generic POS tempalte like {{tl|mn-head}} is being processed (i.e. `generic_pos_template` is set). In such a case, it will have the value of {1} or {2}, depending on whether the language code is being fetched from {{para|1}} (see `lang_in_1`). Otherwise it will be {nil}. * `head_param`: The parameter holding the explicit headword. If `numbered_head` was specified (as for Mongolian headword templates), this has the value {1}, {2} or {3} depending on whether the language code is being fetched from {{para|1}} (see `lang_in_1`) and whether a generic POS template like {{tl|mn-head}} is being processed (see `generic_pos_template`). Otherwise, it has the value {"head"}. Also see the `lang` and `numbered_head` properties in the `data` structure sent to `process_headword()`. * `is_suffix`: True if the current term is a suffix. This is set when processing the `suffix`, `nosuffix` and `clitic` parameters; it is always {false} beforehand (i.e. during `augment_params` and processing of the general `infls` structure). * `insert_specs`: This is a table mapping parameter names to the return value of `Headdata:insert_inflection()`, filled out as parameter values are processed. This lets a given parameter processing function in `infls` gain access to the result of calling `insert_inflection()` on previous parameters (which indicates the number of items inserted as well as whether `-` was specified). The `pos_functions` table contains an entry for each part of speech needing special handling, where the key is the canonical plural part of speech (e.g. {"adverbs"} or {"proper nouns"}). The value associated with each key is a table normally containing a field `infls`, listing the extra part-of-speech-specific inflection and other parameters along with how to handle them. The specs in `infls` are used in three ways: # to augment the `params` object passed to the `process()` function in [[Module:parameters]], specifying how to parse the appropriate inflectional parameters; # to specify how to process any inflectional parameters given and insert them into the `headdata` object passed to `full_headword()` in [[Module:headword]]; # to generate appropriate documentation for the parameters and other changes made by the headword template (e.g. inserting categories). Alternatively, you can separately control the augmentation of the `params` object and the procesing of the resulting arguments. This is done by specifying two fields in place of `infls`, named `params` and `func`. `params` is a table containing extra parameters to add to the overall `params` object passed to the `process()` function in [[Module:parameters]]. `func` is a function of two arguments, normally called `data` (the headword data structure `headdata`) and `args` (the processed arguments table). However, this alternative method is not normally recommended because it leads to duplication between the `params and `func` fields and the documentation, which must be manually specified. A simple example, as used to handle pronouns for Turoyo, is { local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"} pos_functions["pronouns"] = { infls = { {2, type = "genders", validate = valid_genders}, {"f", label = "feminine"}, {"pl", label = "plural"}, }, } } The equivalent using `params` and `func` is { local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"} pos_functions["pronouns"] = { params = { [2] = {type = "genders"}, f = true, pl = true, }, func = function(data, args) data:validate_genders(args[2], valid_genders) data.genders = args[2] data:parse_and_insert_inflection("f", "feminine") data:parse_and_insert_inflection("pl", "plural") end } } Note how the version with separate `params and `func` is longer and splits information on the parameters between the two fields. The `params` structure sets extra user-specifiable parameters {{para|2}} for genders (since the headword is in {{para|1}}) as well a {{para|f}} and {{para|pl}}, and the `func` handler processes those parameters. Note how this is done by calling methods on the headword `data` structure. Each such parameter can have multiple comma-separated values, and each value can have inline modifiers attached to it to specify further properties of the value. The `infls` version ends up making the same method calls, but does it for you instead of you having to do it yourself. These methods are implemented through a metatable set on the headword `data` structure, which is removed before calling `full_headword()` in [[Module:headword]]. The methods access extra information related to headword processing (such as the `args` table) that is stored in the `process_props` field of the headword `data` strucuture. This field is also removed prior to calling `full_headword()`. The methods available on the headword `data` structure are as follows. Each one also has its own documentation. * {parse_inflection(field, props)}: Parse value(s) specified in `field` (a user-specified parameter in the `args` table) and return a list of term objects. Optional `props` specifies additional properties controlling the parsing. * {insert_inflection(terms, label, props)}: Insert the terms in `terms` (a list of term objects as returned by `parse_inflection()`) into the `inflections` list in the headword `data` structure, giving the inflection the label as specified in `label`. Optional `props` specifies additional properties controlling the parsing. * {parse_and_insert_inflection(field, label, props)}: A combination of `parse_inflection()` and `insert_inflection()`, if no further processing of the parsed values needs to be done before insertion. * {insert_fixed_inflection(label, props)}: Insert a "fixed" inflection (a label without associated values) into the `inflections` table. An example (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"}, specifying that the noun belongs to the hidden-''g'' declension. * {resolve_special(terms, handle_special, props)}: Resolve "special" indicators as specified by the user in an inflection parameter. A typical example is {"+"}, requesting a default value. `terms` is the list of parsed term objects and `handle_special` is a handler function to process special indicators and convert them to their actual values. * {validate_genders(genders, valid_genders, props)}: Validate that the user-specified genders in `genders` all belong to the list given in `valid_genders`, throwing an error if not. * {insert_category(category)}: Insert a category into the `categories` list in the headword `data` structure. `category` is normally a string naming the category, which will have the language prepended to it and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech. ]==] function export.process_headword(data) local lang, frame, pos_functions, validate_lang, numbered_head, include_tr, include_ts, include_sc, force_cat, enable_auto_translit, infls, augment_params, augment_headdata = data.lang, data.frame, data.pos_functions, data.validate_lang, data.numbered_head, data.include_tr, data.include_ts, data.include_sc, data.force_cat, data.enable_auto_translit, data.infls, data.augment_params, data.augment_headdata local iparams = { [1] = true, def = true, } local iargs = require(parameters_module).process(frame.args, iparams) local parargs = frame:getParent().args local langcode if not lang then error("Internal error: `data.lang` must be specified; either a language object or `true` for a user-specified language") end local lang_in_1 if lang == true then lang_in_1 = true langcode = ine(parargs[1]) if langcode then langcode = mw.text.trim(langcode) lang = require(languages_module).getByCode(langcode, 1, true) if validate_lang then validate_lang(lang) end else error("Language code (see [[WT:Language codes]]) must be specified in 1=") end else langcode = lang:getCode() if validate_lang then error("Internal error: `data.validate_lang` must not be specified if a language code is given in `data.lang`") end end local poscat = iargs[1] local generic_pos_template = not poscat local pos_param if generic_pos_template then pos_param = lang_in_1 and 2 or 1 poscat = ine(parargs[pos_param]) or mw.title.getCurrentTitle().fullText == ("Templat:%s-head"):format(langcode) and "interjection" or error(("Part of speech must be specified in %s="):format(pos_param)) poscat = require(headword_module).canonicalize_pos(poscat) end local head_param = numbered_head and (generic_pos_template and lang_in_1 and 3 or (generic_pos_template or lang_in_1) and 2 or 1) or "head" local indexing_poscat = generic_pos_template and "head" or poscat local namespace = mw.loadData(headword_data_module).page.namespace -- Partly initialize headdata now for use in generic infls callbacks. Will be further initialized later after -- processing parameters. local headdata = { lang = lang, langcode = langcode, langfullcode = lang:getFullCode(), langname = lang:getCanonicalName(), langfullname = lang:getFullName(), process_props = { namespace = namespace, data = data, indexing_poscat = indexing_poscat, generic_pos_template = generic_pos_template, lang_in_1 = lang_in_1, pos_param = pos_param, head_param = head_param, is_suffix = false, insert_specs = {}, }, pos_category = poscat, orig_poscat = poscat, -- preserve user-specified poscat in case pos_category is changed to 'suffixes' categories = {}, inflections = {enable_auto_translit = enable_auto_translit}, force_cat_output = force_cat, no_redundant_head_cat = true, } setmetatable(headdata, {__index = Headdata}) local params = { [head_param] = {template_default = iargs.def}, head2 = {replaced_by = false, instead = ("use comma-separated |%s="):format(head_param)}, id = true, sort = true, cat = true, nolink = boolean_param, nolinkhead = {type = "boolean", alias_of = "nolink"}, suffix = boolean_param, nosuffix = boolean_param, clitic = true, addlpos = true, var = {type = "boolean", allow = {"both"}}, json = boolean_param, pagename = true, -- for testing } if include_sc then params.sc = {type = "script"} end if include_tr then params.tr = true params.tr2 = {replaced_by = false, instead = "use comma-separated |tr= or <tr:...> inline modifier on head"} end if include_ts then params.ts = true params.ts2 = {replaced_by = false, instead = "use comma-separated |ts= or <ts:...> inline modifier on head"} end if lang_in_1 then params[1] = {required = true} -- required but ignored as already processed above end if generic_pos_template then params[pos_param] = {required = true} -- required but ignored as already processed above end local function resolve_prop(prop, ...) if type(prop) == "function" then prop = prop(headdata, ...) end return prop end local function augment_params_from_infls(infls) infls = resolve_prop(infls) for _, infl in ipairs(infls) do local function interr(txt) error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl))) end local param = infl[1] if param then param = resolve_prop(param) if type(param) ~= "string" and type(param) ~= "number" then interr(("Parameter name %s must be a string or number"):format(dump(param))) end -- We handle defaults as well as validation ourselves. local typ = resolve_prop(infl.type) or "string" if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then -- FIXME: Handle more types. interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format( dump(typ))) end params[param] = {type = typ, required = resolve_prop(infl.required), template_default = resolve_prop(infl.template_default)} if typ ~= "boolean" and type(param) == "string" then params[param .. "2"] = {replaced_by = false, instead = ("use comma-separated |%s="):format(param)} end end end end if infls then augment_params_from_infls(infls) end if augment_params then augment_params(headdata, params) end if pos_functions[indexing_poscat] then local pos_infls = pos_functions[indexing_poscat].infls if pos_infls then augment_params_from_infls(pos_infls) end local pos_params = pos_functions[indexing_poscat].params if pos_params then for key, val in pairs(pos_params) do params[key] = val end end end local args = require("Module:parameters").process(parargs, params) local pagename = args.pagename or mw.loadData(headword_data_module).pagename local sc = args.sc or lang:findBestScript(pagename) headdata.pagename = pagename headdata.process_props.args = args headdata.sc = sc headdata.id = args.id headdata.sort = args.sort -- No redundant script cat unless the user explicitly gave sc= headdata.no_script_code_cat = not args.sc headdata.var = args.var local extra_term_mods = {} if include_tr then insert(extra_term_mods, "tr") end if include_ts then insert(extra_term_mods, "ts") end if include_sc then insert(extra_term_mods, "sc") end if not extra_term_mods[1] then extra_term_mods = nil end local trs = args.tr and split_on_comma(args.tr) or {} local num_trs = #trs local tss = args.ts and split_on_comma(args.ts) or {} local num_tss = #tss local heads = args[head_param] and export.parse_term_with_modifiers { val = args[head_param], paramname = head_param, splitchar = ",", is_head = true, include_mods = extra_term_mods, } or {} local num_heads = #heads if num_heads > 0 and num_trs > 0 and num_heads ~= num_trs then error(("%s head%s specified explicitly but %s translit%s; they must match; use '+' to stand for the default head (the pagename) or default automatic translit and '-' to stand for no translit"):format( num_heads, num_heads > 1 and "s" or "", num_trs, num_trs > 1 and "s" or "")) end if num_heads > 0 and num_tss > 0 and num_heads ~= num_tss then error(("%s head%s specified explicitly but %s transcription%s; they must match; use '+' to stand for the default head (the pagename) and '-' to stand for no transcription"):format( num_heads, num_heads > 1 and "s" or "", num_tss, num_tss > 1 and "s" or "")) end if num_trs > 0 and num_tss > 0 and num_trs ~= num_tss then error(("%s translit%s specified explicitly but %s transcription%s; they must match; use '+' to stand for default automatic translit and '-' to stand for no translit or transcription"):format( num_trs, num_trs > 1 and "s" or "", num_tss, num_tss > 1 and "s" or "")) end -- Be careful here not to overwrite user_specified_heads if it's empty so we can later check user_specified_heads -- to see if the user provided any heads. local max_tr_ts = math.max(num_trs, num_tss) if num_heads == 0 and max_tr_ts > 0 then heads = {} for i = 1, max_tr_ts do heads[i] = {term = "+"} end end if not heads[1] then heads = {{term = "+"}} end for i, headobj in ipairs(heads) do if headobj.tr and trs[i] then if headobj.tr ~= trs[i] then error(("Saw two different translits '%s' and '%s' for head #%s"):format( headobj.tr, trs[i], i)) end else headobj.tr = headobj.tr or trs[i] end if headobj.tr == "+" then headobj.tr = nil end if headobj.ts and tss[i] then if headobj.ts ~= tss[i] then error(("Saw two different transcriptions '%s' and '%s' for head #%s"):format( headobj.ts, tss[i], i)) end else headobj.ts = headobj.ts or tss[i] end if headobj.ts == "-" then headobj.ts = nil end if headobj.term == "+" then headobj.term = args.nolink and pagename or nil if headobj.term and namespace == "Reconstruction" then headobj.term = "*" .. headobj.term end end end headdata.heads = heads local function pagename_is_suffix() if sc:getCode() == "Latn" then -- shortcut Latin terms to avoid unnecessarily loading [[Module:affix]] return pagename:find("^%-") and not pagename:find("%-$") else local affix_type, _, _, _ = require(affix_module).parse_term_for_affixes(pagename, lang, sc) return affix_type == "suffix" end end local clitic_label if args.clitic then clitic_label = require(yesno_module)(args.clitic, args.clitic) end if clitic_label == true then clitic_label = "klitik" end if clitic_label then headdata:insert_category("Klitik") headdata:insert_fixed_inflection(clitic_label) elseif args.suffix or ( not args.nosuffix and pagename_is_suffix() and poscat ~= "Akhiran" and poscat ~= "Bentuk akhiran" ) then headdata.process_props.is_suffix = true local function handle_suffix_pos(pos, is_first) local form_type = pos:match("^Bentuk (.*)$") local actual_poscat if form_type then headdata:insert_category(("Bentuk akhiran %s"):format(form_type)) headdata:insert_fixed_inflection("Bentuk akhiran " .. form_type) else local singular_pos = require(en_utilities_module).singularize(pos) headdata:insert_category(("Akhiran membentuk %s"):format(singular_pos)) headdata:insert_fixed_inflection("Akhiran membentuk " .. singular_pos) end local postype = require(headword_module).pos_lemma_or_nonlemma(pos) if not postype then error(("Unrecognized canonicalized part of speech '%s' in addlpos=, cannot determine whether lemma or non-lemma form"):format( pos )) end actual_poscat = postype == "Lema" and "Akhiran" or "Bentuk akhiran" if is_first then headdata.pos_category = actual_poscat elseif headdata.pos_category ~= actual_poscat then error(("Cannot mix suffixes and suffix forms using addlpos=; '%s' is a %s while overall POS '%s' is a %s; use separate POS headers for the two"): format(pos, actual_poscat, poscat, headdata.pos_category)) end end handle_suffix_pos(poscat, true) if args.addlpos then for _, addlpos in ipairs(split(args.addlpos, "%s*,%s*")) do addlpos = require(headword_module).canonicalize_pos(addlpos) handle_suffix_pos(addlpos, false) end end end if args.cat then for _, cat in ipairs(split_on_comma(args.cat)) do headdata:insert_category(cat) end end local function augment_headdata_from_infls(infls) infls = resolve_prop(infls) for _, infl in ipairs(infls) do local function interr(txt) error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl))) end local function process_labelobjs(labelobjs, originating_term, handle_labelobj) if labelobjs == nil then return end if type(labelobjs) ~= "string" and type(labelobjs) ~= "table" then interr(("Wrong type '%s' for label object(s) %s, expected string or table"):format( type(labelobjs), dump(labelobjs) )) end if type(labelobjs) == "string" or type(labelobjs) == "table" and not labelobjs[1] then labelobjs = {labelobjs} end for _, labelobj in ipairs(labelobjs) do local label, termobj if type(labelobj) == "string" then label = labelobj termobj = originating_term elseif type(labelobj) ~= "table" then interr(("Wrong type '%s' for label object %s, expected string or table"):format( type(labelobj), dump(labelobj) )) label = labelobj.term if type(label) ~= "string" then interr(("Wrong type '%s' for label %s from label object %s, expected string"):format( type(label), dump(label), dump(labelobj) )) end termobj = labelobj end handle_labelobj(label, termobj) end end local param = infl[1] if param then local vals -- Fetch the param and make sure it's a string or number. param = resolve_prop(param) if type(param) ~= "string" and type(param) ~= "number" then interr(("Parameter name %s must be a string or number"):format(dump(param))) end -- Fetch the type and validate. local typ = resolve_prop(infl.type) if typ == nil then typ = "string" end if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then -- FIXME: Handle more types. interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format( dump(typ))) end -- Fetch the value(s). if typ == "genders" or typ == "boolean" then vals = args[param] elseif typ == "string" then local parse_inflection_props = resolve_prop(infl.parse_inflection_props) local include_mods = resolve_prop(infl.include_mods) local no_augment_include_mods = resolve_prop(infl.no_augment_include_mods) if include_mods ~= nil or no_augment_include_mods ~= nil then if parse_inflection_props == nil then parse_inflection_props = {} else parse_inflection_props = shallow_copy(parse_inflection_props) end if include_mods ~= nil then parse_inflection_props.include_mods = include_mods end if no_augment_include_mods ~= nil then parse_inflection_props.no_augment_include_mods = no_augment_include_mods end end vals = headdata:parse_inflection(param, parse_inflection_props) -- Convert an empty list to nil for consistent checking below. if not vals[1] then vals = nil end else interr(("Unrecognized type '%s"):format(typ)) end -- If value(s) nil, fetch the default. if vals == nil and infl.default ~= nil then local default = resolve_prop(infl.default) if typ == "genders" then vals = export.canonicalize_termobj_list(default, "spec", "default") elseif typ == "boolean" then vals = default elseif typ == "string" then vals = export.canonicalize_termobj_list(default, "term", "default") else interr(("Unrecognized type '%s"):format(typ)) end end -- Resolve "special" values (special signals a string values, such as requesting the default with "+"). if vals ~= nil and infl.resolve_special then if typ ~= "string" then interr(("Cannot specify resolve_special= for type %s"):format(dump(typ))) end local resolve_special_props = resolve_prop(infl.resolve_special_props, vals) if infl.is_special ~= nil then if resolve_special_props == nil then resolve_special_props = {} else resolve_special_props = shallow_copy(resolve_special_props) end resolve_special_props.is_special = infl.is_special end vals = headdata:resolve_special(vals, infl.resolve_special, resolve_special_props) end -- Validate the value(s). if vals ~= nil and infl.validate ~= nil then if typ == "boolean" then interr('Cannot specify validate= when type is "boolean"') elseif type(infl.validate) == "function" then infl.validate(headdata, vals) elseif typ == "genders" then headdata:validate_genders(vals, infl.validate) elseif typ == "string" then validate_items { items = vals, field = "term", valid_items = infl.validate, item_type = ("values in |%s="):format(param), } else interr(("Unrecognized type '%s"):format(typ)) end end -- Run the process_after_parse handler, if it exists. if vals ~= nil and infl.process_after_parse ~= nil then local intentionally_nil vals, intentionally_nil = infl.process_after_parse(headdata, vals) if vals == nil and not intentionally_nil then interr("If you return nil from process_after_parse, you must return a second non-nil return " .. "value to indicate that the nil return value was intentional") end end -- "Implement" the values, if non-falsy (i.e. we don't want to fire on boolean false or empty list). -- If a fixed label is specified, insert it. Then, depending on the type, attach the values to a label -- as an inflection, set the `genders` field, or do nothing if boolean (throwing an error if there was -- no fixed label). if vals == true or type(vals) == "table" and vals[1] then if infl.fixed_label and infl.all_fixed_label then interr("Cannot specify both fixed_label= and all_fixed_label=; specify one or the other") end local function check_fixed_label_references_val(label) if type(label) == "table" and label[1] then for _, lab in ipairs(label) do if check_fixed_label_references_val(lab) then return true end end return false end if type(label) == "table" then if not label.term then interr(("Fixed label structure %s does not have a value for `.term`"):format(dump(label))) end label = label.term end if type(label) ~= "string" then interr(("Wrong type for fixed label %s, should be string"):format(type(label))) end return not not label:find("{val}") end local fixed_label = infl.fixed_label local all_fixed_label = infl.all_fixed_label -- If the value being processed is boolean, there's only one value so treat a fixed_label as an -- all_fixed_label and output only once; likewise if the caller specified a fixed_label without -- {val} in it. if fixed_label and (typ == "boolean" or type(fixed_label) ~= "function" and not check_fixed_label_references_val(fixed_label)) then all_fixed_label = fixed_label fixed_label = nil end local inserted_fixed_label if fixed_label then if typ == "boolean" then interr("Boolean fixed_label values should have been converted to all_fixed_label") end for _, valobj in ipairs(vals) do local labelobjs = resolve_prop(fixed_label, valobj) process_labelobjs(labelobjs, valobj, function(label, termobj) if label:find("{val}") then if typ ~= "string" then interr(('Cannot specify {val} in fixed_label %s when type is "%s"'):format(dump(label), typ)) end label = label:gsub("{val}", replacement_escape(valobj.term)) end headdata:insert_fixed_inflection(label, { originating_term = termobj }) inserted_fixed_label = true end) end elseif all_fixed_label then local labelobjs = resolve_prop(all_fixed_label, vals) process_labelobjs(labelobjs, nil, function(label, termobj) if label:find("{vals}") then if typ ~= "string" then interr(('Cannot specify {vals} in all_fixed_label %s when type is "%s"'):format(dump(label), typ)) end local formatted_labels = {} for _, valobj in ipairs(vals) do insert(formatted_labels, add_decorations(valobj.term, valobj, lang)) end label = label:gsub("{vals}", replacement_escape(serial_comma_join(formatted_labels))) end headdata:insert_fixed_inflection(label, termobj) inserted_fixed_label = true end) end local inserted_vals if infl.label ~= nil then if typ ~= "string" then interr(("label=%s can only be specified for type 'string', not '%s'"):format( dump(infl.label), typ )) end local label = resolve_prop(infl.label, vals) if label ~= nil then local insert_inflection_props = resolve_prop(infl.insert_inflection_props, vals) local no_auto_cats = resolve_prop(infl.no_auto_cats, vals) if no_auto_cats ~= nil then if insert_inflection_props == nil then insert_inflection_props = {} else insert_inflection_props = shallow_copy(insert_inflection_props) end insert_inflection_props.no_auto_cats = infl.no_auto_cats end local insert_spec = headdata:insert_inflection(vals, label, insert_inflection_props) headdata.process_props.insert_specs[param] = insert_spec inserted_vals = true end end if typ == "genders" then headdata.genders = vals end local inserted_cat if infl.cat then local allcats = {} for _, valobj in ipairs(vals) do local cats = resolve_prop(infl.cat, valobj) if type(cats) == "string" then cats = {cats} end if cats ~= nil then for _, cat in ipairs(cats) do if cat:find("{val}") then cat = cat:gsub("{val}", replacement_escape(valobj.term)) end insert_if_not(allcats, cat) end end end for _, cat in ipairs(allcats) do headdata:insert_category(cat) inserted_cat = true end end if typ == "boolean" then if not inserted_fixed_label and not inserted_cat then interr(("User set boolean setting for %s= but no fixed label added and no category " .. "inserted; if you took action in process_after_parse(), make sure to return " .. "`nil, true`"):format(param)) end elseif typ == "string" then if not inserted_vals and not inserted_fixed_label then interr(("User set value(s) %s for %s= but no inflection inserted and no fixed label " .. "added; if you took action in process_after_parse(), make sure to return " .. "`nil, true`"):format(dump(vals), param)) end end end else -- no param specified if infl.label or infl.all_fixed_label then interr("Cannot have label= or all_fixed_label= without specifying a param") end if infl.fixed_label then local labelobjs = resolve_prop(infl.fixed_label) process_labelobjs(labelobjs, nil, function(label, termobj) if label:find("{val}") then interr("Cannot specify {val} in a fixed_label= value without specifying a param") end headdata:insert_fixed_inflection(label, { originating_term = termobj }) end) end if infl.cat then local cats = resolve_prop(infl.cat) if type(cats) == "string" then cats = {cats} end if cats ~= nil then for _, cat in ipairs(cats) do if cat:find("{val}") then interr("Cannot specify {val} in a cat= value without specifying a param") end headdata:insert_category(cat) end end end end end end if infls then augment_headdata_from_infls(infls) end if augment_headdata then augment_headdata(headdata, args) end if pos_functions[indexing_poscat] then local pos_infls = pos_functions[indexing_poscat].infls if pos_infls then augment_headdata_from_infls(pos_infls) end local func = pos_functions[indexing_poscat].func if func then func(headdata, args) end end setmetatable(headdata, nil) if args.json then return require("Module:JSON").toJSON(headdata) end headdata.process_props = nil return require(headword_module).full_headword(headdata) end return export nrwzxukyx15w9bs1qysq8d6jicifovd Modul:parameter utilities 828 55604 375382 280880 2026-09-22T07:04:25Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708725|92708725]]) 375382 Scribunto text/plain local export = {} local debug_track_module = "Module:debug/track" local functions_module = "Module:fun" local parameters_module = "Module:parameters" local parse_interface_module = "Module:parse interface" local parse_utilities_module = "Module:parse utilities" local table_module = "Module:table" local dump = mw.dumpObject local error = error local insert = table.insert local ipairs = ipairs local next = next local pairs = pairs local require = require local tonumber = tonumber local type = type --[==[ Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls. ]==] local function debug_track(...) debug_track = require(debug_track_module) return debug_track(...) end local function is_callable(...) is_callable = require(functions_module).is_callable return is_callable(...) end local function list_to_set(...) list_to_set = require(table_module).listToSet return list_to_set(...) end local function parse_inline_modifiers(...) parse_inline_modifiers = require(parse_interface_module).parse_inline_modifiers return parse_inline_modifiers(...) end local function process_params(...) process_params = require(parameters_module).process return process_params(...) end local function shallow_copy(...) shallow_copy = require(table_module).shallowCopy return shallow_copy(...) end local function table_len(...) table_len = require(table_module).length return table_len(...) end ----------------- end loaders ---------------- local function track(page, track_module) return debug_track((track_module or "parameter utilities") .. "/" .. page) end -- Throw an error prefixed with the words "Internal error" (and suffixed with a dumped version of `spec`, if provided). -- This is for logic errors in the code itself rather than template user errors. local function internal_error(msg, spec) if spec then msg = ("%s: %s"):format(msg, dump(spec)) end error(("Internal error: %s"):format(msg)) end -- Table listing the default recognized special separator arguments and how they display. export.default_special_separators = { [";"] = "; ", ["_"] = " ", ["~"] = " ~ ", ["→"] = " → ", } -- Table listing how subitem delimiters display. Unlike for `default_special_separators`, the presence of an item in -- this table does not mean that the delimiter is recognized; only those specified by `data.splitchar` are recognized. export.default_subitem_separator_map = { [";"] = "; ", [","] = ", ", ["/"] = "/", ["_"] = " ", ["~"] = " ~ ", ["→"] = " → ", } --[==[ intro: The purpose of this module is to facilitate implementation of templates that can have arguments specified either through inline modifiers or separate parameters. There are two types of templates supported: those that take a list of items with associated properties, which can be specified either through indexed separate parameters (e.g. {{para|t2}}, {{para|pos3}}) or inline modifiers (`<t:...>`, `<pos:...>`, etc.); and those that take a single term, whose properties can be specified through non-indexed separate parameters (e.g. {{para|t}} or {{para|pos}}) or inline modifiers. Both types of templates can optionally have subitems in the term parameter(s), where the subitems are typically (but not necessarily) separated with commas and each subitem can have its own inline modifiers. Some examples of templates that take a list of items are {{tl|alter}}/{{tl|alt}}; {{tl|synonyms}}/{{tl|syn}}, {{tl|antonyms}}/{{tl|ant}}, and other "nyms" templates; {{tl|col}}, {{tl|col2}}, {{tl|col3}}, {{tl|col4}} and other column templates; {{tl|descendant}}/{{tl|desc}}; {{tl|affix}}/{{tl|af}}, {{tl|prefix}}/{{tl|pre}} and related *fix templates; {{tl|affixusex}}/{{tl|afex}} and related templates; {{tl|IPA}}; {{tl|homophones}}; {{tl|rhymes}}; and several others. Examples of templates that take a single item are form-of templates ({{tl|inflection of}}/{{tl|infl of}}, {{tl|form of}}, and specific templates such as {{tl|alt form}}/{{tl|alternative form of}}, {{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others); for etymology templates ({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates like {{tl|ellipsis}}, {{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and other templates that take an argument structure similar to {{tl|l}} or {{tl|m}}. This module can be thought of as a combination of [[Module:parameters]] (which parses template parameters, and in particular handles the separate parameter versions of the properties) and `parse_inline_modifiers()` in [[Module:parse utilities]] (which parses inline modifiers). The two main entry points are `parse_list_with_inline_modifiers_and_separate_params()` (for templates that take a list of items) and `parse_term_with_inline_modifiers_and_separate_params()` (for templates that take a single item). However, there are other functions provided, e.g. to initialize the `param_mods` structure that is passed to the two entry points. The typical workflow for using `parse_list_with_inline_modifiers_and_separate_params()` looks as follows (a slightly simplified version of the code in [[Module:nyms]]): { local export = {} local parameter_utilities_module = "Module:parameter utilities" ... -- Entry point to be invoked from a template. function export.show(frame) local parent_args = frame:getParent().args -- Parameters that don't have corresponding inline modifiers. Note in particular that the parameter corresponding to -- the items themselves must be specified this way, and must specify either `allow_holes = true` (if the user can -- omit terms, typically by specifying the term using |altN= or <alt:...> so that they remain unlinked) or -- `disallow_holes = true` (if omitting terms is not allowed). (If neither `allow_holes` nor `disallow_holes` is -- specified, an error is thrown in parse_list_with_inline_modifiers_and_separate_params().) local params = { [1] = {required = true, type = "language", default = "und"}, [2] = {list = true, allow_holes = true, required = true, default = "term"}, } local m_param_utils = require(parameter_utilities_module) -- This constructs the `param_mods` structure by adding well-known groups of parameters (such as all the parameters -- associated with based on full_link() in [[Module:links]], with default properties that can be overridden. This is -- easier and less error-prone than manually specifying the `param_mods` structure (see below for how this would -- look). Here, we specify the group "link" (consisting of all the link parameters for use with full_link()), group -- "ref" (which adds the "ref" parameter for specifying references), group "l" (which adds the "l" and "ll" -- parameters for specifying labels) and group "q" (which adds the "q" and "qq" parameters for specifying regular -- qualifiers). By default, labels and qualifiers have `separate_no_index` set so that e.g. |q1= is distinct from -- |q=, the former specifying the left qualifier for the first item and the latter specifying the overall left -- qualifier. For compatibility, we override the `separate_no_index` setting for the group "q", which causes |q= and -- |q1= to be the same, and likewise for |qq= and |qq1=. Finally, also for compatibility, we add an "lb" parameter -- that is an alias of "ll" (in all respects; |lb= is the same as |ll=, |lb1= is the same as |ll1=, <lb:...> is the -- same as <ll:...>, etc.). local param_mods = m_param_utils.construct_param_mods { {group = {"link", "ref", "l"}}, {group = "q", separate_no_index = false}, {param = "lb", alias_of = "ll"}, } -- This processes the raw arguments in `parent_args`, parses inline modifiers and creates corresponding objects -- containing the property values specified either through inline modifiers or separate parameters. local items, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params { params = params, param_mods = param_mods, raw_args = parent_args, termarg = 2, parse_lang_prefix = true, track_module = "nyms", lang = 1, sc = "sc.default", } local lang = args[1] -- Now do the actual implementation of the template. Generally this should be split into a separate function, often -- in a separate module (if the implementation goes in [[Module:foo]], the template interface code goes in -- [[Module:foo/templates]]). ... } The `param_mods` structure controls the properties that can be specified by the user for a given item, and is conceptually very similar to the `param_mods` structure used by `parse_inline_modifiers()`. The key is the name of the parameter (e.g. {"t"}, {"pos"}) and the value is a table with optional elements as follows: * `item_dest`, `store`: Same as the corresponding fields in the `param_mods` structure passed to `parse_inline_modifiers()`. * `type`, `set`, `sublist`, `convert` and associated fields such as `family` and `method`: These control parsing and conversion of the raw values specified by the user and have the same meaning as in [[Module:parameters]] and also in `parse_inline_modifiers()` (which delegates the actual conversion to [[Module:parameters]]). These fields — and for that matter, all fields other than `item_dest`, `store` and `overall` — are forwarded to the `process()` function in [[Module:parameters]]. * `alias_of`: This parameter is an alias of some other parameter. This spec is recognized only by `process()` in [[Module:parameters]], and not by `parse_inline_modifiers()`; to set up an alias in `parse_inline_modifiers()`, you need to make sure (using `item_dest`) that both the alias and aliasee modifiers store their values in the same location, and you need to copy the remaining properties from the aliasee's spec to the aliasing modifier's spec. All of this happens automatically if you generate the `param_mods` structure using `construct_param_mods()`. * `require_index`: This means that the non-indexed parameter version of the property is not recognized. E.g. in the case of the {"sc"} property, use of the {{para|sc}} parameter would result in an error, while {{para|sc1}} is recognized and specifies the {"sc"} property for the first item. The default, if neither `require_index` nor `separate_no_index` is given, is for {{para|sc}} and {{para|sc1}} to mean the same thing (both would specify the {"sc"} property of the first item). Note that `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `separate_no_index`: This means that e.g. the {{para|sc}} parameter is distinct from the {{para|sc1}} parameter (and thus from the `<sc:...>` inline modifier on the first item). This is typically used to distinguish an overall version of a property from the corresponding item-specific property on the first item. (In this case, for example, {{para|sc}} overrides the script code for all items, while {{para|sc1}} overrides the script code only for the first item.) If not given, and if `require_index` is not given, {{para|sc}} and {{para|sc1}} would have the same meaning and refer to the item-specific property on the first item. When this is given, the overall value can be accessed using the `.default` field of the property value in `args`, e.g. in this case `args.sc.default`. Note that (as mentioned above) `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off. * `list`, `allow_holes`, `disallow_holes`: These should '''not''' be given. `list` and `allow_holes` are automatically set for all parameter specs added to the `params` structure used by `process()` in [[Module:parameters]], and `disallow_holes` clashes with `allow_holes`. For the above workflow example, the call to `construct_param_mods()` generates the following `param_mods` structure: { local param_mods = { -- the parameters generated by group "link" alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", type = "genders", }, pos = {}, ng = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, -- the parameters generated by group "ref" ref = { item_dest = "refs", type = "references", }, -- the parameters generated by group "l" l = { type = "labels", separate_no_index = true, }, ll = { type = "labels", separate_no_index = true, }, -- the parameters generated by group "q"; note that `separate_no_index = true` would be set, but is overridden -- (specifying `separate_no_index = false` in the `param_mods` structure is equivalent to not specifying it at all) q = { type = "qualifier", separate_no_index = false, }, qq = { type = "qualifier", separate_no_index = false, }, infl = { type = "form of tags", separate_no_index = true, }, -- the parameter generated by the individual "lb" parameter spec; note that only `alias_of` was explicitly given, -- while `item_dest` is automatically set so that inline modifier <lb:...> stores into the same place as <ll:...>, -- and the other specs are copied from the `ll` spec so `lb` works like `ll` in all regards lb = { alias_of = "ll", item_dest = "ll", type = "labels", separate_no_index = true, }, } } ]==] local qualifier_spec = { type = "qualifier", separate_no_index = true, } local label_spec = { type = "labels", separate_no_index = true, } local form_of_spec = { type = "form of tags", separate_no_index = true, } local recognized_param_mod_groups = { link = { alt = {}, t = { -- [[Module:links]] expects the gloss in "gloss". item_dest = "gloss", }, gloss = { alias_of = "t", }, tr = {}, ts = {}, g = { -- [[Module:links]] expects the genders in "genders". item_dest = "genders", type = "genders", }, pos = {}, ng = {}, lit = {}, id = {}, sc = { separate_no_index = true, type = "script", }, }, lang = { lang = { require_index = true, type = "language", }, }, q = { q = qualifier_spec, qq = qualifier_spec, }, a = { a = label_spec, aa = label_spec, }, l = { l = label_spec, ll = label_spec, }, infl = { infl = form_of_spec, }, ref = { ref = { item_dest = "refs", type = "references", }, }, } local function merge_param_mod_settings(orig, additions) local merged = shallow_copy(orig) for k, v in pairs(additions) do merged[k] = v if k == "require_index" then merged.separate_no_index = nil elseif k == "separate_no_index" then merged.require_index = nil end end merged.default = nil merged.group = nil merged.param = nil merged.exclude = nil merged.include = nil return merged end local function verify_type(spec, param, typ1, typ2) if not spec[param] then return end local val = spec[param] if type(val) ~= typ1 and (not typ2 or type(val) ~= typ2) then internal_error(("Parameter `%s` must be a %s%s but saw a %s"):format(param, typ1, typ2 and " or " .. typ2 or "", type(val)), spec) end end local function verify_well_constructed_spec(spec) local num_control = (spec.default and 1 or 0) + (spec.group and 1 or 0) + (spec.param and 1 or 0) if num_control == 0 then internal_error( "Spec passed to construct_param_mods() must have either the `default`, `group` or `param` keys set", spec) end if num_control > 1 then internal_error( "Exactly one of `default`, `group` or `param` must be set in construct_param_mods() spec", spec) end if spec.list or spec.allow_holes then -- FIXME: We need to support list = "foo" for list parameters that are stored in e.g. 2=, foo2=, foo3=, etc. internal_error("`list` and `allow_holes` may not be set; they are automatically set when constructing the " .. "corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.disallow_holes then internal_error("`disallow_holes` may not be set; it conflicts with `allow_holes`, which is automatically " .. "set when constructing the corresponding spec in the `params` object passed to [[Module:parameters]]", spec) end if spec.include and spec.exclude then internal_error("Saw both `include` and `exclude` in the same spec", spec) end if (spec.include or spec.exclude) and not spec.group then internal_error( "`include` and `exclude` can only be specified along with `group`, not with `default` or `param`", spec) end verify_type(spec, "group", "string", "table") verify_type(spec, "param", "string", "table") verify_type(spec, "include", "table") verify_type(spec, "exclude", "table") end --[==[ Construct the `param_mods` structure used in parsing arguments and inline modifiers from a list of specifications. A sample invocation (a slightly simplified version of the actual invocation associated with {{tl|affix}} and related templates) looks like this: { local param_mods = require("Module:parameter utilities").construct_param_mods { -- We want to require an index for all params (or use separate_no_index, which also requires an index for the -- param corresponding to the first item). {default = true, require_index = true}, {group = {"link", "ref", "lang", "q", "l"}}, -- Override these two to have separate_no_index. {param = {"lit", "pos"}, separate_no_index = true}, } } Each specification either sets the default value for further parameter specs or adds one or more parameters. Parameters can be added directly using `param`, or groups of predefined parameters can be added using `group`. Specifications are one of three types: # Those that set the default properties for future-added parameters. These contain {default = true} as one of the properties of the spec. Specs are processed in order and you can change the defaults mid-way through. # Those that add the parameters associated with one or more pre-defined groups. These contain {group = "group"} or {group = {"group1", "group2", ...}}. The pre-defined parameter groups and their associated properties are listed below. The pre-defined properties of parameters in a group override properties associated with a {default = true} spec, and are in turn overridden by any properties given directly in the spec itself. Note as well that setting the `separate_no_index` property will automatically cause the `require_index` property to be unset and vice-versa, as the two are mutually exclusive. (This happens in the example above, where the {separate_no_index = true} setting associated with the params {"lit"} and {"pos"} cancels out the {require_index = true} default setting, as well as less obviously with the pre-defined {"sc"} property of the {"link"} group, the {"q"} and {"qq"} properties of the {"q"} group, and the {"l"} and {"ll"} properties of the {"l"} group, all of which have an associated pre-defined property {separate_no_index = true}, which overrides and cancels out the {require_index = true} default setting. Finally, when adding the parameters of a group, you can request the only a subset of the parameters be added using either the `include` or `exclude` properties, each of whose values is a list of parameters that specify (respectively) the parameters to include (all other parameters of the group are excluded) or to exclude (all other parameters of the group are included). This is used, for example, in [[Module:romance etymology]] and [[Module:it-etymology]], which specify {group = "link", exclude = {"tr", "ts", "sc"}} to exclude link parameters that aren't relevant to Latin-script languages such as the Romance languages, and conversely in [[Module:IPA/templates]], which specifies {group = "link", include = {"t", "gloss", "pos"}} to include only the specified parameters for use with {{tl|IPA}}. # Those that add individual parameters. These contain {param = "param"} or {param = {"param1", "param2", ...}}, the latter syntax used to control a set of parameters together. The resulting spec is formed by initializing the parameter's settings with any previously-specified default properties (using a spec containing {default = true}) if the parameter hasn't already been initialized, and then overriding the resulting settings with any settings given directly in the specification. In the above example, the {"lit"} and {"pos"} parameters were previously initialized through the {"link"} group (specified in the second of the three specifications) but ended up with {require_index = true} due to the {default = true} spec (the first of the three specifications). We override these two parameters to have {separate_no_index = true} (which, as mentioned above, cancels out {require_index = true}). This is done so that {{tl|affix}} and related templates have {{para|pos}} and {{para|lit}} parameters distinct from {{para|pos1}} and {{para|lit1}}, which are used to specify an overall part of speech (which applies to all parts of the affix, as opposed to applying to just one element of the expression) or a literal definition for the entire expression (instead of just for one element of the expression). The built-in parameter groups are as follows: {|class="wikitable" ! Group !! Group meaning !! Parameter !! Parameter meaning !! Default properties |- | rowspan=11| `link` | rowspan=11| link parameters; same as those available on {{tl|l}}, {{tl|m}} and other linking templates | `alt` || display text, overriding the term's display form || — |- | `t` || gloss (translation) of a non-English term || {item_dest = "gloss"} |- | `gloss` || gloss (translation); same as `t` || {alias_of = "t"} |- | `tr` || transliteration of a non-Latin-script term; only needed if the automatic transliteration is incorrect or unavailable (e.g. in Hebrew, which doesn't have automatic transliteration) || — |- | `ts` || transcription of a non-Latin-script term, if the transliteration is markedly different from the actual pronunciation; should not be used for IPA pronunciations || — |- | `g` || comma-separated list of genders; whitespace may surround the comma and will be ignored || {item_dest = "genders", type = "genders"} |- | `pos` || part of speech for the term || — |- | `ng` || arbitrary non-gloss descriptive text for the term || — |- | `lit` || literal meaning (translation) of the term || — |- | `id` || a sense ID for the term, which links to anchors on the page set by the {{tl|senseid}} template || — |- | `sc` || the script code (see [[Wiktionary:Scripts]]) for the script that the term is written in; rarely necessary, as the script is autodetected (in most cases, correctly) || {separate_no_index = true, type = "script"} |- | rowspan=2| `q` | rowspan=2| left and right normal qualifiers (as displayed using {{tl|q}}) | `q` || left normal qualifier || {separate_no_index = true, type = "qualifier"} |- | `qq` || right normal qualifier || {separate_no_index = true, type = "qualifier"} |- | rowspan=2| `a` | rowspan=2| left and right accent qualifiers (as displayed using {{tl|a}}) | `a` || comma-separated list of left accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `aa` || comma-separated list of right accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | rowspan=2| `l` | rowspan=2| left and right labels (as displayed using {{tl|lb}}, but without categorizing) | `l` || comma-separated list of left labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ll` || comma-separated list of right labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"} |- | `ref` | reference(s) (in the format accepted by [[Module:references]]; see also the documentation for the {{para|ref}} parameter to {{tl|IPA}}) | `ref` || one or more references, in the format accepted by [[Module:references]] || {item_dest = "refs", type = "references"} |- | `lang` | language for an individual term (provided for compatibility; it is preferred to specify languages for individual terms using language prefixes instead) | `lang` || language code (see [[Wiktionary:Languages]]) for the term || {require_index = true, type = "language"} |} ]==] function export.construct_param_mods(specs) local param_mods = {} local default_specs = {} for _, spec in ipairs(specs) do verify_well_constructed_spec(spec) if spec.default then -- This will have an extra `default` field in it, but it will be erased by merge_param_mod_settings() default_specs = spec else if spec.group then local groups = spec.group if type(groups) ~= "table" then groups = {groups} end local include_set if spec.include then include_set = list_to_set(spec.include) end local exclude_set if spec.exclude then exclude_set = list_to_set(spec.exclude) end for _, group in ipairs(groups) do local group_specs = recognized_param_mod_groups[group] if not group_specs then internal_error(("Unrecognized built-in param mod group '%s'"):format(group), spec) end for group_param, group_param_settings in pairs(group_specs) do local include_param if include_set then include_param = include_set[group_param] elseif exclude_set then include_param = not exclude_set[group_param] else include_param = true end if include_param then local merged_settings = merge_param_mod_settings(merge_param_mod_settings( param_mods[group_param] or default_specs, group_param_settings), spec) param_mods[group_param] = merged_settings end end end end if spec.param then local params = spec.param if type(params) ~= "table" then params = {params} end for _, param in ipairs(params) do local settings = merge_param_mod_settings(param_mods[param] or default_specs, spec) -- If this parameter is an alias of another parameter, we need to copy the specs from the other -- parameter, since parse_inline_modifiers() doesn't know about `alias_of` and having the specs -- duplicated won't cause problems for [[Module:parameters]]. We also need to set `item_dest` to -- point to the `item_dest` of the aliasee (defaulting to the aliasee's value itself), so that -- both modifiers write to the same location. Note that this works correctly in the common case of -- <t:...> with `item_dest = "gloss"` and <gloss:...> with `alias_of = "t"`, because both will end -- up with `item_dest = "gloss"`. local aliasee = settings.alias_of if aliasee then local aliasee_settings = param_mods[aliasee] if not aliasee_settings then internal_error(("Undefined aliasee '%s'"):format(aliasee), spec) end for k, v in pairs(aliasee_settings) do if settings[k] == nil then settings[k] = v end end if settings.item_dest == nil then settings.item_dest = aliasee end end param_mods[param] = settings end end end end return param_mods end -- Return true if `k` is a "built-in" (specially recognized) key in a `param_mod` specification. All other keys -- are forwarded to the structure passed to [[Module:parameters]]. local function param_mod_spec_key_is_builtin(k) return k == "item_dest" or k == "overall" or k == "store" end --[==[ Convert the properties in `param_mods` into the appropriate structures for use by `process()` in [[Module:parameters]] and store them in `params`. If `overall_only` is given, only store the properties in `param_mods` that correspond to overall (non-item-specific) parameters. Currently this only happens when `separate_no_index` is specified. ]==] function export.augment_params_with_modifiers(params, param_mods, overall_only) if overall_only then for param_mod, param_mod_spec in pairs(param_mods) do if overall_only == "always" or param_mod_spec.separate_no_index then local param_spec = {} for k, v in pairs(param_mod_spec) do if k ~= "separate_no_index" and k ~= "require_index" and not param_mod_spec_key_is_builtin(k) then param_spec[k] = v end end params[param_mod] = param_spec end end else local list_with_holes -- Add parameters for each term modifier. for param_mod, param_mod_spec in pairs(param_mods) do local param_spec for k, v in pairs(param_mod_spec) do if not param_mod_spec_key_is_builtin(k) then if param_spec == nil then param_spec = {list = true} end param_spec[k] = v end end if param_spec == nil then if list_with_holes == nil then list_with_holes = {list = true, allow_holes = true} end param_spec = list_with_holes elseif param_spec.alias_of == nil then param_spec.allow_holes = true end params[param_mod] = param_spec end end end --[==[ Return true if `k`, a key in an item, refers to a property of the item (is not one of the specially stored values). Note that `lang` and `sc` are considered properties of the item, although `lang` is set when there's a language prefix and both `lang` and `sc` may be set from default values specified in the `data` structure passed into `parse_list_with_inline_modifiers_and_separate_params()` and `parse_term_with_inline_modifiers_and_separate_params()`. If you don't want these treated as property keys, you need to check for them yourself. ]==] function export.item_key_is_property(k) return k ~= "term" and k ~= "termlang" and k ~= "termlangs" and k ~= "itemno" and k ~= "orig_index" and k ~= "separator" end -- Fetch the argument in `args` corresponding to `index_or_value`, which may be a string of the form "foo.default" -- (requesting the value of `args["foo"].default`); a string or number (requesting the value at that key); a function of -- one argument (`args`), which returns the argument value; or the value itself. Return the resulting value and the -- parameter in `args` that the value came from, or nil if unknown (i.e. a function or direct value was specified). local function fetch_argument(args, index_or_value) if not index_or_value then return index_or_value, nil end local index_or_value_type = type(index_or_value) if index_or_value_type == "string" then if index_or_value:sub(-8) == ".default" then local index_without_default = index_or_value:sub(1, -9) local arg_obj = fetch_argument(args, index_without_default) if type(arg_obj) ~= "table" then internal_error(("Requested that the '.default' key of argument `%s` be fetched, but argument value is undefined or not a table"): format(index_without_default), arg_obj) end return arg_obj.default, index_without_default end if index_or_value:match("^%d+$") then index_or_value = tonumber(index_or_value) end return args[index_or_value], index_or_value elseif index_or_value_type == "number" then return args[index_or_value], index_or_value elseif is_callable(index_or_value) then return index_or_value(args), nil end return index_or_value, nil end function export.generate_obj_maybe_parsing_lang_prefix(data) return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix(data) end -- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and -- parse_term_with_inline_modifiers_and_separate_params(), validating certain argument-related fields that are shared -- among the two functions. local function validate_argument_related_fields(data) if not data.termarg then internal_error("`data.termarg` must be given, indicating which argument contains the terms to be parsed", data) end if not data.param_mods then internal_error("`data.param_mods` must be given, indicating the allowed inline modifiers and separate " .. "parameters to copy", data) end local subitem_param_handling = data.subitem_param_handling or "only" if subitem_param_handling ~= "only" and subitem_param_handling ~= "first" and subitem_param_handling ~= "last" then internal_error("Unrecognized value for `data.subitem_param_handling`, should be 'first', 'last' or 'only'", subitem_param_handling) end if data.raw_args then if data.processed_args then internal_error("Only one of `data.raw_args` and `data.processed_args` can be specified", data) end if not data.params then internal_error("When `data.raw_args` is specified, so must `data.params`, so that the raw arguments " .. "can be parsed", data) end if data.params[data.termarg] == nil then internal_error("There must be a spec in `data.params` corresponding to `data.termarg`", data) end else if not data.processed_args then internal_error("Either `data.raw_args` or `data.processed_args` must be specified", data) end if data.params then internal_error("When `data.processed_args` is specified, `data.params` should not be specified", data) end end end local function argval_missing(val) return val == nil or type(val) == "table" and next(val) == nil end -- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and -- parse_term_with_inline_modifiers_and_separate_params(). After parsing inline modifiers, copy the separate parameters -- to the generated object (or to the appropriate subobject if there are multiple). `data` contains the following -- fields: -- -- `args`: The separate-parameter argument structure. -- `param_mods`: The structure describing the inline modifiers. -- `itemno`: The logical item number of the term being processed, or nil if there's only a single term. -- `termobj`: The object to store the inline modifiers into. If there are subitems, they are in the `terms` field; -- otherwise the properties are stored directly into `termobj`. -- `has_subitems`: True if there are subitems. -- `subitem_separator_map`: If `has_subitems` and this is specified, controls the assignment of the `separator` field -- in subitems. If not specified or a delimiter is not in the map, it is copied unchanged. -- `lang`: Language object to store into all items. -- `sc`: Script object to store into all items, or nil. -- `subitem_param_handling`: "only", "first" or "last", indicating what to do if there are multiple subitems. -- `allow_conflicting_inline_mods_and_separate_params`: If true, specifying a value for both an inline modifier and -- corresponding separate parameter is allowed, and the inline modifier takes precedence. Otherwise, an error -- occurs. -- `postprocess_termobj`: Optional function called on all items at the end, to do any postprocessing. Called with one -- argument, the object to postprocess. -- `no_show_decorations`: If true, don't automatically set {show_decorations = true} on the object or subobject if there -- are decorations (i.e. qualifiers, labels or references) specified for the object. local function copy_separate_params_to_termobj_and_postprocess(data) local args, param_mods, itemno, termobj = data.args, data.param_mods, data.itemno, data.termobj local function set_lang_and_sc(termobj) -- Set these after parsing inline modifiers, not in generate_obj(), otherwise we'll get an error in -- parse_inline_modifiers() if we try to use <lang:...> or <sc:...> as inline modifiers. termobj.lang = termobj.lang or data.lang termobj.sc = termobj.sc or data.sc end local function set_show_decorations(termobj) -- Need to set after parsing inline modifiers. if not data.no_show_decorations and (termobj.q or termobj.qq or termobj.a or termobj.aa or termobj.l or termobj.ll or termobj.refs) then termobj.show_decorations = true end end local function fetch_separate_param(args, paramkey, itemno) local argval = args[paramkey] -- Careful with argument values that may be `false`. if argval and itemno then argval = argval[itemno] end return argval end -- Copy separate parameters to a given object. local function copy_separate_params_to_termobj(fetch_destobj) for param_mod, param_mod_spec in pairs(param_mods) do local dest = param_mod_spec.item_dest or param_mod -- Don't do anything with the `sc` param, which will get overwritten below; we don't -- want it to cause an error if there are multiple subitems. if dest ~= "sc" then local argval = fetch_separate_param(args, param_mod, itemno) if not argval_missing(argval) then local destobj = fetch_destobj(param_mod, param_mod_spec, dest) -- Don't overwrite a value already set by an inline modifier. if argval_missing(destobj[dest]) then destobj[dest] = argval elseif not data.allow_conflicting_inline_mods_and_separate_params then error(("Can't specify a value for separate parameter %s%s= because there is " .. "already an inline modifier <%s:...> specifying a value for the term"):format( param_mod, itemno or "", param_mod)) end end end end end if data.has_subitems then -- If there are any separate indexed parameters, we need to copy them to the first, last or only -- subitem, depending on the value of `data.subitem_param_handling` (which defaults to 'only', -- meaning it's an error if there are multiple subitems). Do this before calling -- postprocess_termobj() because the latter sets .lang and .sc and we want the user to be able to -- set separate langN= and scN= parameters. -- If there was no term, `termobj.terms` will not exist; make it exist to make the callers' lives easier. if not termobj.terms then termobj.terms = {} end -- Compute whether any of the separate indexed params exist for this index. local any_param_at_index for param_mod in pairs(param_mods) do local argval = fetch_separate_param(args, param_mod, itemno) if not argval_missing(argval) then any_param_at_index = true break end end -- If there was no term, but there's a separate parameter, we need to create an empty subitem. if any_param_at_index and not termobj.terms[1] then termobj.terms[1] = {} end local function fetch_destobj(param_mod, param_mod_spec, dest) if param_mod_spec.overall then return termobj end if data.subitem_param_handling == "only" and termobj.terms[2] then error(("Can't specify a value for separate parameter %s%s= because there are " .. "multiple subitems (%s) in the term; use an inline modifier"):format( param_mod, itemno or "", #termobj.terms)) end local termind -- q/a/l need to go at the beginning and qq/aa/ll/refs at the end, regardless; otherwise, respect -- `data.subitem_param_handling`. if dest == "q" or dest == "a" or dest == "l" then termind = 1 elseif dest == "qq" or dest == "aa" or dest == "ll" or dest == "refs" then termind = #termobj.terms elseif data.subitem_param_handling == "only" or data.subitem_param_handling == "first" then termind = 1 else termind = #termobj.terms end return termobj.terms[termind] end copy_separate_params_to_termobj(fetch_destobj) for i, subitem in ipairs(termobj.terms) do set_lang_and_sc(subitem) set_show_decorations(subitem) if subitem.delimiter then subitem.separator = i == 1 and "" or data.subitem_separator_map and data.subitem_separator_map[subitem.delimiter] or subitem.delimiter end if data.postprocess_termobj then data.postprocess_termobj(subitem, data) end end else -- Copy all the parsed term-specific parameters into `termobj`. copy_separate_params_to_termobj(function(param_mod, dest) return termobj end) set_lang_and_sc(termobj) set_show_decorations(termobj) if data.postprocess_termobj then data.postprocess_termobj(termobj, data) end end end local function postprocess_termobj(item, data) if not (data.disallow_custom_separators or data.use_semicolon) then if data.has_subitems and item.separator and item.separator:find(",", nil, true) then data.use_semicolon = true else -- If the displayed term (from .term/etc. or .alt) has an embedded comma, use a semicolon to -- join the terms. local term_text = item[data.term_dest] or item.alt if term_text and term_text:find(",", nil, true) then data.use_semicolon = true end end end end --[==[ Parse a list of terms, each of which may have properties specified using inline modifiers or separate parameters. This function is intended for parsing the arguments of templates like {{tl|syn}}, {{tl|ant}} and related ''*nym'' templates; alternative-form templates {{tl|alt}}/{{tl|alter}}; affix templates like {{tl|af}}/{{tl|affix}}, {{tl|com}}/{{tl|compound}}, etc.; affix usex templates like {{tl|afex}}/{{tl|affixusex}}; name templates like {{tl|name translit}}; column templates like {{tl|col}}; pronunciation templates like {{tl|rhyme}}/{{tl|rhymes}} and {{tl|hmp}}/{{tl|homophones}}; etc. In these templates there are one or more terms specified using numeric parameters, and associated separate parameters specifying per-term properties such as {{para|t1}}, {{para|t2}}, {{para|t3}}, ... for the gloss of the first, second, third, ... term respectively. All such properties can also be specified through inline modifiers attached directly to each term (`<t:...>`, `<pos:...>`, etc.). Normally it is an error if both an inline modifier and separate parameter for the same value are given, but this can be overridden (in which case inline modifiers take precedence over separate parameters when both occur). For an example of a typical workflow involving this function, see the comment at the top of this file. Some notable properties of this function: # Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]], containing only the term list itself along with any other parameters that are '''not''' term properties (such as a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is ''augmented'' with list parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the latter case where raw argument processing is done by the caller, they must build the partial `params` structure; augment it themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the processed arguments. In both cases, the return value of this function contains three values: a list of objects, one per term, specifying the term and all properties; the processed arguments structure, so that the non-term-property arguments can be processed as appropriate; and an object containing miscellaneous global computed properties (currently only `use_semicolon`; see below). # Optionally, each term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality is used, for example, by {{tl|col}} and variants, which allow each row to have comma-separated or tilde-separated subitems. When this feature is invoked, the format of the per-term object changes; instead of directly being an object describing the term and its properties, it is an object with a `terms` field containing a list of per-subitem objects along with other top-level fields describing per-term properties. By default, if there are separate parameters specified along with multiple subitems, an error occurs, but this is controllable; currently, you can request that the parameters be assigned to the first or last subitem. # By default, special ''separator'' arguments may be present, mixed in among regular term arguments. Examples of such separator arguments are (by default; this can be overridden) a bare semicolon, specifying that the terms on either side should be separated by a semicolon instead of a comma (indicating a higher-level grouping); a bare tilde, replacing the comma separator with a tilde (indicating that the terms on either side are alternants); and a bare underscore, replacing the comma separator with a space. Separator arguments are ignored when numbering the separate parameters. You disable the separator argument handling entirely if it doesn't make sense to have this (e.g. in {{tl|af}}/{{tl|affix}}, where the separator is always a {{cd|+}} sign). `data` is an object containing several possible fields. 1. Fields that are required or recommended (usually related to argument processing): * `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from {frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw arguments. * `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both) of `raw_args` and `processed_args` must be set. * `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified manually. * `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters, '''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically "augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the "augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by [[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be correctly handled.) * `termarg` ('''required'''): The argument containing the first item with attached inline modifiers to be parsed. Usually a numeric value such as {1} or {2}. * `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used internally to track pages containing template invocations with certain properties. Example properties tracked are missing items with corresponding properties as well as missing items without corresponding properties (which are skipped entirely). To find out the exact properties tracked and the name of the tracking pages, read the code. * `lang` ('''recommended'''): The language object for the language of the items, or the name of the argument to fetch the object from. It is not strictly necessary to specify this, as this function only initializes items based on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used for certain purposes: *# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language prefix). *# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same language code may appear in several items, such as with {{tl|col}} and related templates). The value of `lang` can be any of the following: * If a string of the form "foo.default", it is assumed to be requesting the value of `args["foo"].default`. * Otherwise, if a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching (this also happens with a spec like "2.default"). * Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed to the function as its only argument. * Otherwise, it is used directly. * `sc` ('''recommended'''): The script object for the items, or the name of the argument to fetch the object from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not strictly necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned items if not otherwise set (e.g. by the {{para|sc<var>N</var>}} parameter or `<sc:...>` inline modifier). The most common value is {"sc.default"}. 2. Other argument-related fields: * `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed arguments. It should make modifications in-place. * `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language prefixes are stripped off. Defaults to {"term"}. * `pre_normalize_modifiers`: As in `parse_inline_modifiers()`. * `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence. 3. Fields related to language prefixes: * `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in [[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the `termlangs` field, and the first specified object is stored in the `lang` field). * `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above. * `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an error and will tend to remain unfixed. 4. Fields related to custom/special separators: * `disallow_custom_separators`: If specified, disallow specifying custom separators (semicolon, underscore, tilde; see the internal `default_special_separators` table, or the `special_separators` field) as an item value to override the default separator. By default, the previous separator of each item is considered to be an empty string (for the first item) and otherwise the value of the field `default_separator` (normally a comma + space), unless either the preceding item is one of the values listed in `special_separators`, such as a bare semicolon (which causes the following item's previous separator to be a semicolon + space) or an item has an embedded comma in it (which causes ''all'' items other than the first to have their previous separator be a semicolon + space). The previous separator of each item is set on the item's `separator` property. Bare semicolons and other separator arguments do not count when indexing items using separate parameters. For example, the following is correct: ** {{tl|template|lang|item 1|q1=qualifier 1|;|item 2|q2=qualifier 2}} If `disallow_custom_separators` is specified, however, the `separator` property is not set and separator arguments are not recognized. * `default_separator`: Override the default separator (normally {", "}). * `special_separators`: Table giving the special/custom separators that can be given, and how they should display. If not specified, the default in `default_special_separators` is used. This is a table mapping separator values (such as {"~"}) to the corresponding display string (such as {" ~ "}). 5. Fields related to multiple subitems in a given term: * `splitchar`: A Lua pattern. If specified, each user-specified argument can consist of multiple delimiter-separated subitems, each of which may be followed by inline modifiers. In this case, each element in the returned list of items is no longer an object describing an item, but instead an object with a `terms` field, whose value is a list describing the subitems (whose format is the same as the normal format of an item in the top-level list when `splitchar` is not specified). Each subitem object will have a `delimiter` field holding the actual delimiter occurring before the subitem, which is useful in the case where `splitchar` matches multiple possible characters. In this case, it is possible to specify that a given modifier can only occur after the last subitem and effectively modifies the whole collection of subitems by setting {overall = true} on the modifier. In this case, the modifier's value will be stored in the top-level object (the object with the `terms` field specifying the subitems). Note that splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). In addition, the algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers or inside of square brackets, which do not trigger splitting (whether or not contained within protected sequences). * `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in [[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified (see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that comma+whitespace sequences won't be split. * `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems. The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored into all subitems. * `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator` field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar` are recognized. Delimiters not in this map display as-is. 6. Other fields: * `dont_skip_items`: Normally, items that are completely unspecified (have no term and no properties) are skipped and not inserted into the returned list of items. (Such items cannot occur if {disallow_holes = true} is set on the term specification in the `params` structure passed to `process()` in [[Module:parameters]]. It is generally recommended to do so unless a specific meaning is associated the term value being missing.) If `dont_skip_items` is set, however, items are never skipped, and completely unspecified items will be returned along with others. (They will not have the term or any properties set, but will have the normal non-property fields set; see below.) * `stop_when`: If specified, a function to determine when to prematurely stop processing items. It is passed a single argument, an object containing the following fields: ** `term`: The raw term, prior to parsing off language prefixes and inline modifiers (since the processing of `stop_when` happens before parsing the term). ** `any_param_at_index`: True if any separate property parameters exist for this item. ** `orig_index`: Same as `orig_index` below. ** `itemno`: Same as `itemno` below. ** `stored_itemno`: The index where this item will be stored into the returned items table. This may differ from `itemno` due to skipped items (it will never be different if `dont_skip_items` is set). The function should return true to stop processing items and return the ones processed so far (not including the item currently being processed). This is used, for example, in [[Module:alternative forms]], where an unspecified item signal the end of items and the start of labels. * `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set, causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem. If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations. Three values are returned: the list of items; the processed `args` structure; and an object of miscellaneous computed global values (currently only `use_semicolon`, indicating that commas were found in individual arguments and so the default separator should be a semicolon). In each returned item, there will be one field set for each specified property (either through inline modifiers or separate parameters). If subitems are not allowed, each item directly has fields set on it for the specified properties. If subitems ''are'' allowed, each item contains a `terms` field, which is a list of subitem objects, each of which has fields set on it for the specified properties of that subitem. In addition, the following fields may be set on each item or subitem: * `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given. * `orig_index`: The original index into the item in the items table returned by `process()` in [[Module:parameters]]. This may differ from `itemno` if there are raw semiclons and `disallow_custom_separators` is not given. * `itemno`: The logical index of the item. The index of separate parameters corresponds to this index. This may be different from `orig_index` in the presence of raw semicolons; see above. * `termlang`: If there is a language prefix, the corresponding language object is stored here (only if `parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set). * `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are set, the list of corresponding language objects is stored here. * `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified; (c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value. * `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b) `sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value. * `delimiter`: If subitems are allowed, this is set on subitems and specifies the delimiter used prior to the given subitem (e.g. {","}). * `separator`: The separator to display before the item. Always set on subitems, and set on top-level items if `disallow_custom_separators` is not given. Controlled by `special_separators` (for top-level items) and `subitem_separator_map` (for subitems). * `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified, {show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is called. Use {no_show_decorations = true} to prevent this. ]==] function export.parse_list_with_inline_modifiers_and_separate_params(data) validate_argument_related_fields(data) local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods if raw_args then local params = data.params local termarg_spec = params[termarg] if termarg_spec == true or not termarg_spec.list then internal_error("Term spec in `data.params` must have `list` set", termarg_spec) end if termarg_spec == true or not (termarg_spec.allow_holes or termarg_spec.disallow_holes) then internal_error("Term spec in `data.params` must have either `allow_holes` or `disallow_holes` set", termarg_spec) end export.augment_params_with_modifiers(params, param_mods) args = process_params(raw_args, params) else args = data.processed_args end local process_args_before_parsing = data.process_args_before_parsing if process_args_before_parsing then process_args_before_parsing(args) end -- Find the maximum index among any of the list parameters. local term_args = args[termarg] -- As a special case, the term args might not have a `maxindex` field because they might have -- been declared with `disallow_holes = true`, so fall back to the actual length of the list -- using the table_len function, since # can be unpredictable with arbitrary tables. local maxmaxindex = term_args.maxindex or table_len(term_args) for _, v in pairs(args) do if type(v) == "table" and v.maxindex and v.maxindex > maxmaxindex then maxmaxindex = v.maxindex end end local special_separators = data.special_separators or export.default_special_separators local items, lang_cache, use_semicolon = {}, data.lang_cache or {} local lang = fetch_argument(args, data.lang) if lang then lang_cache[lang:getCode()] = lang end local sc = fetch_argument(args, data.sc) local term_dest = data.term_dest or "term" -- FIXME: this is vulnerable to abusive inputs like 1000000=. local itemno = 0 for i = 1, maxmaxindex do local term = term_args[i] if data.disallow_custom_separators or not special_separators[term] then itemno = itemno + 1 -- Compute whether any of the separate indexed params exist for this index. local any_param_at_index for param_mod in pairs(param_mods) do local argval = args[param_mod] -- Careful with argument values that may be `false`. if argval then argval = argval[itemno] end if not argval_missing(argval) then any_param_at_index = true break end end if data.stop_when and data.stop_when{ term = term, -- FIXME, we should just pass in `any_param_at_index` directly. any_param_at_index = term ~= nil or any_param_at_index, orig_index = i, itemno = itemno, stored_itemno = #items + 1, } then break end -- If any of the params used for formatting this term is present, create a term and add it to the list. if not data.dont_skip_items and term == nil and not any_param_at_index then track("skipped-term", data.track_module) else if not term then track("missing-term", data.track_module) end local termobj = { itemno = itemno, orig_index = i, } if not data.disallow_custom_separators then termobj.separator = i == 1 and "" or special_separators[term_args[i - 1]] end -- Add 1 because first term index starts at 2. local paramname = termarg + i - 1 if term then local function generate_obj(term, parse_err) return export.generate_obj_maybe_parsing_lang_prefix { term = term, termobj = data.splitchar and {} or termobj, term_dest = term_dest, paramname = paramname, parse_lang_prefix = data.parse_lang_prefix, parse_err = parse_err, allow_bad_lang_prefix = data.allow_bad_lang_prefix, allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes, lang_cache = lang_cache, } end parse_inline_modifiers(term, { paramname = paramname, param_mods = param_mods, generate_obj = generate_obj, splitchar = data.splitchar, preserve_splitchar = true, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, outer_container = data.splitchar and termobj or nil, pre_normalize_modifiers = data.pre_normalize_modifiers, }) end -- FIXME: Make into an error, then remove after a month. if data.no_show_qualifiers then track("no_show_qualifiers") end local term_data = { args = args, param_mods = param_mods, itemno = itemno, termobj = termobj, term_dest = term_dest, has_subitems = not not data.splitchar, lang = lang, -- As a special case, if the caller defined a scN= separate param, set it on all subitems if there -- are multiple, falling back to the overall sc= param. sc = args.sc and args.sc[itemno] or sc, subitem_param_handling = data.subitem_param_handling, subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map, allow_conflicting_inline_mods_and_separate_params = data.allow_conflicting_inline_mods_and_separate_params, postprocess_termobj = postprocess_termobj, disallow_custom_separators = data.disallow_custom_separators, use_semicolon = use_semicolon, no_show_decorations = data.no_show_decorations or data.no_show_qualifiers, } copy_separate_params_to_termobj_and_postprocess(term_data) use_semicolon = term_data.use_semicolon insert(items, termobj) end end end if not data.disallow_custom_separators then -- Set the default separator of all those items for which a separator wasn't explicitly given to the default -- separator, defaulting to comma + space; but if any items have embedded commas, set the separator to -- semicolon + space. for _, item in ipairs(items) do if not item.separator then item.separator = use_semicolon and "; " or data.default_separator or ", " end end end return items, args, {use_semicolon = use_semicolon} end --[==[ Parse a single term that may have properties specified through inline modifiers or separate parameters. This differs from `parse_list_with_inline_modifiers_and_separate_params()` in that the latter is for parsing a list of terms, each of which may have properties specified through inline modifiers or separate parameters. Both functions optionally support having multiple subitems in a single term. This function is used e.g. for form-of templates ({{tl|inflection of}}/{{tl|infl of}}, {{tl|form of}}, and specific templates such as {{tl|alt form}}/{{tl|alternative form of}}, {{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others); for etymology templates ({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates like {{tl|ellipsis}}, {{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and for other templates with an argument structure similar to {{tl|l}} or {{tl|m}}. In these templates there is a term specified using a numeric parameter and associated separate parameters specifying term properties such as {{para|t}} for the gloss or {{para|tr}} for manual transliteration. All such properties can also be specified through inline modifiers attached directly to each term (`<t:...>`, `<tr:...>`, etc.). Normally it is an error if both an inline modifier and separate parameter for the same value are given, but this can be overridden (in which case inline modifiers take precedence over separate parameters when both occur). Some notable properties of this function: # Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]], containing only the term list itself along with any other parameters that are '''not''' term properties (such as a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is ''augmented'' with parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the latter case where raw argument processing is done by the caller, they must build the partial `params` structure; augment it themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the processed arguments. In both cases, the return value of this function contains two values, an object specifying the term and all properties; and the processed arguments structure, so that the non-term-property arguments can be processed as appropriate. # Optionally, the term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality is used, for example, by form-of templates. When this feature is invoked, the format of the term object changes; instead of directly being an object describing the term and its properties, it is an object with a `terms` field containing a list of per-subitem objects along with other top-level fields describing per-term properties. By default, if there are separate parameters specified along with multiple subitems, an error occurs, but this is controllable; currently, you can request that the parameters be assigned to the first or last subitem. `data` is an object containing several possible fields. 1. Fields that are required or recommended (usually related to argument processing): * `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from {frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw arguments. * `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both) of `raw_args` and `processed_args` must be set. * `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified manually. * `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters, '''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically "augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the "augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by [[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be correctly handled.) * `termarg` ('''required'''): The argument containing the item with attached inline modifiers to be parsed. Usually a numeric value such as {1} or {2}. * `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used internally to track pages containing template invocations with certain properties. * `lang` ('''recommended'''): The language object for the language of the item or subitems, or the name of the argument to fetch the object from. It is not strictly necessary to specify this, as this function only initializes items based on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used for certain purposes: *# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language prefix). *# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same language code may appear in several subitems). The value of `lang` can be any of the following: * If a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching. * Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed to the function as its only argument. * Otherwise, it is used directly. * `sc` ('''recommended'''): The script object for the item or subitems, or the name of the argument to fetch the object from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not strictly necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned items if not otherwise set (e.g. by the {{para|sc}} parameter or `<sc:...>` inline modifier). The most common value is {"sc"}. * `make_separate_g_into_list`: Set this to {true} if separate gender parameters exist are are specified using {{para|g}}, {{para|g2}}, etc. instead of using a single comma-separated {{para|g}} field. 2. Other argument-related fields: * `adjust_params_before_arg_processing`: An optional function to further adjust the `params` structure prior to calling `process()` in [[Module:parameters]]. This should be used when there are mismatches between the format of a given property as an inline modifier and the corresponding property as a separate parameter (as with the {{para|g}} parameter and {{cd|<g:...>}} modifier, but this particular case is handled by the `make_separate_g_into_list` field). * `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed arguments. It should make modifications in-place. * `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language prefixes are stripped off. Defaults to {"term"}. * `pre_normalize_modifiers`: As in `parse_inline_modifiers()`. * `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence. * `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set, causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem. If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations. 3. Fields related to language prefixes: * `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in [[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the `termlangs` field, and the first specified object is stored in the `lang` field). * `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above. * `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an error and will tend to remain unfixed. 4. Fields related to multiple subitems in the term: * `splitchar`: A Lua pattern. If specified, the user-specified argument can consist of multiple delimiter-separated subitems, each of which may be followed by inline modifiers. In this case, the first returned value is no longer an object describing the item, but instead an object with a `terms` field, whose value is a list describing the subitems (whose format is the same as the normal format of the item when `splitchar` is not specified). Each subitem object will have a `delimiter` field holding the actual delimiter occurring before the subitem, which is useful in the case where `splitchar` matches multiple possible characters. In this case, it is possible to specify that a given modifier can only occur after the last subitem and effectively modifies the whole collection of subitems by setting `overall = true` on the modifier. In this case, the modifier's value will be stored in the top-level object (the object with the `terms` field specifying the subitems). Note that splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). In addition, the algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers or inside of square brackets, which do not trigger splitting (whether or not contained within protected sequences). * `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in [[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified (see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that comma+whitespace sequences won't be split. * `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems. The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored into all subitems. * `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator` field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar` are recognized. Delimiters not in this map display as-is. Two values are returned, an object describing the item (or subitems) and the processed `args` structure. In the returned item, there will be one field set for each specified property (either through inline modifiers or separate parameters). If subitems are not allowed, the item directly has fields set on it for the specified properties. If subitems ''are'' allowed, the item contains a `terms` field, which is a list of subitem objects, each of which has fields set on it for the specified properties of that subitem. In addition, the following fields may be set on the item or each subitem: * `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given. * `termlang`: If there is a language prefix, the corresponding language object is stored here (only if `parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set). * `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are set, the list of corresponding language objects is stored here. * `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified; (c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value. * `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b) `sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value. * `delimiter`: If subitems are allowed, this specifies the delimiter used prior to the given subitem (e.g. {","}). * `separator`: If subitems are allowed, this specifies the displayed form of the delimiter to be shown before a given subitem. The mapping from user-specified delimiters to displayed separators is handled by `subitem_separator_map`; see above. The first subitem always has a blank string in the `separator` field. * `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified, {show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is called. Use {no_show_decorations = true} to prevent this. ]==] function export.parse_term_with_inline_modifiers_and_separate_params(data) validate_argument_related_fields(data) local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods if raw_args then local params = data.params local termarg_spec = params[termarg] if type(termarg_spec) == "table" and termarg_spec.list then internal_error("Term spec in `data.params` must not have `list` set", termarg_spec) end export.augment_params_with_modifiers(params, param_mods, "always") if data.make_separate_g_into_list then -- HACK: g= is a list for compatibility, but sublist as an inline parameter. params.g = {list = true, item_dest = "genders", type = "genders", flatten = true} end local adjust_params_before_arg_processing = data.adjust_params_before_arg_processing if adjust_params_before_arg_processing then adjust_params_before_arg_processing(params) end args = process_params(raw_args, params) else args = data.processed_args end local process_args_before_parsing = data.process_args_before_parsing if process_args_before_parsing then process_args_before_parsing(args) end local term, lang_cache = args[termarg], data.lang_cache local lang = fetch_argument(args, data.lang) if lang and lang_cache then lang_cache[lang:getCode()] = lang end local sc = fetch_argument(args, data.sc) local term_dest = data.term_dest or "term" if not term then track("missing-term", data.track_module) end local termobj, splitchar = {}, data.splitchar if term then local function generate_obj(term, parse_err) return export.generate_obj_maybe_parsing_lang_prefix { term = term, termobj = splitchar and {} or termobj, term_dest = term_dest, paramname = termarg, parse_lang_prefix = data.parse_lang_prefix, parse_err = parse_err, allow_bad_lang_prefix = data.allow_bad_lang_prefix, allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes, lang_cache = lang_cache, } end parse_inline_modifiers(term, { paramname = termarg, param_mods = param_mods, generate_obj = generate_obj, splitchar = splitchar, preserve_splitchar = true, escape_fun = data.escape_fun, unescape_fun = data.unescape_fun, outer_container = splitchar and termobj or nil, pre_normalize_modifiers = data.pre_normalize_modifiers, }) end -- FIXME: Make into an error, then remove after a month. if data.no_show_qualifiers then track("no_show_qualifiers") end copy_separate_params_to_termobj_and_postprocess { args = args, param_mods = param_mods, termobj = termobj, has_subitems = not not splitchar, lang = lang, sc = sc, subitem_param_handling = data.subitem_param_handling, subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map, allow_conflicting_inline_mods_and_separate_params = data.allow_conflicting_inline_mods_and_separate_params, no_show_decorations = data.no_show_decorations or data.no_show_qualifiers, } if splitchar and termobj.terms[2] then track("parse-term-multiple-subitems", data.track_module) track("parse-term-multiple-subitems") end return termobj, args end return export 1kp3k2x2ejn7z1nhgqsp5b8fxr2qs7w Modul:etymon 828 57903 375383 373456 2026-09-22T07:06:30Z Hakimi97 2668 Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92736184|92736184]]) 375383 Scribunto text/plain --[=[ This module implements the {{etymon}} template for structured etymology data on Wiktionary. It enables the creation of etymology trees and text by parsing etymon chains, scraping linked pages for their own {{etymon}} data, and recursively building a tree of derivational relationships. Authors: - Original implementation: [[User:Ioaxxere]] - Full refactor (September 2025): [[User:Fenakhay]] ([[Special:Diff/86717746]]) Modules: - [[Module:etymon]]: main module handling parsing, validation, tree building, and page scraping - [[Module:etymon/data]]: keyword definitions, configuration, and status constants - [[Module:etymon/tree]]: etymology tree rendering - [[Module:etymon/text]]: etymology text generation - [[Module:etymon/categories]]: category generation logic - [[Module:etymon/tracking]]: tracking ]=] local export = {} local __state = { cached_etymon_args = {}, cached_etymon_pages = {}, cached_descendants_checks = {}, senseid_parent_etymon = {}, available_etymon_ids = {}, single_etymons = {}, entry_title = nil, entry_lang_code = nil, current_page_has_inline_etymology = false, current_page_has_redundant_etymology = false, used_idless_etymon = false, toplevel_has_inline_etymology = false, toplevel_redundant_etymology = false, toplevel_idless_etymon = false, has_mismatched_id = false, linked_page_multiple_etymons_idless = false, linked_page_partial_etymology_sections = false, partial_etymology_targets = {}, skip_partial_etymology_category = false, max_depth_reached = 0, total_nodes = 0, language_count = {}, toplevel_keyword_stats = {}, id_stats = nil, warnings = {}, } local function reset_invocation_state() __state.current_page_has_inline_etymology = false __state.current_page_has_redundant_etymology = false __state.used_idless_etymon = false __state.toplevel_has_inline_etymology = false __state.toplevel_redundant_etymology = false __state.toplevel_idless_etymon = false __state.has_mismatched_id = false __state.linked_page_multiple_etymons_idless = false __state.linked_page_partial_etymology_sections = false __state.max_depth_reached = 0 __state.total_nodes = 0 __state.language_count = {} __state.toplevel_keyword_stats = {} __state.warnings = {} end local M = require("Module:module loader").init({ require = { data = "Module:etymon/data", tree = "Module:etymon/tree", text = "Module:etymon/text", categories = "Module:etymon/categories", tracking = "Module:etymon/tracking", descendants = "Module:etymon/descendants", anchors = "Module:anchors", etydate = "Module:etydate", etymology = "Module:etymology", families = "Module:families", languages = "Module:languages", languages_errorgetby = "Module:languages/errorGetBy", links = "Module:links", pages = "Module:pages", parameters = "Module:parameters", string_utilities = "Module:string utilities", template_parser = "Module:template parser", utilities = "Module:utilities", debug = "Module:debug", en_utilities = "Module:en-utilities", parameter_utilities = "Module:parameter utilities", parse_utilities = "Module:parse utilities", template_styles = "Module:TemplateStyles", script_utilities = "Module:script utilities", JSON = "Module:JSON", yesno = "Module:yesno", }, loadData = { headword_data = "Module:headword/data", parameters_data = "Module:parameters/data", text_allowed = "Module:etymon/data/text_allowed", }, }) local Util = {} function Util.format_error(message, preview_only) if preview_only and not M.pages.is_preview() then return nil end return '<span class="error">' .. message .. '</span>' end function Util.add_warning(message, preview_only) local formatted = Util.format_error(message, preview_only) if formatted then table.insert(__state.warnings, formatted) end end function Util.is_text_param_allowed_for_lang(lang) if not lang or type(lang) ~= "table" then return false end local types = lang.getTypes and lang:getTypes() if types and types.family then local code = lang.getCode and lang:getCode() return code and M.text_allowed.families[code] == true end local full_code = lang.getFullCode and lang:getFullCode() if full_code and M.text_allowed.langs[full_code] then return true end if lang.inFamily then for family_code in pairs(M.text_allowed.families) do if lang:inFamily(family_code) then return true end end end return false end function Util.get_lang(code, no_error) if no_error then return M.languages.getByCode(code, nil, true) end return M.languages.getByCode(code, nil, true) or M.languages_errorgetby.code(code, true, true) end -- Match a term language against a text=:lang stop target (supports etymology-only codes). function Util.lang_matches_stop_code(term_lang, stop_code) if not term_lang or not stop_code or stop_code == "" then return false end local stop_lang = Util.get_lang(stop_code, true) if not stop_lang then return false end if term_lang:getCode() == stop_lang:getCode() then return true end if stop_lang:getFullCode() == stop_lang:getCode() then return term_lang:getFullCode() == stop_lang:getCode() end return false end function Util.get_family(code) return M.families.getByCode(code) end function Util.get_lang_exception(lang) -- Families have no language-specific exceptions if lang.getTypes and lang:getTypes().family then return nil end local code = lang:getCode() local lang_exceptions = M.data.config.lang_exceptions if lang_exceptions[code] then return lang_exceptions[code] end for norm_code, exc in pairs(lang_exceptions) do if exc.normalize_to and code == exc.normalize_to then return exc end if exc.normalize_from_families then local should_normalize = false for _, family in ipairs(exc.normalize_from_families) do if lang:inFamily(family) then should_normalize = true break end end if should_normalize and exc.normalize_exclude_families then for _, family in ipairs(exc.normalize_exclude_families) do if lang:inFamily(family) then should_normalize = false break end end end if should_normalize then local ret = {} for k, v in pairs(exc) do ret[k] = v end ret.suppress_tr = nil return ret end end end return nil end function Util.get_norm_lang(lang) local exc = Util.get_lang_exception(lang) if exc and exc.normalize_to then return M.languages.getByCode(exc.normalize_to) end return lang end function Util.resolve_context_lang(lang, node_args) if type(node_args) ~= "table" then return lang end if node_args.status == M.data.STATUS.INLINE then return lang end if not (lang.hasType and lang:hasType("etymology-only")) then return lang end local full = lang.getFull and lang:getFull() if not full or full:getCode() == lang:getCode() then return lang end if full.hasAncestor and full:hasAncestor(lang) then return lang end return full end -- Add default values for boolean modifiers (e.g., <unc> becomes <unc:1>) -- This is needed because Module:parse utilities expects boolean modifiers to have explicit values function Util.add_boolean_defaults(str, param_mods) local result = str for name, spec in pairs(param_mods) do if spec.type == "boolean" then -- Replace <name> with <name:1> (but not <name:...> which already has a value) result = result:gsub("<" .. name .. ">", "<" .. name .. ":1>") end end return result end local REQUEST_TEMPLATE_PARAM_MODS = { rfe = M.parameter_utilities.construct_param_mods { { param = { "sort", "y", "m", "fragment", "section" } }, { param = { "nocat", "box", "noes" }, type = "boolean" }, }, etystub = M.parameter_utilities.construct_param_mods { { param = "sort" }, { param = { "nocat", "nocap", "nodot" }, type = "boolean" }, }, } function Util.expand_request_template(frame, template_name, param_value, lang_code) local param_mods = REQUEST_TEMPLATE_PARAM_MODS[template_name] local with_defaults = Util.add_boolean_defaults(param_value, param_mods) local parsed = M.parse_utilities.parse_inline_modifiers(with_defaults, { param_mods = param_mods, generate_obj = function(text) if M.yesno(text, false) then return { is_boolean = true } end return { text = text } end, }) local template_args = { [1] = lang_code } for name in pairs(param_mods) do template_args[name] = parsed[name] end if not parsed.is_boolean then template_args[2] = parsed.text end return " " .. frame:expandTemplate({ title = template_name, args = template_args, }) end -- Centralized term formatting: handles suppress_term (-), unknown_term (empty/+), and regular terms function Util.format_term(term, is_toplevel, opts) opts = opts or {} -- suppress_term (-) returns nil if term.suppress_term then return nil end local lang = term.lang local exc = Util.get_lang_exception(lang) if is_toplevel then local display_text = term.alt or term.title or "" local sc = term.sc or lang:findBestScript(display_text) local bold_text = tostring(mw.html.create("strong") :addClass("selflink") :wikitext(display_text)) return M.script_utilities.tag_text(bold_text, lang, sc, "term") end local link_params = { lang = lang } link_params.term = not term.unknown_term and term.title or nil link_params.alt = term.alt link_params.id = (not term.unknown_term and term.id and term.id ~= "") and term.id or nil if not (exc and exc.suppress_tr) then link_params.tr = term.tr link_params.ts = term.ts else link_params.suppress_tr = true end link_params.lit = (opts.lit ~= "suppress") and term.lit or nil if opts.gloss ~= "suppress" then link_params.gloss = term.gloss end link_params.genders = term.genders if opts.pos ~= "suppress" then link_params.pos = term.pos link_params.ng = term.ng link_params.infl = term.infl end if exc and exc.suppress_tr then link_params.lit = nil end if opts.tree_ql ~= "suppress" then if term.q then link_params.q = term.q end if term.qq then link_params.qq = term.qq end if term.l then link_params.l = term.l end if term.ll then link_params.ll = term.ll end link_params.show_decorations = term.q or term.qq or term.l or term.ll end return M.links.full_link(link_params, "term") end local __is_content_page_cached function Util.is_content_page() if __is_content_page_cached == nil then __is_content_page_cached = M.pages.is_content_page(mw.title.getCurrentTitle()) end return __is_content_page_cached end local __page_data_cached function Util.get_page_data() if not __page_data_cached then __page_data_cached = M.headword_data.page end return __page_data_cached end -- Extract base keyword from param (without modifiers) local function get_keyword_base(param) if type(param) ~= "string" then return nil end local base = param:match("^:?([^<]+)") or param:gsub("^:", "") return base end local function is_keyword(param, allow_colon_less) if type(param) ~= "string" then return false end local keywords = M.data.keywords if param:sub(1, 1) == ":" then local base = get_keyword_base(param) return keywords[base] ~= nil end if allow_colon_less then local base = get_keyword_base(param) return keywords[base] ~= nil end return false end local function get_keyword(param, allow_colon_less) if type(param) ~= "string" then return nil end local keywords = M.data.keywords if param:sub(1, 1) == ":" then return get_keyword_base(param) end if allow_colon_less then local base = get_keyword_base(param) if keywords[base] then return base end end return nil end local function normalize_keyword(keyword) if keyword:sub(1, 1) == ":" then return keyword end return ":" .. keyword end -- Resolve keyword (possibly an alias) to its canonical form. Used only at input boundaries local function get_canonical_keyword(keyword) if not keyword then return keyword end return M.data.keyword_canonical[keyword] or keyword end local function is_affix_group_keyword(keyword) local config = keyword and M.data.keywords[keyword] return config and config.affix_categories or false end local function reject_removed_surf_keyword(param) local base = get_keyword_base(param) if base == "surf" then error("The `:surf` keyword has been removed. Use `<surf>` on a formation keyword instead (e.g. `:af<surf>`, `:bor<surf>`).") end end local function copy_keyword_info(source) local copy = {} for k, v in pairs(source) do copy[k] = v end return copy end local function lowercase_glossary_display(text) return text:gsub("(%[%[Appendix:Glossary#[^|]+|)([^%]])([^%]]*)%]%]", function(prefix, first, rest) return prefix .. mw.ustring.lower(first) .. rest .. "]]" end) end local function surf_should_keep_formation_phrase(base) if not base.phrase then return false end if base.glossary then return true end return not (base.phrase == "from" and (base.text == "From" or base.text == "from")) end -- Runtime overrides when <surf> is present on a keyword. local function get_effective_keyword_info(keyword, modifiers) local base = M.data.keywords[keyword] if not base or not modifiers or not modifiers.surf then return base end local effective = copy_keyword_info(base) local surf_text = "By [[Appendix:Glossary#surface_analysis|surface analysis]]," local surf_phrase = "by surface analysis," effective.new_sentence = true effective.invisible = "tree" if surf_should_keep_formation_phrase(base) then effective.phrase = surf_phrase .. " " .. base.phrase if base.text then effective.text = surf_text .. " " .. lowercase_glossary_display(base.text) else effective.text = surf_text .. " " .. base.phrase end else effective.text = surf_text effective.phrase = surf_phrase end return effective end -- Build text/phrase for nominalization with <g:code> (uses data module for codes only). local function get_nominalization_label_for_g(code) if not code or code == "" then return nil end local codes = M.data.nominalization_g_codes local adj = codes[code] if not adj and #code == 2 then local gender_adj = codes[code:sub(1, 1)] local number_adj = codes[code:sub(2, 2)] if gender_adj and number_adj then adj = gender_adj .. " " .. number_adj end end if not adj then return nil end local text = adj:gsub("^%l", function(c) return string.upper(c) end) .. " [[Appendix:Glossary#nominalization|nominalization]] of" local phrase = M.en_utilities.add_indefinite_article(adj .. " [[Appendix:Glossary#nominalization|nominalization]] of", false) return { text = text, phrase = phrase } end local EtymonParser = {} -- Keyword modifier definitions EtymonParser.keyword_param_mods = M.parameter_utilities.construct_param_mods { { group = "ref" }, { param = "conj" }, -- conjunction for alternatives: "and", "or", "and/or", etc. { param = { "unc", "surf" }, type = "boolean" }, { param = "text", restrict = { keywords = { "from", "derived" } } }, { param = "lit", restrict = { affix_group = true } }, { param = "g", restrict = { keywords = { "nominalization" } } }, { param = "senseid", restrict = { keywords = { "semantic loan" } } }, } -- Term modifier definitions EtymonParser.etymon_param_mods = M.parameter_utilities.construct_param_mods { {group = {"link", "q", "l", "ref", "infl"}, exclude = {"sc"}}, {param = {"ety", "postype"}}, {param = "unc", type = "boolean"}, {param = "aftype", restrict = {affix_group = true}}, {param = {"bor", "slbor", "lbor"}, type = "boolean", restrict = {affix_group = true}}, } local function get_clean_param_mods(param_mods) local clean = {} for mod_name, mod_def in pairs(param_mods) do clean[mod_name] = {} for key, value in pairs(mod_def) do if key ~= "restrict" then clean[mod_name][key] = value end end end return clean end function EtymonParser.check_modifier_restrictions(modifiers, current_keyword, param_mods) for mod_name, mod_value in pairs(modifiers) do -- Only check restrictions if the modifier has a non-false/nil value if mod_value then local mod_def = param_mods[mod_name] if mod_def and mod_def.restrict then if mod_def.restrict.affix_group then if not is_affix_group_keyword(current_keyword) then local mod_display = mod_value == true and "<" .. mod_name .. ">" or "<" .. mod_name .. ":" .. tostring(mod_value) .. ">" error("The modifier `" .. mod_display .. "` is only allowed for affix-group keywords (e.g. `:af`, `:blend`, `:univ`).") end elseif mod_def.restrict.keywords then local allowed_keywords = mod_def.restrict.keywords local is_allowed = false for _, allowed_keyword in ipairs(allowed_keywords) do if current_keyword == allowed_keyword then is_allowed = true break end end if not is_allowed then local keyword_list = {} for _, kw in ipairs(allowed_keywords) do table.insert(keyword_list, ":" .. kw) end local keyword_str = table.concat(keyword_list, #keyword_list == 2 and " or " or ", ") if #keyword_list > 2 then -- Replace last comma with "or" keyword_str = keyword_str:gsub(", ([^,]+)$", " or %1") end local mod_display = mod_value == true and "<" .. mod_name .. ">" or "<" .. mod_name .. ":" .. tostring(mod_value) .. ">" error("The modifier `" .. mod_display .. "` is only allowed for the keyword" .. (#keyword_list > 1 and "s " or " ") .. keyword_str .. ".") end end end end end end local TERM_RULE_DISALLOW = { suppress = { field = "suppress_term", label = "suppressed" }, unknown = { field = "unknown_term", label = "unknown" }, family = { field = "is_family", label = "family" }, } function EtymonParser.check_etymon_limits(count, limits, label, opts) if not limits then return end opts = opts or {} local min_etymons = limits.min_etymons if min_etymons == nil and not opts.skip_default_min then min_etymons = 1 end if min_etymons and count < min_etymons then if min_etymons > 1 then error("Detected " .. label .. " group with fewer than " .. min_etymons .. " etymons.") else error("Detected " .. label .. " with no etymons.") end end if limits.max_etymons and count > limits.max_etymons then local unit = (limits.max_etymons == 1) and "etymon" or "etymons" error("Detected " .. label .. " with more than " .. limits.max_etymons .. " " .. unit .. ".") end end function EtymonParser.check_term_rules(etymon_data, entry_lang, rules, label) label = label or "term" if rules and rules.disallow then local disallowed = {} for _, typ in ipairs(rules.disallow) do local spec = TERM_RULE_DISALLOW[typ] if spec and etymon_data[spec.field] then table.insert(disallowed, spec.label) end end if #disallowed > 0 then error(label .. " does not support " .. mw.text.listToText(disallowed, "or") .. " etymons.") end end if etymon_data.is_family then if rules and rules.family == "disallowed" then error(label .. " does not support family codes" .. (rules.family_suffix or ".")) elseif not etymon_data.suppress_term then error("Family codes require suppressed term (use family:-).") end end if rules then if rules.require_term and (not etymon_data.term or etymon_data.term == "") then error(label .. " requires a term for each listed form.") end if rules.entry_lang then if Util.get_norm_lang(etymon_data.lang):getFullCode() ~= Util.get_norm_lang(entry_lang):getFullCode() then error(label .. " terms must be in the entry language (" .. entry_lang:getFullCode() .. "), got '" .. etymon_data.lang:getFullCode() .. "'.") end end if rules.ancestor_check then M.etymology.check_ancestor(entry_lang, etymon_data.lang) end elseif etymon_data.is_family and not etymon_data.suppress_term then error("Family codes require suppressed term (use family:-).") end end function EtymonParser.check_keyword_term(etymon_data, entry_lang, keyword) local config = M.data.keywords[keyword] EtymonParser.check_term_rules(etymon_data, entry_lang, config and config.term_rules, "`:" .. keyword .. "`") end function EtymonParser.check_supplement_term(etymon_data, entry_lang, supplement_type) local config = M.data.supplements[supplement_type] EtymonParser.check_term_rules(etymon_data, entry_lang, config and config.term_rules, "|" .. supplement_type .. "=") end -- Parse keyword with modifiers (e.g., ":bor<unc>" or ":bor<ref:{{R:example}}>") function EtymonParser.parse_keyword_modifiers(param) if type(param) ~= "string" then return nil, {} end local base_keyword = get_keyword_base(param) if not base_keyword then return nil, {} end local canonical_keyword = get_canonical_keyword(base_keyword) -- Check if there are any modifiers if not param:find("<", 1, true) then return canonical_keyword, {} end -- Parse modifiers using the same mechanism as etymon parsing local rest_with_defaults = Util.add_boolean_defaults(param, EtymonParser.keyword_param_mods) local function generate_obj(ignored) return {} end local parsed = M.parse_utilities.parse_inline_modifiers(rest_with_defaults:gsub("^:?[^<]+", ""), { param_mods = get_clean_param_mods(EtymonParser.keyword_param_mods), generate_obj = generate_obj }) local modifiers = { unc = parsed.unc or false, refs = parsed.refs, text = parsed.text, lit = parsed.lit, conj = parsed.conj, g = parsed.g, surf = parsed.surf or false, senseid = parsed.senseid, } -- Validate modifiers against restrictions EtymonParser.check_modifier_restrictions(modifiers, canonical_keyword, EtymonParser.keyword_param_mods) return canonical_keyword, modifiers end local function normalize_keyword_param(keyword_with_mods) local trimmed = M.string_utilities.trim(keyword_with_mods) reject_removed_surf_keyword(trimmed:match("^:") and trimmed or (":" .. trimmed)) local base = get_keyword_base(trimmed) if not base or not M.data.keywords[base] then error("Invalid keyword '" .. trimmed .. "' in inline etymology") end local canonical_base = get_canonical_keyword(base) local without_colon = trimmed:gsub("^:", "") local mods_part = without_colon:sub(#base + 1) local kw_param = normalize_keyword(canonical_base .. mods_part) EtymonParser.parse_keyword_modifiers(kw_param) return kw_param end local function get_keyword_mod_names() local names = {} for mod_name in pairs(EtymonParser.keyword_param_mods) do names[mod_name] = true end return names end local function parse_inline_ety_run(ety_string) local body = ety_string or "" if body == "" then error("Empty inline etymology") end local keyword_mod_names = get_keyword_mod_names() local pos = 1 local len = #body local function parse_err(msg) error(msg .. " in inline etymology: '" .. body .. "'") end local function peek_double() return body:sub(pos, pos + 1) == "<<" end local function mod_name_from_unwrapped(unwrapped) return unwrapped:match("^<([^:>]+)") end local function is_keyword_mod(unwrapped) local name = mod_name_from_unwrapped(unwrapped) return name and keyword_mod_names[name] or false end local function read_double_bracket() if not peek_double() then return nil end local start = pos pos = pos + 2 while pos <= len - 1 do if body:sub(pos, pos + 1) == ">>" then local token = body:sub(start, pos + 1) pos = pos + 2 return token, token:sub(2, -2) end pos = pos + 1 end parse_err("Unmatched <<") end local function read_angle_cell() if body:sub(pos, pos) ~= "<" or peek_double() then return nil end local open = pos pos = pos + 1 local depth = 1 local i = pos while i <= len do local ch = body:sub(i, i) if ch == "<" then depth = depth + 1 elseif ch == ">" then depth = depth - 1 if depth == 0 then local inner = body:sub(open + 1, i - 1) pos = i + 1 return inner end end i = i + 1 end parse_err("Unmatched <") end local function read_bare_run() local start = pos while pos <= len and body:sub(pos, pos) ~= "<" do pos = pos + 1 end return body:sub(start, pos - 1) end local function absorb_double_keyword_mods(keyword_str) while peek_double() do local saved = pos local _, unwrapped = read_double_bracket() if is_keyword_mod(unwrapped) then keyword_str = keyword_str .. unwrapped else pos = saved break end end return keyword_str end local kw_start = pos while pos <= len and body:sub(pos, pos) ~= "<" do pos = pos + 1 end local keyword = body:sub(kw_start, pos - 1) if keyword:match("^%s*$") then parse_err("Missing keyword") end keyword = absorb_double_keyword_mods(keyword) local cells = {} while pos <= len do if peek_double() then local _, unwrapped = read_double_bracket() if is_keyword_mod(unwrapped) then parse_err("Unexpected keyword modifier " .. unwrapped .. " outside of a keyword") end table.insert(cells, "+" .. unwrapped) elseif body:sub(pos, pos) == "<" then local inner = read_angle_cell() if inner ~= "" then table.insert(cells, inner) end else local bare = read_bare_run() if bare ~= "" then if bare:sub(1, 1) ~= ":" then parse_err("Unexpected bare text '" .. bare .. "' (use :keyword for nested keywords in inline etymology)") end if not is_keyword(bare, true) then parse_err("Invalid keyword '" .. bare .. "' in inline etymology") end table.insert(cells, absorb_double_keyword_mods(bare)) end end end return { keyword = keyword, cells = cells, } end function EtymonParser.inline_ety_to_pipe(ety_string) local run = parse_inline_ety_run(ety_string) if not run.keyword or run.keyword:match("^%s*$") then return "|" end local pipe_parts = { normalize_keyword_param(M.string_utilities.trim(run.keyword)) } for _, segment in ipairs(run.cells) do if is_keyword(segment, true) then table.insert(pipe_parts, normalize_keyword_param(segment)) else table.insert(pipe_parts, segment) end end return "|" .. table.concat(pipe_parts, "|") .. "|" end function EtymonParser.pipe_to_inline_ety(pipe_string) local cells = {} for cell in pipe_string:gmatch("([^|]+)") do if cell ~= "" then table.insert(cells, cell) end end if #cells == 0 then return "" end local inline_parts = {} for index, cell in ipairs(cells) do local base = get_keyword_base(cell) if base and M.data.keywords[base] then local without_colon = cell:gsub("^:", "") local kw_base, mods = without_colon:match("^([^<]+)(.*)$") local inline_kw = (kw_base or without_colon) .. (mods or ""):gsub("<([^>]+)>", "<<%1>>") if index > 1 then inline_kw = ":" .. inline_kw end table.insert(inline_parts, inline_kw) elseif cell:sub(1, 1) == "+" then local mod = cell:sub(2) if mod:match("^<.->$") then mod = mod:sub(2, -2) end table.insert(inline_parts, "<<" .. mod .. ">>") else table.insert(inline_parts, "<" .. cell .. ">") end end return table.concat(inline_parts, "") end function EtymonParser.parse_inline_ety(ety_string, context_lang) local run = parse_inline_ety_run(ety_string) local keyword = M.string_utilities.trim(run.keyword) reject_removed_surf_keyword(":" .. keyword) if not is_keyword(keyword, true) then error("Invalid keyword '" .. keyword .. "' in inline etymology <ety:" .. keyword .. "...>") end local args = { context_lang:getCode(), normalize_keyword_param(keyword) } for _, segment in ipairs(run.cells) do if is_keyword(segment, true) then table.insert(args, normalize_keyword_param(segment)) else table.insert(args, segment) end end return args end function EtymonParser.parse_etymon(param, context_lang) if is_keyword(param) then return nil end if type(param) ~= "string" then return nil end local lang, rest local is_family = false local before_bracket = param:match("^([^<]*)") or param local lang_code, rest_match = before_bracket:match("^([a-zA-Z][a-zA-Z0-9._-]*):(.*)$") if lang_code then local potential_lang = Util.get_lang(lang_code, true) if potential_lang then lang = potential_lang rest = param:sub(#lang_code + 2) else local potential_family = Util.get_family(lang_code) if potential_family then lang = potential_family rest = param:sub(#lang_code + 2) is_family = true else lang = context_lang rest = param end end else lang = context_lang rest = param end M.tracking.track_term(rest) if rest == "" or rest == "+" then return { lang = lang, term = nil, unknown_term = true, is_family = is_family, } end if rest == "-" then return { lang = lang, term = nil, suppress_term = true, is_family = is_family, } end if not rest:find("<", 1, true) then return { lang = lang, term = M.string_utilities.trim(rest), is_family = is_family, } end local term_text = rest:match("^([^<]*)") or "" local is_unknown = (term_text == "" or term_text == "+") local is_suppress = (term_text == "-") local function generate_obj(ignored_term) return { term = (is_unknown or is_suppress) and nil or M.string_utilities.trim(term_text) } end local rest_with_defaults = Util.add_boolean_defaults(rest, EtymonParser.etymon_param_mods) local parsed_obj = M.parse_utilities.parse_inline_modifiers(rest_with_defaults, { param_mods = get_clean_param_mods(EtymonParser.etymon_param_mods), generate_obj = generate_obj }) if parsed_obj.id and parsed_obj.id:match("^!") then parsed_obj.id = parsed_obj.id:sub(2) parsed_obj.override = true end parsed_obj.lang = lang parsed_obj.is_family = is_family if is_unknown then parsed_obj.unknown_term = true elseif is_suppress then parsed_obj.suppress_term = true end return parsed_obj end function EtymonParser.validate(lang, args, id, title, pos, starts_with_lang_code) -- id is now optional, so only validate if provided if id then if mw.ustring.len(id) < 2 then error("The `id` parameter must have at least two characters.") end if id == title or id == Util.get_page_data().pagename then error("The `id` parameter must not be the same as the page title.") end end local valid_pos = { prefix = true, suffix = true, interfix = true, infix = true, root = true, word = true } if pos and not valid_pos[pos] then error("Unknown value provided for `pos`. Valid values: " .. table.concat(require("Module:table").keysToList(valid_pos), ", ") .. ".") end local current_keyword = "from" local current_keyword_explicit = false local keyword_etymons = {} local keywords = M.data.keywords local function checkKeyword() local config = keywords[current_keyword] if current_keyword == "from" and not current_keyword_explicit and #keyword_etymons == 0 then keyword_etymons = {} return end EtymonParser.check_etymon_limits(#keyword_etymons, config, "`:" .. current_keyword .. "`") keyword_etymons = {} end local start_index = starts_with_lang_code and 2 or 1 for i = start_index, #args do local param = args[i] if type(param) ~= "string" then elseif param:sub(1, 1) == ":" and not is_keyword(param) then reject_removed_surf_keyword(param) error("Invalid keyword '" .. param .. "'. Did you mean a valid keyword like ':bor', ':inh', etc.?") elseif is_keyword(param) then checkKeyword() current_keyword = get_canonical_keyword(get_keyword(param)) current_keyword_explicit = true else local etymon_data = EtymonParser.parse_etymon(param, lang) if etymon_data then table.insert(keyword_etymons, param) EtymonParser.check_keyword_term(etymon_data, lang, current_keyword) -- Check modifier restrictions EtymonParser.check_modifier_restrictions(etymon_data, current_keyword, EtymonParser.etymon_param_mods) -- postype must be "root" or "word" local VALID_POSTYPES = { root = true, word = true } if etymon_data.postype and not VALID_POSTYPES[etymon_data.postype] then error("Invalid <postype:" .. etymon_data.postype .. ">; must be \"root\" or \"word\".") end if etymon_data.ety then local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang) EtymonParser.validate(etymon_data.lang, inline_args, nil, nil, nil, true) end else table.insert(keyword_etymons, param) end end end checkKeyword() end local DataRetriever = {} local function format_etymon_id_hint(id_data, idx) local id = type(id_data) == "table" and id_data.id or id_data local pos = type(id_data) == "table" and id_data.pos if id and id ~= "" and id ~= "*" then return '"' .. id .. '"' end if pos and pos ~= "" then return "unnamed (|pos=" .. pos .. "|)" end return "etymon #" .. idx .. " (no |id= on page)" end local function etymon_target_page_link(page, norm_lang) return M.links.full_link({ term = page, lang = norm_lang, no_generate_alternants = true, }, "term") end -- Summarize {{etymon}} id slots on a linked page for preview warnings. local function summarize_available_etymon_ids(ids) local id_list = {} local all_idless = true local target_has_idless = false local any_pos = false for i, id_data in ipairs(ids) do local id = type(id_data) == "table" and id_data.id or id_data local pos = type(id_data) == "table" and id_data.pos if id and id ~= "" and id ~= "*" then all_idless = false else target_has_idless = true end if pos and pos ~= "" then any_pos = true end table.insert(id_list, format_etymon_id_hint(id_data, i)) end return { id_list = id_list, all_idless = all_idless, target_has_idless = target_has_idless, any_pos = any_pos, count = #ids, options_text = mw.text.listToText(id_list), } end local function ambiguous_etymon_suggestion(page_link, summary) if summary.all_idless then if summary.any_pos then return " None set `|id=` yet; add a unique `|id=` to each on " .. page_link .. ", then `<id:identifier>` after the term here. Section order / hints: " .. summary.options_text .. "." end return " None set `|id=` yet; add a unique `|id=` to each {{etymon}} in that section from top to bottom, then `<id:identifier>` after the term here (same value as `|id=`)." end return " Specify which one with `<id:identifier>` after the term. Options: " .. summary.options_text .. "." end local function warn_ambiguous_etymon_link(page, norm_lang, ids, is_toplevel) local page_link = etymon_target_page_link(page, norm_lang) local summary = summarize_available_etymon_ids(ids) if is_toplevel and summary.target_has_idless then __state.linked_page_multiple_etymons_idless = true end local lang_name = norm_lang:getCanonicalName() local lead = "Etymology link to " .. page_link .. " is ambiguous (" .. summary.count .. " {{etymon}} templates for " .. lang_name .. ")." Util.add_warning(lead .. ambiguous_etymon_suggestion(page_link, summary), true) end local function is_mismatched_explicit_id(base_key, cached_args, parent_etymon) return cached_args == M.data.STATUS.MISSING and not parent_etymon and #(__state.available_etymon_ids[base_key] or {}) > 0 end local function maybe_flag_partial_etymology_reference(base_key, etymon_data, cached_args, is_toplevel) if not is_toplevel or __state.skip_partial_etymology_category then return end if not __state.partial_etymology_targets[base_key] then return end if etymon_data.id and type(cached_args) == "table" then return end __state.linked_page_partial_etymology_sections = true end local function is_nonlemma_etymon_template(template_args) return template_args and M.yesno(template_args.nl, false) end local function warn_mismatched_explicit_id(page, norm_lang, base_key, etymon_id) local page_link = etymon_target_page_link(page, norm_lang) local summary = summarize_available_etymon_ids(__state.available_etymon_ids[base_key] or {}) local lang_name = norm_lang:getCanonicalName() local lead = "Etymology link to " .. page_link .. " uses `<id:" .. etymon_id .. ">`, but no {{etymon}} on that page has `|id=" .. etymon_id .. "|` for " .. lang_name .. "." Util.add_warning(lead .. " Valid IDs: " .. summary.options_text .. ".", true) end -- Given an etymon data, scrape its page and cache the result in the global state object. function DataRetriever.cache_page_etymons(etymon_page, etymon_title, key, etymon_lang, etymon_id, redirected_from, descendants_is_toplevel) local content = etymon_title:getContent() if not content then __state.cached_etymon_args[key] = M.data.STATUS.REDLINK return end -- Check if the linked page is a redirect. If it is, the template parsing -- code below will be effectively skipped, and `scrape_page` will be called -- again on the redirect target (see the bottom of this function) local lang_section_for_descendants = nil local redirect_target = etymon_title.redirect_target if not redirect_target then content = M.pages.get_section(content, etymon_lang:getFullName(), 2) if not content then __state.cached_etymon_args[key] = M.data.STATUS.MISSING return end lang_section_for_descendants = content end local etymon_lang_code = etymon_lang:getFullCode() local lang_page_key = etymon_lang_code .. ":" .. etymon_page local found_templates_for_lang = {} local found_ids = {} local get_node_class = M.template_parser.class_else_type -- Look for all {{etymon}} templates within the page content using the template parser -- This way the same page is never parsed more than once -- Build a map from senseids to their parent etymonids. local active_etymon_args = nil local etymology_section_count = 0 local etymology_sections_with_etymon = 0 local current_etymology_has_etymon = false local current_etymology_has_nonlemma = false local function finalize_current_etymology_section() if etymology_section_count == 0 then return end if current_etymology_has_etymon or current_etymology_has_nonlemma then etymology_sections_with_etymon = etymology_sections_with_etymon + 1 end current_etymology_has_etymon = false current_etymology_has_nonlemma = false end for node in M.template_parser.parse(content):iterate_nodes() do local node_class = get_node_class(node) if node_class == "heading" then -- A new L2 or etymology section acts as a barrier: an {{etymon}} usage -- used previously cannot be the parent of any subsequent senseids. -- Note that we don't have to check for L2s due to the usage of `M.pages.get_section` above. if node:get_name():find("^Etymology") then finalize_current_etymology_section() etymology_section_count = etymology_section_count + 1 active_etymon_args = nil end elseif node_class == "template" then local template_name = node:get_name() if template_name == "etymon" then local template_args = node:get_arguments() -- Check if this etymon is for our language if template_args[1] == etymon_lang_code then if is_nonlemma_etymon_template(template_args) then if etymology_section_count > 0 then current_etymology_has_nonlemma = true end else if etymology_section_count > 0 then current_etymology_has_etymon = true end table.insert(found_templates_for_lang, template_args) if template_args.id then local etymon_key = lang_page_key .. ":" .. template_args.id __state.cached_etymon_args[etymon_key] = template_args __state.cached_etymon_pages[etymon_key] = tostring(etymon_page) table.insert(found_ids, template_args.id) active_etymon_args = template_args else -- Store idless etymon with default key local etymon_key = lang_page_key .. ":*" __state.cached_etymon_args[etymon_key] = template_args __state.cached_etymon_pages[etymon_key] = tostring(etymon_page) table.insert(found_ids, "*") active_etymon_args = template_args end end end elseif active_etymon_args and template_name == "senseid" then local template_args = node:get_arguments() -- This should always be true for proper usages of {{senseid}}. if template_args[1] == etymon_lang_code and template_args[2] then local sense_id_key = lang_page_key .. ":" .. template_args[2] __state.senseid_parent_etymon[sense_id_key] = active_etymon_args __state.cached_etymon_pages[sense_id_key] = tostring(etymon_page) end end end end finalize_current_etymology_section() if lang_section_for_descendants and etymology_section_count > 1 and etymology_sections_with_etymon > 0 and etymology_sections_with_etymon < etymology_section_count then __state.partial_etymology_targets[lang_page_key] = true end if descendants_is_toplevel and lang_section_for_descendants and #found_templates_for_lang > 0 then M.descendants.cache_page_checks({ lang_section = lang_section_for_descendants, etymon_lang_code = etymon_lang_code, found_templates_for_lang = found_templates_for_lang, entry_title = __state.entry_title, entry_lang_code = __state.entry_lang_code, entry_lang = __state.entry_lang_code and Util.get_lang(__state.entry_lang_code, true) or nil, cached_descendants_checks = __state.cached_descendants_checks, lang_page_key = lang_page_key, redirected_from = redirected_from, }) end local id_data_list = {} for _, args in ipairs(found_templates_for_lang) do local id = args.id or "*" table.insert(id_data_list, { id = id, pos = args.pos }) end __state.available_etymon_ids[lang_page_key] = id_data_list if #found_templates_for_lang == 1 then __state.single_etymons[lang_page_key] = found_templates_for_lang[1] end if redirected_from and __state.available_etymon_ids[lang_page_key] then __state.available_etymon_ids[redirected_from] = __state.available_etymon_ids[redirected_from] or {} for _, id_data in ipairs(__state.available_etymon_ids[lang_page_key]) do table.insert(__state.available_etymon_ids[redirected_from], id_data) end end if __state.cached_etymon_args[key] ~= nil or __state.senseid_parent_etymon[key] ~= nil then -- All done! return elseif redirect_target and not redirected_from then -- Try scraping the redirect. etymon_page = redirect_target.prefixedText DataRetriever.cache_page_etymons(etymon_page, redirect_target, lang_page_key .. ":" .. etymon_id, etymon_lang, etymon_id, lang_page_key, descendants_is_toplevel) __state.cached_etymon_args[key] = __state.cached_etymon_args[etymon_lang_code .. ":" .. etymon_page .. ":" .. etymon_id] else __state.cached_etymon_args[key] = M.data.STATUS.MISSING end end local function has_linkable_term(etymon_data) if etymon_data.is_family or etymon_data.suppress_term or etymon_data.unknown_term then return false end local term = etymon_data.term if term == nil or term == "" then return false end return M.string_utilities.trim(term) ~= "" end local function record_term_id_tracking(etymon_data) if not has_linkable_term(etymon_data) then return end local term_page = M.links.get_link_page(etymon_data.term, etymon_data.lang) M.tracking.record_term_id_usage(__state.id_stats, etymon_data, term_page) end -- Given an etymon object, scrape its page (if necessary) and return its own etymon arguments as well as the page name. function DataRetriever.get_etymon_args(etymon_data, is_toplevel) if not has_linkable_term(etymon_data) then return M.data.STATUS.MISSING, nil, nil, nil end local page = M.links.get_link_page(etymon_data.term, etymon_data.lang) local norm_lang = Util.get_norm_lang(etymon_data.lang) local base_key = norm_lang:getFullCode() .. ":" .. page if etymon_data.id then local key = base_key .. ":" .. etymon_data.id local cached_args = __state.cached_etymon_args[key] or __state.senseid_parent_etymon[key] if cached_args == nil then local title = mw.title.new(page) if not title then error('Invalid page title "' .. page .. '" encountered.') end DataRetriever.cache_page_etymons(page, title, key, norm_lang, etymon_data.id, nil, is_toplevel) end cached_args = __state.cached_etymon_args[key] or __state.senseid_parent_etymon[key] -- refresh -- Get etymon_id from parent if this was resolved via senseid local parent_etymon = __state.senseid_parent_etymon[key] local resolved_etymon_id = parent_etymon and parent_etymon.id local descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = is_toplevel, base_key = base_key, lookup = { explicit_id = etymon_data.id, parent_etymon = parent_etymon, }, }) if is_toplevel and descendants_check == nil then local title = mw.title.new(page) if title then DataRetriever.cache_page_etymons(page, title, key, norm_lang, etymon_data.id, nil, true) descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = true, base_key = base_key, lookup = { explicit_id = etymon_data.id, parent_etymon = parent_etymon, }, }) end end local mismatched_id = is_mismatched_explicit_id(base_key, cached_args, parent_etymon) if mismatched_id and is_toplevel then __state.has_mismatched_id = true M.tracking.record_mismatched_id_usage(__state.id_stats, norm_lang, page, etymon_data.id) warn_mismatched_explicit_id(page, norm_lang, base_key, etymon_data.id) end maybe_flag_partial_etymology_reference(base_key, etymon_data, cached_args, is_toplevel) return cached_args, __state.cached_etymon_pages[key], resolved_etymon_id, descendants_check else __state.used_idless_etymon = true if is_toplevel then __state.toplevel_idless_etymon = true end if __state.available_etymon_ids[base_key] == nil then local title = mw.title.new(page) if not title then error('Invalid page title "' .. page .. '" encountered.') end DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, is_toplevel) end local ids = __state.available_etymon_ids[base_key] or {} local count = #ids -- Try to filter by postype if available and we have multiple candidates if count > 1 and etymon_data.postype then local matching_ids = {} for _, id_data in ipairs(ids) do if id_data.pos == etymon_data.postype then table.insert(matching_ids, id_data) end end if #matching_ids == 1 then local matched_id = matching_ids[1].id local matched_key = base_key .. ":" .. matched_id M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "postype") local descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = is_toplevel, base_key = base_key, lookup = { id = matched_id }, }) if is_toplevel and descendants_check == nil then local title = mw.title.new(page) if title then DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, true) descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = true, base_key = base_key, lookup = { id = matched_id }, }) end end local matched_args = __state.cached_etymon_args[matched_key] maybe_flag_partial_etymology_reference(base_key, etymon_data, matched_args, is_toplevel) return matched_args, __state.cached_etymon_pages[matched_key], nil, descendants_check end end if count == 1 then local only_id_data = ids[1] local only_id = (type(only_id_data) == "table" and only_id_data.id) or only_id_data or "*" M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "single") local descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = is_toplevel, base_key = base_key, lookup = { id_data = only_id_data }, }) if is_toplevel and descendants_check == nil then local title = mw.title.new(page) if title then DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, true) descendants_check = M.descendants.get_lookup_check({ cached_descendants_checks = __state.cached_descendants_checks, is_toplevel = true, base_key = base_key, lookup = { id_data = only_id_data }, }) end end local single_args = __state.single_etymons[base_key] maybe_flag_partial_etymology_reference(base_key, etymon_data, single_args, is_toplevel) return single_args, __state.cached_etymon_pages[base_key .. ":" .. only_id], nil, descendants_check elseif count > 1 then M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "ambiguous") warn_ambiguous_etymon_link(page, norm_lang, ids, is_toplevel) maybe_flag_partial_etymology_reference(base_key, etymon_data, M.data.STATUS.AMBIGUOUS, is_toplevel) return M.data.STATUS.AMBIGUOUS, nil, nil, nil else M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "missing") maybe_flag_partial_etymology_reference(base_key, etymon_data, M.data.STATUS.MISSING, is_toplevel) return M.data.STATUS.MISSING, nil, nil, nil end end end local function keyword_invisible_in_tree(keyword_info) if not keyword_info then return false end local inv = keyword_info.invisible return inv == "all" or inv == true or inv == "tree" end -- True when the node has at least one top-level child container visible in the tree. local function node_has_visible_tree_children(node) for _, container in ipairs(node.children or {}) do if not keyword_invisible_in_tree(container.keyword_info) then return true end end return false end -- Count visible term nodes in the tree. local function get_visible_tree_depth(node, skip_child_rendering) local max_depth = 1 if skip_child_rendering or not node then return max_depth end for _, container in ipairs(node.children or {}) do local keyword_info = container.keyword_info if not keyword_invisible_in_tree(keyword_info) then local skip_grandchildren = keyword_info and keyword_info.no_child_categories for _, term in ipairs(container.terms or {}) do if term.is_duplicate then if term.original_has_children then max_depth = math.max(max_depth, 2) end else max_depth = math.max(max_depth, 1 + get_visible_tree_depth(term, skip_grandchildren)) end end end end return max_depth end local function as_param_list(val) if val == nil then return {} end if type(val) == "table" then return val end if type(val) == "string" and val ~= "" then return { val } end return {} end local TreeBuilder = {} -- Build a unique key for deduplication in the seen table function TreeBuilder.build_key(lang, title, args) local norm_lang_code = Util.get_norm_lang(lang):getFullCode() local is_table = type(args) == "table" local id = (is_table and args.id) or "" if title then return norm_lang_code .. ":" .. M.links.get_link_page(title, lang) .. ":" .. id end if is_table and args.status == M.data.STATUS.INLINE then local content_parts = {} for i = 1, #args do content_parts[i] = tostring(args[i]) end return norm_lang_code .. ":*:" .. id .. "\0" .. table.concat(content_parts, "\0") end return norm_lang_code .. ":*:" .. id end -- Copy parsed etymon modifiers onto a tree/supplement term node. function TreeBuilder.apply_etymon_fields(term, etymon_data) term.id = etymon_data.id term.gloss = etymon_data.gloss term.tr = etymon_data.tr term.ts = etymon_data.ts term.alt = etymon_data.alt term.genders = etymon_data.genders term.pos = etymon_data.pos term.ng = etymon_data.ng term.infl = etymon_data.infl term.refs = etymon_data.refs term.is_uncertain = etymon_data.unc term.lit = etymon_data.lit term.q = etymon_data.q term.qq = etymon_data.qq term.l = etymon_data.l term.ll = etymon_data.ll term.suppress_term = etymon_data.suppress_term term.unknown_term = etymon_data.unknown_term term.is_family = etymon_data.is_family term.override = etymon_data.override term.aftype = etymon_data.aftype term.postype = etymon_data.postype term.bor = etymon_data.bor term.lbor = etymon_data.lbor term.slbor = etymon_data.slbor end function TreeBuilder.build_supplement_term(etymon_data, entry_lang, supplement_type) EtymonParser.check_supplement_term(etymon_data, entry_lang, supplement_type) local term = { lang = etymon_data.lang, title = etymon_data.term, children = {}, status = M.data.STATUS.OK, } TreeBuilder.apply_etymon_fields(term, etymon_data) return term end function TreeBuilder.build_supplement_terms(entry_lang, supplement_type, param_value) local terms = {} for _, term_param in ipairs(as_param_list(param_value)) do if type(term_param) == "string" and term_param ~= "" then local etymon_data = EtymonParser.parse_etymon(term_param, entry_lang) if etymon_data then table.insert(terms, TreeBuilder.build_supplement_term(etymon_data, entry_lang, supplement_type)) end end end return terms end -- Attach a |param= supplement defined in etymon_data.supplements (e.g. doublet=). function TreeBuilder.append_term_supplement(data_tree, entry_lang, supplement_type, param_value) local config = M.data.supplements[supplement_type] if not config then error("Unknown supplement '" .. tostring(supplement_type) .. "'.") end local terms = TreeBuilder.build_supplement_terms(entry_lang, supplement_type, param_value) if #terms == 0 then return end data_tree.supplements = data_tree.supplements or {} table.insert(data_tree.supplements, { type = supplement_type, config = config, terms = terms, }) M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, supplement_type, entry_lang, entry_lang, true) end function TreeBuilder.build(lang, title, args, seen, depth, stop_recursion) seen = seen or {} depth = depth or 0 local is_toplevel = (depth == 0) if depth > __state.max_depth_reached then __state.max_depth_reached = depth end __state.total_nodes = __state.total_nodes + 1 local lang_code = lang:getCode() __state.language_count[lang_code] = (__state.language_count[lang_code] or 0) + 1 local current_id = (type(args) == "table" and args.id) or "" local key = TreeBuilder.build_key(lang, title, args) local node = { lang = lang, title = title, id = current_id, args = args, children = {}, status = M.data.STATUS.OK } if type(args) ~= "table" or seen[key] then node.status = args or M.data.STATUS.MISSING -- Mark as duplicate if we've seen this node before if seen[key] then node.is_duplicate = true node.duplicate_key = key local original_node = seen[key] if type(original_node) == "table" and original_node.children and #original_node.children > 0 then node.original_has_children = true end end return node end node.status = args.status or M.data.STATUS.OK seen[key] = node -- If stop_recursion is set, skip parsing children but check for visible children if stop_recursion then local keywords = M.data.keywords local has_visible_children = false for i = 2, #args do local param = args[i] if type(param) == "string" then local keyword_base = get_keyword_base(param) if keyword_base and keywords[keyword_base] then local _, kw_modifiers = EtymonParser.parse_keyword_modifiers(param:sub(1, 1) == ":" and param or (":" .. param)) if not keyword_invisible_in_tree(get_effective_keyword_info(keyword_base, kw_modifiers)) then has_visible_children = true break end elseif param:sub(1, 1) ~= ":" then -- It's a term (not a keyword), so there are visible children has_visible_children = true break end end end node.has_visible_children = has_visible_children return node end -- Parse args into keyword containers local current_keyword = "from" local current_keyword_modifiers = {} local current_container = nil local function ensure_container() if not current_container or current_container.keyword ~= current_keyword then local keyword_info = get_effective_keyword_info(current_keyword, current_keyword_modifiers) current_container = { keyword = current_keyword, keyword_info = keyword_info, keyword_modifiers = current_keyword_modifiers, terms = {}, } table.insert(node.children, current_container) -- Override keyword text/phrase for nominalization with <g:code> if current_keyword_modifiers.g and current_keyword == "nominalization" then local labels = get_nominalization_label_for_g(current_keyword_modifiers.g) if not labels then local codes = {} for c in pairs(M.data.nominalization_g_codes) do table.insert(codes, c) end table.sort(codes) error("Invalid <g:" .. tostring(current_keyword_modifiers.g) .. ">. Supported codes for nominalization: " .. table.concat(codes, ", ")) end current_container.keyword_info = copy_keyword_info(keyword_info) current_container.keyword_info.text = labels.text current_container.keyword_info.phrase = labels.phrase end end return current_container end local parse_context_lang = Util.resolve_context_lang(lang, args) for i = 2, #args do local param = args[i] if is_keyword(param) then local keyword, modifiers = EtymonParser.parse_keyword_modifiers(param) if not keyword then error("Invalid keyword '" .. param .. "'.") end current_keyword = keyword current_keyword_modifiers = modifiers current_container = nil -- Force new container for new keyword elseif type(param) == "string" and param:sub(1, 1) == ":" then reject_removed_surf_keyword(param) error("Invalid keyword '" .. param .. "'. Did you mean a valid keyword like ':bor', ':inh', etc.?") elseif type(param) == "string" then local etymon_data = EtymonParser.parse_etymon(param, parse_context_lang) if etymon_data then -- Track keyword usage at top level M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, current_keyword, lang, etymon_data.lang, is_toplevel) local term_node = {} local container -- Handle suppress_term (-) and unknown_term (empty or +) directly if etymon_data.suppress_term or etymon_data.unknown_term then container = ensure_container() if etymon_data.ety then local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang) inline_args.id = etymon_data.id inline_args.status = M.data.STATUS.INLINE term_node = TreeBuilder.build(etymon_data.lang, nil, inline_args, seen, depth + 1) else term_node = { lang = etymon_data.lang, children = {}, status = M.data.STATUS.OK, } end TreeBuilder.apply_etymon_fields(term_node, etymon_data) else -- Regular term: fetch arguments from page record_term_id_tracking(etymon_data) local etymon_args, page_of, resolved_etymon_id, descendants_check = DataRetriever.get_etymon_args(etymon_data, is_toplevel) -- Check for <ety> inline parameter doesn't override the scraped arguments, unless the latter are missing if etymon_data.ety then if etymon_args == M.data.STATUS.REDLINK or etymon_args == M.data.STATUS.MISSING then __state.current_page_has_inline_etymology = true if is_toplevel then __state.toplevel_has_inline_etymology = true end local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang) -- Track inline ety keywords too local inline_keyword = get_keyword(inline_args[2], true) if inline_keyword and #inline_args >= 3 then local inline_etymon = EtymonParser.parse_etymon(inline_args[3], etymon_data.lang) if inline_etymon then M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, inline_keyword, etymon_data.lang, inline_etymon.lang, is_toplevel) end end inline_args.id = etymon_data.id inline_args.status = M.data.STATUS.INLINE etymon_args = inline_args term_node.page_of = __state.cached_etymon_pages[key] -- term node is on the same page as the parent else -- Scraped arguments exist, <ety> is redundant and ignored __state.current_page_has_redundant_etymology = true if is_toplevel then __state.toplevel_redundant_etymology = true end end end -- Ensure container exists before checking keyword info container = ensure_container() -- Check if current keyword has no_child_categories - if so, stop recursion local keyword_info = container.keyword_info local should_stop_recursion = (stop_recursion or (keyword_info and keyword_info.no_child_categories)) term_node = TreeBuilder.build(etymon_data.lang, etymon_data.term, etymon_args, seen, depth + 1, should_stop_recursion) term_node.target_key = Util.get_norm_lang(etymon_data.lang):getFullCode() .. ":" .. M.links.get_link_page(etymon_data.term, etymon_data.lang) term_node.etymon_id = resolved_etymon_id -- The actual etymon id when resolved via senseid term_node.page_of = page_of TreeBuilder.apply_etymon_fields(term_node, etymon_data) term_node.missing_descendants_header, term_node.missing_descendants_entry = M.descendants.get_term_sync_flags(current_keyword, term_node.status, descendants_check) end table.insert(container.terms, term_node) end end end return node end -- Convert etymology tree to JSON-serializable table local function tree_to_json(node) local obj = { term = node.title, lang = node.lang:getCode(), lang_name = node.lang:getCanonicalName(), id = (node.id and node.id ~= "") and node.id or nil, status = node.status, is_uncertain = node.is_uncertain or nil, is_duplicate = node.is_duplicate or nil, gloss = node.gloss, transliteration = node.tr, transcription = node.ts, alt = node.alt, g = node.genders, pos = node.pos, ng = node.ng, infl = node.infl, children = {}, } for _, container in ipairs(node.children or {}) do local keyword_info = container.keyword_info if keyword_info then local container_obj = { keyword = container.keyword, keyword_label = keyword_info.text, keyword_abbrev = keyword_info.abbrev, is_group = keyword_info.is_group or nil, is_invisible = keyword_info.invisible or nil, is_uncertain = (container.keyword_modifiers and container.keyword_modifiers.unc) or nil, terms = {}, } for _, term in ipairs(container.terms or {}) do table.insert(container_obj.terms, tree_to_json(term)) end table.insert(obj.children, container_obj) end end return obj end -- Build and return the etymology data tree for a given term. function export.get_tree(lang, title, args, options) options = options or {} __state.entry_title = title __state.entry_lang_code = lang:getCode() __state.id_stats = M.tracking.new_id_stats() __state.skip_partial_etymology_category = options.skip_partial_etymology_category == true if options.validate then EtymonParser.validate(lang, args, options.id, title, options.pos, false) end local lang_code = lang:getCode() local start_index = (args[1] == lang_code) and 2 or 1 local tree_args = { [1] = lang_code, id = options.id or args.id } for i = start_index, #args do table.insert(tree_args, args[i]) end __state.cached_etymon_args[lang_code .. ":" .. title .. ":" .. (tree_args.id or "")] = tree_args local ety_data_tree = TreeBuilder.build(lang, title, tree_args) if options.json then return M.JSON.toJSON(tree_to_json(ety_data_tree)) end return ety_data_tree end -- Given a language code, page name and optionally the id= parameter, -- render the tree and only the etymology tree for the relevant page. -- Fetches and parses the corresponding {{etymon}} from the requested page, -- and any further pages needed to render the tree. -- Parameters can be passed either through the #invoke or as -- template parameters *through* an #invoke. function export.render_tree_for_etymon_on_page(frame) local frame_args = frame.args local parent_args = frame:getParent().args local langcode = frame_args[1] or parent_args[1] local pagename = frame_args[2] or parent_args[2] local id = frame_args["id"] or parent_args["id"] local display_title = frame_args["title"] or parent_args["title"] local parsed_title = mw.title.new(pagename, 0) local title if parsed_title.namespace == 0 then title = M.pages.safe_page_name(parsed_title) elseif parsed_title.namespace == 118 then title = "*" .. M.pages.safe_page_name(parsed_title) else error("Unsupported namespace for render_tree_for_etymon_on_page: " .. parsed_title.namespace) end local lang = Util.get_lang(langcode) __state.entry_title = title __state.entry_lang_code = lang:getCode() __state.id_stats = M.tracking.new_id_stats() -- Construct etymon_data for DataRetriever.get_args. local etymon_data = { lang = lang, term = title, id = id } local args, pagename = DataRetriever.get_etymon_args(etymon_data, true) if args == M.data.STATUS.MISSING then error("The etymon template was not found (language " .. langcode .. ", title '" .. title .. "'" .. (id and ", ID '" .. id .. "'" or ", no ID given") .. "). Page contents may have changed in the interim.") end local tree_title = display_title or title if lang:stripDiacritics(M.links.remove_links(tree_title)) ~= lang:stripDiacritics(M.links.remove_links(title)) then M.tracking.track_title_pagename_mismatch(lang) end reset_invocation_state() local ety_data_tree = export.get_tree(lang, tree_title, args, { validate = true, id = id, }) local output = {} table.insert(output, M.template_styles("Module:etymon/styles.css")) table.insert(output, M.tree.render({ data_tree = ety_data_tree, format_term_func = function(term, is_toplevel) return Util.format_term(term, is_toplevel, { gloss = "suppress", pos = "suppress", lit = "suppress", tree_ql = "suppress", }) end, })) return table.concat(output) end function export.main(frame) local parent_args = frame:getParent().args local args = M.parameters.process(parent_args, M.parameters_data.etymon) local lang = args[1] local etymon_args = args[2] local id = args.id local title = args.title local text = args.text local tree = args.tree local etydate = args.etydate local doublet = args.doublet local rfe = args.rfe local etystub = args.etystub local is_nonlemma = M.yesno(args.nl, false) local page_data = Util.get_page_data() if not title then title = page_data.pagename if page_data.namespace == "Reconstruction" then title = "*" .. title end end local entry_pagename = page_data.pagename if page_data.namespace == "Reconstruction" then entry_pagename = "*" .. entry_pagename end if lang:stripDiacritics(M.links.remove_links(title)) ~= lang:stripDiacritics(M.links.remove_links(entry_pagename)) then M.tracking.track_title_pagename_mismatch(lang) end local current_L2 = M.pages.get_current_L2() if current_L2 then local norm_lang = Util.get_norm_lang(lang) local norm_name = norm_lang:getCanonicalName() if current_L2 ~= norm_name then local lang_desc = lang:getCode() .. " (" .. lang:getCanonicalName() .. ")" if norm_lang:getCode() ~= lang:getCode() then lang_desc = lang_desc .. ", normalized to " .. norm_lang:getCode() .. " (" .. norm_name .. ")" end error("Language '" .. lang_desc .. "' does not match the L2 header (" .. current_L2 .. ").") end end reset_invocation_state() local ety_data_tree = export.get_tree(lang, title, etymon_args, { validate = true, pos = args.pos, id = id, json = args.json, skip_partial_etymology_category = is_nonlemma, }) if args.json then return ety_data_tree end local output = {} local text_allowlist_mode = M.text_allowed.default_mode or "off" if text and text_allowlist_mode ~= "off" and not Util.is_text_param_allowed_for_lang(lang) then local msg = "Etymology texts (parameter <code>text=</code>) are not allowed for " .. lang:getFullName() .. "; see [[Template:etymon#Text allowlist|Template:etymon § Text allowlist]] for the list of languages that may use the <code>text=</code> parameter." if text_allowlist_mode == "error" then error(msg) else Util.add_warning(msg, true) end end local lang_exc = Util.get_lang_exception(lang) if lang_exc and lang_exc.disallow then local disallow = lang_exc.disallow local error_text = " for " .. lang:getFullName() if disallow.ref then error_text = error_text .. "; see " .. disallow.ref else error_text = error_text .. "." end if tree and disallow.tree then error("Etymology trees are not allowed" .. error_text) end if text and disallow.text then error("Etymology texts are not allowed" .. error_text) end end if etydate then local etydate_param_mods = M.parameter_utilities.construct_param_mods { { group = "ref" }, { param = "refn" }, { param = "nocap", type = "boolean" }, } local function generate_etydate_obj(etydate_text) local etydate_specs = {} for spec in etydate_text:gmatch("[^,]+") do table.insert(etydate_specs, mw.text.trim(spec)) end return { [1] = etydate_specs } end local parsed_etydate = M.parse_utilities.parse_inline_modifiers(etydate, { param_mods = etydate_param_mods, generate_obj = generate_etydate_obj }) local etydate_args = { [1] = parsed_etydate[1], nocap = parsed_etydate.nocap or false, } ety_data_tree.supplements = ety_data_tree.supplements or {} table.insert(ety_data_tree.supplements, { type = "etydate", etydate_text = M.etydate.format_etydate(etydate_args, { omit_refs = true }), etydate_refs = (parsed_etydate.refs and #parsed_etydate.refs > 0) and parsed_etydate.refs or nil, }) end TreeBuilder.append_term_supplement(ety_data_tree, lang, "doublet", doublet) local has_visible_children = node_has_visible_tree_children(ety_data_tree) -- Suppress trees for multiword entries and one-step chains local visible_tree_depth = get_visible_tree_depth(ety_data_tree) local is_trivial_tree = visible_tree_depth <= 2 local is_multiword = title:find("%s") ~= nil or title:find("_") ~= nil if tree and (is_multiword or is_trivial_tree) then tree = false end if tree then table.insert(output, M.template_styles("Module:etymon/styles.css")) table.insert(output, M.tree.render({ data_tree = ety_data_tree, format_term_func = function(term, is_toplevel) return Util.format_term(term, is_toplevel, { gloss = "suppress", pos = "suppress", lit = "suppress", tree_ql = "suppress", }) end, })) end local tree_disallowed = lang_exc and lang_exc.disallow and lang_exc.disallow.tree local ety_tree_json = M.JSON.toJSON(tree_to_json(ety_data_tree)) local anchor = M.anchors.etymonid(lang, id, { no_tree = args.notree, title = title, empty_tree = (not has_visible_children) or tree_disallowed, ety_tree_json = ety_tree_json, }) table.insert(output, anchor) local text_stop_lang_missing = nil if text then local max_depth, stop_at_blue_link, stop_at_lang, stop_at_lang_or_bluelink if text == "++" then max_depth, stop_at_blue_link = false, false elseif text == "+" then max_depth, stop_at_blue_link = 1, false elseif text == "*" then max_depth, stop_at_blue_link = false, true elseif text:match("^:[^*]+%*$") then -- Stop at a specific language OR first bluelink after it, e.g., ":ota*" -- If the target language is a redlink, continue to the first bluelink local lang_code = text:match("^:([^*]+)%*$") if lang_code and lang_code ~= "" then local lang_obj = Util.get_lang(lang_code, true) if lang_obj then stop_at_lang_or_bluelink = lang_code else Util.add_warning('Invalid language code "' .. lang_code .. '" in text parameter. Showing full chain instead.') max_depth, stop_at_blue_link = false, false end else Util.add_warning('Empty language code in text parameter. Showing full chain instead.') max_depth, stop_at_blue_link = false, false end elseif text:sub(1, 1) == ":" then -- Stop at a specific language, e.g., ":ar" stops at first Arabic term local lang_code = text:sub(2) if lang_code ~= "" then -- Validate the language code local lang_obj = Util.get_lang(lang_code, true) if lang_obj then stop_at_lang = lang_code else Util.add_warning('Invalid language code "' .. lang_code .. '" in text parameter. Showing full chain instead.') max_depth, stop_at_blue_link = false, false -- default to ++ end else Util.add_warning('Empty language code in text parameter. Showing full chain instead.') max_depth, stop_at_blue_link = false, false -- default to ++ end else local num = tonumber(text) if num and num >= 1 then max_depth, stop_at_blue_link = num, false else error('Invalid text value "' .. text .. '". Valid values are: "++" (full chain), "+" (first step only), "*" (until first blue link), a number (max steps), ":lang" (stop at language), or ":lang*" (stop at language or first bluelink if redlink)') end end local text_output, text_render_meta = M.text.render({ data_tree = ety_data_tree, format_term_func = Util.format_term, lang_matches_stop_code = Util.lang_matches_stop_code, max_depth = max_depth, stop_at_blue_link = stop_at_blue_link, curr_page = page_data.pagename, nodot = args.nodot, dot = args.dot, stop_at_lang = stop_at_lang, stop_at_lang_or_bluelink = stop_at_lang_or_bluelink, }) table.insert(output, text_output) if stop_at_lang and text_render_meta and not text_render_meta.stop_lang_reached then M.tracking.track_text_stop_lang_missing(lang, stop_at_lang) text_stop_lang_missing = stop_at_lang end end if rfe then table.insert(output, Util.expand_request_template(frame, "rfe", rfe, lang:getCode())) end if etystub then table.insert(output, Util.expand_request_template(frame, "etystub", etystub, lang:getCode())) end if is_nonlemma then table.insert(output, " " .. frame:expandTemplate({ title = "nonlemma", args = {}, })) end local categories = {} if Util.is_content_page() then M.tracking.track_tree_metrics({ max_depth_reached = __state.max_depth_reached, total_nodes = __state.total_nodes, language_count = __state.language_count, lang = lang, }) categories = M.categories.build({ data_tree = ety_data_tree, page_lang = lang, available_etymon_ids = __state.available_etymon_ids, senseid_parent_etymon = __state.senseid_parent_etymon, get_norm_lang_func = Util.get_norm_lang, lang_exc = lang_exc, suppress_categories = lang_exc and lang_exc.suppress_categories, nocat = args.nocat, tree = tree, text = text, exnihilo = args.exnihilo, toplevel_has_inline_etymology = __state.toplevel_has_inline_etymology, toplevel_redundant_etymology = __state.toplevel_redundant_etymology, toplevel_idless_etymon = __state.toplevel_idless_etymon, has_mismatched_id = __state.has_mismatched_id, linked_page_multiple_etymons_idless = __state.linked_page_multiple_etymons_idless, linked_page_partial_etymology_sections = __state.linked_page_partial_etymology_sections, text_stop_lang_missing = text_stop_lang_missing, }) M.tracking.track_keywords(__state.toplevel_keyword_stats, lang) M.tracking.track_page_id(lang, id) M.tracking.track_ids(__state.id_stats, lang) end if #categories > 0 then table.insert(output, M.categories.format(categories, lang)) end if __state.warnings then for i, warning in ipairs(__state.warnings) do table.insert(output, (i == 1 and "\n" or "") .. warning .. "\n") end end return table.concat(output) end return export 83f0fr5rgmxbm610bma1l6w6fdin4eo Wikikamus:Penyelia/Pengundian/Lynumiss untuk penyelia 4 September 2026 4 141417 375360 375335 2026-09-22T03:18:06Z PeaceSeekers 3334 /* Keputusan */ Balas 375360 wikitext text/x-wiki <big>'''[[Pengguna:Lynumiss|Lynumiss]]'''</big> ([[Perbincangan Pengguna:Lynumiss |Perbualan]] • [[Khas:Sumbangan/Lynumiss|Sumbangan]] • [[Khas:Pusat_pengesahan/Lynumiss|Sejagat]] • <span class="plainlinks">[https://xtools.wmcloud.org/ec/ms.wiktionary.org/Lynumiss Statistik]</span>) :{| class=wikitable |- | '''Diusulkan oleh''' | [https://ms.wiktionary.org/wiki/Wikikamus:Penyelia/Permohonan#Lynumiss_(Perbincangan_-_Sumbangan) Rombituon] |- | '''Jumlah suntingan''' | 1,297 setakat 03:33, 4 September 2026 (UTC) |- | '''Mekanisme''' | Waktu pengundian: * Bermula pada 4 September 2026. * Berakhir pada <strike>18 September 2026 (2 minggu).</strike> 20 September 2026. Syarat pengundian: * Pengguna yang menyokong minimum sebanyak tiga (3) orang. * Dipersetujui minimum 70% pengundi. Undian 'berkecuali' tidak diambil kira. * Semua [[Wikikamus:Penyelia/Polisi#Takrifan dan istilah umum|pengguna berdaftar sah]] boleh mengundi. |} == Undian == <!-- Guna {{Undi|Y}} atau {{Undi|T}} diikuti dengan ~~~~ untuk menandatangan undian--> # {{Undi|Y}} [[Pengguna:Ultron90|Ultron90]] ([[Perbincangan pengguna:Ultron90|bincang]]) 13:49, 4 September 2026 (UTC) # {{Undi|Y}} - [[Pengguna:Rulwarih|Rulwarih]] ([[Perbincangan pengguna:Rulwarih|bincang]]) 13:26, 15 September 2026 (UTC) # {{Undi|Y}} - [[Pengguna:SNN95|SNN95]] ([[Perbincangan pengguna:SNN95|bincang]]) 12:25, 19 September 2026 (UTC) # {{Undi|Y}} - '''[[Pengguna:Hakimi97|محمد حكيمي]]''' ([[Perbincangan pengguna:Hakimi97|bincang]]) 12:28, 19 September 2026 (UTC) == Komen == == Keputusan == Memandangkan undi belum cukup setakat 14 hari, saya panjangkan sedikit tempoh pengundian hingga 20 haribulan. [[Pengguna:PeaceSeekers|PeaceSeekers]] ([[Perbincangan pengguna:PeaceSeekers|bincang]]) 12:26, 19 September 2026 (UTC) :Undian tamat dengan jumlah persetujuan mencukupi. Pencalonan diterima dan akan dibawa ke Meta bagi urusan lanjut. [[Pengguna:PeaceSeekers|PeaceSeekers]] ([[Perbincangan pengguna:PeaceSeekers|bincang]]) 03:18, 22 September 2026 (UTC) ek78zcm0kkxnl1w9wh63xphveggaml7 Wikikamus:bdr/lawa 4 145784 375343 2026-09-21T14:06:26Z Sipatung 10839 Mencipta laman 375343 wikitext text/x-wiki ==Bahasa {{bahasa|bdr}}== ===Kata sifat=== {{inti|bdr|kata sifat}} # {{label|1=bdr|2=dialek|3=Sabah}} lawa {{cp|bdr|Lukisan kekanak a '''lawa''' bana.|Lukisan budak itu sangat '''[[cantik]]'''.}} mmsu7s3eipb4pyfjer1r5e6sjpceffg