Wikikamus
mswiktionary
https://ms.wiktionary.org/wiki/Wikikamus:Laman_Utama
MediaWiki 1.47.0-wmf.20
case-sensitive
Media
Khas
Perbincangan
Pengguna
Perbincangan pengguna
Wikikamus
Perbincangan Wikikamus
Fail
Perbincangan fail
MediaWiki
Perbincangan MediaWiki
Templat
Perbincangan templat
Bantuan
Perbincangan bantuan
Kategori
Perbincangan kategori
Lampiran
Perbincangan lampiran
Rima
Perbincangan rima
Tesaurus
Perbincangan tesaurus
Indeks
Perbincangan indeks
Petikan
Perbincangan petikan
Rekonstruksi
Perbincangan rekonstruksi
Padanan isyarat
Perbincangan padanan isyarat
Konkordans
Perbincangan konkordans
TimedText
TimedText talk
Modul
Perbincangan modul
Acara
Perbincangan acara
Modul:languages
828
8666
375364
375316
2026-09-22T04:33:42Z
Hakimi97
2668
Membatalkan semakan [[Special:Diff/375316|375316]] oleh [[Special:Contributions/Hakimi97|Hakimi97]] ([[User talk:Hakimi97|bincang]])
375364
Scribunto
text/plain
--[==[ intro:
This module implements fetching of language-specific information and processing text in a given language.
===Types of languages===
There are two types of languages: full languages and etymology-only languages. The essential difference is that only
full languages appear in L2 headings in vocabulary entries, and hence categories like [[:Category:French nouns]] exist
only for full languages. Etymology-only languages have either a full language or another etymology-only language as
their parent (in the parent-child inheritance sense), and for etymology-only languages with another etymology-only
language as their parent, a full language can always be derived by following the parent links upwards. For example,
"Canadian French", code `fr-CA`, is an etymology-only language whose parent is the full language "French", code `fr`.
An example of an etymology-only language with another etymology-only parent is "Northumbrian Old English", code
`ang-nor`, which has "Anglian Old English", code `ang-ang` as its parent; this is an etymology-only language whose
parent is "Old English", code `ang`, which is a full language. (This is because Northumbrian Old English is considered
a variety of Anglian Old English.) Sometimes the parent is the "Undetermined" language, code `und`; this is the case,
for example, for "substrate" languages such as "Pre-Greek", code `qsb-grc`, and "the BMAC substrate", code `qsb-bma`.
It is important to distinguish language ''parents'' from language ''ancestors''. The parent-child relationship is one
of containment, i.e. if X is a child of Y, X is considered a variety of Y. On the other hand, the ancestor-descendant
relationship is one of descent in time. For example, "Classical Latin", code `la-cla`, and "Late Latin", code `la-lat`,
are both etymology-only languages with "Latin", code `la`, as their parents, because both of the former are varieties
of Latin. However, Late Latin does *NOT* have Classical Latin as its parent because Late Latin is *not* a variety of
Classical Latin; rather, it is a descendant. There is in fact a separate `ancestors` field that is used to express the
ancestor-descendant relationship, and Late Latin's ancestor is given as Classical Latin. It is also important to note
that sometimes an etymology-only language is actually the conceptual ancestor of its parent language. This happens,
for example, with "Old Italian" (code `roa-oit`), which is an etymology-only variant of full language "Italian" (code
`it`), and with "Old Latin" (code `itc-ola`), which is an etymology-only variant of Latin. In both cases, the full
language has the etymology-only variant listed as an ancestor. This allows a Latin term to inherit from Old Latin
using the {{tl|inh}} template (where in this template, "inheritance" refers to ancestral inheritance, i.e. inheritance
in time, rather than in the parent-child sense); likewise for Italian and Old Italian.
Full languages come in three subtypes:
* {regular}: This indicates a full language that is attested according to [[WT:CFI]] and therefore permitted in the
main namespace. There may also be reconstructed terms for the language, which are placed in the
{Reconstruction} namespace and must be prefixed with * to indicate a reconstruction. Most full languages
are natural (not constructed) languages, but a few constructed languages (e.g. Esperanto and Volapük,
among others) are also allowed in the mainspace and considered regular languages.
* {reconstructed}: This language is not attested according to [[WT:CFI]], and therefore is allowed only in the
{Reconstruction} namespace. All terms in this language are reconstructed, and must be prefixed with
*. Languages such as Proto-Indo-European and Proto-Germanic are in this category.
* {appendix-constructed}: This language is attested but does not meet the additional requirements set out for
constructed languages ([[WT:CFI#Constructed languages]]). Its entries must therefore be in
the Appendix namespace, but they are not reconstructed and therefore should not have *
prefixed in links. Most constructed languages are of this subtype.
Both full languages and etymology-only languages have a {Language} object associated with them, which is fetched using
the {getByCode} function in [[Modul:languages]] to convert a language code to a {Language} object. Depending on the
options supplied to this function, etymology-only languages may or may not be accepted, and family codes may be
accepted (returning a {Family} object as described in [[Modul:families]]). There are also separate {getByCanonicalName}
functions in [[Modul:languages]] and [[Modul:etymology languages]] to convert a language's canonical name to a
{Language} object (depending on whether the canonical name refers to a full or etymology-only language).
===Textual representations===
Textual strings belonging to a given language come in several different ''text variants'':
# The ''input text'' is what the user supplies in wikitext, in the parameters to {{tl|m}}, {{tl|l}}, {{tl|ux}},
{{tl|t}}, {{tl|lang}} and the like.
# The ''corrected input text'' is the input text with some corrections and/or normalizations applied, such as
bad-character replacements for certain languages, like replacing `l` or `1` to [[palochka]] in some languages written
in Cyrillic. (FIXME: This currently goes under the name ''display text'' but that will be repurposed below. Also,
[[User:Surjection]] suggests renaming this to ''normalized input text'', but "normalized" is used in a different sense
in [[Modul:usex]].)
# The ''display text'' is the text in the form as it will be displayed to the user. This is what appears in headwords,
in usexes, in displayed internal links, etc. This can include accent marks that are removed to form the stripped
display text (see below), as well as embedded bracketed links that are variously processed further. The display text
is generated from the corrected input text by applying language-specific transformations; for most languages, there
will be no such transformations. The general reason for having a difference between input and display text is to allow
for extra information in the input text that is not displayed to the user but is sent to the transliteration module.
Note that having different display and input text is only supported currently through special-casing but will be
generalized. Examples of transformations are: (1) Removing the {{cd|^}} that is used in certain East Asian (and
possibly other unicameral) languages to indicate capitalization of the transliteration (which is currently
special-cased); (2) for Korean, removing or otherwise processing hyphens (which is currently special-cased); (3) for
Arabic, removing a ''sukūn'' diacritic placed over a ''tāʔ marbūṭa'' (like this: ةْ) to indicate that the
''tāʔ marbūṭa'' is pronounced and transliterated as /t/ instead of being silent [NOTE, NOT IMPLEMENTED YET]; (4) for
Thai and Khmer, converting space-separated words to bracketed words and resolving respelling substitutions such as
`[กรีน/กฺรีน]`, which indicate how to transliterate given words [NOTE, NOT IMPLEMENTED YET except in language-specific
templates like {{tl|th-usex}}].
## The ''right-resolved display text'' is the result of removing brackets around one-part embedded links and resolving
two-part embedded links into their right-hand components (i.e. converting two-part links into the displayed form).
The process of right-resolution is what happens when you call {{cd|remove_links()}} in [[Modul:links]] on some text.
When applied to the display text, it produces exactly what the user sees, without any link markup.
# The ''stripped display text'' is the result of applying diacritic-stripping to the display text.
## The ''left-resolved stripped display text'' [NEED BETTER NAME] is the result of applying left-resolution to the
stripped display text, i.e. similar to right-resolution but resolving two-part embedded links into their left-hand
components (i.e. the linked-to page). If the display text refers to a single page, the resulting of applying
diacritic stripping and left-resolution produces the ''logical pagename''.
# The ''physical pagename text'' is the result of converting the stripped display text into physical page links. If the
stripped display text contains embedded links, the left side of those links is converted into physical page links;
otherwise, the entire text is considered a pagename and converted in the same fashion. The conversion does three
things: (1) converts characters not allowed in pagenames into their "unsupported title" representation, e.g.
{{cd|Unsupported titles/`gt`}} in place of the logical name {{cd|>}}; (2) handles certain special-cased
unsupported-title logical pagenames, such as {{cd|Unsupported titles/Space}} in place of {{cd|[space]}} and
{{cd|Unsupported titles/Ancient Greek dish}} in place of a very long Greek name for a gourmet dish as found in
Aristophanes; (3) converts "mammoth" pagenames such as [[a]] into their appropriate split component, e.g.
[[a/languages A to L]].
# The ''source translit text'' is the text as supplied to the language-specific {{cd|transliterate()}} method. The form
of the source translit text may need to be language-specific, e.g Thai and Khmer will need the corrected input text,
whereas other languages may need to work off the display text. [FIXME: It's still unclear to me how embedded bracketed
links are handled in the existing code.] In general, embedded links need to be right-resolved (see above), but when
this happens is unclear to me [FIXME]. Some languages have a chop-up-and-paste-together scheme that sends parts of the
text through the transliterate mechanism, and for others (those listed with "cont" in {{cd|substitution}} in
[[Modul:languages/data]]) they receive the full input text, but preprocessed in certain ways. (The wisdom of this is
still unclear to me.)
# The ''transliterated text'' (or ''transliteration'') is the result of transliterating the source translit text. Unlike
for all the other text variants except the transcribed text, it is always in the Latin script.
# The ''transcribed text'' (or ''transcription'') is the result of transcribing the source translit text, where
"transcription" here means a close approximation to the phonetic form of the language in languages (e.g. Akkadian,
Sumerian, Ancient Egyptian, maybe Tibetan) that have a wide difference between the written letters and spoken form.
Unlike for all the other text variants other than the transliterated text, it is always in the Latin script.
Currently, the transcribed text is always supplied manually be the user; there is no such thing as a
{{cd|transcribe()}} method on language objects.
# The ''sort key'' is the text used in sort keys for determining the placing of pages in categories they belong to. The
sort key is generated from the pagename or a specified ''sort base'' by lowercasing, doing language-specific
transformations and then uppercasing the result. If the sort base is supplied and is generated from input text, it
needs to be converted to display text, have embedded links removed through right-resolution and have
diacritic-stripping applied.
# There are other text variants that occur in usexes (specifically, there are normalized variants of several of the
above text variants), but we can skip them for now.
The following methods exist on {Language} objects to convert between different text variants:
# {correctInputText} (currently called {makeDisplayText}): This converts input text to corrected input text.
# {stripDiacritics}: This converts to stripped display text. [FIXME: This needs some rethinking. In particular,
{stripDiacritics} is sometimes called on input text, corrected input text or display text (in various paths inside of
[[Modul:links]], and, in the case of input text, usually from other modules). We need to make sure we don't try to
convert input text to display text twice, but at the same time we need to support calling it directly on input text
since so many modules do this. This means we need to add a parameter indicating whether the passed-in text is input,
corrected input, or display text; if the former two, we call {correctInputText} ourselves.]
# {logicalToPhysical}: This converts logical pagenames to physical pagenames.
# {transliterate}: This appears to convert input text with embedded brackets removed into a transliteration.
[FIXME: This needs some rethinking. In particular, it calls {processDisplayText} on its input, which won't work
for Thai and Khmer, so we may need language-specific flags indicating whether to pass the input text directly to the
language transliterate method. In addition, I'm not sure how embedded links are handled in the existing translit code;
a lot of callers remove the links themselves before calling {transliterate()}, which I assume is wrong.]
# {makeSortKey}: This converts display text (?) to a sort key. [FIXME: Clarify this.]
]==]
local export = {}
local debug_track_module = "Modul:debug/track"
local etymology_languages_data_module = "Modul:etymology languages/data"
local families_module = "Modul:families"
local headword_page_module = "Modul:headword/page"
local json_module = "Modul:JSON"
local language_like_module = "Modul:language-like"
local languages_data_module = "Modul:languages/data"
local languages_data_patterns_module = "Modul:languages/data/patterns"
local links_data_module = "Modul:links/data"
local load_module = "Modul:load"
local scripts_module = "Modul:scripts"
local scripts_data_module = "Modul:scripts/data"
local string_encode_entities_module = "Modul:string/encode entities"
local string_pattern_escape_module = "Modul:string/patternEscape"
local string_replacement_escape_module = "Modul:string/replacementEscape"
local string_utilities_module = "Modul:string utilities"
local table_module = "Modul:table"
local utilities_module = "Modul:utilities"
local wikimedia_languages_module = "Modul:wikimedia languages"
local mw = mw
local string = string
local table = table
local char = string.char
local concat = table.concat
local find = string.find
local floor = math.floor
local get_by_code -- Defined below.
local get_data_module_name -- Defined below.
local get_extra_data_module_name -- Defined below.
local getmetatable = getmetatable
local gmatch = string.gmatch
local gsub = string.gsub
local insert = table.insert
local ipairs = ipairs
local is_known_language_tag = mw.language.isKnownLanguageTag
local make_object -- Defined below.
local match = string.match
local next = next
local pairs = pairs
local remove = table.remove
local require = require
local select = select
local setmetatable = setmetatable
local sub = string.sub
local type = type
local unstrip = mw.text.unstrip
-- Loaded as needed by findBestScript.
local Hans_chars
local Hant_chars
local function check_object(...)
check_object = require(utilities_module).check_object
return check_object(...)
end
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function decode_entities(...)
decode_entities = require(string_utilities_module).decode_entities
return decode_entities(...)
end
local function decode_uri(...)
decode_uri = require(string_utilities_module).decode_uri
return decode_uri(...)
end
local function deep_copy(...)
deep_copy = require(table_module).deepCopy
return deep_copy(...)
end
local function encode_entities(...)
encode_entities = require(string_encode_entities_module)
return encode_entities(...)
end
local function get_L2_sort_key(...)
get_L2_sort_key = require(headword_page_module).get_L2_sort_key
return get_L2_sort_key(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function find_best_script_without_lang(...)
find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang
return find_best_script_without_lang(...)
end
local function get_family(...)
get_family = require(families_module).getByCode
return get_family(...)
end
local function get_plaintext(...)
get_plaintext = require(utilities_module).get_plaintext
return get_plaintext(...)
end
local function get_wikimedia_lang(...)
get_wikimedia_lang = require(wikimedia_languages_module).getByCode
return get_wikimedia_lang(...)
end
local function keys_to_list(...)
keys_to_list = require(table_module).keysToList
return keys_to_list(...)
end
local function list_to_set(...)
list_to_set = require(table_module).listToSet
return list_to_set(...)
end
local function load_data(...)
load_data = require(load_module).load_data
return load_data(...)
end
local function make_family_object(...)
make_family_object = require(families_module).makeObject
return make_family_object(...)
end
local function pattern_escape(...)
pattern_escape = require(string_pattern_escape_module)
return pattern_escape(...)
end
local function replacement_escape(...)
replacement_escape = require(string_replacement_escape_module)
return replacement_escape(...)
end
local function safe_require(...)
safe_require = require(load_module).safe_require
return safe_require(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local function to_json(...)
to_json = require(json_module).toJSON
return to_json(...)
end
local function u(...)
u = require(string_utilities_module).char
return u(...)
end
local function ugsub(...)
ugsub = require(string_utilities_module).gsub
return ugsub(...)
end
local function ulen(...)
ulen = require(string_utilities_module).len
return ulen(...)
end
local function ulower(...)
ulower = require(string_utilities_module).lower
return ulower(...)
end
local function umatch(...)
umatch = require(string_utilities_module).match
return umatch(...)
end
local function uupper(...)
uupper = require(string_utilities_module).upper
return uupper(...)
end
local function track(page)
debug_track("languages/" .. page)
return true
end
local function normalize_code(code)
return load_data(languages_data_module).aliases[code] or code
end
local function check_inputs(self, check, default, ...)
local n = select("#", ...)
if n == 0 then
return false
end
local ret = check(self, (...))
if ret ~= nil then
return ret
elseif n > 1 then
local inputs = {...}
for i = 2, n do
ret = check(self, inputs[i])
if ret ~= nil then
return ret
end
end
end
return default
end
local function make_link(self, target, display)
local prefix, main
if self:getFamilyCode() == "qfa-sub" then
prefix, main = display:match("^(the )(.*)")
if not prefix then
prefix, main = display:match("^(a )(.*)")
end
end
return (prefix or "") .. "[[" .. target .. "|" .. (main or display) .. "]]"
end
-- Convert risky characters to HTML entities, which minimizes interference once returned (e.g. for "sms:a", "<!-- -->" etc.).
local function escape_risky_characters(text)
-- Spacing characters in isolation generally need to be escaped in order to be properly processed by the MediaWiki
-- software.
if umatch(text, "^%s*$") then
return encode_entities(text, text)
end
return encode_entities(text, "!#%&*+/:;<=>?@[\\]_{|}")
end
-- Temporarily convert various formatting characters to PUA to prevent them from being disrupted by the substitution process.
local function doTempSubstitutions(text, subbedChars, keepCarets, noTrim)
-- Clone so that we don't insert any extra patterns into the table in package.loaded. For some reason, using require
-- seems to keep memory use down; probably because the table is always cloned.
local patterns = shallow_copy(require(languages_data_patterns_module))
if keepCarets then
insert(patterns, "((\\+)%^)")
insert(patterns, "((%^))")
end
-- Ensure any whitespace at the beginning and end is temp substituted, to prevent it from being accidentally
-- trimmed. We only want to trim any final spaces added during the substitution process (e.g. by a module), which
-- means we only do this during the first round of temp substitutions.
if not noTrim then
insert(patterns, "^([\128-\191\244]*(%s+))")
insert(patterns, "((%s+)[\128-\191\244]*)$")
end
-- Pre-substitution, of "[[" and "]]", which makes pattern matching more accurate.
text = gsub(text, "%f[%[]%[%[", "\1"):gsub("%f[%]]%]%]", "\2")
local i = #subbedChars
for _, pattern in ipairs(patterns) do
-- Patterns ending in \0 stand are for things like "[[" or "]]"), so the inserted PUA are treated as breaks
-- between terms by modules that scrape info from pages.
local term_divider
pattern = gsub(pattern, "%z$", function(divider)
term_divider = divider == "\0"
return ""
end)
text = gsub(text, pattern, function(...)
local m = {...}
local m1New = m[1]
for k = 2, #m do
local n = i + k - 1
subbedChars[n] = m[k]
local byte2 = floor(n / 4096) % 64 + (term_divider and 128 or 136)
local byte3 = floor(n / 64) % 64 + 128
local byte4 = n % 64 + 128
m1New = gsub(m1New, pattern_escape(m[k]), "\244" .. char(byte2) .. char(byte3) .. char(byte4), 1)
end
i = i + #m - 1
return m1New
end)
end
text = gsub(text, "\1", "%[%["):gsub("\2", "%]%]")
return text, subbedChars
end
-- Reinsert any formatting that was temporarily substituted.
local function undoTempSubstitutions(text, subbedChars)
for i = 1, #subbedChars do
local byte2 = floor(i / 4096) % 64 + 128
local byte3 = floor(i / 64) % 64 + 128
local byte4 = i % 64 + 128
text = gsub(text, "\244[" .. char(byte2) .. char(byte2+8) .. "]" .. char(byte3) .. char(byte4),
replacement_escape(subbedChars[i]))
end
text = gsub(text, "\1", "%[%["):gsub("\2", "%]%]")
return text
end
-- Check if the raw text is an unsupported title, and if so return that. Otherwise, remove HTML entities. We do the
-- pre-conversion to avoid loading the unsupported title list unnecessarily.
local function checkNoEntities(text)
local textNoEnc = decode_entities(text)
if textNoEnc ~= text and load_data(links_data_module).unsupported_titles[text] then
return text
else
return textNoEnc
end
end
-- If no script object is provided (or if it's invalid or None), get one.
local function checkScript(text, self, sc)
if not check_object("script", true, sc) or sc:getCode() == "None" then
return self:findBestScript(text)
end
return sc
end
local function normalize(text, sc)
text = sc:fixDiscouragedSequences(text)
return sc:toFixedNFD(text)
end
--[=[
Subfunction of iterateSectionSubstitutions(). Process an individual chunk of text according to the specifications in
`substitution_data`. The input parameters are all as in the documentation of iterateSectionSubstitutions() except for
`recursed`, which is set to true if we called ourselves recursively to process a script-specific setting or
script-wide fallback. Returns two values: the processed text and the actual substitution data used to do the
substitutions (same as the `actual_substitution_data` return value to iterateSectionSubstitutions()).
]=]
local function doSubstitutions(self, text, sc, substitution_data, data_field, function_name, recursed)
-- BE CAREFUL in this function because the value at any level can be `false`, which causes no processing to be done
-- and blocks any further fallback processing.
local actual_substitution_data = substitution_data
-- If there are language-specific substitutes given in the data module, use those.
if type(substitution_data) == "table" then
-- If a script is specified, run this function with the script-specific data before continuing.
local sc_code = sc:getCode()
local has_substitution_data = false
if substitution_data[sc_code] ~= nil then
has_substitution_data = true
if substitution_data[sc_code] then
text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data[sc_code], data_field,
function_name, true)
end
-- Hant, Hans and Hani are usually treated the same, so add a special case to avoid having to specify each one
-- separately.
elseif sc_code:match("^Han") and substitution_data.Hani ~= nil then
has_substitution_data = true
if substitution_data.Hani then
text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data.Hani, data_field,
function_name, true)
end
-- Substitution data with key 1 in the outer table may be given as a fallback.
elseif substitution_data[1] ~= nil then
has_substitution_data = true
if substitution_data[1] then
text, actual_substitution_data = doSubstitutions(self, text, sc, substitution_data[1], data_field,
function_name, true)
end
end
-- Iterate over all strings in the "from" subtable, and gsub with the corresponding string in "to". We work with
-- the NFD decomposed forms, as this simplifies many substitutions.
if substitution_data.from then
has_substitution_data = true
for i, from in ipairs(substitution_data.from) do
-- Normalize each loop, to ensure multi-stage substitutions work correctly.
text = sc:toFixedNFD(text)
text = ugsub(text, sc:toFixedNFD(from), substitution_data.to[i] or "")
end
end
if substitution_data.remove_diacritics then
has_substitution_data = true
text = sc:toFixedNFD(text)
-- Convert exceptions to PUA.
local remove_exceptions, substitutes = substitution_data.remove_exceptions
if remove_exceptions then
substitutes = {}
local i = 0
for _, exception in ipairs(remove_exceptions) do
exception = sc:toFixedNFD(exception)
text = ugsub(text, exception, function(m)
i = i + 1
local subst = u(0x80000 + i)
substitutes[subst] = m
return subst
end)
end
end
-- Strip diacritics.
text = ugsub(text, "[" .. substitution_data.remove_diacritics .. "]", "")
-- Convert exceptions back.
if remove_exceptions then
text = text:gsub("\242[\128-\191]*", substitutes)
end
end
if not has_substitution_data and sc._data[data_field] then
-- If language-specific sort key (etc.) is nil, fall back to script-wide sort key (etc.).
text, actual_substitution_data = doSubstitutions(self, text, sc, sc._data[data_field], data_field,
function_name, true)
end
elseif type(substitution_data) == "string" then
-- If there is a dedicated function module, use that.
local module = safe_require("Modul:" .. substitution_data)
if module then
-- TODO: translit functions should take objects, not codes.
-- TODO: translit functions should be called with form NFD.
if function_name == "tr" then
if not module[function_name] then
error(("Internal error: Module [[%s]] has no function named 'tr'"):format(substitution_data))
end
text = module[function_name](text, self._code, sc:getCode())
elseif function_name == "stripDiacritics" then
-- FIXME, get rid of this arm after renaming makeEntryName -> stripDiacritics.
if module[function_name] then
text = module[function_name](sc:toFixedNFD(text), self, sc)
elseif module.makeEntryName then
text = module.makeEntryName(sc:toFixedNFD(text), self, sc)
else
error(("Internal error: Module [[%s]] has no function named 'stripDiacritics' or 'makeEntryName'"
):format(substitution_data))
end
else
if not module[function_name] then
error(("Internal error: Module [[%s]] has no function named '%s'"):format(
substitution_data, function_name))
end
text = module[function_name](sc:toFixedNFD(text), self, sc)
end
else
error("Substitution data '" .. substitution_data .. "' does not match an existing module.")
end
elseif substitution_data == nil and sc._data[data_field] then
-- If language-specific sort key (etc.) is nil, fall back to script-wide sort key (etc.).
text, actual_substitution_data = doSubstitutions(self, text, sc, sc._data[data_field], data_field,
function_name, true)
end
-- Don't normalize to NFC if this is the inner loop or if a module returned nil.
if recursed or not text then
return text, actual_substitution_data
end
-- Fix any discouraged sequences created during the substitution process, and normalize into the final form.
return sc:toFixedNFC(sc:fixDiscouragedSequences(text)), actual_substitution_data
end
--[=[
Split the text into sections, based on the presence of temporarily substituted formatting characters, then iterate over
each section to apply substitutions (e.g. transliteration or diacritic stripping). This avoids putting PUA (Private Use
Area) characters through language-specific modules, which may be unequipped for them. This function is passed the
following values:
* `self` (the Language object);
* `text` (the text to process);
* `sc` (the script of the text, which must be specified; callers should call checkScript() as needed to autodetect the
script of the text if not given explicitly by the user);
* `subbedChars` (an array of the same length as the text, indicating which characters have been substituted and by
what, or {nil} if no substitutions are to happen);
* `keepCarets` (DOCUMENT ME);
* `substitution_data` (the data indicating which substitutions to apply, taken directly from `data_field` in the
language's data structure in a submodule of [[Modul:languages/data]]);
* `data_field` (the field from which `substitution_data` was fetched, such as {"sort_key"} or {"strip_diacritics"});
* `function_name` (the name of the function to call to do the substitution, in case `substitution_data` specifies a
module to do the substitution);
* `notrim` (don't trim whitespace at the edges of `text`; set when computing the sort key, because whitespace at the
beginning of a sort key is significant and causes the resulting page to be sorted at the beginning of the category
it's in).
Return three values:
# the processed text;
# the value of `subbedChars` that was passed in, possibly modified with additional character substitutions; will be
{nil} if {nil} was passed in;
# the actual substitution data that was used to apply substitutions to `text`; this may be different from the value
of `substitution_data` passed in if that value recursively specified script-specific substitutions or if no
substitution data could be found in the language-specific data (e.g. {nil} was passed in or a structure was passed
in that had no setting for the script given in `sc`), but a script-wide fallback value was set; currently it is
only used by {makeSortKey()}.
]=]
local function iterateSectionSubstitutions(self, text, sc, subbedChars, keepCarets, substitution_data, data_field,
function_name, notrim)
local sections
-- See [[Modul:languages/data]].
if not find(text, "\244") or load_data(languages_data_module).substitution[self._code] == "cont" then
sections = {text}
else
sections = split(text, "\244[\128-\143][\128-\191]*", true)
end
local actual_substitution_data
for _, section in ipairs(sections) do
-- Don't bother processing empty strings or whitespace (which may also not be handled well by dedicated
-- modules).
if gsub(section, "%s+", "") ~= "" then
local sub, this_actual_substitution_data = doSubstitutions(self, section, sc, substitution_data, data_field,
function_name)
actual_substitution_data = this_actual_substitution_data
-- Second round of temporary substitutions, in case any formatting was added by the main substitution
-- process. However, don't do this if the section contains formatting already (as it would have had to have
-- been escaped to reach this stage, and therefore should be given as raw text).
if sub and subbedChars then
local noSub
for _, pattern in ipairs(require(languages_data_patterns_module)) do
if match(section, pattern .. "%z?") then
noSub = true
end
end
if not noSub then
sub, subbedChars = doTempSubstitutions(sub, subbedChars, keepCarets, true)
end
end
if not sub then
text = sub
break
end
text = sub and gsub(text, pattern_escape(section), replacement_escape(sub), 1) or text
end
end
if not notrim then
-- Trim, unless there are only spacing characters, while ignoring any final formatting characters.
-- Do not trim sort keys because spaces at the beginning are significant.
text = text and text:gsub("^([\128-\191\244]*)%s+(%S)", "%1%2"):gsub("(%S)%s+([\128-\191\244]*)$", "%1%2") or
nil
end
return text, subbedChars, actual_substitution_data
end
-- Process carets (and any escapes). Default to simple removal, if no pattern/replacement is given.
local function processCarets(text, pattern, repl)
local rep
repeat
text, rep = gsub(text, "\\\\(\\*^)", "\3%1")
until rep == 0
return (text:gsub("\\^", "\4")
:gsub(pattern or "%^", repl or "")
:gsub("\3", "\\")
:gsub("\4", "^"))
end
-- Remove carets if they are used to capitalize parts of transliterations (unless they have been escaped).
local function removeCarets(text, sc)
if not sc:hasCapitalization() and sc:isTransliterated() and text:find("^", 1, true) then
return processCarets(text)
else
return text
end
end
local Language = {}
--[==[
Return the language code of the language. Example: {fr"} for French.
]==]
function Language:getCode()
return self._code
end
--[==[
Return the canonical name of the language. This is the name used to represent that language on Wiktionary, and is
guaranteed to be unique to that language alone. Example: {"French"} for French.
]==]
function Language:getCanonicalName()
local name = self._name
if name == nil then
name = self._data[1]
self._name = name
end
return name
end
--[==[
Return the display form of the language. The display form of a language, family or script is the form it takes when
appearing as the <code><var>source</var></code> in categories such as <code>English terms derived from
<var>source</var></code> or <code>English given names from <var>source</var></code>, and is also the displayed text
in `##makeCategoryLink()` links. For full and etymology-only languages, this is the same as the canonical name, but
for families, it reads <code>"<var>name</var> languages"</code> (e.g. {"Indo-Iranian languages"}), and for scripts,
it reads <code>"<var>name</var> script"</code> (e.g. {"Arabic script"}).
]==]
function Language:getDisplayForm()
local form = self._displayForm
if form == nil then
form = self:getCanonicalName()
-- Add article and " substrate" to substrates that lack them.
if self:getFamilyCode() == "qfa-sub" then
if not (sub(form, 1, 4) == "the " or sub(form, 1, 2) == "a ") then
form = "a " .. form
end
if not match(form, " [Ss]ubstrate") then
form = form .. " substrate"
end
end
self._displayForm = form
end
return form
end
--[==[
Return the value which should be used in the HTML `lang=` attribute for tagged text in the language.
]==]
function Language:getHTMLAttribute(sc, region)
local code = self._code
if not find(code, "-", 1, true) then
return code .. "-" .. sc:getCode() .. (region and "-" .. region or "")
end
local parent = self:getParent()
region = region or match(code, "%f[%u][%u-]+%f[%U]")
if parent then
return parent:getHTMLAttribute(sc, region)
end
-- TODO: ISO family codes can also be used.
return "mis-" .. sc:getCode() .. (region and "-" .. region or "")
end
--[==[
Return a list of the aliases that the language is known by, excluding the canonical name. Aliases are synonyms for the
language in question. The names are not guaranteed to be unique, in that sometimes more than one language is known by
the same name. Example: { {"High German", "New High German", "Deutsch"}} for {{lnl|de}}.
]==]
function Language:getAliases()
self:loadInExtraData()
return require(language_like_module).getAliases(self)
end
--[==[
Return a list of the known subvarieties of a given language, excluding subvarieties that have been given explicit
etymology-only language codes. The names are not guaranteed to be unique, in that sometimes a given name refers to a
subvariety of more than one language. Example: { {"Southern Aymara", "Central Aymara"}} for {{lnl|ay}}. Note that the
returned value can have nested tables in it, when a subvariety goes by more than one name. Example:
{ {"North Azerbaijani", "South Azerbaijani", {"Afshar", "Afshari", "Afshar Azerbaijani", "Afchar"}, {"Qashqa'i", "Qashqai", "Kashkay"}, "Sonqor"}}
for {{lnl|az}}. Here, for example, Afshar, Afshari, Afshar Azerbaijani and Afchar all refer to the same subvariety,
whose preferred name is Afshar (the one listed first). To avoid a return value with nested tables in it, specify a
non-{nil} value for the `flatten` parameter; in that case, the return value would be
{ {"North Azerbaijani", "South Azerbaijani", "Afshar", "Afshari", "Afshar Azerbaijani", "Afchar", "Qashqa'i", "Qashqai", "Kashkay", "Sonqor"}}.
]==]
function Language:getVarieties(flatten)
self:loadInExtraData()
return require(language_like_module).getVarieties(self, flatten)
end
--[==[
Return a table of the "other names" that the language is known by, which are listed in the `other_names` field in the
language's extra-data module. It should be noted that the `other_names` field itself is deprecated, and entries listed
there should eventually be moved to either `aliases` or `varieties`, or removed if they refer to larger entities which
the language in question is actually a part of.
]==]
function Language:getOtherNames() -- To be eventually removed, once there are no more uses of the `other_names` field.
self:loadInExtraData()
return require(language_like_module).getOtherNames(self)
end
--[==[
Return a combined table of the canonical name, aliases, varieties and other names of a given language.
]==]
function Language:getAllNames()
self:loadInExtraData()
return require(language_like_module).getAllNames(self)
end
--[==[
Return a table of types as a set (with the types as keys). The possible types are
* {language}: This is a language, either full or etymology-only.
* {full}: This is a "full" (not etymology-only) language, i.e. the union of {regular}, {reconstructed} and
{appendix-constructed}. Note that the types {full} and {etymology-only} also exist for families, so if you want to
check specifically for a full language and you have an object that might be a family, you should use
{hasType("language", "full")} and not simply {hasType("full")}.
* {etymology-only}: This is an etymology-only (not full) language, whose parent is another etymology-only language or a
full language. Note that the types {full} and {etymology-only} also exist for families, so if you want to check
specifically for an etymology-only language and you have an object that might be a family, you should use
{hasType("language", "etymology-only")} and not simply {hasType("etymology-only")}.
* {regular}: This indicates a full language that is attested according to [[WT:CFI]] and therefore permitted in the main
namespace. There may also be reconstructed terms for the language, which are placed in the {Reconstruction}
namespace and must be prefixed with `*` to indicate a reconstruction. Most full languages are natural
(not constructed) languages, but a few constructed languages (e.g. Esperanto and Volapük, among others) are also
allowed in the mainspace and considered regular languages.
* {reconstructed}: This language is not attested according to [[WT:CFI]], and therefore is allowed only in the
{Reconstruction} namespace. All terms in this language are reconstructed, and must be prefixed with `*`. Languages
such as Proto-Indo-European and Proto-Germanic are in this category. '''Exception:''' Reconstructed languages can
have entries in the mainspace using the ''anti-asterisk'' feature. This requires that both the headword for the
mainspace term (as specified using {{para|head}}) and links to the term are prefixed with the anti-asterisk
indicator `!!`.
* {appendix-constructed}: This language is attested but does not meet the additional requirements set out for
constructed languages ([[WT:CFI#Constructed languages]]). Its entries must therefore be in the Appendix namespace,
but they are not reconstructed and therefore should not have * prefixed in links.
]==]
function Language:getTypes()
local types = self._types
if types == nil then
types = {language = true}
if self:getFullCode() == self._code then
types.full = true
else
types["etymology-only"] = true
end
for t in gmatch(self._data.type, "[^,]+") do
types[t] = true
end
self._types = types
end
return types
end
--[==[
Given a list of types as strings, return {true} if the language has all of them.
]==]
function Language:hasType(...)
Language.hasType = require(language_like_module).hasType
return self:hasType(...)
end
--[==[
Return a list containing `WikimediaLanguage` objects (see [[Modul:wikimedia languages]]), which represent languages
and their codes as they are used in Wikimedia projects for interwiki linking and such. More than one object may be
returned, as a single Wiktionary language may correspond to multiple Wikimedia languages. For example, Wiktionary's
single code `sh` (Serbo-Croatian) maps to four Wikimedia codes: `sh` (Serbo-Croatian), `bs` (Bosnian), `hr` (Croatian)
and `sr` (Serbian). The code for the Wikimedia language is retrieved from the `wikimedia_codes` property in the data
modules. If that property is not present, the code of the current language is used. If none of the available codes is
actually a valid Wikimedia code, an empty list is returned.
]==]
function Language:getWikimediaLanguages()
local wm_langs = self._wikimediaLanguageObjects
if wm_langs == nil then
local codes = self:getWikimediaLanguageCodes()
wm_langs = {}
for i = 1, #codes do
wm_langs[i] = get_wikimedia_lang(codes[i])
end
self._wikimediaLanguageObjects = wm_langs
end
return wm_langs
end
function Language:getWikimediaLanguageCodes()
local wm_langs = self._wikimediaLanguageCodes
if wm_langs == nil then
wm_langs = self._data.wikimedia_codes
if wm_langs then
wm_langs = split(wm_langs, ",", true, true)
else
local code = self._code
if is_known_language_tag(code) then
wm_langs = {code}
else
-- Inherit, but only if no codes are specified in the data *and*
-- the language code isn't a valid Wikimedia language code.
local parent = self:getParent()
wm_langs = parent and parent:getWikimediaLanguageCodes() or {}
end
end
self._wikimediaLanguageCodes = wm_langs
end
return wm_langs
end
--[==[
Return the name of the Wikipedia article for the language. `project` specifies the language and project to retrieve
the article from, defaulting to {"enwiki"} for the English Wikipedia. Normally if specified it should be the project
code for a specific-language Wikipedia e.g. "zhwiki" for the Chinese Wikipedia, but it can be any project, including
non-Wikipedia ones. If the project is the English Wikipedia and the property {wikipedia_article} is present in the data
module it will be used first. In all other cases, a sitelink will be generated from {:getWikidataItem} (if set). The
resulting value (or lack of value) is cached so that subsequent calls are fast. If no value could be determined, and
`noCategoryFallback` is {false}, `##Language:getCategoryName` is used as fallback; otherwise, {nil} is returned. Note
that if `noCategoryFallback` is {nil} or omitted, it defaults to {false} if the project is the English Wikipedia,
otherwise to {true}. In other words, under normal circumstances, if the English Wikipedia article couldn't be retrieved,
the return value will fall back to a link to the language's category, but this won't normally happen for any other
project.
]==]
function Language:getWikipediaArticle(noCategoryFallback, project)
Language.getWikipediaArticle = require(language_like_module).getWikipediaArticle
return self:getWikipediaArticle(noCategoryFallback, project)
end
function Language:makeWikipediaLink()
return make_link(self, "w:" .. self:getWikipediaArticle(), self:getCanonicalName())
end
--[==[
Return the name of the Wikimedia Commons category page for the language.
]==]
function Language:getCommonsCategory()
Language.getCommonsCategory = require(language_like_module).getCommonsCategory
return self:getCommonsCategory()
end
--[==[
Return the Wikidata item ID for the language or {nil}. This corresponds to the the second field in the data modules.
]==]
function Language:getWikidataItem()
Language.getWikidataItem = require(language_like_module).getWikidataItem
return self:getWikidataItem()
end
--[==[
Return a list of `Script` objects for all scripts that the language is written in. See [[Modul:scripts]].
]==]
function Language:getScripts()
local scripts = self._scriptObjects
if scripts == nil then
local codes = self:getScriptCodes()
if codes[1] == "All" then
scripts = load_data(scripts_data_module)
else
scripts = {}
for i = 1, #codes do
scripts[i] = get_script(codes[i])
end
end
self._scriptObjects = scripts
end
return scripts
end
--[==[
Return the list of script codes in the language's data file.
]==]
function Language:getScriptCodes()
local scripts = self._scriptCodes
if scripts == nil then
scripts = self._data[4]
if scripts then
local codes, n = {}, 0
for code in gmatch(scripts, "[^,]+") do
n = n + 1
-- Special handling of "Hants", which represents "Hani", "Hant" and "Hans" collectively.
if code == "Hants" then
codes[n] = "Hani"
codes[n + 1] = "Hant"
codes[n + 2] = "Hans"
n = n + 2
else
codes[n] = code
end
end
scripts = codes
else
scripts = {"None"}
end
self._scriptCodes = scripts
end
return scripts
end
--[==[
Given some text, iterate through the scripts of a given language trying to find the script that best matches the text.
Return a {Script} object representing the script. If no match is found at all, it returns the {None} script object.
]==]
function Language:findBestScript(text, forceDetect)
if not text or text == "" or text == "-" then
return get_script("None")
end
-- Differs from table returned by getScriptCodes, as Hants is not normalized into its constituents.
local codes = self._bestScriptCodes
if codes == nil then
codes = self._data[4]
codes = codes and split(codes, ",", true, true) or {"None"}
self._bestScriptCodes = codes
end
local first_sc = codes[1]
if first_sc == "All" then
return find_best_script_without_lang(text)
end
local codes_len = #codes
if not (forceDetect or first_sc == "Hants" or codes_len > 1) then
first_sc = get_script(first_sc)
local charset = first_sc.characters
return charset and umatch(text, "[" .. charset .. "]") and first_sc or get_script("None")
end
-- Remove all formatting characters.
text = get_plaintext(text)
-- Remove all spaces and any ASCII punctuation. Some non-ASCII punctuation is script-specific, so can't be removed.
text = ugsub(text, "[%s!\"#%%&'()*,%-./:;?@[\\%]_{}]+", "")
if #text == 0 then
return get_script("None")
end
-- Try to match every script against the text,
-- and return the one with the most matching characters.
local bestcount, bestscript, length = 0
for i = 1, codes_len do
local sc = codes[i]
-- Special case for "Hants", which is a special code that represents whichever of "Hant" or "Hans" best matches,
-- or "Hani" if they match equally. This avoids having to list all three. In addition, "Hants" will be treated
-- as the best match if there is at least one matching character, under the assumption that a Han script is
-- desirable in terms that contain a mix of Han and other scripts (not counting those which use Jpan or Kore).
if sc == "Hants" then
local Hani = get_script("Hani")
if not Hant_chars then
Hant_chars = load_data("Modul:zh/data/ts")
Hans_chars = load_data("Modul:zh/data/st")
end
local t, s, found = 0, 0
-- This is faster than using mw.ustring.gmatch directly.
for ch in gmatch((ugsub(text, "[" .. Hani.characters .. "]", "\255%0")), "\255(.[\128-\191]*)") do
found = true
if Hant_chars[ch] then
t = t + 1
if Hans_chars[ch] then
s = s + 1
end
elseif Hans_chars[ch] then
s = s + 1
else
t, s = t + 1, s + 1
end
end
if found then
if t == s then
return Hani
end
return get_script(t > s and "Hant" or "Hans")
end
else
sc = get_script(sc)
if not length then
length = ulen(text)
end
-- Count characters by removing everything in the script's charset and comparing to the original length.
local charset = sc.characters
local count = charset and length - ulen((ugsub(text, "[" .. charset .. "]+", ""))) or 0
if count >= length then
return sc
elseif count > bestcount then
bestcount = count
bestscript = sc
end
end
end
-- Return best matching script, or otherwise None.
return bestscript or get_script("None")
end
--[==[
Return a `Family` object for the language family that the language belongs to. See [[Modul:families]].
]==]
function Language:getFamily()
local family = self._familyObject
if family == nil then
family = self:getFamilyCode()
-- If the value is nil, it's cached as false.
family = family and get_family(family) or false
self._familyObject = family
end
return family or nil
end
--[==[
Return the family code in the language's data file.
]==]
function Language:getFamilyCode()
local family = self._familyCode
if family == nil then
-- If the value is nil, it's cached as false.
family = self._data[3] or false
self._familyCode = family
end
return family or nil
end
function Language:getFamilyName()
local family = self._familyName
if family == nil then
family = self:getFamily()
-- If the value is nil, it's cached as false.
family = family and family:getCanonicalName() or false
self._familyName = family
end
return family or nil
end
do
local function check_family(self, family)
if type(family) == "table" then
family = family:getCode()
end
if self:getFamilyCode() == family then
return true
end
local self_family = self:getFamily()
if self_family:inFamily(family) then
return true
-- If the family isn't a real family (e.g. creoles) check any ancestors.
elseif self_family:inFamily("qfa-not") then
local ancestors = self:getAncestors()
for _, ancestor in ipairs(ancestors) do
if ancestor:inFamily(family) then
return true
end
end
end
end
--[==[
Check whether the language belongs to `family` (which can be a family code or object). A list of objects can be given in
place of `family`; in that case, return true if the language belongs to any of the specified families. Note that some
languages (in particular, certain creoles) can have multiple immediate ancestors potentially belonging to different
families; in that case, return true if the language belongs to any of the specified families.
]==]
function Language:inFamily(...)
if self:getFamilyCode() == nil then
return false
end
return check_inputs(self, check_family, false, ...)
end
end
function Language:getParent()
local parent = self._parentObject
if parent == nil then
parent = self:getParentCode()
-- If the value is nil, it's cached as false.
parent = parent and get_by_code(parent, nil, true, true) or false
self._parentObject = parent
end
return parent or nil
end
function Language:getParentCode()
local parent = self._parentCode
if parent == nil then
-- If the value is nil, it's cached as false.
parent = self._data.parent or false
self._parentCode = parent
end
return parent or nil
end
function Language:getParentName()
local parent = self._parentName
if parent == nil then
parent = self:getParent()
-- If the value is nil, it's cached as false.
parent = parent and parent:getCanonicalName() or false
self._parentName = parent
end
return parent or nil
end
function Language:getParentChain()
local chain = self._parentChain
if chain == nil then
chain = {}
local parent, n = self:getParent(), 0
while parent do
n = n + 1
chain[n] = parent
parent = parent:getParent()
end
self._parentChain = chain
end
return chain
end
do
local function check_lang(self, lang)
for _, parent in ipairs(self:getParentChain()) do
if (type(lang) == "string" and lang or lang:getCode()) == parent:getCode() then
return true
end
end
end
function Language:hasParent(...)
return check_inputs(self, check_lang, false, ...)
end
end
--[==[
If the language is etymology-only, iterate through parents until a full language or family is found, and the
corresponding object is returned. If the language is a full language, return that language.
]==]
function Language:getFull()
local full = self._fullObject
if full == nil then
full = self:getFullCode()
full = full == self._code and self or get_by_code(full)
self._fullObject = full
end
return full
end
--[==[
If the language is an etymology-only language, iterate through parents until a full language or family is found, and
the corresponding code is returned. If the language is a full language, return that language's code.
]==]
function Language:getFullCode()
return self._fullCode or self._code
end
--[==[
If the language is an etymology-only language, iterate through parents until a full language or family is found, and the
corresponding canonical name is returned. If the language is a full language, return that language's canonical name.
]==]
function Language:getFullName()
local full = self._fullName
if full == nil then
full = self:getFull():getCanonicalName()
self._fullName = full
end
return full
end
--[==[
Return a list of {Language} objects for all languages that this language is directly descended from. Generally this is
only a single language, but creoles, pidgins and mixed languages can have multiple ancestors.
]==]
function Language:getAncestors()
local ancestors = self._ancestorObjects
if ancestors == nil then
ancestors = {}
local ancestor_codes = self:getAncestorCodes()
if #ancestor_codes > 0 then
for _, ancestor in ipairs(ancestor_codes) do
insert(ancestors, get_by_code(ancestor, nil, true))
end
else
local fam = self:getFamily()
local protoLang = fam and fam:getProtoLanguage() or nil
-- For the cases where the current language is the proto-language
-- of its family, or an etymology-only language that is ancestral to that
-- proto-language, we need to step up a level higher right from the
-- start.
if protoLang and (
protoLang:getCode() == self._code or
(self:hasType("etymology-only") and protoLang:hasAncestor(self))
) then
fam = fam:getFamily()
protoLang = fam and fam:getProtoLanguage() or nil
end
while not protoLang and not (not fam or fam:getCode() == "qfa-not") do
fam = fam:getFamily()
protoLang = fam and fam:getProtoLanguage() or nil
end
insert(ancestors, protoLang)
end
self._ancestorObjects = ancestors
end
return ancestors
end
do
-- Avoid a language being its own ancestor via class inheritance. We only need to check for this if the language has
-- inherited an ancestor table from its parent, because we never want to drop ancestors that have been explicitly
-- set in the data. Recursively iterate over ancestors until we either find self or run out. If self is found,
-- return true.
local function check_ancestor(self, lang)
local codes = lang:getAncestorCodes()
if not codes then
return nil
end
for i = 1, #codes do
local code = codes[i]
if code == self._code then
return true
end
local anc = get_by_code(code, nil, true)
if check_ancestor(self, anc) then
return true
end
end
end
--[==[
Return a list of `Language` codes for all languages that this language is directly descended from. Generally this is
only a single language, but creoles, pidgins and mixed languages can have multiple ancestors.
]==]
function Language:getAncestorCodes()
if self._ancestorCodes then
return self._ancestorCodes
end
local data = self._data
local codes = data.ancestors
if codes == nil then
codes = {}
self._ancestorCodes = codes
return codes
end
codes = split(codes, ",", true, true)
self._ancestorCodes = codes
-- If there are no codes or the ancestors weren't inherited data, there's nothing left to check.
if #codes == 0 or self:getData(false, "raw").ancestors ~= nil then
return codes
end
local i = 1
while i <= #codes do
if check_ancestor(self, self) then
remove(codes, i)
else
i = i + 1
end
end
return codes
end
end
--[==[
Given a list of language objects or codes, return true if at least one of them is an ancestor. This includes any
etymology-only children of that ancestor. If the language's ancestor(s) are etymology-only languages, it will also
return true for those language parent(s) (e.g. if Vulgar Latin is the ancestor, it will also return true for its parent,
Latin). However, a parent is excluded from this if the ancestor is also ancestral to that parent (e.g. if Classical
Persian is the ancestor, Persian would return {false}, because Classical Persian is also ancestral to Persian).
]==]
function Language:hasAncestor(...)
local function iterateOverAncestorTree(node, func, parent_check)
local ancestors = node:getAncestors()
local ancestorsParents = {}
for _, ancestor in ipairs(ancestors) do
-- When checking the parents of the other language, and the ancestor is also a parent, skip to the next
-- ancestor, so that we exclude any etymology-only children of that parent that are not directly related
-- (see below).
local ret = (parent_check or not node:hasParent(ancestor)) and
func(ancestor) or iterateOverAncestorTree(ancestor, func, parent_check)
if ret then
return ret
end
end
-- Check the parents of any ancestors. We don't do this if checking the parents of the other language, so that
-- we exclude any etymology-only children of those parents that are not directly related (e.g. if the ancestor
-- is Vulgar Latin and we are checking New Latin, we want it to return false because they are on different
-- ancestral branches. As such, if we're already checking the parent of New Latin (Latin) we don't want to
-- compare it to the parent of the ancestor (Latin), as this would be a false positive; it should be one or the
-- other).
if not parent_check then
return nil
end
for _, ancestor in ipairs(ancestors) do
local ancestorParents = ancestor:getParentChain()
for _, ancestorParent in ipairs(ancestorParents) do
if ancestorParent:getCode() == self._code or ancestorParent:hasAncestor(ancestor) then
break
else
insert(ancestorsParents, ancestorParent)
end
end
end
for _, ancestorParent in ipairs(ancestorsParents) do
local ret = func(ancestorParent)
if ret then
return ret
end
end
end
local function do_iteration(otherlang, parent_check)
-- otherlang can't be self
if (type(otherlang) == "string" and otherlang or otherlang:getCode()) == self._code then
return false
end
repeat
if iterateOverAncestorTree(
self,
function(ancestor)
return ancestor:getCode() == (type(otherlang) == "string" and otherlang or otherlang:getCode())
end,
parent_check
) then
return true
elseif type(otherlang) == "string" then
otherlang = get_by_code(otherlang, nil, true)
end
otherlang = otherlang:getParent()
parent_check = false
until not otherlang
end
local parent_check = true
for _, otherlang in ipairs{...} do
local ret = do_iteration(otherlang, parent_check)
if ret then
return true
end
end
return false
end
do
local function construct_node(lang, memo)
local branch, ancestors = {lang = lang:getCode()}
memo[lang:getCode()] = branch
for _, ancestor in ipairs(lang:getAncestors()) do
if ancestors == nil then
ancestors = {}
end
insert(ancestors, memo[ancestor:getCode()] or construct_node(ancestor, memo))
end
branch.ancestors = ancestors
return branch
end
function Language:getAncestorChain()
local chain = self._ancestorChain
if chain == nil then
chain = construct_node(self, {})
self._ancestorChain = chain
end
return chain
end
end
function Language:getAncestorChainOld()
local chain = self._ancestorChain
if chain == nil then
chain = {}
local step = self
while true do
local ancestors = step:getAncestors()
step = #ancestors == 1 and ancestors[1] or nil
if not step then
break
end
insert(chain, step)
end
self._ancestorChain = chain
end
return chain
end
local function fetch_descendants(self, fmt)
local descendants, family = {}, self:getFamily()
-- Iterate over all three datasets.
for _, data in ipairs{
require("Modul:languages/code to canonical name"),
require("Modul:etymology languages/code to canonical name"),
require("Modul:families/code to canonical name"),
} do
for code in pairs(data) do
local lang = get_by_code(code, nil, true, true)
if not lang then
error(("Internal error: code '%s' can't be fetched"):format(code))
end
-- Test for a descendant. Earlier tests weed out most candidates, while the more intensive tests are only used sparingly.
if (
code ~= self._code and -- Not self.
lang:inFamily(family) and -- In the same family.
(
family:getProtoLanguageCode() == self._code or -- Self is the protolanguage.
self:hasDescendant(lang) or -- Full hasDescendant check.
(lang:getFullCode() == self._code and not self:hasAncestor(lang)) -- Etymology-only child which isn't an ancestor.
)
) then
if fmt == "object" then
insert(descendants, lang)
elseif fmt == "code" then
insert(descendants, code)
elseif fmt == "name" then
insert(descendants, lang:getCanonicalName())
end
end
end
end
return descendants
end
function Language:getDescendants()
local descendants = self._descendantObjects
if descendants == nil then
descendants = fetch_descendants(self, "object")
self._descendantObjects = descendants
end
return descendants
end
function Language:getDescendantCodes()
local descendants = self._descendantCodes
if descendants == nil then
descendants = fetch_descendants(self, "code")
self._descendantCodes = descendants
end
return descendants
end
function Language:getDescendantNames()
local descendants = self._descendantNames
if descendants == nil then
descendants = fetch_descendants(self, "name")
self._descendantNames = descendants
end
return descendants
end
do
local function check_lang(self, lang)
if type(lang) == "string" then
lang = get_by_code(lang, nil, true)
end
if lang:hasAncestor(self) then
return true
end
end
function Language:hasDescendant(...)
return check_inputs(self, check_lang, false, ...)
end
end
local function fetch_children(self, fmt)
local m_etym_data = require(etymology_languages_data_module)
local self_code, children = self._code, {}
for code, lang in pairs(m_etym_data) do
local _lang = lang
repeat
local parent = _lang.parent
if parent == self_code then
if fmt == "object" then
insert(children, get_by_code(code, nil, true))
elseif fmt == "code" then
insert(children, code)
elseif fmt == "name" then
insert(children, lang[1])
end
break
end
_lang = m_etym_data[parent]
until not _lang
end
return children
end
function Language:getChildren()
local children = self._childObjects
if children == nil then
children = fetch_children(self, "object")
self._childObjects = children
end
return children
end
function Language:getChildrenCodes()
local children = self._childCodes
if children == nil then
children = fetch_children(self, "code")
self._childCodes = children
end
return children
end
function Language:getChildrenNames()
local children = self._childNames
if children == nil then
children = fetch_children(self, "name")
self._childNames = children
end
return children
end
function Language:hasChild(...)
local lang = ...
if not lang then
return false
elseif type(lang) == "string" then
lang = get_by_code(lang, nil, true)
end
if lang:hasParent(self) then
return true
end
return self:hasChild(select(2, ...))
end
--[==[
Return the name of the main category of that language. Example: {French language"} for French, whose category is at
[[:Category:French language]]. Unless optional argument `nocap` is given, the language name at the beginning of the
returned value will be capitalized. This capitalization is correct for category names, but not if the language name is
lowercase and the returned value of this function is used in the middle of a sentence.
]==]
function Language:getCategoryName(nocap)
local name = self._categoryName
if name == nil then
name = self:getCanonicalName()
-- If a substrate, omit any leading article.
if self:getFamilyCode() == "qfa-sub" then
name = name:gsub("^sebuah ", ""):gsub("^suatu ", "")
end
-- Only add " language" if a full language.
if self:hasType("full") then
-- Unless the canonical name already ends with "language", "lect" or their derivatives, add " language".
if not (match(name, "^[Bb]ahasa") or match(name, "^[Ll]ek")) then
name = "Bahasa " .. name
end
end
self._categoryName = name
end
if nocap then
return name
end
return mw.getContentLanguage():ucfirst(name)
end
--[==[
Create a link to the category; the link text is the canonical name.
]==]
function Language:makeCategoryLink()
return make_link(self, ":Kategori:" .. self:getCategoryName(), self:getDisplayForm())
end
function Language:getStandardCharacters(sc)
local standard_chars = self._data.standard_chars
if type(standard_chars) ~= "table" then
return standard_chars
elseif sc and type(sc) ~= "string" then
check_object("script", nil, sc)
sc = sc:getCode()
end
if (not sc) or sc == "None" then
local scripts = {}
for _, script in pairs(standard_chars) do
insert(scripts, script)
end
return concat(scripts)
end
if standard_chars[sc] then
return standard_chars[sc] .. (standard_chars[1] or "")
end
end
--[==[
Strip diacritics from display text `text` (in a language-specific fashion), which is in the script `sc`. If `sc` is
omitted or {nil}, the script is autodetected. This also strips certain punctuation characters from the end and (in the
case of Spanish upside-down question mark and exclamation points) from the beginning; strips any whitespace at the
end of the text or between the text and final stripped punctuation characters; and applies some language-specific
Unicode normalizations to replace discouraged characters with their prescribed alternatives. Return the stripped text.
]==]
function Language:stripDiacritics(text, sc)
if (not text) or text == "" then
return text
end
sc = checkScript(text, self, sc)
text = normalize(text, sc)
-- FIXME, rename makeEntryName to stripDiacritics and get rid of second and third return values
-- everywhere
local _
text, _, _ = iterateSectionSubstitutions(self, text, sc, nil, nil,
self._data.strip_diacritics or self._data.entry_name, "strip_diacritics", "stripDiacritics")
text = umatch(text, "^[¿¡]?(.-[^%s%p].-)%s*[؟?!;՛՜ ՞ ՟?!︖︕।॥။၊་།]?$") or text
return text
end
--[==[
Convert a ''logical'' pagename (the pagename as it appears to the user, after diacritics and punctuation have been
stripped) to a ''physical'' pagename (the pagename as it appears in the MediaWiki database). Reasons for a difference
between the two are (a) unsupported titles such as `[ ]` (with square brackets in them), `#` (pound/hash sign) and
`¯\_(ツ)_/¯` (with underscores), as well as overly long titles of various sorts; (b) "mammoth" pages that are split into
parts (e.g. `a`, which is split into physical pagenames `a/languages A to L` and `a/languages M to Z`). For almost all
purposes, you should work with logical and not physical pagenames. But there are certain use cases that require physical
pagenames, such as checking the existence of a page or retrieving a page's contents.
`pagename` is the logical pagename to be converted. `is_reconstructed_or_appendix` indicates whether the page is in the
`Reconstruction` or `Appendix` namespaces. If it is omitted or has the value {nil}, the pagename is checked for an
initial asterisk, and if found, the page is assumed to be a `Reconstruction` page. Setting a value of `false` or `true`
to `is_reconstructed_or_appendix` disables this check and allows for mainspace pagenames that begin with an asterisk.
]==]
function Language:logicalToPhysical(pagename, is_reconstructed_or_appendix)
-- FIXME: This probably shouldn't happen but it happens when makeEntryName() receives nil.
if pagename == nil then
track("nil-passed-to-logicalToPhysical")
return nil
end
local initial_asterisk
if is_reconstructed_or_appendix == nil then
local pagename_minus_initial_asterisk
initial_asterisk, pagename_minus_initial_asterisk = pagename:match("^(%*)(.*)$")
if pagename_minus_initial_asterisk then
is_reconstructed_or_appendix = true
pagename = pagename_minus_initial_asterisk
elseif self:hasType("appendix-constructed") then
is_reconstructed_or_appendix = true
end
end
if not is_reconstructed_or_appendix then
-- Check if the pagename is a listed unsupported title.
local unsupportedTitles = load_data(links_data_module).unsupported_titles
if unsupportedTitles[pagename] then
return "Tajuk tidak disokong/" .. unsupportedTitles[pagename]
end
end
-- Set `unsupported` as true if certain conditions are met.
local unsupported
-- Check if there's an unsupported character. \239\191\189 is the replacement character U+FFFD, which can't be typed
-- directly here due to an abuse filter. Unix-style dot-slash notation is also unsupported, as it is used for
-- relative paths in links, as are 3 or more consecutive tildes. Note: match is faster with magic
-- characters/charsets; find is faster with plaintext.
if (
match(pagename, "[#<>%[%]_{|}]") or
find(pagename, "\239\191\189") or
match(pagename, "%f[^%z/]%.%.?%f[%z/]") or
find(pagename, "~~~")
) then
unsupported = true
-- If it looks like an interwiki link.
elseif find(pagename, ":") then
local prefix = gsub(pagename, "^:*(.-):.*", ulower)
if (
load_data("Modul:data/namespaces")[prefix] or
load_data("Modul:data/interwikis")[prefix]
) then
unsupported = true
end
end
-- Escape unsupported characters so they can be used in titles. ` is used as a delimiter for this, so a raw use of
-- it in an unsupported title is also escaped here to prevent interference; this is only done with unsupported
-- titles, though, so inclusion won't in itself mean a title is treated as unsupported (which is why it's excluded
-- from the earlier test).
if unsupported then
-- FIXME: This conversion needs to be different for reconstructed pages with unsupported characters. There
-- aren't any currently, but if there ever are, we need to fix this e.g. to put them in something like
-- Reconstruction:Proto-Indo-European/Unsupported titles/`lowbar``num`.
local unsupported_characters = load_data(links_data_module).unsupported_characters
pagename = pagename:gsub("[#<>%[%]_`{|}\239]\191?\189?", unsupported_characters)
:gsub("%f[^%z/]%.%.?%f[%z/]", function(m)
return (gsub(m, "%.", "`period`"))
end)
:gsub("~~~+", function(m)
return (gsub(m, "~", "`tilde`"))
end)
pagename = "Tajuk tidak disokong/" .. pagename
elseif not is_reconstructed_or_appendix then
-- Check if this is a mammoth page. If so, which subpage should we link to?
local m_links_data = load_data(links_data_module)
local mammoth_page_type = m_links_data.mammoth_pages[pagename]
if mammoth_page_type then
local canonical_name = self:getFullName()
if canonical_name ~= "Rentas bahasa" and canonical_name ~= "Bahasa Melayu" then
local this_subpage
local L2_sort_key = get_L2_sort_key(canonical_name)
for _, subpage_spec in ipairs(m_links_data.mammoth_page_subpage_types[mammoth_page_type]) do
-- unpack() fails utterly on data loaded using mw.loadData() even if offsets are given
local subpage, pattern = subpage_spec[1], subpage_spec[2]
if pattern == true or L2_sort_key:match(pattern) then
this_subpage = subpage
break
end
end
if not this_subpage then
error(("Internal error: Bad data in mammoth_page_subpage_pages in [[Modul:links/data]] for mammoth page %s, type %s; last entry didn't have 'true' in it"):format(
pagename, mammoth_page_type))
end
pagename = pagename .. "/" .. this_subpage
end
end
end
return (initial_asterisk or "") .. pagename
end
--[==[
Strip the diacritics from a display pagename and convert the resulting logical pagename into a physical pagename.
This allows you, for example, to retrieve the contents of the page or check its existence. WARNING: This is deprecated
and will be going away. It is a simple composition of `##Language:stripDiacritics` and `##Language:logicalToPhysical`;
most callers only want the former, and if you need both, call them both yourself.
`text` and `sc` are as in `##Language:stripDiacritics`, and `is_reconstructed_or_appendix` is as in
`##Language:logicalToPhysical`.
]==]
function Language:makeEntryName(text, sc, is_reconstructed_or_appendix)
track("makeEntryName called")
return self:logicalToPhysical(self:stripDiacritics(text, sc), is_reconstructed_or_appendix)
end
--[==[
Generate term alternants (e.g. simplified versions of traditional Chinese terms) using a language-specific method, and
return them as a table. If the language provides no method for generating alternants, return a table containing only the
input term.
]==]
function Language:generateAlternants(text, sc)
local generate_alternants = self._data.generate_alternants
if generate_alternants == nil then
return {text}
end
sc = checkScript(text, self, sc)
return require("Modul:" .. generate_alternants).generateAlternants(text, self, sc)
end
--[==[
Create a sort key for the given stripped text, following the rules appropriate for the language. This removes
diacritical marks from the stripped text if they are not considered significant for sorting, and may perform some other
changes. Any initial hyphen is also removed, and anything in parentheses is removed as well. The `sort_key` setting for
each language in the data modules defines the replacements made by this function, or it gives the name of the module
that takes the stripped text and returns a sortkey.
]==]
function Language:makeSortKey(text, sc)
if (not text) or text == "" then
return text
end
if match(text, "<[^<>]+>") then
track("track HTML tag")
end
-- Remove directionality/control characters, bold, italics, soft hyphens, strip markers and HTML tags.
-- FIXME: Partly duplicated with remove_formatting() in [[Modul:links]].
text = ugsub(text, "[\194\173\216\156\226\128\142\226\128\143\226\128\170-\226\128\174\226\129\166-\226\129\175]", "")
text = text:gsub("('*)'''(.-'*)'''", "%1%2"):gsub("('*)''(.-'*)''", "%1%2")
text = gsub(unstrip(text), "<[^<>]+>", "")
text = decode_uri(text, "PATH")
text = checkNoEntities(text)
-- Remove initial hyphens and * unless the term only consists of spacing + punctuation characters.
text = ugsub(text, "^([-]*)[-־ـ᠊*]+([-]*)(.*[^%s%p].*)", "%1%2%3")
sc = checkScript(text, self, sc)
text = normalize(text, sc)
text = removeCarets(text, sc)
-- For languages with dotted dotless i, ensure that "İ" is sorted as "i", and "I" is sorted as "ı".
if self:hasDottedDotlessI() then
text = gsub(text, "I\204\135", "i") -- decomposed "İ"
:gsub("I", "ı")
text = sc:toFixedNFD(text)
end
-- Convert to lowercase, make the sortkey, then convert to uppercase. Where the language has dotted dotless i, it is
-- usually not necessary to convert "i" to "İ" and "ı" to "I" first, because "I" will always be interpreted as
-- conventional "I" (not dotless "İ") by any sorting algorithms, which will have been taken into account by the
-- sortkey substitutions themselves. However, if no sortkey substitutions have been specified, then conversion is
-- necessary so as to prevent "i" and "ı" both being sorted as "I".
--
-- An exception is made for scripts that (sometimes) sort by scraping page content, as that means they are sensitive
-- to changes in capitalization (as it changes the target page).
if not sc:sortByScraping() then
text = ulower(text)
end
local actual_substitution_data, _
-- Don't trim whitespace here because it's significant at the beginning of a sort key or sort base.
text, _, actual_substitution_data = iterateSectionSubstitutions(self, text, sc, nil, nil,
self._data.sort_key, "sort_key", "makeSortKey", "notrim")
if not sc:sortByScraping() then
if self:hasDottedDotlessI() and not actual_substitution_data then
text = text:gsub("ı", "I"):gsub("i", "İ")
text = sc:toFixedNFC(text)
end
text = uupper(text)
end
-- Remove parentheses, as long as they are either preceded or followed by something.
text = gsub(text, "(.)[()]+", "%1"):gsub("[()]+(.)", "%1")
text = escape_risky_characters(text)
return text
end
--[==[
Create the form used as as a basis for display text and transliteration. FIXME: Rename to correctInputText().
]==]
local function processDisplayText(text, self, sc, keepCarets, keepPrefixes)
local subbedChars = {}
text, subbedChars = doTempSubstitutions(text, subbedChars, keepCarets)
text = decode_uri(text, "PATH")
text = checkNoEntities(text)
sc = checkScript(text, self, sc)
text = normalize(text, sc)
text, subbedChars = iterateSectionSubstitutions(self, text, sc, subbedChars, keepCarets, self._data.display_text,
"display_text", "makeDisplayText")
text = removeCarets(text, sc)
-- Remove any interwiki link prefixes (unless they have been escaped or this has been disabled).
if find(text, ":") and not keepPrefixes then
local rep
repeat
text, rep = gsub(text, "\\\\(\\*:)", "\3%1")
until rep == 0
text = gsub(text, "\\:", "\4")
while true do
local prefix = gsub(text, "^(.-):.+", function(m1)
return (gsub(m1, "\244[\128-\191]*", ""))
end)
-- Check if the prefix is an interwiki, though ignore capitalised Wiktionary:, which is a namespace.
if not prefix or prefix == text or prefix == "Wikikamus"
or not (load_data("Modul:data/interwikis")[ulower(prefix)] or prefix == "") then
break
end
text = gsub(text, "^(.-):(.*)", function(m1, m2)
local ret = {}
for subbedChar in gmatch(m1, "\244[\128-\191]*") do
insert(ret, subbedChar)
end
return concat(ret) .. m2
end)
end
text = gsub(text, "\3", "\\"):gsub("\4", ":")
end
return text, subbedChars
end
--[==[
Make the display text (i.e. what is displayed on the page).
]==]
function Language:makeDisplayText(text, sc, keepPrefixes)
if not text or text == "" then
return text
end
local subbedChars
text, subbedChars = processDisplayText(text, self, sc, nil, keepPrefixes)
text = escape_risky_characters(text)
return undoTempSubstitutions(text, subbedChars)
end
--[==[
Transliterate the text from the given script into Latin script (see [[Wiktionary:Transliteration and romanization]]).
The language must have the `translit` property for this to work; if it is not present, {nil} is returned.
The `sc` parameter is handled by the transliteration module, and how it is handled is specific to that module. Some
transliteration modules may tolerate {nil} as the script, others require it to be one of the possible scripts that the
module can transliterate, and will throw an error if it's not one of them. For this reason, the `sc` parameter should
always be provided when writing non-language-specific code.
The `module_override` parameter is used to override the default module that is used to provide the transliteration.
This is useful in cases where you need to demonstrate a particular module in use, but there is no default module yet,
or you want to demonstrate an alternative version of a transliteration module before making it official. It should not
be used in real modules or templates, only for testing. All uses of this parameter are tracked by
[[Wiktionary:Tracking/languages/module_override]].
'''Known bugs''':
* This function assumes {tr(s1) .. tr(s2) == tr(s1 .. s2)}. When this assertion fails, wikitext markups like
<nowiki>'''</nowiki> can cause wrong transliterations.
* HTML entities like `&apos;`, often used to escape wikitext markups, do not work.
]==]
function Language:transliterate(text, sc, module_override)
-- If there is no text, or the language doesn't have transliteration data and there's no override, return nil.
if not text or text == "" or text == "-" then
return text
end
-- If the script is not transliteratable (and no override is given), return nil.
sc = checkScript(text, self, sc)
if not (sc:isTransliterated() or module_override) then
-- temporary tracking to see if/when this gets triggered
track("non-transliterable")
track("non-transliterable/" .. self._code)
track("non-transliterable/" .. sc:getCode())
track("non-transliterable/" .. sc:getCode() .. "/" .. self._code)
return nil
end
-- Remove any strip markers.
text = unstrip(text)
-- Do not process the formatting into PUA characters for certain languages.
local processed = load_data(languages_data_module).substitution[self._code] ~= "none"
-- Get the display text with the keepCarets flag set.
local subbedChars
if processed then
text, subbedChars = processDisplayText(text, self, sc, true)
end
-- Transliterate (using the module override if applicable).
text, subbedChars = iterateSectionSubstitutions(self, text, sc, subbedChars, true, module_override or
self._data.translit, "translit", "tr")
if not text then
return nil
end
-- Incomplete transliterations return nil.
local charset = sc.characters
if charset and umatch(text, "[" .. charset .. "]") then
-- Remove any characters in Latin, which includes Latin characters also included in other scripts (as these are
-- false positives), as well as any PUA substitutions. Anything remaining should only be script code "None"
-- (e.g. numerals).
local check_text = ugsub(text, "[" .. get_script("Latn").characters .. "-]+", "")
-- Set none_is_last_resort_only flag, so that any non-None chars will cause a script other than "None" to be
-- returned.
if find_best_script_without_lang(check_text, true):getCode() ~= "None" then
return nil
end
end
if processed then
text = escape_risky_characters(text)
text = undoTempSubstitutions(text, subbedChars)
end
-- If the script does not use capitalization, then capitalize any letters of the transliteration which are
-- immediately preceded by a caret (and remove the caret).
if text and not sc:hasCapitalization() and text:find("^", 1, true) then
text = processCarets(text, "%^([\128-\191\244]*%*?)([^\128-\191\244][\128-\191]*)", function(m1, m2)
return m1 .. uupper(m2)
end)
end
-- Track module overrides.
if module_override ~= nil then
track("module_override")
end
return text
end
do
local function handle_language_spec(self, spec, sc)
local ret = self["_" .. spec]
if ret == nil then
ret = self._data[spec]
if type(ret) == "string" then
ret = list_to_set(split(ret, ",", true, true))
end
self["_" .. spec] = ret
end
if type(ret) == "table" then
ret = ret[sc:getCode()]
end
return not not ret
end
function Language:overrideManualTranslit(sc)
return handle_language_spec(self, "override_translit", sc)
end
function Language:link_tr(sc)
return handle_language_spec(self, "link_tr", sc)
end
end
--[==[
Return {true} if the language has a transliteration module, or {false} if it doesn't.
]==]
function Language:hasTranslit()
return not not self._data.translit
end
--[==[
Return {true} if the language uses the letters I/ı and İ/i, or {false} if it doesn't.]==]
function Language:hasDottedDotlessI()
return not not self._data.dotted_dotless_i
end
function Language:toJSON(opts)
local strip_diacritics, strip_diacritics_patterns, strip_diacritics_remove_diacritics = self._data.strip_diacritics
if strip_diacritics then
if strip_diacritics.from then
strip_diacritics_patterns = {}
for i, from in ipairs(strip_diacritics.from) do
insert(strip_diacritics_patterns, {from = from, to = strip_diacritics.to[i] or ""})
end
end
strip_diacritics_remove_diacritics = strip_diacritics.remove_diacritics
end
-- mainCode should only end up non-nil if dontCanonicalizeAliases is passed to make_object().
-- props should either contain zero-argument functions to compute the value, or the value itself.
local props = {
ancestors = function() return self:getAncestorCodes() end,
canonicalName = function() return self:getCanonicalName() end,
categoryName = function() return self:getCategoryName("nocap") end,
code = self._code,
mainCode = self._mainCode,
parent = function() return self:getParentCode() end,
full = function() return self:getFullCode() end,
stripDiacriticsPatterns = strip_diacritics_patterns,
stripDiacriticsRemoveDiacritics = strip_diacritics_remove_diacritics,
family = function() return self:getFamilyCode() end,
aliases = function() return self:getAliases() end,
varieties = function() return self:getVarieties() end,
otherNames = function() return self:getOtherNames() end,
scripts = function() return self:getScriptCodes() end,
type = function() return keys_to_list(self:getTypes()) end,
wikimediaLanguages = function() return self:getWikimediaLanguageCodes() end,
wikidataItem = function() return self:getWikidataItem() end,
wikipediaArticle = function() return self:getWikipediaArticle(true) end,
}
local ret = {}
for prop, val in pairs(props) do
if not opts.skip_fields or not opts.skip_fields[prop] then
if type(val) == "function" then
ret[prop] = val()
else
ret[prop] = val
end
end
end
-- Use `deep_copy` when returning a table, so that there are no editing restrictions imposed by `mw.loadData`.
return opts and opts.lua_table and deep_copy(ret) or to_json(ret, opts)
end
function export.getDataModuleName(code)
local letter = match(code, "^(%l)%l%l?$")
return "Modul:" .. (
letter == nil and "languages/data/exceptional" or
#code == 2 and "languages/data/2" or
"languages/data/3/" .. letter
)
end
get_data_module_name = export.getDataModuleName
function export.getExtraDataModuleName(code)
return get_data_module_name(code) .. "/extra"
end
get_extra_data_module_name = export.getExtraDataModuleName
do
local function make_stack(data)
local key_types = {
[2] = "unique",
aliases = "unique",
otherNames = "unique",
type = "append",
varieties = "unique",
wikipedia_article = "unique",
wikimedia_codes = "unique"
}
local function __index(self, k)
local stack, key_type = getmetatable(self), key_types[k]
-- Data that isn't inherited from the parent.
if key_type == "unique" then
local v = stack[stack[make_stack]][k]
if v == nil then
local layer = stack[0]
if layer then -- Could be false if there's no extra data.
v = layer[k]
end
end
return v
-- Data that is appended by each generation.
elseif key_type == "append" then
local parts, offset, n = {}, 0, stack[make_stack]
for i = 1, n do
local part = stack[i][k]
if part == nil then
offset = offset + 1
else
parts[i - offset] = part
end
end
return offset ~= n and concat(parts, ",") or nil
end
local n = stack[make_stack]
while true do
local layer = stack[n]
if not layer then -- Could be false if there's no extra data.
return nil
end
local v = layer[k]
if v ~= nil then
return v
end
n = n - 1
end
end
local function __newindex()
error("table is read-only")
end
local function __pairs(self)
-- Iterate down the stack, caching keys to avoid duplicate returns.
local stack, seen = getmetatable(self), {}
local n = stack[make_stack]
local iter, state, k, v = pairs(stack[n])
return function()
repeat
repeat
k = iter(state, k)
if k == nil then
n = n - 1
local layer = stack[n]
if not layer then -- Could be false if there's no extra data.
return nil
end
iter, state, k = pairs(layer)
end
until not (k == nil or seen[k])
-- Get the value via a lookup, as the one returned by the
-- iterator will be the raw value from the current layer,
-- which may not be the one __index will return for that
-- key. Also memoize the key in `seen` (even if the lookup
-- returns nil) so that it doesn't get looked up again.
-- TODO: store values in `self`, avoiding the need to create
-- the `seen` table. The iterator will need to iterate over
-- `self` with `next` first to find these on future loops.
v, seen[k] = self[k], true
until v ~= nil
return k, v
end
end
local __ipairs = require(table_module).indexIpairs
function make_stack(data)
local stack = {
data,
[make_stack] = 1, -- stores the length and acts as a sentinel to confirm a given metatable is a stack.
__index = __index,
__newindex = __newindex,
__pairs = __pairs,
__ipairs = __ipairs,
}
stack.__metatable = stack
return setmetatable({}, stack), stack
end
return make_stack(data)
end
local function get_stack(data)
local stack = getmetatable(data)
return stack and type(stack) == "table" and stack[make_stack] and stack or nil
end
--[==[
<span style="color: var(--wikt-palette-red,#BA0000)">This function is not for use in entries or other content pages.</span>
Return a blob of data about the language. The format of this blob is undocumented, and perhaps unstable; it's intended
for things like the module's own unit-tests, which are "close friends" with the module and will be kept up-to-date as
the format changes. If `extra` is set, any extra data in the relevant `/extra` module will be included. (Note that it
will be included anyway if it has already been loaded into the language object.) If `raw` is set, then the returned data
will not contain any data inherited from parent objects.
-- Do NOT use these methods!
-- All uses should be pre-approved on the talk page!
]==]
function Language:getData(extra, raw)
if extra then
self:loadInExtraData()
end
local data = self._data
-- If raw is not set, just return the data.
if not raw then
return data
end
local stack = get_stack(data)
-- If there isn't a stack or its length is 1, return the data. Extra data (if any) will be included, as it's
-- stored at key 0 and doesn't affect the reported length.
if stack == nil then
return data
end
local n = stack[make_stack]
if n == 1 then
return data
end
extra = stack[0]
-- If there isn't any extra data, return the top layer of the stack.
if extra == nil then
return stack[n]
end
-- If there is, return a new stack which has the top layer at key 1 and the extra data at key 0.
data, stack = make_stack(stack[n])
stack[0] = extra
return data
end
function Language:loadInExtraData()
-- Only full languages have extra data.
if not self:hasType("language", "full") then
return
end
local data = self._data
-- If there's no stack, create one.
local stack = get_stack(self._data)
if stack == nil then
data, stack = make_stack(data)
-- If already loaded, return.
elseif stack[0] ~= nil then
return
end
self._data = data
-- Load extra data from the relevant module and add it to the stack at key 0, so that the __index and __pairs
-- metamethods will pick it up, since they iterate down the stack until they run out of layers.
local code = self._code
local modulename = get_extra_data_module_name(code)
-- No data cached as false.
stack[0] = modulename and load_data(modulename)[code] or false
end
--[==[
Return the name of the module containing the language's data, e.g. [[Modul:languages/data/3/k]] for three-letter codes
beginning with `k`.
]==]
function Language:getDataModuleName()
local name = self._dataModuleName
if name == nil then
name = self:hasType("etymology-only") and etymology_languages_data_module or
get_data_module_name(self._mainCode or self._code)
self._dataModuleName = name
end
return name
end
--[==[
Return the name of the module containing the language's extra data, e.g. [[Modul:languages/data/3/k/extra]] for
three-letter codes beginning with `k`.
]==]
function Language:getExtraDataModuleName()
local name = self._extraDataModuleName
if name == nil then
name = not self:hasType("etymology-only") and get_extra_data_module_name(self._mainCode or self._code) or false
self._extraDataModuleName = name
end
return name or nil
end
function export.makeObject(code, data, dontCanonicalizeAliases)
local data_type = type(data)
if data_type ~= "table" then
error(("bad argument #2 to 'makeObject' (table expected, got %s)"):format(data_type))
end
-- Convert any aliases.
local input_code = code
code = normalize_code(code)
input_code = dontCanonicalizeAliases and input_code or code
local parent
if data.parent then
parent = get_by_code(data.parent, nil, true, true)
else
parent = Language
end
parent.__index = parent
local lang = {_code = input_code}
-- This can only happen if dontCanonicalizeAliases is passed to make_object().
if code ~= input_code then
lang._mainCode = code
end
local parent_data = parent._data
if parent_data == nil then
-- Full code is the same as the code.
lang._fullCode = parent._code or code
else
-- Copy full code.
lang._fullCode = parent._fullCode
local stack = get_stack(parent_data)
if stack == nil then
parent_data, stack = make_stack(parent_data)
end
-- Insert the input data as the new top layer of the stack.
local n = stack[make_stack] + 1
data, stack[n], stack[make_stack] = parent_data, data, n
end
lang._data = data
return setmetatable(lang, parent)
end
make_object = export.makeObject
end
--[==[
Find the language whose code matches the one provided. If it exists, return a `Language` object representing the
language. Otherwise, return {nil}, unless `paramForError` is given, in which case an error is generated. If
`paramForError` is {true}, a generic error message mentioning the bad code is generated; otherwise `paramForError`
should be a string or number specifying the parameter that the code came from, and this parameter will be mentioned
in the error message along with the bad code. If `allowEtymLang` is specified, etymology-only language codes are allowed
and looked up along with full language codes. If `allowFamily` is specified, language family codes are allowed and
looked up along with normal language codes.
]==]
function export.getByCode(code, paramForError, allowEtymLang, allowFamily)
-- [[ Track uses of paramForError, ultimately so it can be removed, as error-handling should be done by
-- [[Modul:parameters]], not here. ]] -- I (Benwing) disagree; that is an idealistic goal unlikely to be achievable
-- realistically.
if paramForError ~= nil then
track("paramForError")
end
if type(code) ~= "string" then
local typ
if not code then
typ = "nil"
elseif check_object("language", true, code) then
typ = "sebuah objek bahasa"
elseif check_object("family", true, code) then
typ = "sebuah objek keluarga"
else
typ = "sebuah " .. type(code)
end
error("The function getByCode expects a string as its first argument, but received " .. typ .. ".")
end
local m_data = load_data(languages_data_module)
if m_data.aliases[code] or m_data.track[code] then
track(code)
end
local norm_code = normalize_code(code)
-- Get the data, checking for etymology-only languages if allowEtymLang is set.
local data = load_data(get_data_module_name(norm_code))[norm_code] or
allowEtymLang and load_data(etymology_languages_data_module)[norm_code]
-- If no data was found and allowFamily is set, check the family data. If the main family data was found, make the
-- object with [[Modul:families]] instead, as family objects have different methods. However, if it's an
-- etymology-only family, use make_object in this module (which handles object inheritance), and the
-- family-specific methods will be inherited from the parent object.
if data == nil and allowFamily then
data = load_data("Modul:families/data")[norm_code]
if data ~= nil then
if data.parent == nil then
return make_family_object(norm_code, data)
elseif not allowEtymLang then
data = nil
end
end
end
local retval = code and data and make_object(code, data)
if not retval and paramForError then
require("Modul:languages/errorGetBy").code(code, paramForError, allowEtymLang, allowFamily)
end
return retval
end
get_by_code = export.getByCode
--[==[
Find the language whose canonical name (the name used to represent that language on Wiktionary) or other name matches
the one provided. If it exists, return a `Language` object representing the language. Otherwise, return {nil}, unless
`paramForError` is given, in which case an error is generated. If `allowEtymLang` is specified, etymology-only language
codes are allowed and looked up along with full language codes. If `allowFamily` is specified, language family codes are
allowed and looked up along with normal language codes. The canonical name of languages should always be unique (it is
an error for two languages on Wiktionary to share the same canonical name), so this is guaranteed to give at most one
result. This function is powered by [[Modul:languages/canonical names]], which contains a pre-generated mapping of
full-language canonical names to codes. It is generated by going through the [[:Category:Language data modules]] for
full languages. When `allowEtymLang` is specified for the above function, [[Modul:etymology languages/canonical names]]
may also be used, and when `allowFamily` is specified for the above function, [[Modul:families/canonical names]] may
also be used.
]==]
function export.getByCanonicalName(name, errorIfInvalid, allowEtymLang, allowFamily)
local byName = load_data("Modul:languages/canonical names")
local code = byName and byName[name]
if not code and allowEtymLang then
byName = load_data("Modul:etymology languages/canonical names")
code = byName and byName[name] or
byName[gsub(name, "^[Ss]ubstratum", "")] or
byName[gsub(name, "^sebuah ", "")] or
byName[gsub(name, "^sebuah [Ss]ubstratum ", ""):gsub("", "")] or
-- For etymology families like "ira-pro".
-- FIXME: This is not ideal, as it allows " languages" to be appended to any etymology-only language, too.
byName[match(name, "^[Bb]ahasa%-bahasa (.*)$")]
end
if not code and allowFamily then
byName = load_data("Modul:families/canonical names")
code = byName[name] or byName[match(name, "^^[Bb]ahasa%-bahasa (.*)$")]
end
local retval = code and get_by_code(code, errorIfInvalid, allowEtymLang, allowFamily)
if not retval and errorIfInvalid then
require("Modul:languages/errorGetBy").canonicalName(name, allowEtymLang, allowFamily)
end
return retval
end
--[==[
Used by [[Modul:languages/data/2]] (et al.) and [[Modul:etymology languages/data]], [[Modul:families/data]],
[[Modul:scripts/data]] and [[Modul:writing systems/data]] to finalize the data into the format that is actually returned.
]==]
function export.finalizeData(data, main_type, variety)
local fields = {"type"}
if main_type == "language" then
insert(fields, 4) -- script codes
insert(fields, "ancestors")
insert(fields, "link_tr")
insert(fields, "override_translit")
insert(fields, "wikimedia_codes")
elseif main_type == "script" then
insert(fields, 3) -- writing system codes
end -- Families and writing systems have no extra fields to process.
local fields_len = #fields
for _, entity in next, data do
if variety then
-- Move parent from 3 to "parent" and family from "family" to 3. These are different for the sake of
-- convenience, since very few varieties have the family specified, whereas all of them have a parent.
entity.parent, entity[3] = entity[3], entity.family
entity.family = nil
-- Give the type "regular" iff not a variety and no other types are assigned.
elseif not (entity.type or entity.parent) then
entity.type = "regular"
end
for i = 1, fields_len do
local key = fields[i]
local field = entity[key]
if field and type(field) == "string" then
entity[key] = gsub(field, "%s*,%s*", ",")
end
end
end
return data
end
--[==[
For backwards compatibility only; modules should require the error themselves.
]==]
function export.err(lang_code, param, code_desc, template_tag, not_real_lang)
return require("Modul:languages/error")(lang_code, param, code_desc, template_tag, not_real_lang)
end
return export
dbicznkmimzt0t4bckchxj0kvhhiklz
Modul:headword
828
9757
375353
369309
2026-09-22T03:13:28Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714451|92714451]])
375353
Scribunto
text/plain
local export = {}
-- Named constants for all modules used, to make it easier to swap out sandbox versions.
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local gender_and_number_module = "Module:gender and number"
local headword_data_module = "Module:headword/data"
local headword_page_module = "Module:headword/page"
local links_module = "Module:links"
local load_module = "Module:load"
local pages_module = "Module:pages"
local palindromes_module = "Module:palindromes"
local scripts_module = "Module:scripts"
local scripts_data_module = "Module:scripts/data"
local script_utilities_module = "Module:script utilities"
local script_utilities_data_module = "Module:script utilities/data"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local dump = mw.dumpObject
local insert = table.insert
local ipairs = ipairs
local max = math.max
local new_title = mw.title.new
local pairs = pairs
local require = require
local toNFC = mw.ustring.toNFC
local toNFD = mw.ustring.toNFD
local type = type
local ufind = mw.ustring.find
local ugmatch = mw.ustring.gmatch
local ugsub = mw.ustring.gsub
local umatch = mw.ustring.match
--[==[
Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==]
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function contains(...)
contains = require(table_module).contains
return contains(...)
end
local function encode_entities(...)
encode_entities = require(string_utilities_module).encode_entities
return encode_entities(...)
end
local function extend(...)
extend = require(table_module).extend
return extend(...)
end
local function find_best_script_without_lang(...)
find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang
return find_best_script_without_lang(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function format_genders(...)
format_genders = require(gender_and_number_module).format_genders
return format_genders(...)
end
local function format_decorations(...)
format_decorations = require(decorations_module).format_decorations
return format_decorations(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_current_L2(...)
get_current_L2 = require(pages_module).get_current_L2
return get_current_L2(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function is_palindrome(...)
is_palindrome = require(palindromes_module).is_palindrome
return is_palindrome(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function load_data(...)
load_data = require(load_module).load_data
return load_data(...)
end
local function pattern_escape(...)
pattern_escape = require(string_utilities_module).pattern_escape
return pattern_escape(...)
end
local function pluralize(...)
pluralize = require(en_utilities_module).pluralize
return pluralize(...)
end
local function process_page(...)
process_page = require(headword_page_module).process_page
return process_page(...)
end
local function remove_links(...)
remove_links = require(links_module).remove_links
return remove_links(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function tag_text(...)
tag_text = require(script_utilities_module).tag_text
return tag_text(...)
end
local function tag_transcription(...)
tag_transcription = require(script_utilities_module).tag_transcription
return tag_transcription(...)
end
local function tag_translit(...)
tag_translit = require(script_utilities_module).tag_translit
return tag_translit(...)
end
local function trim(...)
trim = require(string_utilities_module).trim
return trim(...)
end
local function ulen(...)
ulen = require(string_utilities_module).len
return ulen(...)
end
local function ucfirst(...)
ucfirst = require(string_utilities_module).ucfirst
return ucfirst(...)
end
--[==[
Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==]
local m_data
local function get_data()
m_data = load_data(headword_data_module)
return m_data
end
local script_data
local function get_script_data()
script_data = load_data(scripts_data_module)
return script_data
end
local script_utilities_data
local function get_script_utilities_data()
script_utilities_data = load_data(script_utilities_data_module)
return script_utilities_data
end
-- If set to true, categories always appear, even in non-mainspace pages
local test_force_categories = false
-- Add a tracking category to track entries with certain (unusually undesirable) properties. `track_id` is an identifier
-- for the particular property being tracked and goes into the tracking page. Specifically, this adds a link in the
-- page text to [[Wiktionary:Tracking/headword/TRACK_ID]], meaning you can find all entries with the `track_id` property
-- by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID]].
--
-- If `lang` (a language object) is given, an additional tracking page [[Wiktionary:Tracking/headword/TRACK_ID/CODE]] is
-- linked to where CODE is the language code of `lang`, and you can find all entries in the combination of `track_id`
-- and `lang` by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID/CODE]]. This makes it possible to
-- isolate only the entries with a specific tracking property that are in a given language. Note that if `lang`
-- references at etymology-only language, both that language's code and its full parent's code are tracked.
local function track(track_id, lang)
local tracking_page = "headword/" .. track_id
if lang and lang:hasType("etymology-only") then
debug_track{tracking_page, tracking_page .. "/" .. lang:getCode(),
tracking_page .. "/" .. lang:getFullCode()}
elseif lang then
debug_track{tracking_page, tracking_page .. "/" .. lang:getCode()}
else
debug_track(tracking_page)
end
return true
end
local function text_in_script(text, script_code)
local sc = get_script(script_code)
if not sc then
error("Internal error: Bad script code " .. script_code)
end
local characters = sc.characters
local out
if characters then
text = ugsub(text, "%W", "")
out = ufind(text, "[" .. characters .. "]")
end
if out then
return true
else
return false
end
end
local spacingPunctuation = "[%s%p]+"
--[[ List of punctuation or spacing characters that are found inside of words.
Used to exclude characters from the regex above. ]]
local wordPunc = "-#%%&@־׳״'.·*’་•:᠊"
local notWordPunc = "[^" .. wordPunc .. "]+"
--[=[
Format a term (either a head term or an inflection term) along with any decorations (left or right qualifiers, labels,
references or customized separator). `part` is the object specifying the term (and `lang` the language of the term),
which should optionally contain:
* left qualifiers in `q`, an array of strings;
* right qualifiers in `qq`, an array of strings;
* left labels in `l`, an array of strings;
* right labels in `ll`, an array of strings;
* references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text`
(formatted reference text) and optionally `name` and/or `group`;
* a separator in `separator`, defaulting to " <i>or</i> " if this is not the first term (j > 1), otherwise "".
`formatted` is the formatted version of the term itself, and `j` is the index of the term.
]=]
local function format_term_with_decorations(lang, part, formatted, j)
local function part_non_empty(field)
local list = part[field]
if not list then
return nil
end
if type(list) ~= "table" then
error(("Internal error: Wrong type for `part.%s`=%s, should be \"table\""):format(field, dump(list)))
end
return list[1]
end
if part_non_empty("q") or part_non_empty("qq") or part_non_empty("l") or
part_non_empty("ll") or part_non_empty("refs") then
formatted = format_decorations {
lang = lang,
text = formatted,
q = part.q,
qq = part.qq,
l = part.l,
ll = part.ll,
refs = part.refs,
}
end
local separator = part.separator or j > 1 and " <i>atau</i> " -- use "" to request no separator
if separator then
formatted = separator .. formatted
end
return formatted
end
--[==[Return true if the given head is multiword according to the algorithm used in full_headword().]==]
function export.head_is_multiword(head)
for possibleWordBreak in ugmatch(head, spacingPunctuation) do
if umatch(possibleWordBreak, notWordPunc) then
return true
end
end
return false
end
do
local function workaround_to_exclude_chars(s)
return (ugsub(s, notWordPunc, "\2%1\1"))
end
--[==[
Add appropriate links to `head`, correctly handling multiword terms. This is intended for multiword terms but can
be used for any term if you want links added to single-word terms as well. If you want to only add
links to multiword terms, first check that the term is multiword using `head_is_multiword`.
If `default` is specified, this will escape colons so that they don't get interpreted as interwiki links. This
should generally only be used when `head` is an actual pagename or is taken from a {{para|pagename}} parameter, not
when taken from a {{para|head}} parameter.
]==]
function export.add_multiword_links(head, default)
head = "\1" .. ugsub(head, spacingPunctuation, workaround_to_exclude_chars) .. "\2"
if default then
head = head
:gsub("(\1[^\2]*)\\([:#][^\2]*\2)", "%1\\\\%2")
:gsub("(\1[^\2]*)([:#][^\2]*\2)", "%1\\%2")
end
--Escape any remaining square brackets to stop them breaking links (e.g. "[citation needed]").
head = encode_entities(head, "[]", true, true)
--[=[
use this when workaround is no longer needed:
head = "[[" .. ugsub(head, WORDBREAKCHARS, "]]%1[[") .. "]]"
Remove any empty links, which could have been created above
at the beginning or end of the string.
]=]
return (head
:gsub("\1\2", "")
:gsub("[\1\2]", {["\1"] = "[[", ["\2"] = "]]"}))
end
end
local function non_categorizable(full_raw_pagename)
return full_raw_pagename:find("^Lampiran:Gerak isyarat/") or
-- Unsupported titles with descriptive names.
(full_raw_pagename:find("^Tajuk tidak disokong/") and not full_raw_pagename:find("`"))
end
local function tag_text_and_add_decorations(data, head, formatted, j)
-- Add language and script wrapper.
formatted = tag_text(formatted, data.lang, head.sc, "head", nil, j == 1 and data.id or nil)
-- Add decorations (qualifiers, labels, references and separator).
return format_term_with_decorations(data.lang, head, formatted, j)
end
-- Format a headword with transliterations.
local function format_headword(data)
-- Are there non-empty transliterations?
local has_translits = false
local has_manual_translits = false
------ Format the headwords. ------
local head_parts = {}
local unique_head_parts = {}
local has_multiple_heads = not not data.heads[2]
for j, head in ipairs(data.heads) do
if head.tr or head.ts then
has_translits = true
end
if head.tr and head.tr_manual or head.ts then
has_manual_translits = true
end
local formatted
-- Apply processing to the headword, for formatting links and such.
if head.term:find("[[", nil, true) and head.sc:getCode() ~= "Image" then
formatted = language_link{term = head.term, lang = data.lang}
else
formatted = data.lang:makeDisplayText(head.term, head.sc, true)
end
local head_part = tag_text_and_add_decorations(data, head, formatted, j)
insert(head_parts, head_part)
-- If multiple heads, try to determine whether all heads display the same. To do this we need to effectively
-- rerun the text tagging and addition of decorations, using 1 for all indices.
if has_multiple_heads then
local unique_head_part
if j == 1 then
unique_head_part = head_part
else
unique_head_part = tag_text_and_add_decorations(data, head, formatted, 1)
end
unique_head_parts[unique_head_part] = true
end
end
local set_size = 0
if has_multiple_heads then
for _ in pairs(unique_head_parts) do
set_size = set_size + 1
end
end
if set_size == 1 then
head_parts = head_parts[1]
else
head_parts = concat(head_parts)
end
if has_manual_translits then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr/LANGCODE]]
track("manual-tr", data.lang)
end
------ Format the transliterations and transcriptions. ------
local translits_formatted
if has_translits then
local translit_parts = {}
for _, head in ipairs(data.heads) do
if head.tr or head.ts then
local this_parts = {}
if head.tr then
insert(this_parts, tag_translit(head.tr, data.lang:getCode(), "head", nil, head.tr_manual))
if head.ts then
insert(this_parts, " ")
end
end
if head.ts then
insert(this_parts, "/" .. tag_transcription(head.ts, data.lang:getCode(), "head") .. "/")
end
insert(translit_parts, concat(this_parts))
end
end
translits_formatted = " (" .. concat(translit_parts, " <i>atau</i> ") .. ")"
local function format_transliteration_page_link(langname)
local transliteration_pagename = "Transliterasi bahasa " .. langname
local transliteration_page = new_title(transliteration_pagename, "Wikikamus")
if transliteration_page and transliteration_page:getContent() then
return ("[[Wikikamus:%s|•]]"):format(transliteration_pagename)
end
return nil
end
local translit_page_link = format_transliteration_page_link(data.lang:getCanonicalName())
-- If data.lang is an etymology-only language and we didn't find a translation page for it, fall back to the
-- full parent.
if not translit_page_link and data.lang:hasType("etymology-only") then
translit_page_link = format_transliteration_page_link(data.lang:getFullName())
end
if translit_page_link then
translits_formatted = " " .. translit_page_link .. translits_formatted
end
else
translits_formatted = ""
end
------ Paste heads and transliterations/transcriptions. ------
local lemma_gloss
if data.gloss then
lemma_gloss = ' <span class="ib-content qualifier-content">' .. data.gloss .. '</span>'
else
lemma_gloss = ""
end
return head_parts .. translits_formatted .. lemma_gloss
end
local function format_headword_genders(data, is_varform_only)
local retval = ""
if data.genders and data.genders[1] then
if data.gloss then
retval = ","
end
local pos_for_cat
if not data.nogendercat and not is_varform_only then
local no_gender_cat = (m_data or get_data()).no_gender_cat
if not (no_gender_cat[data.lang:getCode()] or no_gender_cat[data.lang:getFullCode()]) then
pos_for_cat = (m_data or get_data()).pos_for_gender_number_cat[data.pos_category:gsub("^reconstructed ", "")]
end
end
local text, cats = format_genders(data.genders, data.lang, pos_for_cat)
if cats then
extend(data.categories, cats)
end
retval = retval .. " " .. text
end
return retval
end
-- Forward reference
local format_inflections
local function format_inflection_parts(data, parts)
for j, part in ipairs(parts) do
if type(part) ~= "table" then
part = {term = part}
end
local partaccel = part.accel
local face = part.face or "bold"
if face ~= "bold" and face ~= "plain" and face ~= "hypothetical" then
error("The face `" .. face .. "` " .. (
(script_utilities_data or get_script_utilities_data()).faces[face] and
"should not be used for non-headword terms on the headword line." or
"is invalid."
))
end
-- Here the final part 'or data.nolinkinfl' allows to have 'nolinkinfl=true'
-- right into the 'data' table to disable inflection links of the entire headword
-- when inflected forms aren't entry-worthy, e.g.: in Vulgar Latin
local nolinkinfl = part.face == "hypothetical" or (part.nolink and track("nolink") or part.nolinkinfl) or (
data.nolink and track("nolink") or data.nolinkinfl)
local formatted
if part.label then
-- FIXME: There should be a better way of italicizing a label. As is, this isn't customizable.
formatted = "<i>" .. part.label .. "</i>"
else
-- Convert the term into a full link. Don't show a transliteration here unless enable_auto_translit is
-- requested, either at the `parts` level (i.e. per inflection) or at the `data.inflections` level (i.e.
-- specified for all inflections). This is controllable in {{head}} using autotrinfl=1 for all inflections,
-- or fNautotr=1 for an individual inflection (remember that a single inflection may be associated with
-- multiple terms). The reason for doing this is to avoid clutter in headword lines by default in languages
-- where the script is relatively straightforward to read by learners (e.g. Greek, Russian), but allow it
-- to be enabled in languages with more complex scripts (e.g. Arabic).
--
-- FIXME: With nested inflections, should we also respect `enable_auto_translit` at the top level of the
-- nested inflections structure?
local tr = part.tr or not (parts.enable_auto_translit or data.inflections.enable_auto_translit) and "-" or nil
local postprocess_annotations
if part.inflections then
postprocess_annotations = function(infldata)
insert(infldata.annotations, format_inflections(data, part.inflections))
end
end
formatted = full_link(
{
term = not nolinkinfl and part.term or nil,
alt = part.alt or (nolinkinfl and part.term or nil),
lang = part.lang or data.lang,
sc = part.sc or parts.sc or nil,
gloss = part.gloss,
pos = part.pos,
lit = part.lit,
id = part.id,
genders = part.genders,
tr = tr,
ts = part.ts,
accel = partaccel or parts.accel,
postprocess_annotations = postprocess_annotations,
},
face
)
end
parts[j] = format_term_with_decorations(part.lang or data.lang, part,
formatted, j)
end
local parts_output
if parts[1] then
parts_output = (parts.label and " " or "") .. concat(parts)
elseif parts.request then
parts_output = " <small>[sila nyatakan]</small>"
insert(data.categories, "Permohonan fleksi dalam entri bahasa " .. data.lang:getFullName())
else
parts_output = ""
end
local parts_label = parts.label and ("<i>" .. parts.label .. "</i>") or ""
return format_term_with_decorations(data.lang, parts, parts_label .. parts_output, 1)
end
-- Format the inflections following the headword or nested after a given inflection. Declared local above.
function format_inflections(data, inflections)
if inflections and inflections[1] then
-- Format each inflection individually.
for key, infl in ipairs(inflections) do
inflections[key] = format_inflection_parts(data, infl)
end
return concat(inflections, ", ")
else
return ""
end
end
-- Format the top-level inflections following the headword. Currently this just adds parens around the
-- formatted comma-separated inflections in `data.inflections`.
local function format_top_level_inflections(data)
local result = format_inflections(data, data.inflections)
if result ~= "" then
return " (" .. result .. ")"
else
return result
end
end
-- Forward reference
local check_red_link_inflections
-- Check a single inflection (which consists of a label and zero or more terms, each possibly with nested inflections)
-- for red links. If so, insert a red-link category based on `plpos` (the plural part of speech to insert in the
-- category), stop further processing, and return true. If no red links found, return false.
local function check_red_link_inflection_parts(data, parts, plpos)
for _, part in ipairs(parts) do
if type(part) ~= "table" then
part = {term = part}
end
local term = part.term
if term and not term:find("%[%[") then
local stripped_physical_term = get_link_page(term, data.lang, part.sc or parts.sc or nil)
if stripped_physical_term then
local title = mw.title.new(stripped_physical_term)
if title and not title:getContent() then
insert(data.categories, data.lang:getFullName() .. " " .. plpos .. " with red links in their headword lines")
return true
end
end
end
if part.inflections then
if check_red_link_inflections(data, part.inflections, plpos) then
return true
end
end
end
return false
end
-- Check a set of inflections (each of which describes a single inflection of the term, such as feminine or plural, and
-- consists of a label and zero or more terms, each possibly with nested inflections) for red links. If so, insert a
-- red-link category based on `plpos` (the plural part of speech to insert in the category), stop further processing,
-- and return true. If no red links found, return false.
function check_red_link_inflections(data, inflections, plpos)
if inflections and inflections[1] then
-- Check each inflection individually.
for key, infl in ipairs(inflections) do
if check_red_link_inflection_parts(data, infl, plpos) then
return true
end
end
end
return false
end
-- Check the top-level inflections in `data.inflections`, along with any nested inflections, for red links. If so,
-- insert a red-link category based on `plpos` (the plural part of speech to insert in the category), stop further
-- processing, and return true. If no red links found, return false.
local function check_red_link_inflections_top_level(data, plpos)
return check_red_link_inflections(data, data.inflections, plpos)
end
--[==[
Returns the plural form of `pos`, a raw part of speech input, which could be singular or
plural. Irregular plural POS are taken into account (e.g. "kanji" pluralizes to
"kanji").
]==]
function export.pluralize_pos(pos)
-- Make the plural form of the part of speech
return (m_data or get_data()).irregular_plurals[pos] or
pos:sub(-1) == "s" and pos or
pluralize(pos)
end
--[==[
Return "lemma" if the given POS is a lemma, "non-lemma form" if a non-lemma form, or nil
if unknown. The POS passed in must be in its plural form ("nouns", "prefixes", etc.).
If you have a POS in its singular form, call {export.pluralize_pos()} above to pluralize it
in a smart fashion that knows when to add "-s" and when to add "-es", and also takes
into account any irregular plurals.
If `best_guess` is given and the POS is in neither the lemma nor non-lemma list, guess
based on whether it ends in " forms"; otherwise, return nil.
]==]
function export.pos_lemma_or_nonlemma(plpos, best_guess)
local m_headword_data = m_data or get_data()
local isLemma = m_headword_data.lemmas
-- Is it a lemma category?
if isLemma[plpos] then
return "Lema"
end
local plpos_no_recon = plpos:gsub("^reconstructed ", "")
if isLemma[plpos_no_recon] then
return "Lema"
end
-- Is it a nonlemma category?
local isNonLemma = m_headword_data.nonlemmas
if isNonLemma[plpos] or isNonLemma[plpos_no_recon] then
return "Bentuk bukan lema"
end
local plpos_no_mut = plpos:gsub("^mutated ", "")
if isLemma[plpos_no_mut] or isNonLemma[plpos_no_mut] then
return "Bentuk bukan lema"
elseif best_guess then
return plpos:find("^Bentuk ") and "Bentuk bukan lema" or "Lema"
else
return nil
end
end
--[==[
Canonicalize a part of speech as specified in 2= in {{tl|head}}. This checks for POS aliases and non-lemma form
aliases ending in 'f', and then pluralizes if the POS term does not have an invariable plural.
]==]
function export.canonicalize_pos(pos)
-- FIXME: Temporary code to throw an error for alias 'pre' (= preposition) that will go away.
if pos == "pre" then
-- Don't throw error on 'pref' as it's an alias for "prefix".
error("POS 'pre' for 'preposition' no longer allowed as it's too ambiguous; use 'prep'")
end
-- Likewise for pro = pronoun.
if pos == "pro" or pos == "prof" then
error("POS 'pro' for 'pronoun' no longer allowed as it's too ambiguous; use 'pron'")
end
local m_headword_data = m_data or get_data()
if m_headword_data.pos_aliases[pos] then
pos = m_headword_data.pos_aliases[pos]
elseif pos:sub(-1) == "f" then
pos = pos:sub(1, -2)
pos = "Bentuk " .. (m_headword_data.pos_aliases[pos] or pos)
end
return export.pluralize_pos(pos)
end
-- Find and return the maximum index in the array `data[element]` (which may have gaps in it), and initialize it to a
-- zero-length array if unspecified. Check to make sure all keys are numeric (other than "maxindex", which is set by
-- [[Module:parameters]] for list parameters), all values are strings, and unless `allow_blank_string` is given,
-- no blank (zero-length) strings are present.
local function init_and_find_maximum_index(data, element, allow_blank_string)
local maxind = 0
if not data[element] then
data[element] = {}
end
local typ = type(data[element])
if typ ~= "table" then
error(("Internal error: In full_headword(), `data.%s` must be an array but is a %s"):format(element, typ))
end
for k, v in pairs(data[element]) do
if k ~= "maxindex" then
if type(k) ~= "number" then
error(("Internal error: Unrecognized non-numeric key '%s' in `data.%s`"):format(k, element))
end
if k > maxind then
maxind = k
end
if v then
if type(v) ~= "string" then
error(("Internal error: For key '%s' in `data.%s`, value should be a string but is a %s"):format(k, element, type(v)))
end
if not allow_blank_string and v == "" then
error(("Internal error: For key '%s' in `data.%s`, blank string not allowed; use 'false' for the default"):format(k, element))
end
end
end
end
return maxind
end
--[==[
-- Add the page to various maintenance categories for the language and the
-- whole page. These are placed in the headword somewhat arbitrarily, but
-- mainly because headword templates are mandatory for entries (meaning that
-- in theory it provides full coverage).
--
-- This is provided as an external entry point so that modules which transclude
-- information from other entries (such as {{tl|ja-see}}) can take advantage
-- of this feature as well, because they are used in place of a conventional
-- headword template.]==]
do
-- Handle any manual sortkeys that have been specified in raw categories
-- by tracking if they are the same or different from the automatically-
-- generated sortkey, so that we can track them in maintenance
-- categories.
local function handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
sortkey = sortkey or lang:makeSortKey(page.pagename)
-- If there are raw categories with no sortkey, then they will be
-- sorted based on the default MediaWiki sortkey, so we check against
-- that.
if tbl == true then
if page.raw_defaultsort ~= sortkey then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik")
end
return
end
local redundant, different
for k in pairs(tbl) do
if k == sortkey then
redundant = true
else
different = true
end
end
if redundant then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih lewah")
end
if different then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik")
end
return sortkey
end
function export.maintenance_cats(page, lang, lang_cats, page_cats)
extend(page_cats, page.cats)
lang = lang:getFull() -- since we are just generating categories
local canonical = lang:getCanonicalName()
local tbl = page.wikitext_topic_cat[lang:getCode()]
local sortkey = nil
if tbl then
sortkey = handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori topik yang menggunakan penanda mentah")
end
tbl = page.wikitext_langname_cat[canonical]
if tbl then
handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori nama bahasa yang menggunakan penanda mentah")
end
if get_current_L2() ~= "Bahasa " .. canonical then
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan pengepala bahasa tidak betul")
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header/LANGCODE]]
track("pengepala bahasa tidak betul", lang)
end
end
end
--[==[This is the primary external entry point.
{{lua|full_headword(data)}}
This is used by {{temp|head}} and various language-specific headword templates (e.g. {{temp|ru-adj}} for Russian adjectives, {{temp|de-noun}} for German nouns, etc.) to display an entire headword line.
See [[#Further explanations for full_headword()]]
]==]
function export.full_headword(data)
-- Prevent data from being destructively modified.
data = shallow_copy(data)
------------ 1. Basic checks for old-style (multi-arg) calling convention. ------------
if data.getCanonicalName then
error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) of properties, not a language object")
end
if not data.lang or type(data.lang) ~= "table" or not data.lang.getCode then
error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) and `data.lang` must be a language object")
end
if data.id and type(data.id) ~= "string" then
error("Internal error: The id in the data table should be a string.")
end
------------ 2. Initialize pagename etc. ------------
local langcode = data.lang:getCode()
local full_langcode = data.lang:getFullCode()
local langname = data.lang:getCanonicalName()
local full_langname = data.lang:getFullName()
local raw_pagename = data.pagename
local page
local m_headword_data = m_data or get_data()
if raw_pagename and raw_pagename ~= m_headword_data.pagename then -- for testing, doc pages, etc.
-- data.pagename is often set on documentation and test pages through the pagename= parameter of various
-- templates, to emulate running on that page. Having a large number of such test templates on a single
-- page often leads to timeouts, because we fetch and parse the contents of each page in turn. However,
-- we don't really need to do that and can function fine without fetching and parsing the contents of a
-- given page, so turn off content fetching/parsing (and also setting the DEFAULTSORT key through a parser
-- function, which is *slooooow*) in certain namespaces where test and documentation templates are likely to
-- be found and where actual content does not live (User, Template, Module).
local actual_namespace = m_headword_data.page.namespace
local no_fetch_content = actual_namespace == "User" or actual_namespace == "Template" or
actual_namespace == "Module"
page = process_page(raw_pagename, no_fetch_content)
else
page = m_headword_data.page
end
local namespace = page.namespace
if data.altform then
-- Temporary tracking for use of old altform=
track("altform", data.lang)
end
local is_varform_only = data.var and data.var ~= "both"
local is_varform_both = data.var == "both"
------------ 3. Initialize `data.heads` table; if old-style, convert to new-style. ------------
if type(data.heads) == "table" and type(data.heads[1]) == "table" then
-- new-style
if data.translits or data.transcriptions then
error("Internal error: In full_headword(), if `data.heads` is new-style (array of head objects), `data.translits` and `data.transcriptions` cannot be given")
end
else
-- convert old-style `heads`, `translits` and `transcriptions` to new-style
local maxind = max(
init_and_find_maximum_index(data, "heads"),
init_and_find_maximum_index(data, "translits", true),
init_and_find_maximum_index(data, "transcriptions", true)
)
for i = 1, maxind do
data.heads[i] = {
term = data.heads[i],
tr = data.translits[i],
ts = data.transcriptions[i],
}
end
end
-- Make sure there's at least one head.
if not data.heads[1] then
data.heads[1] = {}
end
------------ 4. Initialize and validate `data.categories` and `data.whole_page_categories`, and determine `pos_category` if not given, and add basic categories. ------------
init_and_find_maximum_index(data, "categories")
init_and_find_maximum_index(data, "whole_page_categories")
local pos_category_already_present = false
if data.categories[1] then
local escaped_langname = pattern_escape(full_langname)
local matches_lang_pattern = "^" .. escaped_langname .. " "
for _, cat in ipairs(data.categories) do
-- Does the category begin with the language name? If not, tag it with a tracking category.
if not cat:find(matches_lang_pattern) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category/LANGCODE]]
track("no lang category", data.lang)
end
end
-- If `pos_category` not given, try to infer it from the first specified category. If this doesn't work, we
-- throw an error below.
if not data.pos_category and data.categories[1]:find(matches_lang_pattern) then
data.pos_category = data.categories[1]:gsub(matches_lang_pattern, "")
-- Optimization to avoid inserting category already present.
pos_category_already_present = true
end
end
if not data.pos_category then
error("Internal error: `data.pos_category` not specified and could not be inferred from the categories given in "
.. "`data.categories`. Either specify the plural part of speech in `data.pos_category` "
.. "(e.g. \"proper nouns\") or ensure that the first category in `data.categories` is formed from the "
.. "language's canonical name plus the plural part of speech (e.g. \"Norwegian Bokmål proper nouns\")."
)
end
-- Insert a category at the beginning for the part of speech unless it's already present or `data.noposcat` given.
if not pos_category_already_present and not data.noposcat and not is_varform_only then
local pos_category = ucfirst(data.pos_category) .. " bahasa " .. full_langname
-- FIXME: [[User:Theknightwho]] Why is this special case here? Please add an explanatory comment.
if pos_category ~= "Aksara Han rentas bahasa" then
insert(data.categories, 1, pos_category)
end
end
-- Try to determine whether the part of speech refers to a lemma or a non-lemma form; if we can figure this out,
-- add an appropriate category.
local postype = export.pos_lemma_or_nonlemma(data.pos_category)
if not postype then
-- We don't know what this category is, so tag it with a tracking category.
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/LANGCODE]]
track("unrecognized pos", data.lang)
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS/LANGCODE]]
track("unrecognized pos/pos/" .. data.pos_category, data.lang)
elseif not data.noposcat and not is_varform_only then
insert(data.categories, 1, ucfirst(postype) .. " bahasa " .. full_langname)
end
-- Categorize variant forms into 'variant lemmas' or 'variant non-lemma forms'. Originally proposed in
-- [[Wiktionary:Beer parlour/2024/June#Decluttering the altform mess]] as 'alternative forms'; renamed in
-- [[Wiktionary:Beer parlour/2026/July#Renaming the "alternative forms" categories]].
if (is_varform_only or is_varform_both) and postype then
insert(data.categories, 1, postype .. " kelaianan bahasa " .. full_langname)
end
------------ 5. Create a default headword, and add links to multiword page names. ------------
-- Determine if this is an "anti-asterisk" term, i.e. an attested term in a language that must normally be
-- reconstructed.
local is_anti_asterisk = data.heads[1].term and data.heads[1].term:find("^!!")
local lang_reconstructed = data.lang:hasType("reconstructed")
if is_anti_asterisk then
if not lang_reconstructed then
error("Anti-asterisk feature (head= beginning with !!) can only be used with reconstructed languages")
end
lang_reconstructed = false
end
-- Determine if term is reconstructed
local is_reconstructed = namespace == "Rekonstruksi" or lang_reconstructed
-- Create a default headword based on the pagename, which is determined in
-- advance by the data module so that it only needs to be done once.
local default_head = page.pagename
-- Add links to multi-word page names when appropriate
if not (is_reconstructed or data.nolinkhead) then
local no_links = m_headword_data.no_multiword_links
if not (no_links[langcode] or no_links[full_langcode]) and export.head_is_multiword(default_head) then
default_head = export.add_multiword_links(default_head, true)
end
end
if is_reconstructed then
default_head = "*" .. default_head
end
------------ 6. Check the namespace against the language type. ------------
if namespace == "" then
if lang_reconstructed then
error("Entries in " .. langname .. " must be placed in the Rekonstruksi: namespace")
elseif data.lang:hasType("appendix-constructed") then
error("Entries in " .. langname .. " must be placed in the Lampiran: namespace")
end
elseif namespace == "Petikan" or namespace == "Tesaurus" then
error("Headword templates should not be used in the " .. namespace .. ": namespace.")
end
------------ 7. Fill in missing values in `data.heads`. ------------
-- True if any script among the headword scripts has spaces in it.
local any_script_has_spaces = false
-- True if any term has a redundant head= param.
local has_redundant_head_param = false
for _, head in ipairs(data.heads) do
------ 7a. If missing head, replace with default head.
if not head.term then
head.term = default_head
elseif head.term == default_head then
has_redundant_head_param = true
elseif is_anti_asterisk and head.term == "!!" then
-- If explicit head=!! is given, it's an anti-asterisk term and we fill in the default head.
head.term = "!!" .. default_head
elseif head.term:find("^[!?]$") then
-- If explicit head= just consists of ! or ?, add it to the end of the default head.
head.term = default_head .. head.term
end
head.term_no_initial_bang_bang = is_anti_asterisk and head.term:sub(3) or head.term
if is_reconstructed then
local head_term = head.term
if head_term:find("%[%[") then
head_term = remove_links(head_term)
end
if head_term:sub(1, 1) ~= "*" then
error("The headword '" .. head_term .. "' must begin with '*' to indicate that it is reconstructed.")
end
end
------ 7b. Try to detect the script(s) if not provided. If a per-head script is provided, that takes precedence,
------ otherwise fall back to the overall script if given. If neither given, autodetect the script.
local auto_sc = data.lang:findBestScript(head.term)
if (
auto_sc:getCode() == "None" and
find_best_script_without_lang(head.term):getCode() ~= "None"
) then
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dalam bentuk tulisan tidak piawai")
end
if not (head.sc or data.sc) then -- No script code given, so use autodetected script.
head.sc = auto_sc
else
if not head.sc then -- Overall script code given.
head.sc = data.sc
end
-- Track uses of sc parameter.
if head.sc:getCode() == auto_sc:getCode() then
track("redundant script code", data.lang)
if not data.no_script_code_cat then
insert(data.categories, "Perkataan dengan kod tulisan lewah bahasa " .. full_langname )
end
else
track("non-redundant manual script code", data.lang)
if not data.no_script_code_cat then
insert(data.categories, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. full_langname )
end
end
end
-- If using a discouraged character sequence, add to maintenance category.
if head.sc:hasNormalizationFixes() == true then
local composed_head = toNFC(head.term)
if head.sc:fixDiscouragedSequences(composed_head) ~= composed_head then
insert(data.whole_page_categories, "Laman menggunakan jujukan aksara tidak digalakkan")
end
end
any_script_has_spaces = any_script_has_spaces or head.sc:hasSpaces()
------ 7c. Create automatic transliterations for any non-Latin headwords without manual translit given
------ (provided automatic translit is available, e.g. not in Persian or Hebrew).
-- Make transliterations
head.tr_manual = nil
-- Try to generate a transliteration if necessary
if head.tr == "-" then
head.tr = nil
else
local notranslit = m_headword_data.notranslit
if not (notranslit[langcode] or notranslit[full_langcode]) and head.sc:isTransliterated() then
head.tr_manual = not not head.tr
local text = head.term_no_initial_bang_bang
if not data.lang:link_tr(head.sc) then
text = remove_links(text)
end
local automated_tr = data.lang:transliterate(text, head.sc)
if automated_tr then
local manual_tr = head.tr
if manual_tr then
if remove_links(manual_tr) == remove_links(automated_tr) then
insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi lewah")
else
insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi manual tidak lewah")
end
end
if not manual_tr then
head.tr = automated_tr
end
end
-- There is still no transliteration?
-- Add the entry to a cleanup category.
if not head.tr then
head.tr = "<small>transliterasi diperlukan</small>"
-- FIXME: No current support for 'Request for transliteration of Classical Persian terms' or similar.
-- Consider adding this support in [[Module:category tree/poscatboiler/data/entry maintenance]].
insert(data.categories, "Permintaan transliterasi perkataan bahasa " .. full_langname)
else
-- Otherwise, trim it.
head.tr = trim(head.tr)
end
end
end
-- Link to the transliteration entry for languages that require this.
if head.tr and data.lang:link_tr(head.sc) then
head.tr = full_link{
term = head.tr,
lang = data.lang,
sc = get_script("Latn"),
tr = "-"
}
end
end
------------ 8. Maybe tag the title with the appropriate script code, using the `display_title` mechanism. ------------
-- Assumes that the scripts in "toBeTagged" will never occur in the Reconstruction namespace.
-- (FIXME: Don't make assumptions like this, and if you need to do so, throw an error if the assumption is violated.)
-- Avoid tagging ASCII as Hani even when it is tagged as Hani in the headword, as in [[check]]. The check for ASCII
-- might need to be expanded to a check for any Latin characters and whitespace or punctuation.
local display_title
-- Where there are multiple headwords, use the script for the first. This assumes the first headword is similar to
-- the pagename, and that headwords that are in different scripts from the pagename aren't first. This seems to be
-- about the best we can do (alternatively we could potentially do script detection on the pagename).
local dt_script = data.heads[1].sc
local dt_script_code = dt_script:getCode()
local page_non_ascii = namespace == "" and not page.pagename:find("^[%z\1-\127]+$")
local unsupported_pagename, unsupported = page.full_raw_pagename:gsub("^Tajuk tidak disokong/", "")
if unsupported == 1 and page.unsupported_titles[unsupported_pagename] then
display_title = 'Tajuk tidak disokong/<span class="' .. dt_script_code .. '">' .. page.unsupported_titles[unsupported_pagename] .. '</span>'
elseif page_non_ascii and m_headword_data.toBeTagged[dt_script_code]
or (dt_script_code == "Jpan" and (text_in_script(page.pagename, "Hira") or text_in_script(page.pagename, "Kana")))
or (dt_script_code == "Kore" and text_in_script(page.pagename, "Hang")) then
display_title = '<span class="' .. dt_script_code .. '">' .. page.full_raw_pagename .. '</span>'
-- Keep Han entries region-neutral in the display title.
elseif page_non_ascii and (dt_script_code == "Hant" or dt_script_code == "Hans") then
display_title = '<span class="Hani">' .. page.full_raw_pagename .. '</span>'
elseif namespace == "Rekonstruksi" then
local matched
display_title, matched = ugsub(
page.full_raw_pagename,
"^(Rekonstruksi:[^/]+/)(.+)$",
function(before, term)
return before .. tag_text(term, data.lang, dt_script)
end
)
if matched == 0 then
display_title = nil
end
end
-- FIXME: Generalize this.
-- If the current language uses Aran (Nastaliq), e.g. Urdu, and there's more than one language on the page, don't
-- set the display title because we don't want Nastaliq for terms that also exist in other languages that don't
-- display in Nastaliq (e.g. Arabic or Persian). Because the word "Urdu" occurs near the end of the alphabet, Urdu
-- fonts tend to override the fonts of other languages. FIXME: This is checking for more than one language on the
-- page but instead needs to check if there are any languages using scripts other than Aran.
if dt_script_code == "Aran" and page.L2_list.n > 1 then
display_title = nil
end
if display_title then
mw.getCurrentFrame():callParserFunction(
"DISPLAYTITLE",
display_title
)
end
------------ 9. Insert additional categories. ------------
if data.force_cat_output then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/force cat output]]
track("force cat output")
end
if has_redundant_head_param then
if not data.no_redundant_head_cat then
-- This is not the right way to go about this; too many exceptions and problems due to language-specific headword
-- handling customization. If we want this, it should be opt-in by a given language passing in the default headword.
-- insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan parameter pengepala lewah")
end
end
-- If the first head is multiword (after removing links), maybe insert into "LANG multiword terms".
if not data.nomultiwordcat and not is_varform_only and any_script_has_spaces and postype == "lemma" then
local no_multiword_cat = m_headword_data.no_multiword_cat
if not (no_multiword_cat[langcode] or no_multiword_cat[full_langcode]) then
-- Check for spaces or hyphens, but exclude prefixes and suffixes.
-- Use the pagename, not the head= value, because the latter may have extra
-- junk in it, e.g. superscripted text that throws off the algorithm.
local no_hyphen = m_headword_data.hyphen_not_multiword_sep
-- Exclude hyphens if the data module states that they should for this language.
local checkpattern = (no_hyphen[langcode] or no_hyphen[full_langcode]) and ".[%s፡]." or ".[%s%-፡]."
local is_multiword = umatch(page.pagename, checkpattern)
if is_multiword and not non_categorizable(page.full_raw_pagename) then
insert(data.categories, "Perkataan berbilang kata bahasa " .. full_langname)
elseif not is_multiword then
local long_word_threshold = m_headword_data.long_word_thresholds[langcode] or
m_headword_data.long_word_thresholds[full_langcode]
if long_word_threshold and ulen(page.pagename) >= long_word_threshold then
insert(data.categories, "Perkataan panjang bahasa " .. full_langname)
end
end
end
end
-- Determine whether to insert a category 'LANGNAME POS in SCRIPT'. If there are multiple heads, we may need to check
-- each head, as the heads may (theoretically) have different scripts.
local default_sccat = m_headword_data.default_sccat
if data.sccat or not is_varform_only and (default_sccat[langcode] or langcode ~= full_langcode and default_sccat[full_langcode]) then
local function needs_sccat(sccat_entry, sc)
if sccat_entry == true or not sccat_entry then
return sccat_entry
end
if type(sccat_entry) == "table" then
local in_list = contains(sccat_entry, sc:getCode())
if sccat_entry[1] == "not" then
in_list = not in_list
end
return in_list
end
return nil
end
for _, head in ipairs(data.heads) do
-- First check the `sccat` specified at the {{head}} level.
local this_needs_sccat = needs_sccat(data.sccat, head.sc)
-- If that wasn't given, check the default sccat at the language level for the lang code.
if this_needs_sccat == nil and not is_varform_only then
this_needs_sccat = needs_sccat(default_sccat[langcode], head.sc)
end
-- If that wasn't found and the lang code is an etym code, check the default sccat at the parent language level.
if this_needs_sccat == nil and not is_varform_only and langcode ~= full_langcode then
this_needs_sccat = needs_sccat(default_sccat[full_langcode], head.sc)
end
if this_needs_sccat then
insert(data.categories, ucfirst(data.pos_category) .. " bahasa " .. full_langname .. " dalam " ..
head.sc:getDisplayForm(data.lang))
end
end
end
-- Reconstructed terms often use weird combinations of scripts and realistically aren't spelled so much as notated.
if namespace ~= "Rekonstruksi" and not is_varform_only then
-- Map from languages to a string containing the characters to ignore when considering whether a term has
-- multiple written scripts in it. Typically these are Greek or Cyrillic letters used for their phonetic
-- values.
local characters_to_ignore = {
["aaq"] = "αάὰ", -- Penobscot (Algonquian)
["acy"] = "δθ", -- Cypriot Arabic
["aez"] = "β", -- Aeka (Trans-New Guinea)
["anc"] = "γ", -- Ngas (Chadic/Afroasiatic)
["aou"] = "χ", -- A'ou (Kra-Dai)
["art-blk"] = "ч", -- Bolak (conlang)
["awg"] = "β", -- Anguthimri (Pama-Nyungan)
["az"] = "ь", -- Azerbaijani (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["ba"] = "ь", -- Bashkir (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["bhp"] = "β", -- Bima (Austronesian)
["bjz"] = "β", -- Baruga (Trans-New Guinea)
["byk"] = "θ", -- Biao (Kra-Dai)
["cdy"] = "θ", -- Chadong (Kra-Dai)
["chp"] = "θ", -- Chipewyan (Athabaskan)
["cjh"] = "χ", -- Upper Chehalis (Salishan)
["clm"] = "χ", -- Klallam (Salishan)
["col"] = "χ", -- Colombia-Wenatchi (Salishan)
["coo"] = "χθ", -- Comox (Salishan)
["crx"] = "θ", -- Carrier (Athabaskan)
["ets"] = "θ", -- Yekhee (Edoid/Niger-Congo)
["ett"] = "χ", -- Etruscan (isolate; in romanizations)
["fla"] = "χ", -- Montana Salish (Salishan)
["grt"] = "་", -- Garo (South Asian Sino-Tibetan)
["gmw-gts"] = "χ", -- Gottscheerish (Bavarian variant spoken in Slovenia)
["hur"] = "χθ", -- Halkomelem (Salishan)
["itc-psa"] = "f", -- Pre-Samnite (Italic; normally written in Greek)
["izh"] = "ь", -- Ingrian (Finnic)
["kic"] = "θ", -- Kickapoo (Algonquian)
["kk"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["ky"] = "ь", -- Kyrgyz (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["lil"] = "χ", -- Lillooet (Salishan)
["lsi"] = "ꓹ", -- Lashi (Lolo-Burmese/Sino-Tibetan; represents a glottal stop)
["mhz"] = "β", -- Mor (Austronesian)
["mqn"] = "β", -- Moronene (Austronesian)
["neg"]= "ӡā", -- Negidal (Tungusic; normally in Cyrillic)
["oka"] = "χ", -- Okanagan (Salishan)
["ole"] = "θ", -- Olekha (Sino-Tibetan)
["oui"] = "γβ", -- Old Uyghur (Turkic; FIXME: others? E.g. Greek delta (δ)?)
["pox"] = "χ", -- Polabian (West Slavic)
["rif"] = "ε", -- Tarifit (Berber)
["rom"] = "Θθ", -- Romani (Indic: International Standard; two different thetas???)
["rpn"] = "β", -- Repanbitip (Austronesian)
["sah"] = "ь", -- Yakut (Turkic; 1929 - 1939 Latin spelling)
["sit-jap"] = "χ", -- Japhug (Sino-Tibetan)
["sjw"] = "θ", -- Shawnee (Algonquian)
["squ"] = "χ", -- Squamish (Salishan)
["str"] = "χθ", -- Saanich (Salishan)
["teh"] = "χ", -- Tehuelche (Chonan; spoken in Argentina)
["tep"] = "η", -- Tepecano (Uto-Aztecan)
["thp"] = "χ", -- Thompson (Salishan)
["tk"] = "ь", -- Turkmen (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["tt"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["twa"] = "χ", -- Twana (Salishan)
["wbl"] = "ы", -- Wakhi (Iranian)
["xbc"] = "ϸ", -- Bactrian (Iranian; represents š; normally written in Greek)
["yha"] = "θ", -- Baha (Kra-Dai)
["za"] = "зч", -- Zhuang (Tai/Kra-Dai); 1957-1982 alphabet used two Cyrillic letters (as well as some others like
-- ƃ, ƅ, ƨ, ɯ and ɵ that look like Cyrillic or Greek but are actually Latin)
["zlw-slv"] = "χђћ", -- Slovincian (West Slavic; FIXME: χ is Greek, the other two are Cyrillic, but I'm not sure
-- the currect characters are being chosen in the entry names)
["zng"] = "θ", -- Mang (Mon-Khmer)
["ztp"] = "θ", -- Loxicha Zapotec (Zapotecan)
}
-- Determine how many real scripts are found in the pagename, where we exclude symbols and such. We exclude
-- scripts whose `character_category` is false as well as Zmth (mathematical notation symbols), which has a
-- category of "Mathematical notation symbols". When counting scripts, we need to elide language-specific
-- variants because e.g. Beng and as-Beng have slightly different characters but we don't want to consider them
-- two different scripts (e.g. [[এৰ]] has two characters which are detected respectively as Beng and as-Beng).
local seen_scripts = {}
local num_seen_scripts = 0
local num_loops = 0
local canon_pagename = page.pagename
local ch_to_ignore = characters_to_ignore[full_langcode]
if ch_to_ignore then
canon_pagename = ugsub(canon_pagename, "[" .. ch_to_ignore .. "]", "")
end
while true do
if canon_pagename == "" or num_seen_scripts >= 2 or num_loops >= 10 then
break
end
-- Make sure we don't get into a loop checking the same script over and over again; happens with e.g. [[ᠪᡳ]]
num_loops = num_loops + 1
local pagename_script = find_best_script_without_lang(canon_pagename, "None only as last resort")
local script_chars = pagename_script.characters
if not script_chars then
-- we are stuck; this happens with None
break
end
local script_code = pagename_script:getCode()
local replaced
canon_pagename, replaced = ugsub(canon_pagename, "[" .. script_chars .. "]", "")
if (
replaced and
script_code ~= "Zmth" and
(script_data or get_script_data())[script_code] and
script_data[script_code].character_category ~= false
) then
script_code = script_code:gsub("^.-%-", "")
if not seen_scripts[script_code] then
seen_scripts[script_code] = true
num_seen_scripts = num_seen_scripts + 1
end
end
end
if num_seen_scripts > 1 then
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dieja dalam berbilang tulisan")
end
end
-- Categorise for unusual characters. Takes into account combining characters, so that we can categorise for characters with diacritics that aren't encoded as atomic characters (e.g. U̠). These can be in two formats: single combining characters (i.e. character + diacritic(s)) or double combining characters (i.e. character + diacritic(s) + character). Each can have any number of diacritics.
local standard = data.lang:getStandardCharacters()
if not is_varform_only and standard and not non_categorizable(page.full_raw_pagename) then
local function char_category(char)
local specials = {
["#"] = "number sign",
["("] = "parentheses",
[")"] = "parentheses",
["<"] = "angle brackets",
[">"] = "angle brackets",
["["] = "square brackets",
["]"] = "square brackets",
["_"] = "underscore",
["{"] = "braces",
["|"] = "vertical line",
["}"] = "braces",
["ß"] = "ẞ",
["\205\133"] = "", -- this is UTF-8 for U+0345 ( ͅ)
["\239\191\189"] = "replacement character",
}
char = toNFD(char)
:gsub(".[\128-\191]*", function(m)
local new_m = specials[m]
new_m = new_m or m:uupper()
return new_m
end)
return toNFC(char)
end
if full_langcode ~= "hi" and full_langcode ~= "lo" then
local standard_chars_scripts = {}
for _, head in ipairs(data.heads) do
standard_chars_scripts[head.sc:getCode()] = true
end
-- Iterate over the scripts, in case there is more than one (as they can have different sets of standard characters).
for code in pairs(standard_chars_scripts) do
local sc_standard = data.lang:getStandardCharacters(code)
if sc_standard then
if page.pagename_len > 1 then
local explode_standard = {}
local function explode(char)
explode_standard[char] = true
return ""
end
local sc_standard_exploded = ugsub(sc_standard, page.comb_chars.combined_double, explode)
-- The following is correct; it relies on side-effecing the explode_standard[] table.
ugsub(sc_standard_exploded, page.comb_chars.combined_single, explode):gsub(".[\128-\191]*", explode)
local num_cat_inserted
for char in pairs(page.explode_pagename) do
if not explode_standard[char] then
if char:find("[0-9]") then
if not num_cat_inserted then
insert(data.categories, "Perkataan dieja dengan nombor bahasa " .. full_langname)
num_cat_inserted = true
end
elseif ufind(char, page.emoji_pattern) then
insert(data.categories, "Perkataan dieja dengan emoji bahasa " .. full_langname)
else
local upper = char_category(char)
if not explode_standard[upper] then
char = upper
end
insert(data.categories, "Perkataan dieja dengan " .. char .. " bahasa " .. full_langname)
end
end
end
end
-- If a diacritic doesn't appear in any of the standard characters, also categorise for it generally.
sc_standard = toNFD(sc_standard)
for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_single) do
if not umatch(sc_standard, diacritic) then
insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname)
end
end
for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_double) do
if not umatch(sc_standard, diacritic) then
insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname)
end
end
end
end
-- Ancient Greek, Hindi and Lao handled the old way for now, as their standard chars still need to be converted to the new format (because there are a lot of them).
elseif ulen(page.pagename) ~= 1 then
for character in ugmatch(page.pagename, "([^" .. standard .. "])") do
local upper = char_category(character)
if not umatch(upper, "[" .. standard .. "]") then
character = upper
end
insert(data.categories, "Perkataan dieja dengan " .. character .. " bahasa " .. full_langname)
end
end
end
if not is_varform_only and data.heads[1].sc:isSystem("alphabet") then
local pagename, i = page.pagename:ulower(), 2
while umatch(pagename, "(%a)" .. ("%1"):rep(i)) do
i = i + 1
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan " .. i .. " contoh huruf yang sama berturut-turut")
end
end
-- Categorise for palindromes
if not is_varform_only and not data.nopalindromecat and namespace ~= "Rekonstruksi" and ulen(page.pagename) > 2
-- FIXME: Use of first script here seems hacky. What is the clean way of doing this in the presence of
-- multiple scripts?
and is_palindrome(page.pagename, data.lang, data.heads[1].sc) then
insert(data.categories, "Palindrom bahasa " .. full_langname)
end
if namespace == "" and not lang_reconstructed then
for _, head in ipairs(data.heads) do
if page.full_raw_pagename ~= get_link_page(remove_links(head.term), data.lang, head.sc) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch/LANGCODE]]
track("pagename spelling mismatch", data.lang)
break
end
end
end
-- Add red link category if called for and we're not a "large" page, where such checks are disabled.
if data.checkredlinks and not m_headword_data.large_pages[m_headword_data.pagename] then
local plposcat = type(data.checkredlinks) == "string" and data.checkredlinks or data.pos_category
check_red_link_inflections_top_level(data, plposcat)
end
-- Add to various maintenance categories.
export.maintenance_cats(page, data.lang, data.categories, data.whole_page_categories)
------------ 10. Format and return headwords, genders, inflections and categories. ------------
-- Format and return all the gathered information. This may add more categories (e.g. gender/number categories),
-- so make sure we do it before evaluating `data.categories`.
local text = '<span class="headword-line">' ..
format_headword(data) ..
format_headword_genders(data, is_varform_only) ..
format_top_level_inflections(data) .. '</span>'
-- Language-specific categories.
local cats = format_categories(
data.categories, data.lang, data.sort_key, page.encoded_pagename,
data.force_cat_output or test_force_categories, data.heads[1].sc
)
-- Language-agnostic categories.
local whole_page_cats = format_categories(
data.whole_page_categories, nil, "-"
)
return text .. cats .. whole_page_cats
end
return export
dsnbdd62bmi4jy62atkjng0gw1w5e6o
375385
375353
2026-09-22T07:10:05Z
Hakimi97
2668
Betulkan ejaan
375385
Scribunto
text/plain
local export = {}
-- Named constants for all modules used, to make it easier to swap out sandbox versions.
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local gender_and_number_module = "Module:gender and number"
local headword_data_module = "Module:headword/data"
local headword_page_module = "Module:headword/page"
local links_module = "Module:links"
local load_module = "Module:load"
local pages_module = "Module:pages"
local palindromes_module = "Module:palindromes"
local scripts_module = "Module:scripts"
local scripts_data_module = "Module:scripts/data"
local script_utilities_module = "Module:script utilities"
local script_utilities_data_module = "Module:script utilities/data"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local dump = mw.dumpObject
local insert = table.insert
local ipairs = ipairs
local max = math.max
local new_title = mw.title.new
local pairs = pairs
local require = require
local toNFC = mw.ustring.toNFC
local toNFD = mw.ustring.toNFD
local type = type
local ufind = mw.ustring.find
local ugmatch = mw.ustring.gmatch
local ugsub = mw.ustring.gsub
local umatch = mw.ustring.match
--[==[
Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==]
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function contains(...)
contains = require(table_module).contains
return contains(...)
end
local function encode_entities(...)
encode_entities = require(string_utilities_module).encode_entities
return encode_entities(...)
end
local function extend(...)
extend = require(table_module).extend
return extend(...)
end
local function find_best_script_without_lang(...)
find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang
return find_best_script_without_lang(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function format_genders(...)
format_genders = require(gender_and_number_module).format_genders
return format_genders(...)
end
local function format_decorations(...)
format_decorations = require(decorations_module).format_decorations
return format_decorations(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_current_L2(...)
get_current_L2 = require(pages_module).get_current_L2
return get_current_L2(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function is_palindrome(...)
is_palindrome = require(palindromes_module).is_palindrome
return is_palindrome(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function load_data(...)
load_data = require(load_module).load_data
return load_data(...)
end
local function pattern_escape(...)
pattern_escape = require(string_utilities_module).pattern_escape
return pattern_escape(...)
end
local function pluralize(...)
pluralize = require(en_utilities_module).pluralize
return pluralize(...)
end
local function process_page(...)
process_page = require(headword_page_module).process_page
return process_page(...)
end
local function remove_links(...)
remove_links = require(links_module).remove_links
return remove_links(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function tag_text(...)
tag_text = require(script_utilities_module).tag_text
return tag_text(...)
end
local function tag_transcription(...)
tag_transcription = require(script_utilities_module).tag_transcription
return tag_transcription(...)
end
local function tag_translit(...)
tag_translit = require(script_utilities_module).tag_translit
return tag_translit(...)
end
local function trim(...)
trim = require(string_utilities_module).trim
return trim(...)
end
local function ulen(...)
ulen = require(string_utilities_module).len
return ulen(...)
end
local function ucfirst(...)
ucfirst = require(string_utilities_module).ucfirst
return ucfirst(...)
end
--[==[
Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==]
local m_data
local function get_data()
m_data = load_data(headword_data_module)
return m_data
end
local script_data
local function get_script_data()
script_data = load_data(scripts_data_module)
return script_data
end
local script_utilities_data
local function get_script_utilities_data()
script_utilities_data = load_data(script_utilities_data_module)
return script_utilities_data
end
-- If set to true, categories always appear, even in non-mainspace pages
local test_force_categories = false
-- Add a tracking category to track entries with certain (unusually undesirable) properties. `track_id` is an identifier
-- for the particular property being tracked and goes into the tracking page. Specifically, this adds a link in the
-- page text to [[Wiktionary:Tracking/headword/TRACK_ID]], meaning you can find all entries with the `track_id` property
-- by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID]].
--
-- If `lang` (a language object) is given, an additional tracking page [[Wiktionary:Tracking/headword/TRACK_ID/CODE]] is
-- linked to where CODE is the language code of `lang`, and you can find all entries in the combination of `track_id`
-- and `lang` by visiting [[Special:WhatLinksHere/Wiktionary:Tracking/headword/TRACK_ID/CODE]]. This makes it possible to
-- isolate only the entries with a specific tracking property that are in a given language. Note that if `lang`
-- references at etymology-only language, both that language's code and its full parent's code are tracked.
local function track(track_id, lang)
local tracking_page = "headword/" .. track_id
if lang and lang:hasType("etymology-only") then
debug_track{tracking_page, tracking_page .. "/" .. lang:getCode(),
tracking_page .. "/" .. lang:getFullCode()}
elseif lang then
debug_track{tracking_page, tracking_page .. "/" .. lang:getCode()}
else
debug_track(tracking_page)
end
return true
end
local function text_in_script(text, script_code)
local sc = get_script(script_code)
if not sc then
error("Internal error: Bad script code " .. script_code)
end
local characters = sc.characters
local out
if characters then
text = ugsub(text, "%W", "")
out = ufind(text, "[" .. characters .. "]")
end
if out then
return true
else
return false
end
end
local spacingPunctuation = "[%s%p]+"
--[[ List of punctuation or spacing characters that are found inside of words.
Used to exclude characters from the regex above. ]]
local wordPunc = "-#%%&@־׳״'.·*’་•:᠊"
local notWordPunc = "[^" .. wordPunc .. "]+"
--[=[
Format a term (either a head term or an inflection term) along with any decorations (left or right qualifiers, labels,
references or customized separator). `part` is the object specifying the term (and `lang` the language of the term),
which should optionally contain:
* left qualifiers in `q`, an array of strings;
* right qualifiers in `qq`, an array of strings;
* left labels in `l`, an array of strings;
* right labels in `ll`, an array of strings;
* references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text`
(formatted reference text) and optionally `name` and/or `group`;
* a separator in `separator`, defaulting to " <i>or</i> " if this is not the first term (j > 1), otherwise "".
`formatted` is the formatted version of the term itself, and `j` is the index of the term.
]=]
local function format_term_with_decorations(lang, part, formatted, j)
local function part_non_empty(field)
local list = part[field]
if not list then
return nil
end
if type(list) ~= "table" then
error(("Internal error: Wrong type for `part.%s`=%s, should be \"table\""):format(field, dump(list)))
end
return list[1]
end
if part_non_empty("q") or part_non_empty("qq") or part_non_empty("l") or
part_non_empty("ll") or part_non_empty("refs") then
formatted = format_decorations {
lang = lang,
text = formatted,
q = part.q,
qq = part.qq,
l = part.l,
ll = part.ll,
refs = part.refs,
}
end
local separator = part.separator or j > 1 and " <i>atau</i> " -- use "" to request no separator
if separator then
formatted = separator .. formatted
end
return formatted
end
--[==[Return true if the given head is multiword according to the algorithm used in full_headword().]==]
function export.head_is_multiword(head)
for possibleWordBreak in ugmatch(head, spacingPunctuation) do
if umatch(possibleWordBreak, notWordPunc) then
return true
end
end
return false
end
do
local function workaround_to_exclude_chars(s)
return (ugsub(s, notWordPunc, "\2%1\1"))
end
--[==[
Add appropriate links to `head`, correctly handling multiword terms. This is intended for multiword terms but can
be used for any term if you want links added to single-word terms as well. If you want to only add
links to multiword terms, first check that the term is multiword using `head_is_multiword`.
If `default` is specified, this will escape colons so that they don't get interpreted as interwiki links. This
should generally only be used when `head` is an actual pagename or is taken from a {{para|pagename}} parameter, not
when taken from a {{para|head}} parameter.
]==]
function export.add_multiword_links(head, default)
head = "\1" .. ugsub(head, spacingPunctuation, workaround_to_exclude_chars) .. "\2"
if default then
head = head
:gsub("(\1[^\2]*)\\([:#][^\2]*\2)", "%1\\\\%2")
:gsub("(\1[^\2]*)([:#][^\2]*\2)", "%1\\%2")
end
--Escape any remaining square brackets to stop them breaking links (e.g. "[citation needed]").
head = encode_entities(head, "[]", true, true)
--[=[
use this when workaround is no longer needed:
head = "[[" .. ugsub(head, WORDBREAKCHARS, "]]%1[[") .. "]]"
Remove any empty links, which could have been created above
at the beginning or end of the string.
]=]
return (head
:gsub("\1\2", "")
:gsub("[\1\2]", {["\1"] = "[[", ["\2"] = "]]"}))
end
end
local function non_categorizable(full_raw_pagename)
return full_raw_pagename:find("^Lampiran:Gerak isyarat/") or
-- Unsupported titles with descriptive names.
(full_raw_pagename:find("^Tajuk tidak disokong/") and not full_raw_pagename:find("`"))
end
local function tag_text_and_add_decorations(data, head, formatted, j)
-- Add language and script wrapper.
formatted = tag_text(formatted, data.lang, head.sc, "head", nil, j == 1 and data.id or nil)
-- Add decorations (qualifiers, labels, references and separator).
return format_term_with_decorations(data.lang, head, formatted, j)
end
-- Format a headword with transliterations.
local function format_headword(data)
-- Are there non-empty transliterations?
local has_translits = false
local has_manual_translits = false
------ Format the headwords. ------
local head_parts = {}
local unique_head_parts = {}
local has_multiple_heads = not not data.heads[2]
for j, head in ipairs(data.heads) do
if head.tr or head.ts then
has_translits = true
end
if head.tr and head.tr_manual or head.ts then
has_manual_translits = true
end
local formatted
-- Apply processing to the headword, for formatting links and such.
if head.term:find("[[", nil, true) and head.sc:getCode() ~= "Image" then
formatted = language_link{term = head.term, lang = data.lang}
else
formatted = data.lang:makeDisplayText(head.term, head.sc, true)
end
local head_part = tag_text_and_add_decorations(data, head, formatted, j)
insert(head_parts, head_part)
-- If multiple heads, try to determine whether all heads display the same. To do this we need to effectively
-- rerun the text tagging and addition of decorations, using 1 for all indices.
if has_multiple_heads then
local unique_head_part
if j == 1 then
unique_head_part = head_part
else
unique_head_part = tag_text_and_add_decorations(data, head, formatted, 1)
end
unique_head_parts[unique_head_part] = true
end
end
local set_size = 0
if has_multiple_heads then
for _ in pairs(unique_head_parts) do
set_size = set_size + 1
end
end
if set_size == 1 then
head_parts = head_parts[1]
else
head_parts = concat(head_parts)
end
if has_manual_translits then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/manual-tr/LANGCODE]]
track("manual-tr", data.lang)
end
------ Format the transliterations and transcriptions. ------
local translits_formatted
if has_translits then
local translit_parts = {}
for _, head in ipairs(data.heads) do
if head.tr or head.ts then
local this_parts = {}
if head.tr then
insert(this_parts, tag_translit(head.tr, data.lang:getCode(), "head", nil, head.tr_manual))
if head.ts then
insert(this_parts, " ")
end
end
if head.ts then
insert(this_parts, "/" .. tag_transcription(head.ts, data.lang:getCode(), "head") .. "/")
end
insert(translit_parts, concat(this_parts))
end
end
translits_formatted = " (" .. concat(translit_parts, " <i>atau</i> ") .. ")"
local function format_transliteration_page_link(langname)
local transliteration_pagename = "Transliterasi bahasa " .. langname
local transliteration_page = new_title(transliteration_pagename, "Wikikamus")
if transliteration_page and transliteration_page:getContent() then
return ("[[Wikikamus:%s|•]]"):format(transliteration_pagename)
end
return nil
end
local translit_page_link = format_transliteration_page_link(data.lang:getCanonicalName())
-- If data.lang is an etymology-only language and we didn't find a translation page for it, fall back to the
-- full parent.
if not translit_page_link and data.lang:hasType("etymology-only") then
translit_page_link = format_transliteration_page_link(data.lang:getFullName())
end
if translit_page_link then
translits_formatted = " " .. translit_page_link .. translits_formatted
end
else
translits_formatted = ""
end
------ Paste heads and transliterations/transcriptions. ------
local lemma_gloss
if data.gloss then
lemma_gloss = ' <span class="ib-content qualifier-content">' .. data.gloss .. '</span>'
else
lemma_gloss = ""
end
return head_parts .. translits_formatted .. lemma_gloss
end
local function format_headword_genders(data, is_varform_only)
local retval = ""
if data.genders and data.genders[1] then
if data.gloss then
retval = ","
end
local pos_for_cat
if not data.nogendercat and not is_varform_only then
local no_gender_cat = (m_data or get_data()).no_gender_cat
if not (no_gender_cat[data.lang:getCode()] or no_gender_cat[data.lang:getFullCode()]) then
pos_for_cat = (m_data or get_data()).pos_for_gender_number_cat[data.pos_category:gsub("^reconstructed ", "")]
end
end
local text, cats = format_genders(data.genders, data.lang, pos_for_cat)
if cats then
extend(data.categories, cats)
end
retval = retval .. " " .. text
end
return retval
end
-- Forward reference
local format_inflections
local function format_inflection_parts(data, parts)
for j, part in ipairs(parts) do
if type(part) ~= "table" then
part = {term = part}
end
local partaccel = part.accel
local face = part.face or "bold"
if face ~= "bold" and face ~= "plain" and face ~= "hypothetical" then
error("The face `" .. face .. "` " .. (
(script_utilities_data or get_script_utilities_data()).faces[face] and
"should not be used for non-headword terms on the headword line." or
"is invalid."
))
end
-- Here the final part 'or data.nolinkinfl' allows to have 'nolinkinfl=true'
-- right into the 'data' table to disable inflection links of the entire headword
-- when inflected forms aren't entry-worthy, e.g.: in Vulgar Latin
local nolinkinfl = part.face == "hypothetical" or (part.nolink and track("nolink") or part.nolinkinfl) or (
data.nolink and track("nolink") or data.nolinkinfl)
local formatted
if part.label then
-- FIXME: There should be a better way of italicizing a label. As is, this isn't customizable.
formatted = "<i>" .. part.label .. "</i>"
else
-- Convert the term into a full link. Don't show a transliteration here unless enable_auto_translit is
-- requested, either at the `parts` level (i.e. per inflection) or at the `data.inflections` level (i.e.
-- specified for all inflections). This is controllable in {{head}} using autotrinfl=1 for all inflections,
-- or fNautotr=1 for an individual inflection (remember that a single inflection may be associated with
-- multiple terms). The reason for doing this is to avoid clutter in headword lines by default in languages
-- where the script is relatively straightforward to read by learners (e.g. Greek, Russian), but allow it
-- to be enabled in languages with more complex scripts (e.g. Arabic).
--
-- FIXME: With nested inflections, should we also respect `enable_auto_translit` at the top level of the
-- nested inflections structure?
local tr = part.tr or not (parts.enable_auto_translit or data.inflections.enable_auto_translit) and "-" or nil
local postprocess_annotations
if part.inflections then
postprocess_annotations = function(infldata)
insert(infldata.annotations, format_inflections(data, part.inflections))
end
end
formatted = full_link(
{
term = not nolinkinfl and part.term or nil,
alt = part.alt or (nolinkinfl and part.term or nil),
lang = part.lang or data.lang,
sc = part.sc or parts.sc or nil,
gloss = part.gloss,
pos = part.pos,
lit = part.lit,
id = part.id,
genders = part.genders,
tr = tr,
ts = part.ts,
accel = partaccel or parts.accel,
postprocess_annotations = postprocess_annotations,
},
face
)
end
parts[j] = format_term_with_decorations(part.lang or data.lang, part,
formatted, j)
end
local parts_output
if parts[1] then
parts_output = (parts.label and " " or "") .. concat(parts)
elseif parts.request then
parts_output = " <small>[sila nyatakan]</small>"
insert(data.categories, "Permohonan fleksi dalam entri bahasa " .. data.lang:getFullName())
else
parts_output = ""
end
local parts_label = parts.label and ("<i>" .. parts.label .. "</i>") or ""
return format_term_with_decorations(data.lang, parts, parts_label .. parts_output, 1)
end
-- Format the inflections following the headword or nested after a given inflection. Declared local above.
function format_inflections(data, inflections)
if inflections and inflections[1] then
-- Format each inflection individually.
for key, infl in ipairs(inflections) do
inflections[key] = format_inflection_parts(data, infl)
end
return concat(inflections, ", ")
else
return ""
end
end
-- Format the top-level inflections following the headword. Currently this just adds parens around the
-- formatted comma-separated inflections in `data.inflections`.
local function format_top_level_inflections(data)
local result = format_inflections(data, data.inflections)
if result ~= "" then
return " (" .. result .. ")"
else
return result
end
end
-- Forward reference
local check_red_link_inflections
-- Check a single inflection (which consists of a label and zero or more terms, each possibly with nested inflections)
-- for red links. If so, insert a red-link category based on `plpos` (the plural part of speech to insert in the
-- category), stop further processing, and return true. If no red links found, return false.
local function check_red_link_inflection_parts(data, parts, plpos)
for _, part in ipairs(parts) do
if type(part) ~= "table" then
part = {term = part}
end
local term = part.term
if term and not term:find("%[%[") then
local stripped_physical_term = get_link_page(term, data.lang, part.sc or parts.sc or nil)
if stripped_physical_term then
local title = mw.title.new(stripped_physical_term)
if title and not title:getContent() then
insert(data.categories, data.lang:getFullName() .. " " .. plpos .. " with red links in their headword lines")
return true
end
end
end
if part.inflections then
if check_red_link_inflections(data, part.inflections, plpos) then
return true
end
end
end
return false
end
-- Check a set of inflections (each of which describes a single inflection of the term, such as feminine or plural, and
-- consists of a label and zero or more terms, each possibly with nested inflections) for red links. If so, insert a
-- red-link category based on `plpos` (the plural part of speech to insert in the category), stop further processing,
-- and return true. If no red links found, return false.
function check_red_link_inflections(data, inflections, plpos)
if inflections and inflections[1] then
-- Check each inflection individually.
for key, infl in ipairs(inflections) do
if check_red_link_inflection_parts(data, infl, plpos) then
return true
end
end
end
return false
end
-- Check the top-level inflections in `data.inflections`, along with any nested inflections, for red links. If so,
-- insert a red-link category based on `plpos` (the plural part of speech to insert in the category), stop further
-- processing, and return true. If no red links found, return false.
local function check_red_link_inflections_top_level(data, plpos)
return check_red_link_inflections(data, data.inflections, plpos)
end
--[==[
Returns the plural form of `pos`, a raw part of speech input, which could be singular or
plural. Irregular plural POS are taken into account (e.g. "kanji" pluralizes to
"kanji").
]==]
function export.pluralize_pos(pos)
-- Make the plural form of the part of speech
return (m_data or get_data()).irregular_plurals[pos] or
pos:sub(-1) == "s" and pos or
pluralize(pos)
end
--[==[
Return "lemma" if the given POS is a lemma, "non-lemma form" if a non-lemma form, or nil
if unknown. The POS passed in must be in its plural form ("nouns", "prefixes", etc.).
If you have a POS in its singular form, call {export.pluralize_pos()} above to pluralize it
in a smart fashion that knows when to add "-s" and when to add "-es", and also takes
into account any irregular plurals.
If `best_guess` is given and the POS is in neither the lemma nor non-lemma list, guess
based on whether it ends in " forms"; otherwise, return nil.
]==]
function export.pos_lemma_or_nonlemma(plpos, best_guess)
local m_headword_data = m_data or get_data()
local isLemma = m_headword_data.lemmas
-- Is it a lemma category?
if isLemma[plpos] then
return "Lema"
end
local plpos_no_recon = plpos:gsub("^reconstructed ", "")
if isLemma[plpos_no_recon] then
return "Lema"
end
-- Is it a nonlemma category?
local isNonLemma = m_headword_data.nonlemmas
if isNonLemma[plpos] or isNonLemma[plpos_no_recon] then
return "Bentuk bukan lema"
end
local plpos_no_mut = plpos:gsub("^mutated ", "")
if isLemma[plpos_no_mut] or isNonLemma[plpos_no_mut] then
return "Bentuk bukan lema"
elseif best_guess then
return plpos:find("^Bentuk ") and "Bentuk bukan lema" or "Lema"
else
return nil
end
end
--[==[
Canonicalize a part of speech as specified in 2= in {{tl|head}}. This checks for POS aliases and non-lemma form
aliases ending in 'f', and then pluralizes if the POS term does not have an invariable plural.
]==]
function export.canonicalize_pos(pos)
-- FIXME: Temporary code to throw an error for alias 'pre' (= preposition) that will go away.
if pos == "pre" then
-- Don't throw error on 'pref' as it's an alias for "prefix".
error("POS 'pre' for 'preposition' no longer allowed as it's too ambiguous; use 'prep'")
end
-- Likewise for pro = pronoun.
if pos == "pro" or pos == "prof" then
error("POS 'pro' for 'pronoun' no longer allowed as it's too ambiguous; use 'pron'")
end
local m_headword_data = m_data or get_data()
if m_headword_data.pos_aliases[pos] then
pos = m_headword_data.pos_aliases[pos]
elseif pos:sub(-1) == "f" then
pos = pos:sub(1, -2)
pos = "Bentuk " .. (m_headword_data.pos_aliases[pos] or pos)
end
return export.pluralize_pos(pos)
end
-- Find and return the maximum index in the array `data[element]` (which may have gaps in it), and initialize it to a
-- zero-length array if unspecified. Check to make sure all keys are numeric (other than "maxindex", which is set by
-- [[Module:parameters]] for list parameters), all values are strings, and unless `allow_blank_string` is given,
-- no blank (zero-length) strings are present.
local function init_and_find_maximum_index(data, element, allow_blank_string)
local maxind = 0
if not data[element] then
data[element] = {}
end
local typ = type(data[element])
if typ ~= "table" then
error(("Internal error: In full_headword(), `data.%s` must be an array but is a %s"):format(element, typ))
end
for k, v in pairs(data[element]) do
if k ~= "maxindex" then
if type(k) ~= "number" then
error(("Internal error: Unrecognized non-numeric key '%s' in `data.%s`"):format(k, element))
end
if k > maxind then
maxind = k
end
if v then
if type(v) ~= "string" then
error(("Internal error: For key '%s' in `data.%s`, value should be a string but is a %s"):format(k, element, type(v)))
end
if not allow_blank_string and v == "" then
error(("Internal error: For key '%s' in `data.%s`, blank string not allowed; use 'false' for the default"):format(k, element))
end
end
end
end
return maxind
end
--[==[
-- Add the page to various maintenance categories for the language and the
-- whole page. These are placed in the headword somewhat arbitrarily, but
-- mainly because headword templates are mandatory for entries (meaning that
-- in theory it provides full coverage).
--
-- This is provided as an external entry point so that modules which transclude
-- information from other entries (such as {{tl|ja-see}}) can take advantage
-- of this feature as well, because they are used in place of a conventional
-- headword template.]==]
do
-- Handle any manual sortkeys that have been specified in raw categories
-- by tracking if they are the same or different from the automatically-
-- generated sortkey, so that we can track them in maintenance
-- categories.
local function handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
sortkey = sortkey or lang:makeSortKey(page.pagename)
-- If there are raw categories with no sortkey, then they will be
-- sorted based on the default MediaWiki sortkey, so we check against
-- that.
if tbl == true then
if page.raw_defaultsort ~= sortkey then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik")
end
return
end
local redundant, different
for k in pairs(tbl) do
if k == sortkey then
redundant = true
else
different = true
end
end
if redundant then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih lewah")
end
if different then
insert(lang_cats, "Perkataan bahasa " .. lang:getFullName() .. " dengan kunci isih tidak lewah dan tidak automatik")
end
return sortkey
end
function export.maintenance_cats(page, lang, lang_cats, page_cats)
extend(page_cats, page.cats)
lang = lang:getFull() -- since we are just generating categories
local canonical = lang:getCanonicalName()
local tbl = page.wikitext_topic_cat[lang:getCode()]
local sortkey = nil
if tbl then
sortkey = handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori topik yang menggunakan penanda mentah")
end
tbl = page.wikitext_langname_cat[canonical]
if tbl then
handle_raw_sortkeys(tbl, sortkey, page, lang, lang_cats)
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan kategori nama bahasa yang menggunakan penanda mentah")
end
if get_current_L2() ~= "Bahasa " .. canonical then
insert(lang_cats, "Entri bahasa " .. canonical .. " dengan pengepala bahasa tidak betul")
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/incorrect language header/LANGCODE]]
track("pengepala bahasa tidak betul", lang)
end
end
end
--[==[This is the primary external entry point.
{{lua|full_headword(data)}}
This is used by {{temp|head}} and various language-specific headword templates (e.g. {{temp|ru-adj}} for Russian adjectives, {{temp|de-noun}} for German nouns, etc.) to display an entire headword line.
See [[#Further explanations for full_headword()]]
]==]
function export.full_headword(data)
-- Prevent data from being destructively modified.
data = shallow_copy(data)
------------ 1. Basic checks for old-style (multi-arg) calling convention. ------------
if data.getCanonicalName then
error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) of properties, not a language object")
end
if not data.lang or type(data.lang) ~= "table" or not data.lang.getCode then
error("Internal error: In full_headword(), the first argument `data` needs to be a Lua object (table) and `data.lang` must be a language object")
end
if data.id and type(data.id) ~= "string" then
error("Internal error: The id in the data table should be a string.")
end
------------ 2. Initialize pagename etc. ------------
local langcode = data.lang:getCode()
local full_langcode = data.lang:getFullCode()
local langname = data.lang:getCanonicalName()
local full_langname = data.lang:getFullName()
local raw_pagename = data.pagename
local page
local m_headword_data = m_data or get_data()
if raw_pagename and raw_pagename ~= m_headword_data.pagename then -- for testing, doc pages, etc.
-- data.pagename is often set on documentation and test pages through the pagename= parameter of various
-- templates, to emulate running on that page. Having a large number of such test templates on a single
-- page often leads to timeouts, because we fetch and parse the contents of each page in turn. However,
-- we don't really need to do that and can function fine without fetching and parsing the contents of a
-- given page, so turn off content fetching/parsing (and also setting the DEFAULTSORT key through a parser
-- function, which is *slooooow*) in certain namespaces where test and documentation templates are likely to
-- be found and where actual content does not live (User, Template, Module).
local actual_namespace = m_headword_data.page.namespace
local no_fetch_content = actual_namespace == "User" or actual_namespace == "Template" or
actual_namespace == "Module"
page = process_page(raw_pagename, no_fetch_content)
else
page = m_headword_data.page
end
local namespace = page.namespace
if data.altform then
-- Temporary tracking for use of old altform=
track("altform", data.lang)
end
local is_varform_only = data.var and data.var ~= "both"
local is_varform_both = data.var == "both"
------------ 3. Initialize `data.heads` table; if old-style, convert to new-style. ------------
if type(data.heads) == "table" and type(data.heads[1]) == "table" then
-- new-style
if data.translits or data.transcriptions then
error("Internal error: In full_headword(), if `data.heads` is new-style (array of head objects), `data.translits` and `data.transcriptions` cannot be given")
end
else
-- convert old-style `heads`, `translits` and `transcriptions` to new-style
local maxind = max(
init_and_find_maximum_index(data, "heads"),
init_and_find_maximum_index(data, "translits", true),
init_and_find_maximum_index(data, "transcriptions", true)
)
for i = 1, maxind do
data.heads[i] = {
term = data.heads[i],
tr = data.translits[i],
ts = data.transcriptions[i],
}
end
end
-- Make sure there's at least one head.
if not data.heads[1] then
data.heads[1] = {}
end
------------ 4. Initialize and validate `data.categories` and `data.whole_page_categories`, and determine `pos_category` if not given, and add basic categories. ------------
init_and_find_maximum_index(data, "categories")
init_and_find_maximum_index(data, "whole_page_categories")
local pos_category_already_present = false
if data.categories[1] then
local escaped_langname = pattern_escape(full_langname)
local matches_lang_pattern = "^" .. escaped_langname .. " "
for _, cat in ipairs(data.categories) do
-- Does the category begin with the language name? If not, tag it with a tracking category.
if not cat:find(matches_lang_pattern) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/no lang category/LANGCODE]]
track("no lang category", data.lang)
end
end
-- If `pos_category` not given, try to infer it from the first specified category. If this doesn't work, we
-- throw an error below.
if not data.pos_category and data.categories[1]:find(matches_lang_pattern) then
data.pos_category = data.categories[1]:gsub(matches_lang_pattern, "")
-- Optimization to avoid inserting category already present.
pos_category_already_present = true
end
end
if not data.pos_category then
error("Internal error: `data.pos_category` not specified and could not be inferred from the categories given in "
.. "`data.categories`. Either specify the plural part of speech in `data.pos_category` "
.. "(e.g. \"proper nouns\") or ensure that the first category in `data.categories` is formed from the "
.. "language's canonical name plus the plural part of speech (e.g. \"Norwegian Bokmål proper nouns\")."
)
end
-- Insert a category at the beginning for the part of speech unless it's already present or `data.noposcat` given.
if not pos_category_already_present and not data.noposcat and not is_varform_only then
local pos_category = ucfirst(data.pos_category) .. " bahasa " .. full_langname
-- FIXME: [[User:Theknightwho]] Why is this special case here? Please add an explanatory comment.
if pos_category ~= "Aksara Han rentas bahasa" then
insert(data.categories, 1, pos_category)
end
end
-- Try to determine whether the part of speech refers to a lemma or a non-lemma form; if we can figure this out,
-- add an appropriate category.
local postype = export.pos_lemma_or_nonlemma(data.pos_category)
if not postype then
-- We don't know what this category is, so tag it with a tracking category.
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/LANGCODE]]
track("unrecognized pos", data.lang)
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/unrecognized pos/POS/LANGCODE]]
track("unrecognized pos/pos/" .. data.pos_category, data.lang)
elseif not data.noposcat and not is_varform_only then
insert(data.categories, 1, ucfirst(postype) .. " bahasa " .. full_langname)
end
-- Categorize variant forms into 'variant lemmas' or 'variant non-lemma forms'. Originally proposed in
-- [[Wiktionary:Beer parlour/2024/June#Decluttering the altform mess]] as 'alternative forms'; renamed in
-- [[Wiktionary:Beer parlour/2026/July#Renaming the "alternative forms" categories]].
if (is_varform_only or is_varform_both) and postype then
insert(data.categories, 1, postype .. " kelainan bahasa " .. full_langname)
end
------------ 5. Create a default headword, and add links to multiword page names. ------------
-- Determine if this is an "anti-asterisk" term, i.e. an attested term in a language that must normally be
-- reconstructed.
local is_anti_asterisk = data.heads[1].term and data.heads[1].term:find("^!!")
local lang_reconstructed = data.lang:hasType("reconstructed")
if is_anti_asterisk then
if not lang_reconstructed then
error("Anti-asterisk feature (head= beginning with !!) can only be used with reconstructed languages")
end
lang_reconstructed = false
end
-- Determine if term is reconstructed
local is_reconstructed = namespace == "Rekonstruksi" or lang_reconstructed
-- Create a default headword based on the pagename, which is determined in
-- advance by the data module so that it only needs to be done once.
local default_head = page.pagename
-- Add links to multi-word page names when appropriate
if not (is_reconstructed or data.nolinkhead) then
local no_links = m_headword_data.no_multiword_links
if not (no_links[langcode] or no_links[full_langcode]) and export.head_is_multiword(default_head) then
default_head = export.add_multiword_links(default_head, true)
end
end
if is_reconstructed then
default_head = "*" .. default_head
end
------------ 6. Check the namespace against the language type. ------------
if namespace == "" then
if lang_reconstructed then
error("Entries in " .. langname .. " must be placed in the Rekonstruksi: namespace")
elseif data.lang:hasType("appendix-constructed") then
error("Entries in " .. langname .. " must be placed in the Lampiran: namespace")
end
elseif namespace == "Petikan" or namespace == "Tesaurus" then
error("Headword templates should not be used in the " .. namespace .. ": namespace.")
end
------------ 7. Fill in missing values in `data.heads`. ------------
-- True if any script among the headword scripts has spaces in it.
local any_script_has_spaces = false
-- True if any term has a redundant head= param.
local has_redundant_head_param = false
for _, head in ipairs(data.heads) do
------ 7a. If missing head, replace with default head.
if not head.term then
head.term = default_head
elseif head.term == default_head then
has_redundant_head_param = true
elseif is_anti_asterisk and head.term == "!!" then
-- If explicit head=!! is given, it's an anti-asterisk term and we fill in the default head.
head.term = "!!" .. default_head
elseif head.term:find("^[!?]$") then
-- If explicit head= just consists of ! or ?, add it to the end of the default head.
head.term = default_head .. head.term
end
head.term_no_initial_bang_bang = is_anti_asterisk and head.term:sub(3) or head.term
if is_reconstructed then
local head_term = head.term
if head_term:find("%[%[") then
head_term = remove_links(head_term)
end
if head_term:sub(1, 1) ~= "*" then
error("The headword '" .. head_term .. "' must begin with '*' to indicate that it is reconstructed.")
end
end
------ 7b. Try to detect the script(s) if not provided. If a per-head script is provided, that takes precedence,
------ otherwise fall back to the overall script if given. If neither given, autodetect the script.
local auto_sc = data.lang:findBestScript(head.term)
if (
auto_sc:getCode() == "None" and
find_best_script_without_lang(head.term):getCode() ~= "None"
) then
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dalam bentuk tulisan tidak piawai")
end
if not (head.sc or data.sc) then -- No script code given, so use autodetected script.
head.sc = auto_sc
else
if not head.sc then -- Overall script code given.
head.sc = data.sc
end
-- Track uses of sc parameter.
if head.sc:getCode() == auto_sc:getCode() then
track("redundant script code", data.lang)
if not data.no_script_code_cat then
insert(data.categories, "Perkataan dengan kod tulisan lewah bahasa " .. full_langname )
end
else
track("non-redundant manual script code", data.lang)
if not data.no_script_code_cat then
insert(data.categories, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. full_langname )
end
end
end
-- If using a discouraged character sequence, add to maintenance category.
if head.sc:hasNormalizationFixes() == true then
local composed_head = toNFC(head.term)
if head.sc:fixDiscouragedSequences(composed_head) ~= composed_head then
insert(data.whole_page_categories, "Laman menggunakan jujukan aksara tidak digalakkan")
end
end
any_script_has_spaces = any_script_has_spaces or head.sc:hasSpaces()
------ 7c. Create automatic transliterations for any non-Latin headwords without manual translit given
------ (provided automatic translit is available, e.g. not in Persian or Hebrew).
-- Make transliterations
head.tr_manual = nil
-- Try to generate a transliteration if necessary
if head.tr == "-" then
head.tr = nil
else
local notranslit = m_headword_data.notranslit
if not (notranslit[langcode] or notranslit[full_langcode]) and head.sc:isTransliterated() then
head.tr_manual = not not head.tr
local text = head.term_no_initial_bang_bang
if not data.lang:link_tr(head.sc) then
text = remove_links(text)
end
local automated_tr = data.lang:transliterate(text, head.sc)
if automated_tr then
local manual_tr = head.tr
if manual_tr then
if remove_links(manual_tr) == remove_links(automated_tr) then
insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi lewah")
else
insert(data.categories, "Perkataan bahasa ".. full_langname .. " dengan transliterasi manual tidak lewah")
end
end
if not manual_tr then
head.tr = automated_tr
end
end
-- There is still no transliteration?
-- Add the entry to a cleanup category.
if not head.tr then
head.tr = "<small>transliterasi diperlukan</small>"
-- FIXME: No current support for 'Request for transliteration of Classical Persian terms' or similar.
-- Consider adding this support in [[Module:category tree/poscatboiler/data/entry maintenance]].
insert(data.categories, "Permintaan transliterasi perkataan bahasa " .. full_langname)
else
-- Otherwise, trim it.
head.tr = trim(head.tr)
end
end
end
-- Link to the transliteration entry for languages that require this.
if head.tr and data.lang:link_tr(head.sc) then
head.tr = full_link{
term = head.tr,
lang = data.lang,
sc = get_script("Latn"),
tr = "-"
}
end
end
------------ 8. Maybe tag the title with the appropriate script code, using the `display_title` mechanism. ------------
-- Assumes that the scripts in "toBeTagged" will never occur in the Reconstruction namespace.
-- (FIXME: Don't make assumptions like this, and if you need to do so, throw an error if the assumption is violated.)
-- Avoid tagging ASCII as Hani even when it is tagged as Hani in the headword, as in [[check]]. The check for ASCII
-- might need to be expanded to a check for any Latin characters and whitespace or punctuation.
local display_title
-- Where there are multiple headwords, use the script for the first. This assumes the first headword is similar to
-- the pagename, and that headwords that are in different scripts from the pagename aren't first. This seems to be
-- about the best we can do (alternatively we could potentially do script detection on the pagename).
local dt_script = data.heads[1].sc
local dt_script_code = dt_script:getCode()
local page_non_ascii = namespace == "" and not page.pagename:find("^[%z\1-\127]+$")
local unsupported_pagename, unsupported = page.full_raw_pagename:gsub("^Tajuk tidak disokong/", "")
if unsupported == 1 and page.unsupported_titles[unsupported_pagename] then
display_title = 'Tajuk tidak disokong/<span class="' .. dt_script_code .. '">' .. page.unsupported_titles[unsupported_pagename] .. '</span>'
elseif page_non_ascii and m_headword_data.toBeTagged[dt_script_code]
or (dt_script_code == "Jpan" and (text_in_script(page.pagename, "Hira") or text_in_script(page.pagename, "Kana")))
or (dt_script_code == "Kore" and text_in_script(page.pagename, "Hang")) then
display_title = '<span class="' .. dt_script_code .. '">' .. page.full_raw_pagename .. '</span>'
-- Keep Han entries region-neutral in the display title.
elseif page_non_ascii and (dt_script_code == "Hant" or dt_script_code == "Hans") then
display_title = '<span class="Hani">' .. page.full_raw_pagename .. '</span>'
elseif namespace == "Rekonstruksi" then
local matched
display_title, matched = ugsub(
page.full_raw_pagename,
"^(Rekonstruksi:[^/]+/)(.+)$",
function(before, term)
return before .. tag_text(term, data.lang, dt_script)
end
)
if matched == 0 then
display_title = nil
end
end
-- FIXME: Generalize this.
-- If the current language uses Aran (Nastaliq), e.g. Urdu, and there's more than one language on the page, don't
-- set the display title because we don't want Nastaliq for terms that also exist in other languages that don't
-- display in Nastaliq (e.g. Arabic or Persian). Because the word "Urdu" occurs near the end of the alphabet, Urdu
-- fonts tend to override the fonts of other languages. FIXME: This is checking for more than one language on the
-- page but instead needs to check if there are any languages using scripts other than Aran.
if dt_script_code == "Aran" and page.L2_list.n > 1 then
display_title = nil
end
if display_title then
mw.getCurrentFrame():callParserFunction(
"DISPLAYTITLE",
display_title
)
end
------------ 9. Insert additional categories. ------------
if data.force_cat_output then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/force cat output]]
track("force cat output")
end
if has_redundant_head_param then
if not data.no_redundant_head_cat then
-- This is not the right way to go about this; too many exceptions and problems due to language-specific headword
-- handling customization. If we want this, it should be opt-in by a given language passing in the default headword.
-- insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan parameter pengepala lewah")
end
end
-- If the first head is multiword (after removing links), maybe insert into "LANG multiword terms".
if not data.nomultiwordcat and not is_varform_only and any_script_has_spaces and postype == "lemma" then
local no_multiword_cat = m_headword_data.no_multiword_cat
if not (no_multiword_cat[langcode] or no_multiword_cat[full_langcode]) then
-- Check for spaces or hyphens, but exclude prefixes and suffixes.
-- Use the pagename, not the head= value, because the latter may have extra
-- junk in it, e.g. superscripted text that throws off the algorithm.
local no_hyphen = m_headword_data.hyphen_not_multiword_sep
-- Exclude hyphens if the data module states that they should for this language.
local checkpattern = (no_hyphen[langcode] or no_hyphen[full_langcode]) and ".[%s፡]." or ".[%s%-፡]."
local is_multiword = umatch(page.pagename, checkpattern)
if is_multiword and not non_categorizable(page.full_raw_pagename) then
insert(data.categories, "Perkataan berbilang kata bahasa " .. full_langname)
elseif not is_multiword then
local long_word_threshold = m_headword_data.long_word_thresholds[langcode] or
m_headword_data.long_word_thresholds[full_langcode]
if long_word_threshold and ulen(page.pagename) >= long_word_threshold then
insert(data.categories, "Perkataan panjang bahasa " .. full_langname)
end
end
end
end
-- Determine whether to insert a category 'LANGNAME POS in SCRIPT'. If there are multiple heads, we may need to check
-- each head, as the heads may (theoretically) have different scripts.
local default_sccat = m_headword_data.default_sccat
if data.sccat or not is_varform_only and (default_sccat[langcode] or langcode ~= full_langcode and default_sccat[full_langcode]) then
local function needs_sccat(sccat_entry, sc)
if sccat_entry == true or not sccat_entry then
return sccat_entry
end
if type(sccat_entry) == "table" then
local in_list = contains(sccat_entry, sc:getCode())
if sccat_entry[1] == "not" then
in_list = not in_list
end
return in_list
end
return nil
end
for _, head in ipairs(data.heads) do
-- First check the `sccat` specified at the {{head}} level.
local this_needs_sccat = needs_sccat(data.sccat, head.sc)
-- If that wasn't given, check the default sccat at the language level for the lang code.
if this_needs_sccat == nil and not is_varform_only then
this_needs_sccat = needs_sccat(default_sccat[langcode], head.sc)
end
-- If that wasn't found and the lang code is an etym code, check the default sccat at the parent language level.
if this_needs_sccat == nil and not is_varform_only and langcode ~= full_langcode then
this_needs_sccat = needs_sccat(default_sccat[full_langcode], head.sc)
end
if this_needs_sccat then
insert(data.categories, ucfirst(data.pos_category) .. " bahasa " .. full_langname .. " dalam " ..
head.sc:getDisplayForm(data.lang))
end
end
end
-- Reconstructed terms often use weird combinations of scripts and realistically aren't spelled so much as notated.
if namespace ~= "Rekonstruksi" and not is_varform_only then
-- Map from languages to a string containing the characters to ignore when considering whether a term has
-- multiple written scripts in it. Typically these are Greek or Cyrillic letters used for their phonetic
-- values.
local characters_to_ignore = {
["aaq"] = "αάὰ", -- Penobscot (Algonquian)
["acy"] = "δθ", -- Cypriot Arabic
["aez"] = "β", -- Aeka (Trans-New Guinea)
["anc"] = "γ", -- Ngas (Chadic/Afroasiatic)
["aou"] = "χ", -- A'ou (Kra-Dai)
["art-blk"] = "ч", -- Bolak (conlang)
["awg"] = "β", -- Anguthimri (Pama-Nyungan)
["az"] = "ь", -- Azerbaijani (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["ba"] = "ь", -- Bashkir (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["bhp"] = "β", -- Bima (Austronesian)
["bjz"] = "β", -- Baruga (Trans-New Guinea)
["byk"] = "θ", -- Biao (Kra-Dai)
["cdy"] = "θ", -- Chadong (Kra-Dai)
["chp"] = "θ", -- Chipewyan (Athabaskan)
["cjh"] = "χ", -- Upper Chehalis (Salishan)
["clm"] = "χ", -- Klallam (Salishan)
["col"] = "χ", -- Colombia-Wenatchi (Salishan)
["coo"] = "χθ", -- Comox (Salishan)
["crx"] = "θ", -- Carrier (Athabaskan)
["ets"] = "θ", -- Yekhee (Edoid/Niger-Congo)
["ett"] = "χ", -- Etruscan (isolate; in romanizations)
["fla"] = "χ", -- Montana Salish (Salishan)
["grt"] = "་", -- Garo (South Asian Sino-Tibetan)
["gmw-gts"] = "χ", -- Gottscheerish (Bavarian variant spoken in Slovenia)
["hur"] = "χθ", -- Halkomelem (Salishan)
["itc-psa"] = "f", -- Pre-Samnite (Italic; normally written in Greek)
["izh"] = "ь", -- Ingrian (Finnic)
["kic"] = "θ", -- Kickapoo (Algonquian)
["kk"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["ky"] = "ь", -- Kyrgyz (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["lil"] = "χ", -- Lillooet (Salishan)
["lsi"] = "ꓹ", -- Lashi (Lolo-Burmese/Sino-Tibetan; represents a glottal stop)
["mhz"] = "β", -- Mor (Austronesian)
["mqn"] = "β", -- Moronene (Austronesian)
["neg"]= "ӡā", -- Negidal (Tungusic; normally in Cyrillic)
["oka"] = "χ", -- Okanagan (Salishan)
["ole"] = "θ", -- Olekha (Sino-Tibetan)
["oui"] = "γβ", -- Old Uyghur (Turkic; FIXME: others? E.g. Greek delta (δ)?)
["pox"] = "χ", -- Polabian (West Slavic)
["rif"] = "ε", -- Tarifit (Berber)
["rom"] = "Θθ", -- Romani (Indic: International Standard; two different thetas???)
["rpn"] = "β", -- Repanbitip (Austronesian)
["sah"] = "ь", -- Yakut (Turkic; 1929 - 1939 Latin spelling)
["sit-jap"] = "χ", -- Japhug (Sino-Tibetan)
["sjw"] = "θ", -- Shawnee (Algonquian)
["squ"] = "χ", -- Squamish (Salishan)
["str"] = "χθ", -- Saanich (Salishan)
["teh"] = "χ", -- Tehuelche (Chonan; spoken in Argentina)
["tep"] = "η", -- Tepecano (Uto-Aztecan)
["thp"] = "χ", -- Thompson (Salishan)
["tk"] = "ь", -- Turkmen (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["tt"] = "ь", -- Kazakh (Turkic; Yañalif Latin spelling, c. 1928 - 1938)
["twa"] = "χ", -- Twana (Salishan)
["wbl"] = "ы", -- Wakhi (Iranian)
["xbc"] = "ϸ", -- Bactrian (Iranian; represents š; normally written in Greek)
["yha"] = "θ", -- Baha (Kra-Dai)
["za"] = "зч", -- Zhuang (Tai/Kra-Dai); 1957-1982 alphabet used two Cyrillic letters (as well as some others like
-- ƃ, ƅ, ƨ, ɯ and ɵ that look like Cyrillic or Greek but are actually Latin)
["zlw-slv"] = "χђћ", -- Slovincian (West Slavic; FIXME: χ is Greek, the other two are Cyrillic, but I'm not sure
-- the currect characters are being chosen in the entry names)
["zng"] = "θ", -- Mang (Mon-Khmer)
["ztp"] = "θ", -- Loxicha Zapotec (Zapotecan)
}
-- Determine how many real scripts are found in the pagename, where we exclude symbols and such. We exclude
-- scripts whose `character_category` is false as well as Zmth (mathematical notation symbols), which has a
-- category of "Mathematical notation symbols". When counting scripts, we need to elide language-specific
-- variants because e.g. Beng and as-Beng have slightly different characters but we don't want to consider them
-- two different scripts (e.g. [[এৰ]] has two characters which are detected respectively as Beng and as-Beng).
local seen_scripts = {}
local num_seen_scripts = 0
local num_loops = 0
local canon_pagename = page.pagename
local ch_to_ignore = characters_to_ignore[full_langcode]
if ch_to_ignore then
canon_pagename = ugsub(canon_pagename, "[" .. ch_to_ignore .. "]", "")
end
while true do
if canon_pagename == "" or num_seen_scripts >= 2 or num_loops >= 10 then
break
end
-- Make sure we don't get into a loop checking the same script over and over again; happens with e.g. [[ᠪᡳ]]
num_loops = num_loops + 1
local pagename_script = find_best_script_without_lang(canon_pagename, "None only as last resort")
local script_chars = pagename_script.characters
if not script_chars then
-- we are stuck; this happens with None
break
end
local script_code = pagename_script:getCode()
local replaced
canon_pagename, replaced = ugsub(canon_pagename, "[" .. script_chars .. "]", "")
if (
replaced and
script_code ~= "Zmth" and
(script_data or get_script_data())[script_code] and
script_data[script_code].character_category ~= false
) then
script_code = script_code:gsub("^.-%-", "")
if not seen_scripts[script_code] then
seen_scripts[script_code] = true
num_seen_scripts = num_seen_scripts + 1
end
end
end
if num_seen_scripts > 1 then
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dieja dalam berbilang tulisan")
end
end
-- Categorise for unusual characters. Takes into account combining characters, so that we can categorise for characters with diacritics that aren't encoded as atomic characters (e.g. U̠). These can be in two formats: single combining characters (i.e. character + diacritic(s)) or double combining characters (i.e. character + diacritic(s) + character). Each can have any number of diacritics.
local standard = data.lang:getStandardCharacters()
if not is_varform_only and standard and not non_categorizable(page.full_raw_pagename) then
local function char_category(char)
local specials = {
["#"] = "number sign",
["("] = "parentheses",
[")"] = "parentheses",
["<"] = "angle brackets",
[">"] = "angle brackets",
["["] = "square brackets",
["]"] = "square brackets",
["_"] = "underscore",
["{"] = "braces",
["|"] = "vertical line",
["}"] = "braces",
["ß"] = "ẞ",
["\205\133"] = "", -- this is UTF-8 for U+0345 ( ͅ)
["\239\191\189"] = "replacement character",
}
char = toNFD(char)
:gsub(".[\128-\191]*", function(m)
local new_m = specials[m]
new_m = new_m or m:uupper()
return new_m
end)
return toNFC(char)
end
if full_langcode ~= "hi" and full_langcode ~= "lo" then
local standard_chars_scripts = {}
for _, head in ipairs(data.heads) do
standard_chars_scripts[head.sc:getCode()] = true
end
-- Iterate over the scripts, in case there is more than one (as they can have different sets of standard characters).
for code in pairs(standard_chars_scripts) do
local sc_standard = data.lang:getStandardCharacters(code)
if sc_standard then
if page.pagename_len > 1 then
local explode_standard = {}
local function explode(char)
explode_standard[char] = true
return ""
end
local sc_standard_exploded = ugsub(sc_standard, page.comb_chars.combined_double, explode)
-- The following is correct; it relies on side-effecing the explode_standard[] table.
ugsub(sc_standard_exploded, page.comb_chars.combined_single, explode):gsub(".[\128-\191]*", explode)
local num_cat_inserted
for char in pairs(page.explode_pagename) do
if not explode_standard[char] then
if char:find("[0-9]") then
if not num_cat_inserted then
insert(data.categories, "Perkataan dieja dengan nombor bahasa " .. full_langname)
num_cat_inserted = true
end
elseif ufind(char, page.emoji_pattern) then
insert(data.categories, "Perkataan dieja dengan emoji bahasa " .. full_langname)
else
local upper = char_category(char)
if not explode_standard[upper] then
char = upper
end
insert(data.categories, "Perkataan dieja dengan " .. char .. " bahasa " .. full_langname)
end
end
end
end
-- If a diacritic doesn't appear in any of the standard characters, also categorise for it generally.
sc_standard = toNFD(sc_standard)
for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_single) do
if not umatch(sc_standard, diacritic) then
insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname)
end
end
for diacritic in ugmatch(page.decompose_pagename, page.comb_chars.diacritics_double) do
if not umatch(sc_standard, diacritic) then
insert(data.categories, "Perkataan dieja dengan ◌" .. diacritic .. " bahasa " .. full_langname)
end
end
end
end
-- Ancient Greek, Hindi and Lao handled the old way for now, as their standard chars still need to be converted to the new format (because there are a lot of them).
elseif ulen(page.pagename) ~= 1 then
for character in ugmatch(page.pagename, "([^" .. standard .. "])") do
local upper = char_category(character)
if not umatch(upper, "[" .. standard .. "]") then
character = upper
end
insert(data.categories, "Perkataan dieja dengan " .. character .. " bahasa " .. full_langname)
end
end
end
if not is_varform_only and data.heads[1].sc:isSystem("alphabet") then
local pagename, i = page.pagename:ulower(), 2
while umatch(pagename, "(%a)" .. ("%1"):rep(i)) do
i = i + 1
insert(data.categories, "Perkataan bahasa " .. full_langname .. " dengan " .. i .. " contoh huruf yang sama berturut-turut")
end
end
-- Categorise for palindromes
if not is_varform_only and not data.nopalindromecat and namespace ~= "Rekonstruksi" and ulen(page.pagename) > 2
-- FIXME: Use of first script here seems hacky. What is the clean way of doing this in the presence of
-- multiple scripts?
and is_palindrome(page.pagename, data.lang, data.heads[1].sc) then
insert(data.categories, "Palindrom bahasa " .. full_langname)
end
if namespace == "" and not lang_reconstructed then
for _, head in ipairs(data.heads) do
if page.full_raw_pagename ~= get_link_page(remove_links(head.term), data.lang, head.sc) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch]]
-- [[Special:WhatLinksHere/Wiktionary:Tracking/headword/pagename spelling mismatch/LANGCODE]]
track("pagename spelling mismatch", data.lang)
break
end
end
end
-- Add red link category if called for and we're not a "large" page, where such checks are disabled.
if data.checkredlinks and not m_headword_data.large_pages[m_headword_data.pagename] then
local plposcat = type(data.checkredlinks) == "string" and data.checkredlinks or data.pos_category
check_red_link_inflections_top_level(data, plposcat)
end
-- Add to various maintenance categories.
export.maintenance_cats(page, data.lang, data.categories, data.whole_page_categories)
------------ 10. Format and return headwords, genders, inflections and categories. ------------
-- Format and return all the gathered information. This may add more categories (e.g. gender/number categories),
-- so make sure we do it before evaluating `data.categories`.
local text = '<span class="headword-line">' ..
format_headword(data) ..
format_headword_genders(data, is_varform_only) ..
format_top_level_inflections(data) .. '</span>'
-- Language-specific categories.
local cats = format_categories(
data.categories, data.lang, data.sort_key, page.encoded_pagename,
data.force_cat_output or test_force_categories, data.heads[1].sc
)
-- Language-agnostic categories.
local whole_page_cats = format_categories(
data.whole_page_categories, nil, "-"
)
return text .. cats .. whole_page_cats
end
return export
phbggqxeqxeecc32oqns9mjq2wituwm
Modul:links
828
9771
375349
281073
2026-09-22T03:12:32Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92723705|92723705]])
375349
Scribunto
text/plain
local export = {}
--[=[
[[Unsupported titles]], pages with high memory usage,
extraction modules and part-of-speech names are listed
at [[Module:links/data]].
Other modules used:
[[Module:script utilities]]
[[Module:scripts]]
[[Module:languages]] and its submodules
[[Module:gender and number]]
[[Module:debug/track]]
]=]
local anchors_module = "Module:anchors"
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local form_of_module = "Module:form of"
local gender_and_number_module = "Module:gender and number"
local languages_module = "Module:languages"
local load_module = "Module:load"
local memoize_module = "Module:memoize"
local pages_module = "Module:pages"
local scripts_module = "Module:scripts"
local script_utilities_module = "Module:script utilities"
local string_encode_entities_module = "Module:string/encode entities"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local find = string.find
local get_current_title = mw.title.getCurrentTitle
local insert = table.insert
local ipairs = ipairs
local match = string.match
local new_title = mw.title.new
local pairs = pairs
local remove = table.remove
local sub = string.sub
local toNFC = mw.ustring.toNFC
local tostring = tostring
local type = type
local unstrip = mw.text.unstrip
local NAMESPACE = get_current_title().nsText
local function anchor_encode(...)
anchor_encode = require(memoize_module)(mw.uri.anchorEncode, true)
return anchor_encode(...)
end
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function decode_entities(...)
decode_entities = require(string_utilities_module).decode_entities
return decode_entities(...)
end
local function decode_uri(...)
decode_uri = require(string_utilities_module).decode_uri
return decode_uri(...)
end
-- Can't yet replace, as the [[Module:string utilities]] version no longer has automatic double-encoding prevention, which requires changes here to account for.
local function encode_entities(...)
encode_entities = require(string_encode_entities_module)
return encode_entities(...)
end
local function extend(...)
extend = require(table_module).extend
return extend(...)
end
local function find_best_script_without_lang(...)
find_best_script_without_lang = require(scripts_module).findBestScriptWithoutLang
return find_best_script_without_lang(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function format_genders(...)
format_genders = require(gender_and_number_module).format_genders
return format_genders(...)
end
local function format_decorations(...)
format_decorations = require(decorations_module).format_decorations
return format_decorations(...)
end
local function get_current_L2(...)
get_current_L2 = require(pages_module).get_current_L2
return get_current_L2(...)
end
local function get_lang(...)
get_lang = require(languages_module).getByCode
return get_lang(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function language_anchor(...)
language_anchor = require(anchors_module).language_anchor
return language_anchor(...)
end
local function load_data(...)
load_data = require(load_module).load_data
return load_data(...)
end
local function request_script(...)
request_script = require(script_utilities_module).request_script
return request_script(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local function tag_text(...)
tag_text = require(script_utilities_module).tag_text
return tag_text(...)
end
local function tag_translit(...)
tag_translit = require(script_utilities_module).tag_translit
return tag_translit(...)
end
local function trim(...)
trim = require(string_utilities_module).trim
return trim(...)
end
local function u(...)
u = require(string_utilities_module).char
return u(...)
end
local function ulower(...)
ulower = require(string_utilities_module).lower
return ulower(...)
end
local function umatch(...)
umatch = require(string_utilities_module).match
return umatch(...)
end
local m_headword_data
local function get_headword_data()
m_headword_data = load_data("Module:headword/data")
return m_headword_data
end
local function track(page, code)
local tracking_page = "links/" .. page
debug_track(tracking_page)
if code then
debug_track(tracking_page .. "/" .. code)
end
end
local function field_non_empty(list, field)
if not list then
return nil
end
if type(list) ~= "table" then
error(("Internal error: Wrong type for `termobj.%s`=%s, should be %s\"table\""):format(
field, mw.dumpObject(list), field == "q" or field == "qq" and "\"string\" or " or ""))
end
return list[1]
end
--[=[
Add any decorations (left or right regular or accent qualifiers, labels or references) to an item. `text` is the
item's text (to which to add the decorations) and `itemobj` is the object specifying the item's decorations, which
should optionally contain:
* left regular qualifiers in `q` (an array of strings or a single string); an empty array will be ignored;
* right regular qualifiers in `qq` (an array of strings or a single string); an empty array will be ignored;
* left accent qualifiers in `a` (an array of strings); an empty array will be ignored;
* right accent qualifiers in `aa` (an array of strings); an empty array will be ignored;
* left labels in `l` (an array of strings); an empty array will be ignored;
* right labels in `ll` (an array of strings); an empty array will be ignored;
* references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text`
(formatted reference text) and optionally `name` and/or `group`; an empty array will be ignored.
`lang` is a language object and is required if any accent qualifiers or labels are given.
]=]
local function add_text_decorations(text, itemobj, lang)
local q = itemobj.q
if type(q) == "string" then
q = { q }
end
local qq = itemobj.qq
if type(qq) == "string" then
qq = { qq }
end
if field_non_empty(q, "q") or field_non_empty(qq, "qq") or field_non_empty(itemobj.a, "a") or
field_non_empty(itemobj.aa, "aa") or field_non_empty(itemobj.l, "l") or field_non_empty(itemobj.ll, "ll") or
field_non_empty(itemobj.refs, "refs") then
text = format_decorations {
lang = lang,
text = text,
q = q,
qq = qq,
a = itemobj.a,
aa = itemobj.aa,
l = itemobj.l,
ll = itemobj.ll,
refs = itemobj.refs,
}
end
return text
end
local function selective_trim(...)
-- Unconditionally trimmed charset.
local always_trim =
"\194\128-\194\159" .. -- U+0080-009F (C1 control characters)
"\194\173" .. -- U+00AD (soft hyphen)
"\226\128\170-\226\128\174" .. -- U+202A-202E (directionality formatting characters)
"\226\129\166-\226\129\169" -- U+2066-2069 (directionality formatting characters)
-- Standard trimmed charset.
local standard_trim = "%s" .. -- (default whitespace charset)
"\226\128\139-\226\128\141" .. -- U+200B-200D (zero-width spaces)
always_trim
-- If there are non-whitespace characters, trim all characters in `standard_trim`.
-- Otherwise, only trim the characters in `always_trim`.
selective_trim = function(text)
if text == "" then
return text
end
local trimmed = trim(text, standard_trim)
if trimmed ~= "" then
return trimmed
end
return trim(text, always_trim)
end
return selective_trim(...)
end
local function escape(text, str)
local rep
repeat
text, rep = text:gsub("\\\\(\\*" .. str .. ")", "\5%1")
until rep == 0
return (text:gsub("\\" .. str, "\6"))
end
local function unescape(text, str)
return (text
:gsub("\5", "\\")
:gsub("\6", str))
end
-- Remove bold, italics, soft hyphens, strip markers and HTML tags.
local function remove_formatting(str)
str = str
:gsub("('*)'''(.-'*)'''", "%1%2")
:gsub("('*)''(.-'*)''", "%1%2")
:gsub("", "")
return (unstrip(str)
:gsub("<[^<>]+>", ""))
end
--[==[
Split `text` on double slashes (taking account of escaping backslashes). Return a list of split items.
]==]
function export.split_on_slashes(text)
if text:find("\\", nil, true) then
track("escaped", "split_on_slashes")
end
text = split(escape(text, "//"), "//", true) or {}
for i, v in ipairs(text) do
text[i] = unescape(v, "//")
if v == "" then
text[i] = false
end
end
return text
end
--[==[
If `text` is a wikilink (i.e. in the form `<nowiki>[[foo]]</nowiki>` or `<nowiki>[[foo|bar]]</nowiki>`), return the link
target and display text. If the wikilink is single-part, the display text will be the same as the link target. If the
text is not in wikilink form, return {nil} for both link target and display text.
]==]
function export.get_wikilink_parts(text)
-- TODO: replace `allow_bad_target` with `allow_unsupported`, with support for links to unsupported titles, including escape sequences.
if ( -- Filters out anything but "[[...]]" with no intermediate "[[" or "]]".
not match(text, "^()%[%[") or -- Faster than sub(text, 1, 2) ~= "[[".
find(text, "[[", 3, true) or
find(text, "]]", 3, true) ~= #text - 1
) then
return nil, nil
end
local pipe = find(text, "|", 3, true)
local title, display
if pipe then
title, display = sub(text, 3, pipe - 1), sub(text, pipe + 1, -3)
else
title = sub(text, 3, -3)
display = title
end
return title, display
end
-- Does the work of export.get_fragment, but can be called directly to avoid unnecessary checks for embedded links.
local function get_fragment(text)
text = escape(text, "#")
-- Replace numeric character references with the corresponding character (' → '),
-- as they contain #, which causes the numeric character reference to be
-- misparsed (wa'a → wa'a → pagename wa&, fragment 39;a).
text = decode_entities(text)
local target, fragment = text:match("^(.-)#(.+)$")
target = target or text
target = unescape(target, "#")
fragment = fragment and unescape(fragment, "#") or nil
return target, fragment
end
--[==[
Given a link target possibly containing a fragment (i.e. a pound sign and following anchor), return two values, the
actual target (minus the fragment) and the fragment, or {nil} if there is no fragment.
]==]
function export.get_fragment(text)
if text:find("\\", nil, true) then
track("escaped", "get_fragment")
end
-- If there are no embedded links, process input.
local open = find(text, "[[", nil, true)
if not open then
return get_fragment(text)
end
local close = find(text, "]]", open + 2, true)
if not close then
return get_fragment(text)
-- If there is one, but it's redundant (i.e. encloses everything with no pipe), remove and process.
elseif open == 1 and close == #text - 1 and not find(text, "|", 3, true) then
return get_fragment(sub(text, 3, -3))
end
-- Otherwise, return the input.
return text, nil
end
--[==[
Given a link target as passed to `full_link()`, get the actual page that the target refers to. This removes
bold, italics, strip markets and HTML; calls `makeEntryName()` for the language in question; converts targets
beginning with `*` to the Reconstruction namespace; and converts appendix-constructed languages to the Appendix
namespace. Returns up to three values:
# the actual page to link to, or {nil} to not link to anything;
# how the target should be displayed as, if the user didn't explicitly specify any display text; generally the
same as the original target, but minus any anti-asterisk !!;
# the value `true` if the target had a backslash-escaped * in it (FIXME: explain this more clearly).
FIXME: This should not be calling deprecated `makeEntryName()` and should always return the same number of values.
]==]
function export.get_link_page_with_auto_display(target, lang, sc, plain)
local orig_target = target
if not target then
return nil
elseif target:find("\\", nil, true) then
track("escaped", "get_link_page")
end
target = remove_formatting(target)
if target:sub(1, 1) == ":" then
track("initial colon")
-- FIXME, the auto_display (second return value) should probably remove the colon
return target:sub(2), orig_target
end
local prefix = target:match("^(.-):")
-- Convert any escaped colons
target = target:gsub("\\:", ":")
if prefix then
-- If this is an a link to another namespace or an interwiki link, ensure there's an initial colon and then
-- return what we have (so that it works as a conventional link, and doesn't do anything weird like add the term
-- to a category.)
prefix = ulower(trim(prefix))
if prefix ~= "" and (
load_data("Module:data/namespaces")[prefix] or
load_data("Module:data/interwikis")[prefix]
) then
return target, orig_target
end
end
-- Check if the term is reconstructed and remove any asterisk. Also check for anti-asterisk (!!).
-- Otherwise, handle the escapes.
local reconstructed, escaped, anti_asterisk
if not plain then
target, reconstructed = target:gsub("^%*(.)", "%1")
if reconstructed == 0 then
target, anti_asterisk = target:gsub("^!!(.)", "%1")
if anti_asterisk == 1 then
-- Remove !! from original. FIXME! We do it this way because the call to remove_formatting() above
-- may cause non-initial !! to be interpreted as anti-asterisks. We should surely move the
-- remove_formatting() call later.
orig_target = orig_target:gsub("^!!", "")
end
end
end
target, escaped = target:gsub("^(\\-)\\%*", "%1*")
if not (sc and sc:getCode() ~= "None") then
sc = lang:findBestScript(target)
end
-- Remove carets if they are used to capitalize parts of transliterations (unless they have been escaped).
if (not sc:hasCapitalization()) and sc:isTransliterated() and target:match("%^") then
target = escape(target, "^")
:gsub("%^", "")
target = unescape(target, "^")
end
-- Get the entry name for the language.
target = lang:makeEntryName(target, sc, reconstructed == 1 or lang:hasType("appendix-constructed"))
-- If the link contains unexpanded template parameters, then don't create a link.
if target:match("{{{.-}}}") then
-- FIXME: Should we return the original target as the default display value (second return value)?
return nil
end
-- Link to appendix for reconstructed terms and terms in appendix-only languages. Plain links interpret *
-- literally, however.
if reconstructed == 1 then
if lang:getFullCode() == "und" then
-- Return the original target as default display value. If we don't do this, we wrongly get
-- [Term?] displayed instead.
return nil, orig_target
end
target = "Rekonstruksi:Bahasa " .. lang:getFullName() .. "/" .. target
-- Reconstructed languages and substrates require an initial *.
elseif anti_asterisk ~= 1 and (lang:hasType("reconstructed") or lang:getFamilyCode() == "qfa-sub") then
error(("The specified language %s is unattested, while the term '%s' does not begin with '*' to indicate that it is reconstructed.")
:
format(lang:getCanonicalName(), orig_target))
elseif lang:hasType("appendix-constructed") then
target = "Lampiran:Bahasa " .. lang:getFullName() .. "/" .. target
else
target = target
end
return target, orig_target, escaped > 0
end
function export.get_link_page(target, lang, sc, plain)
local target, auto_display, escaped = export.get_link_page_with_auto_display(target, lang, sc, plain)
return target, escaped
end
-- Make a link from a given link's parts.
local function make_link(link, lang, sc, id, isolated, cats, no_alt_ast, plain)
-- Convert percent encoding to plaintext.
link.target = link.target and decode_uri(link.target, "PATH")
link.fragment = link.fragment and decode_uri(link.fragment, "PATH")
-- Find fragments (if one isn't already set).
-- Prevents {{l|en|word#Etymology 2|word}} from linking to [[word#Etymology 2#English]].
-- # can be escaped as \#.
if link.target and link.fragment == nil then
link.target, link.fragment = get_fragment(link.target)
end
-- Process the target
local auto_display, escaped
link.target, auto_display, escaped = export.get_link_page_with_auto_display(link.target, lang, sc, plain)
-- Create a default display form.
-- If the target is "" then it's a link like [[#English]], which refers to the current page.
if auto_display == "" then
auto_display = (m_headword_data or get_headword_data()).pagename
end
-- If the display is the target and the reconstruction * has been escaped, remove the escaping backslash.
if escaped then
auto_display = auto_display:gsub("\\([^\\]*%*)", "%1", 1)
end
-- Process the display form.
if link.display then
local orig_display = link.display
link.display = lang:makeDisplayText(link.display, sc, true)
if cats then
auto_display = lang:makeDisplayText(auto_display, sc)
-- If the alt text is the same as what would have been automatically generated, then the alt parameter is
-- redundant (e.g. {{l|en|foo|foo}}, {{l|en|w:foo|foo}}, but not {{l|en|w:foo|w:foo}}). If they're
-- different, but the alt text could have been entered as the term parameter without it affecting the target
-- page, then the target parameter is redundant (e.g. {{l|ru|фу|фу́}}). If `no_alt_ast` is true, use pcall
-- to catch the error which will be thrown if this is a reconstructed lang and the alt text doesn't have *.
if link.display == auto_display then
insert(cats, "Pautan dengan parameter alt lewah bahasa " .. lang:getFullName())
else
local ok, check
if no_alt_ast then
ok, check = pcall(export.get_link_page, orig_display, lang, sc, plain)
else
ok = true
check = export.get_link_page(orig_display, lang, sc, plain)
end
if ok and link.target == check then
insert(cats, "Pautan dengan parameter sasaran lewah bahasa " .. lang:getFullName())
end
end
end
else
link.display = lang:makeDisplayText(auto_display, sc)
end
if not link.target then
return link.display
end
-- If the target is the same as the current page, there is no sense id
-- and either the language code is "und" or the current L2 is the current
-- language then return a "self-link" like the software does.
if link.target == get_current_title().prefixedText then
local fragment, current_L2 = link.fragment, get_current_L2()
if (
fragment and fragment == current_L2 or
not (id or fragment) and (lang:getFullCode() == "und" or lang:getFullName() == current_L2)
) then
return tostring(mw.html.create("strong")
:addClass("selflink")
:wikitext(link.display))
end
end
-- Add fragment. Do not add a section link to "Undetermined", as such sections do not exist and are invalid.
-- TabbedLanguages handles links without a section by linking to the "last visited" section, but adding
-- "Undetermined" would break that feature. For localized prefixes that make syntax error, please use the
-- format: ["xyz"] = true.
local prefix = link.target:match("^:*([^:]+):")
prefix = prefix and ulower(prefix)
if prefix ~= "category" and not (prefix and load_data("Module:data/interwikis")[prefix]) then
if (link.fragment or link.target:sub(-1) == "#") and not plain then
track("fragment", lang:getFullCode())
if cats then
insert(cats, "Pautan dengan serpihan manual bahasa " .. lang:getFullName())
end
end
if not link.fragment then
if id then
link.fragment = lang:getFullCode() == "und" and anchor_encode(id) or language_anchor(lang, id)
elseif lang:getFullCode() ~= "und" and not (link.target:match("^Lampiran:") or link.target:match("^Rekonstruksi:")) then
link.fragment = anchor_encode("Bahasa " .. lang:getFullName())
end
end
end
-- Put inward-facing square brackets around a link to isolated spacing character(s).
if isolated and link.display[1] and not umatch(decode_entities(link.display), "%S") then
link.display = "]" .. link.display .. "["
end
link.target = link.target:gsub("^(:?)(.*)", function(m1, m2)
return m1 .. encode_entities(m2, "#%&+/:<=>@[\\]_{|}")
end)
link.fragment = link.fragment and encode_entities(remove_formatting(link.fragment), "#%&+/:<=>@[\\]_{|}")
return "[[" ..
link.target:gsub("^[^:]", ":%0") .. (link.fragment and "#" .. link.fragment or "") .. "|" .. link.display .. "]]"
end
-- Split a link into its parts.
local function parse_link(linktext)
local link = { target = linktext }
local target = link.target
link.target, link.display = target:match("^(..-)|(.+)$")
if not link.target then
link.target = target
link.display = target
end
-- There's no point in processing these, as they aren't real links.
local target_lower = link.target:lower()
for _, false_positive in ipairs({ "category", "cat", "file", "image" }) do
if target_lower:match("^" .. false_positive .. ":") then
return nil
end
end
link.display = decode_entities(link.display)
link.target, link.fragment = get_fragment(link.target)
-- So that make_link does not look for a fragment again.
if not link.fragment then
link.fragment = false
end
return link
end
local function check_params_ignored_when_embedded(alt, lang, id, cats)
if alt then
track("alt-ignored")
if cats then
insert(cats, "Pautan dengan parameter alt dihirau bahasa " .. lang:getFullName())
end
end
if id then
track("id-ignored")
if cats then
insert(cats, "Pautan dengan parameter id dihirau bahasa " .. lang:getFullName())
end
end
end
-- Find embedded links and ensure they link to the correct section.
local function process_embedded_links(text, alt, lang, sc, id, cats, no_alt_ast, plain)
-- Process the non-linked text.
text = lang:makeDisplayText(text, sc, true)
-- If the text begins with * and another character, then act as if each link begins with *. However, don't do this if the * is contained within a link at the start. E.g. `|*[[foo]]` would set all_reconstructed to true, while `|[[*foo]]` would not.
local all_reconstructed = false
if not plain then
-- anchor_encode removes links etc.
if anchor_encode(text):sub(1, 1) == "*" then
all_reconstructed = true
end
-- Otherwise, handle any escapes.
text = text:gsub("^(\\-)\\%*", "%1*")
end
check_params_ignored_when_embedded(alt, lang, id, cats)
local function process_link(space1, linktext, space2)
local capture = "[[" .. linktext .. "]]"
local link = parse_link(linktext)
-- Return unprocessed false positives untouched (e.g. categories).
if not link then
return capture
end
if all_reconstructed then
if link.target:find("^!!") then
-- Check for anti-asterisk !! at the beginning of a target, indicating that a reconstructed term
-- wants a part of the term to link to a non-reconstructed term, e.g. Old English
-- {{ang-noun|m|head=*[[!!Crist|Cristes]] [[!!mæsseǣfen]]}}.
link.target = link.target:sub(3)
-- Also remove !! from the display, which may have been copied from the target (as in mæsseǣfen in
-- the example above).
link.display = link.display:gsub("^!!", "")
elseif not link.target:match("^%*") then
link.target = "*" .. link.target
end
end
linktext = make_link(link, lang, sc, id, false, nil, no_alt_ast, plain)
:gsub("^%[%[", "\3")
:gsub("%]%]$", "\4")
return space1 .. linktext .. space2
end
-- Use chars 1 and 2 as temporary substitutions, so that we can use charsets. These are converted to chars 3 and 4 by process_link, which means we can convert any remaining chars 1 and 2 back to square brackets (i.e. those not part of a link).
text = text
:gsub("%[%[", "\1")
:gsub("%]%]", "\2")
-- If the script uses ^ to capitalize transliterations, make sure that any carets preceding links are on the inside, so that they get processed with the following text.
if (
text:find("^", nil, true) and
not sc:hasCapitalization() and
sc:isTransliterated()
) then
text = escape(text, "^")
:gsub("%^\1", "\1%^")
text = unescape(text, "^")
end
text = text:gsub("\1(%s*)([^\1\2]-)(%s*)\2", process_link)
-- Remove the extra * at the beginning of a language link if it's immediately followed by a link whose display begins with * too.
if all_reconstructed then
text = text:gsub("^%*\3([^|\1-\4]+)|%*", "\3%1|*")
end
return (text
:gsub("[\1\3]", "[[")
:gsub("[\2\4]", "]]")
)
end
local function simple_link(term, fragment, alt, lang, sc, id, cats, no_alt_ast, suppress_redundant_wikilink_cat)
local plain
if lang == nil then
lang, plain = get_lang("und"), true
end
-- Get the link target and display text. If the term is the empty string, treat the input as a link to the current page.
if term == "" then
term = get_current_title().prefixedText
elseif term then
local new_term, new_alt = export.get_wikilink_parts(term)
if new_term then
check_params_ignored_when_embedded(alt, lang, id, cats)
-- [[|foo]] links are treated as plaintext "[[|foo]]".
-- FIXME: Pipes should be handled via a proper escape sequence, as they can occur in unsupported titles.
if new_term == "" then
term, alt = nil, term
else
local title = new_title(new_term)
if title then
local ns = title.namespace
-- File: and Category: links should be returned as-is.
if ns == 6 or ns == 14 then
return term
end
end
term, alt = new_term, new_alt
if cats then
if not (suppress_redundant_wikilink_cat and suppress_redundant_wikilink_cat(term, alt)) then
insert(cats, "Pautan bahasa " .. lang:getFullName() .. " dengan pautan wiki lewah")
end
end
end
end
end
if alt then
alt = selective_trim(alt)
if alt == "" then
alt = nil
end
end
-- If there's nothing to process, return nil.
if not (term or alt) then
return nil
end
-- If there is no script, get one.
if not sc then
sc = lang:findBestScript(alt or term)
end
-- Embedded wikilinks need to be processed individually.
if term then
local open = find(term, "[[", nil, true)
if open and find(term, "]]", open + 2, true) then
return process_embedded_links(term, alt, lang, sc, id, cats, no_alt_ast, plain)
end
term = selective_trim(term)
end
-- If not, make a link using the parameters.
return make_link({
target = term,
display = alt,
fragment = fragment
}, lang, sc, id, true, cats, no_alt_ast, plain)
end
--[==[
Create a basic link to the given term. It links to the language section (such as `==English==`), but it does not add
language and script wrappers, so any code that uses this function should call `[[Module:script utilities#tag_text]]`
to add such wrappers itself at some point. The first argument, `data`, may contain the following items, a subset of the
items used in the `data` argument of `##full_link`. If any other items are included, they are ignored.
{ {
term = "entry_to_link_to",
alt = "link_text_or_displayed_text",
lang = language_object,
sc = script_object,
fragment = "link_fragment",
id = "sense_id",
no_alt_ast = boolean,
suppress_redundant_wikilink_cat = function(term, alt) -> boolean,
cats = nil or {}, -- NOTE: If given, will be side-effected to return categories to add the page to.
} }
Specifically:
* `term`: Term to turn into a link. This is generally the name of a page, possibly with extra diacritics added (e.g.
length marks in Latin or Old English terms, accents in Russian terms, vowel diacritics in Arabic terms, etc.), which
are stripped to determine the actual pagename. The term can contain wikilinks already embedded in it. These are
processed individually just like a single link would be. The `alt` argument is ignored in this case.
* `alt`: The alternative display for the link, if different from the linked page. If this is {nil}, the `term` argument
is used instead (much like regular wikilinks). If `term` contains wikilinks in it, this argument is ignored and has no
effect. (Links in which the alt is ignored are tracked with the tracking template
{{whatlinkshere|tracking=links/alt-ignored}}.)
* `lang` ('''required'''): The [[Module:languages#Language objects|language object]] for the term being linked. The link
or links in `term` will normally have their fragment set to point to the canonical name (see
{{tl|language data documentation}}) of `lang` (or, if it is an etymology-only language, to the canonical name of its
L2 parent). (However, if `id` is defined, the fragment will point to a language-specific sense ID corresponding to
this field, which in turn will be overridden by `fragment` if specified.)
* `sc`: The [[Module:scripts#Script objects|script object]] for the term being linked. This rarely needs to be specified
because it is autodetected based on `term` or `alt`, and the detection is usually correct. It is used to determine
how to convert the term into a pagename, possibly by stripping certain diacritics from the term's text.
* `fragment`: If specified, overrides the fragment in the generated link (i.e. the portion after `#`, which determines
where on the page to go to when the link is clicked). If not specified, the fragment is generated from `id` (if given)
or otherwise from `lang`.
* `id`: Sense ID string. If this argument is defined, the link will point to a language-specific sense ID
({{ll|en|identifier|id=HTML}}) created by the template {{temp|senseid}}. The fragment for a sense ID consists of the
language's canonical name, a hyphen (`-`), and the string that was supplied as the `id` argument. This is useful when
a term has more than one sense in a language. If the `term` argument contains wikilinks, this argument is ignored.
(Links in which the sense ID is ignored are tracked with the tracking template
{{whatlinkshere|tracking=links/id-ignored}}.)
* `no_alt_ast`: This is the same as `no_alt_ast` in `##full_link()`. See that function for more information.
* `suppress_redundant_wikilink_cat`: This is the same as `suppress_redundant_wikilink_cat` in `##full_link()`. See
that function for more information.
* `cats`: This should be either {nil} or an empty list. In the latter case, tracking categories will be added to the
list when appropriate. The caller can choose to add the page to those categories (as is done by ##full_link()`).
The following special options are processed for each link (both simple terms and with embedded wikilinks):
* The target page name will be processed by stripping certain diacritics (as mentioned above) and converting the
resulting ''logical'' pagename to a ''physical'' pagename (which will be different from the logical pagename in the
case of pages with unsupported characters in them and certain overly large pages, such as [[a]], that are split into
parts).
* If the term starts with `*`, then it is considered a reconstructed term, and a link to the `Reconstruction:` namespace
will be created. If the text contains embedded wikilinks, then `*` is automatically applied to each one individually,
while preserving the displayed form of each link as it was given. This allows linking to phrases containing multiple
reconstructed terms, while only showing the `*` once at the beginning.
* If the text starts with `:`, then the link is treated as "raw" and the above steps are skipped. This can be used in
rare cases where the page name begins with `*` or if diacritics should not be stripped. For example:
** {{tl|l|en|*nix}} links to the nonexistent page [[Reconstruction:English/nix]] (`*` is interpreted as a
reconstruction), but {{tl|l|en|:*nix}} links to [[*nix]].
** {{tl|l|sl|Franche-Comté}} links to the nonexistent page [[Franche-Comte]] (`é` is converted to `e` by the
diacritic-stripping process), but {{tl|l|sl|:Franche-Comté}} links to [[Franche-Comté]].
]==]
function export.language_link(data)
if type(data) ~= "table" then
error(
"The first argument to the function language_link must be a table. See [[Module:links/documentation]] for more information.")
elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then
track("escaped", "language_link")
end
-- Categorize links to "und".
local lang, cats = data.lang, data.cats
if cats and lang:getCode() == "und" then
insert(cats, "Pautan bahasa tidak ditentukan")
end
return simple_link(
data.term,
data.fragment,
data.alt,
lang,
data.sc,
data.id,
cats,
data.no_alt_ast,
data.suppress_redundant_wikilink_cat
)
end
function export.plain_link(data)
if type(data) ~= "table" then
error(
"The first argument to the function plain_link must be a table. See Module:links/documentation for more information.")
elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then
track("escaped", "plain_link")
end
return simple_link(
data.term,
data.fragment,
data.alt,
nil,
data.sc,
data.id,
data.cats,
data.no_alt_ast,
data.suppress_redundant_wikilink_cat
)
end
--[==[Replace any links with links to the correct section, but don't link the whole text if no embedded links are found. Returns the display text form.]==]
function export.embedded_language_links(data)
if type(data) ~= "table" then
error(
"The first argument to the function embedded_language_links must be a table. See Module:links/documentation for more information.")
elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then
track("escaped", "embedded_language_links")
end
local term, lang, sc = data.term, data.lang, data.sc
-- If we don't have a script, get one.
if not sc then
sc = lang:findBestScript(term)
end
-- Do we have embedded wikilinks? If so, they need to be processed individually.
local open = find(term, "[[", nil, true)
if open and find(term, "]]", open + 2, true) then
return process_embedded_links(term, data.alt, lang, sc, data.id, data.cats, data.no_alt_ast)
end
-- If not, return the display text.
term = selective_trim(term)
-- FIXME: Double-escape any percent-signs, because we don't want to treat non-linked text as having percent-encoded
-- characters. This is a hack: percent-decoding should come out of [[Module:languages]] and only dealt with in this
-- module, as it's specific to links.
term = term:gsub("%%", "%%25")
return lang:makeDisplayText(term, sc, true)
end
function export.mark(text, item_type, face, lang)
local tag = { "", "" }
if item_type == "gloss" then
tag = { '<span class="mention-gloss-double-quote">“</span><span class="mention-gloss">',
'</span><span class="mention-gloss-double-quote">”</span>' }
if type(text) == "string" and text:match("^''[^'].*''$") then
-- Temporary tracking for mention glosses that are entirely italicized or bolded, which is probably
-- wrong. (Note that this will also find bolded mention glosses since they use triple apostrophes.)
track("italicized-mention-gloss", lang and lang:getFullCode() or nil)
end
elseif item_type == "tr" then
if face == "term" then
tag = { '<span lang="' .. lang:getFullCode() .. '" class="tr mention-tr Latn">',
'</span>' }
else
tag = { '<span lang="' .. lang:getFullCode() .. '" class="tr Latn">', '</span>' }
end
elseif item_type == "ts" then
-- \226\129\160 = word joiner (zero-width non-breaking space) U+2060
tag = { '<span class="ts mention-ts Latn">/\226\129\160', '\226\129\160/</span>' }
elseif item_type == "pos" then
tag = { '<span class="ann-pos">', '</span>' }
elseif item_type == "non-gloss" then
tag = { '<span class="ann-non-gloss">', '</span>' }
elseif item_type == "annotations" then
tag = { '<span class="mention-gloss-paren annotation-paren">(</span>',
'<span class="mention-gloss-paren annotation-paren">)</span>' }
elseif item_type == "infl" then
tag = { '<span class="ann-infl">', '</span>' }
end
if type(text) == "string" then
return tag[1] .. text .. tag[2]
else
return ""
end
end
--[=[
Implementation of `format_transliteration` and `format_transcription`. The implementation is identical except that the
field containing the transliteration or transcription may vary and is specified in `field`, and the way a given
transliteration or transcription is tagged may vary and is controlled by `tag_fn`, which is passed three parameters:
`text` (the transliteration or transcription), `lang` (the language passed in) and `face` (the face passed in).
On input, `item` is the transliteration or transcription or list of such objects; `field` is either {"tr"} or {"ts"};
`tag_fn` is a function of three parameters to tag the item, as described above; `lang` is the language object of the
term whose transliteration or transcription is specified; and `face` is a string indicating how to display the item
(generally only the strings {"term"} and {"default"} are recognized).
]=]
local function format_transliteration_or_transcription(item, field, tag_fn, lang, item_face)
if type(item) == "string" then
return tag_fn(item, lang, item_face)
end
local formatted_items = {}
for _, itemobj in ipairs(item) do
local tagged_item = tag_fn(itemobj[field], lang, item_face)
tagged_item = add_text_decorations(tagged_item, item, lang)
insert(formatted_items, tagged_item)
end
if formatted_items[2] then
-- FIXME: This should be customizable.
return concat(formatted_items, " <i>or</i> ")
else
return formatted_items[1]
end
end
--[==[
Format a transliteration string or list of transliteration objects. `tr` contains the transliteration(s), which for
forward compatibility reasons can only be either a single transliteration string or a list of transliteration objects
(each of which has a `tr` field holding the transliteration and optional fields `q`, `qq`, `l`, `ll` and/or `refs`).
This correctly handles multiple transliterations as well as decorations (qualifiers, labels or references) attached to
transliterations.
]==]
function export.format_transliteration(tr, lang, face)
return format_transliteration_or_transcription(tr, "tr", tag_translit, lang, face)
end
local function tag_transcription(ts, _lang, _face)
return export.mark(ts, "ts")
end
--[==[
Format a transcription string or list of transcription objects. `ts` contains the transcription(s), which for forward
compatibility reasons can only be either a single transcription string or a list of transcription objects (each of which
has a `ts` field holding the transcription and optional decoration fields `q`, `qq`, `l`, `ll` and/or `refs`). This
correctly handles multiple transcriptions as well as decorations (qualifiers, labels or references) attached to
transcriptions.
]==]
function export.format_transcription(ts, lang, face)
return format_transliteration_or_transcription(ts, "ts", tag_transcription, lang, face)
end
local pos_tags
--[==[
Format the annotations that are displayed with a link created by `full_link()`. Annotations are the extra bits of
information that are displayed following the linked term, and include things such as gender, transliteration, gloss,
etc. The first argument is a table with some or all of the following keys (all are optional):
* `interwiki`: An interwiki link. This is used for links in translation tables to the corresponding term in another
Wiktionary. If specified, it should be a fully formatted link and is inserted as-is at the beginning of the output.
See the `interwiki()` function in [[Module:translations]].
* `genders`: Table containing a list of gender specifications in the style of [[Module:gender and number]]. If
specified, these are formatted using `format_genders()` in [[Module:gender and number]] and the result inserted at
the beginning of the output, following any interwiki link and (in all cases) directly after a no-break space.
* `tr`: Transliteration or transliterations. Currently, this is always a one-item list. It is a list because of
potential support for per-alternant transliterations due to the multiple alternants (separated by `//`) that can be
specified in `term` or `alt`. The item in the list can be either a string or a list of transliteration objects (see
`format_transliteration()`).
* `ts`: Transcription or transcriptions. Like `tr`, this is currently always a one-item list, with the item being either
a single string or a list of transcription objects, as described in `format_transcription()`.
* `gloss`: Gloss that translates the term in the link.
* `pos`: Part of speech of the linked term. If the given argument matches one of the aliases in `pos_aliases` in
[[Module:headword/data]], or consists of a part of speech or alias followed by `f` (for a non-lemma form), expand it
appropriately. Otherwise, just show the given text as it is.
* `infl`: A list of tags, each a string. Multiple tag sets may be encoded in this list by separating them with an
element consisting of a semicolon. If there are multiple tag sets, they are formatted individually and separated
by a semicolon + space.
* `ng`: Arbitrary non-gloss descriptive text for the link. This should be used in preference to putting descriptive text
in `gloss` or `pos`.
* `lit`: Literal meaning of the term, if the usual meaning is figurative or idiomatic.
* `postprocess_annotations`: A function to postprocess the annotations, after they have been formatted (see below).
The `interwiki` and `genders` properties are formatted specially, and `postprocess_annotations` is a callback function
rather than an item to display; all others are formatted (each in their own way), separated by commas and placed inside
of parentheses (except that if both transliteration and transcription are present, they are separated by a space). The
order of the annotations is `tr`+`ts`, `gloss`, `pos`, `infl`, `ng` and `lit`. The `postprocess_annotations` function,
if supplied, is passed a single object, a table with two keys `data` (the `data` object passed into
`format_link_annotations()`) and `annotations` (the formatted annotations, prior to being concatenated). It should
side-effect the `annotations` list, e.g. by inserting more annotations. (It is used to handle nested inflections in
[[Module:headword]]. FIXME: It should probably be generalized so that it can return the final formatted string, to allow
for e.g. changing the way the annotations are formatted.)
* The second argument is a string controlling the "face" that the terms are displayed in. Currently it only affects
transliteration and transcription and only when the value {"term"} is passed in, in which case those annotations are
displayed italicized.
]==]
function export.format_link_annotations(data, face)
local output = {}
-- Interwiki link
if data.interwiki then
insert(output, data.interwiki)
end
-- Genders
if type(data.genders) ~= "table" then
data.genders = { data.genders }
end
if data.genders and data.genders[1] then
local genders, gender_cats = format_genders(data.genders, data.lang)
insert(output, " " .. genders)
if gender_cats then
local cats = data.cats
if cats then
extend(cats, gender_cats)
end
end
end
local annotations = {}
-- Transliteration and transcription
local tr = data.tr and data.tr[1] or nil
local ts = data.ts and data.ts[1] or nil
if tr or ts then
local item_face
if face == "term" then
item_face = face
else
item_face = "default"
end
local formatted_tr = tr and export.format_transliteration(tr, data.lang, item_face) or nil
local formatted_ts = ts and export.format_transcription(ts, data.lang, item_face) or nil
if formatted_tr and formatted_ts then
insert(annotations, formatted_tr .. " " .. formatted_ts)
else
insert(annotations, formatted_tr or formatted_ts)
end
end
-- Gloss/translation
if data.gloss then
insert(annotations, export.mark(data.gloss, "gloss"))
end
-- Part of speech
if data.pos then
-- debug category for pos= containing transcriptions
if data.pos:match("/[^><]-/") then
data.pos = data.pos .. "[[Kategori:Pautan yang mungkin mengandungi transkripsi dalam pos]]"
end
-- Canonicalize part of speech aliases as well as non-lemma aliases like 'nf' or 'nounf' for "noun form".
pos_tags = pos_tags or (m_headword_data or get_headword_data()).pos_aliases
local pos = pos_tags[data.pos]
if not pos and data.pos:find("f$") then
local pos_form = data.pos:sub(1, -2)
-- We only expand something ending in 'f' if the result is a recognized non-lemma POS.
pos_form = "bentuk " .. (pos_tags[pos_form] or pos_form)
if (m_headword_data or get_headword_data()).nonlemmas[pos_form] then
pos = pos_form
end
end
insert(annotations, export.mark(pos or data.pos, "pos"))
end
-- Inflection data
if data.infl then
local m_form_of = require(form_of_module)
-- Split tag sets manually, since tagged_inflections creates a numbered list, and we do not want that.
local infl_outputs = {}
local tag_sets = m_form_of.split_tag_set(data.infl)
for _, tag_set in ipairs(tag_sets) do
table.insert(infl_outputs,
m_form_of.tagged_inflections({ tags = tag_set, lang = data.lang, nocat = true, nolink = true, nowrap = true }))
end
insert(annotations, export.mark(table.concat(infl_outputs, "; "), "infl"))
end
-- Non-gloss text
if data.ng then
insert(annotations, export.mark(data.ng, "non-gloss"))
end
-- Literal/sum-of-parts meaning
if data.lit then
insert(annotations, "secara harfiah " .. export.mark(data.lit, "gloss"))
end
-- Provide a hook to insert additional annotations such as nested inflections.
if data.postprocess_annotations then
data.postprocess_annotations {
data = data,
annotations = annotations
}
end
if annotations[1] then
insert(output, " " .. export.mark(concat(annotations, ", "), "annotations"))
end
return concat(output)
end
-- Encode certain characters to avoid various delimiter-related issues at various stages. We need to encode < and >
-- because they end up forming part of CSS class names inside of <span ...> and will interfere with finding the end
-- of the HTML tag. I first tried converting them to URL encoding, i.e. %3C and %3E; they then appear in the URL as
-- %253C and %253E, which get mapped back to %3C and %3E when passed to [[Module:accel]]. But mapping them to <
-- and > somehow works magically without any further work; they appear in the URL as < and >, and get passed to
-- [[Module:accel]] as < and >. I have no idea who along the chain of calls is doing the encoding and decoding. If
-- someone knows, please modify this comment appropriately!
local accel_char_map
local function get_accel_char_map()
accel_char_map = {
["%"] = ".",
[" "] = "_",
["_"] = u(0xFFF0),
["<"] = "<",
[">"] = ">",
}
return accel_char_map
end
local function encode_accel_param_chars(param)
return (param:gsub("[%% <>_]", accel_char_map or get_accel_char_map()))
end
local function encode_accel_param(prefix, param)
if not param then
return ""
end
if type(param) == "table" then
local filled_params = {}
-- There may be gaps in the sequence, especially for translit params.
local maxindex = 0
for k in pairs(param) do
if type(k) == "number" and k > maxindex then
maxindex = k
end
end
for i = 1, maxindex do
filled_params[i] = param[i] or ""
end
-- [[Module:accel]] splits these up again.
param = concat(filled_params, "*~!")
end
-- This is decoded again by [[WT:ACCEL]].
return prefix .. encode_accel_param_chars(param)
end
local function insert_if_not_blank(list, item)
if item == "" then
return
end
insert(list, item)
end
local function get_css_classes(lang, tr, accel, nowrap)
if not accel and not nowrap then
return ""
end
local classes = {}
if accel then
insert(classes, "form-of lang-" .. lang:getFullCode())
local form = accel.form
if form then
insert(classes, encode_accel_param_chars(form) .. "-form-of")
end
insert_if_not_blank(classes, encode_accel_param("gender-", accel.gender))
insert_if_not_blank(classes, encode_accel_param("pos-", accel.pos))
insert_if_not_blank(classes, encode_accel_param("transliteration-", accel.translit or (tr ~= "-" and tr or nil)))
insert_if_not_blank(classes, encode_accel_param("target-", accel.target))
insert_if_not_blank(classes, encode_accel_param("origin-", accel.lemma))
insert_if_not_blank(classes, encode_accel_param("origin_transliteration-", accel.lemma_translit))
if accel.no_store then
insert(classes, "form-of-nostore")
end
end
if nowrap then
insert(classes, nowrap)
end
return concat(classes, " ")
end
--[==[
Creates a full link, with annotations (see `##format_link_annotations`), in the style of {{tl|l}} or {{tl|m}}.
The first argument, `data`, must be a table. It contains the various elements that can be supplied as parameters to
{{tl|l}} or {{tl|m}}:
{ {
-- Basic link-related fields
term = "entry_to_link_to",
alt = "link_text_or_displayed_text",
lang = language_object,
sc = script_object,
fragment = "link_fragment",
id = "sense_id",
accel = {accelerated_creation_tags},
-- Link annotation fields
interwiki = "interwiki_link",
genders = {"gender1", "gender2", ...} or {{spec = "gender1", q = {"left qualifier", ...}, qq = {"right qualifier"}, ...}, ...},
tr = "transliteration" or "-" or {{tr = "transliteration", q = {"left_qualifier", ...}, qq = {"right_qualifier", ...}, ..., genders = {gender_spec, ...}}, ...},
ts = "transcription" or {{ts = "transliteration", q = {"left_qualifier", ...}, qq = {"right_qualifier", ...}, ..., genders = {gender_spec, ...}}, ...},
gloss = "gloss",
pos = "part_of_speech_tag",
infl = {"infl1_tag1", "infl1_tag2", ..., ";", "infl2_tag1", "infl2_tag2", ...},
ng = "non-gloss text",
lit = "literal_translation",
postprocess_annotations = function({data = full_link_data, annotations = {"annotation1", "annotation2", ...}}) -> nil,
-- Other transliteration-related fields
respect_link_tr = boolean,
never_call_transliteration_module = boolean,
suppress_tr = boolean,
-- Decoration fields
q = { "left_qualifier1", "left_qualifier2", ...} or "left_qualifier",
qq = { "right_qualifier1", "right_qualifier2", ...} or "right_qualifier",
l = { "left_label1", "left_label2", ...},
ll = { "right_label1", "right_label2", ...},
a = { "left_accent_qualifier1", "left_accent_qualifier2", ...},
aa = { "right_accent_qualifier1", "right_accent_qualifier2", ...},
refs = { "formatted_ref1", "formatted_ref2", ...} or { {text = "text", name = "name", group = "group"}, ... },
pretext = "text_at_beginning",
posttext = "text_at_end",
show_decorations = boolean,
-- Fields controlling tracking categories
track_sc = boolean,
no_nonstandard_sc_cat = boolean,
suppress_redundant_wikilink_cat = function(term, alt) -> boolean,
-- Miscellaneous fields
no_alt_ast = boolean,
no_generate_alternants = boolean,
} }
Any one of the items in the `data` table (except for `lang`) may be {nil}. If none of `term`, `alt` and `tr` is present,
a term request will be shown. Thus, calling {full_link{ term = term, lang = lang, sc = sc }}, where `term` is the page
to link to (which may have diacritics that will be stripped and/or embedded bracketed links) and `lang` is a
[[Module:languages#Language objects|language object]] from [[Module:languages]], will give a plain link similar to the
one produced by the template {{tl|l}}, and calling {full_link( { term = term, lang = lang, sc = sc }, "term" )} will
give a link similar to the one produced by the template {{tl|m}}.
The function will:
* Try to determine the script, based on the characters found in the `term` or `alt` argument, if the script was not
given. If a script is given and `track_sc` is {true}, it will check whether the input script is the same as the one
which would have been automatically generated and add the category ` ``lang`` terms with redundant script codes` if
yes, or ` ``lang`` terms with non-redundant manual script codes` if no. This should be used when the input script
object is directly determined by a template's `sc` parameter.
* Call `simple_link()` on the `term` or `alt` forms, to remove diacritics in the page name, process any embedded
wikilinks and create links to Reconstruction or Appendix pages when necessary. (`simple_link()` is almost exactly the
same as `##language_link()`; the latter is a simple wrapper around the former that adds a bit of extra tracking.)
* Call `[[Module:script utilities#tag_text]]` to add the appropriate language and script tags to the term and
italicize terms written in the Latin script if necessary. Accelerated creation tags, as used by [[WT:ACCEL]], are
included.
* Generate a transliteration, based on the `alt` or `term` arguments, if the script is not Latin, no transliteration was
provided in `tr` and the combination of the term's language and script support automatic transliteration. The
transliteration itself will be linked if both `.respect_link_tr` is specified and the language of the term has the
`link_tr` property set for the script of the term; but not otherwise.
* Add the annotations (transliteration, gender, gloss, etc.) after the link.
* If `no_alt_ast` is specified, then the `alt` text does not need to contain an asterisk if the language is
reconstructed. This should only be used by modules which really need to allow links to reconstructions that don't
display asterisks (e.g. number boxes).
* If `suppress_redundant_wikilink_cat` is specified, it should be a function that indicates whether to suppress the
generation of the ` ``lang`` links with redundant wikilinks` tracking category. It is passed two arguments, the `term`
and `alt` parameters. Normally, this tracking category is added whenever the `term` argument consists entirely of a
one-part or two-part embedded link, which is considered "redundant" in that the link can be rewritten into separate
`term` and `alt` arguments without any embedded links. For certain wrapping templates, however, otherwise "redundant"
embedded links are necessary to prevent interpretation of certain characters as delimiters. For example, {{tl|col}}
and related templates use `~` as a separator, as well as `,` when not followed by a space. In these templates,
embedded links are required to correctly link to terms containing those delimiters, such as [[Micros~1]] and
[[1,6-Cleves acid]], but will incorrectly trigger the addition of the tracking category unless the appropriate
`suppress_redundant_wikilink_cat` function is given.
* If `pretext` or `posttext` is specified, this is text to (respectively) prepend or append to the output, directly
before processing decorations (qualifiers, labels and references). This can be used to add arbitrary extra text inside
of the decorations.
* If `show_decorations` is specified, then decorations specified in `data` (i.e. left and right qualifiers, accent
qualifiers, labels and references) will be displayed, otherwise they will be ignored. (This is because a fair amount
of code stores decorations in these fields and displays them itself, rather than expecting {full_link()} to display
them.)
* ]==]
function export.full_link(data, face, allow_self_link, show_qualifiers)
if type(data) ~= "table" then
error("The first argument to the function full_link must be a table. "
.. "See Module:links/documentation for more information.")
elseif data.term and data.term:find("\\", nil, true) or data.alt and data.alt:find("\\", nil, true) then
track("escaped", "full_link")
end
if show_qualifiers then
-- FIXME: Eventually remove the error code. Added 2026-09-17, remove after 2026-10-17 or so.
error("Can't pass fourth parameter `show_qualifiers` any more. Set `show_decorations = true` on data.")
end
if data.show_qualifiers then
-- FIXME: Eventually remove the error code. Added 2026-09-18, remove after 2026-10-18 or so.
error("Can't set field `show_qualifiers`; use `show_decorations`")
end
if data.no_generate_forms then
-- FIXME: Eventually remove the error code. Added 2026-09-14, remove after 2026-10-14 or so.
error("Can't set field `no_generate_forms`; use `no_generate_alternants`")
end
-- Prevent data from being destructively modified.
data = shallow_copy(data)
data.cats = {}
-- Categorize links to "und".
local lang, cats = data.lang, data.cats
if cats and lang:getCode() == "und" then
insert(cats, "Pautan bahasa tidak ditentukan")
end
local terms = { true }
-- Generate multiple alternants if applicable.
for _, param in ipairs { "term", "alt" } do
if type(data[param]) == "string" and data[param]:find("//", nil, true) then
data[param] = export.split_on_slashes(data[param])
elseif type(data[param]) == "string" and not (type(data.term) == "string" and data.term:find("//", nil, true)) then
if not data.no_generate_alternants then
data[param] = lang:generateAlternants(data[param])
else
data[param] = { data[param] }
end
else
data[param] = {}
end
end
for _, param in ipairs { "sc", "tr", "ts" } do
data[param] = { data[param] }
end
for _, param in ipairs { "term", "alt", "sc", "tr", "ts" } do
for i in pairs(data[param]) do
terms[i] = true
end
end
-- Create the link
local outparts = {}
local id, no_alt_ast, suppress_redundant_wikilink_cat, accel, never_call_transliteration_module =
data.id, data.no_alt_ast, data.suppress_redundant_wikilink_cat, data.accel,
data.never_call_transliteration_module
local link_tr = data.respect_link_tr and lang:link_tr(data.sc[1])
for i in ipairs(terms) do
local link
-- Is there any text to show?
if (data.term[i] or data.alt[i]) then
-- Try to detect the script if it was not provided
local display_term = data.alt[i] or data.term[i]
local best = lang:findBestScript(display_term)
-- no_nonstandard_sc_cat is intended for use in [[Module:interproject]]
if (
not data.no_nonstandard_sc_cat and
best:getCode() == "None" and
find_best_script_without_lang(display_term):getCode() ~= "None"
) then
insert(cats, "Perkataan bahasa " .. lang:getFullName() .. " dalam bentuk tulisan tidak piawai")
end
if not data.sc[i] then
data.sc[i] = best
-- Track uses of sc parameter.
elseif data.track_sc then
if data.sc[i]:getCode() == best:getCode() then
insert(cats, "Perkataan dengan kod tulisan lewah bahasa " .. lang:getFullName())
else
insert(cats, "Perkataan dengan kod tulisan manual tidak lewah bahasa " .. lang:getFullName())
end
end
-- If using a discouraged character sequence, add to maintenance category
if data.sc[i]:hasNormalizationFixes() == true then
if (data.term[i] and data.sc[i]:fixDiscouragedSequences(toNFC(data.term[i])) ~= toNFC(data.term[i])) or (data.alt[i] and data.sc[i]:fixDiscouragedSequences(toNFC(data.alt[i])) ~= toNFC(data.alt[i])) then
insert(cats, "Laman menggunakan jujukan aksara tidak digalakkan")
end
end
link = simple_link(
data.term[i],
data.fragment,
data.alt[i],
lang,
data.sc[i],
id,
cats,
no_alt_ast,
suppress_redundant_wikilink_cat
)
end
-- simple_link can return nil, so check if a link has been generated.
if link then
-- Add "nowrap" class to prefixes in order to prevent wrapping after the hyphen
local nowrap
local display_term = data.alt[i] or data.term[i]
if display_term and (display_term:find("^%-") or display_term:find("^־")) then -- Hebrew maqqef -- FIXME, use hyphens from [[Module:affix]]
nowrap = "nowrap"
end
link = tag_text(link, lang, data.sc[i], face, get_css_classes(lang, data.tr[i], accel, nowrap))
else
--[[ No term to show.
Is there at least a transliteration we can work from? ]]
link = request_script(lang, data.sc[i])
-- No link to show, and no transliteration either. Show a term request (unless it's a substrate, as they rarely take terms).
if (link == "" or (not data.tr[i]) or data.tr[i] == "-") and lang:getFamilyCode() ~= "qfa-sub" then
-- If there are multiple terms, break the loop instead.
if i > 1 then
remove(outparts)
break
elseif NAMESPACE ~= "Templat" then
insert(cats, "Permintaan perkataan bahasa " .. lang:getFullName())
end
link = "<small>[Istilah?]</small>"
end
end
insert(outparts, link)
if i < #terms then insert(outparts, "<span class=\"Zsym mention\" style=\"font-size:100%;\"> / </span>") end
end
-- When suppress_tr is true, do not show or generate any transliteration
if data.suppress_tr then
data.tr[1] = nil
else
-- TODO: Currently only handles the first transliteration, pending consensus on how to handle multiple translits for multiple forms, as this is not always desirable (e.g. traditional/simplified Chinese).
if data.tr[1] == "" or data.tr[1] == "-" then
data.tr[1] = nil
else
local phonetic_extraction = load_data("Module:links/data").phonetic_extraction
phonetic_extraction = phonetic_extraction[lang:getCode()] or phonetic_extraction[lang:getFullCode()]
if phonetic_extraction then
data.tr[1] = data.tr[1] or
require(phonetic_extraction).getTranslit(export.remove_links(data.alt[1] or data.term[1]))
elseif (data.term[1] or data.alt[1]) and data.sc[1]:isTransliterated() then
-- Track whenever there is manual translit. The categories below like 'terms with redundant transliterations'
-- aren't sufficient because they only work with reference to automatic translit and won't operate at all in
-- languages without any automatic translit, like Persian and Hebrew.
if data.tr[1] then
local full_code = lang:getFullCode()
track("manual-tr", full_code)
end
if not never_call_transliteration_module then
-- Try to generate a transliteration.
local text = data.alt[1] or data.term[1]
if not link_tr then
text = export.remove_links(text, true)
end
local automated_tr = lang:transliterate(text, data.sc[1])
if automated_tr then
local manual_tr = data.tr[1]
if manual_tr then
if export.remove_links(manual_tr) == export.remove_links(automated_tr) then
insert(cats, "Perkataan dengan transliterasi lewah bahasa " .. lang:getFullName())
else
-- Prevents Arabic root categories from flooding the tracking categories.
if NAMESPACE ~= "Kategori" then
insert(cats,
"Perkataan dengan transliterasi manual tidak lewah bahasa " .. lang:getFullName())
end
end
end
if not manual_tr or lang:overrideManualTranslit(data.sc[1]) then
data.tr[1] = automated_tr
end
end
end
end
end
end
-- Link to the transliteration entry for languages that require this
if data.tr[1] and link_tr and not data.tr[1]:match("%[%[(.-)%]%]") then
data.tr[1] = simple_link(
data.tr[1],
nil,
nil,
lang,
get_script("Latn"),
nil,
cats,
no_alt_ast,
suppress_redundant_wikilink_cat
)
elseif data.tr[1] and not link_tr then
-- Remove the pseudo-HTML tags added by remove_links.
data.tr[1] = data.tr[1]:gsub("</?link>", "")
end
if data.tr[1] and not umatch(data.tr[1], "[^%s%p]") then data.tr[1] = nil end
insert(outparts, export.format_link_annotations(data, face))
if data.pretext then
insert(outparts, 1, data.pretext)
end
if data.posttext then
insert(outparts, data.posttext)
end
local categories = cats[1] and format_categories(cats, lang, "-", nil, nil, data.sc) or ""
local output = concat(outparts)
if data.show_decorations then
output = add_text_decorations(output, data, lang)
end
return output .. categories
end
--[==[
Strip links by replacing all wikilinks with their displayed text, and remove any categories. This function can be
invoked either from a template or from another module. Specifically, this function deletes category links, the targets
of piped links, and any double square brackets involved in links (other than file links, which are untouched). If `tag`
is set, then any links removed will be given pseudo-HTML tags, which allow the substitution functions in
[[Module:languages]] to properly subdivide the text in order to reduce the chance of substitution failures in modules
which scrape pages like [[Module:zh-translit]]. (FIXME: This is quite hacky. We probably want this to be integrated
into [[Module:languages]], but we can't do that until we know that nothing is pushing pipe linked transliterations
through it for languages which don't have link_tr set.)
* `<nowiki>[[page|displayed text]]</nowiki>` → `displayed text`
* `<nowiki>[[page and displayed text]]</nowiki>` → `page and displayed text`
* `<nowiki>[[Category:English lemmas|WORD]]</nowiki>` → ''(nothing)''
]==]
function export.remove_links(text, tag)
if type(text) == "table" then
text = text.args[1]
end
if not text or text == "" then
return ""
end
text = text
:gsub("%[%[", "\1")
:gsub("%]%]", "\2")
-- Parse internal links for the display text.
text = text:gsub("(\1)([^\1\2]-)(\2)",
function(c1, c2, c3)
-- Don't remove files.
for _, false_positive in ipairs({ "file", "image" }) do
if c2:lower():match("^" .. false_positive .. ":") then return c1 .. c2 .. c3 end
end
-- Remove categories completely.
for _, false_positive in ipairs({ "category", "cat" }) do
if c2:lower():match("^" .. false_positive .. ":") then return "" end
end
-- In piped links, remove all text before the pipe, unless it's the final character (i.e. the pipe trick), in which case just remove the pipe.
c2 = c2:match("^[^|]*|(.+)") or c2:match("([^|]+)|$") or c2
if tag then
return "<link>" .. c2 .. "</link>"
else
return c2
end
end)
text = text
:gsub("\1", "[[")
:gsub("\2", "]]")
return text
end
function export.section_link(link)
if type(link) ~= "string" then
error("The first argument to section_link was a " .. type(link) .. ", but it should be a string.")
elseif link:find("\\", nil, true) then
track("escaped", "section_link")
end
local target, section = get_fragment((link:gsub("_", " ")))
if not section then
error("No \"#\" delineating a section name")
end
return simple_link(
target,
section,
target .. " § " .. section
)
end
return export
q7sfw2xyccv8fplvxsqjqtahr7rn6v9
Modul:gender and number
828
9772
375355
185123
2026-09-22T03:13:44Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92717017|92717017]])
375355
Scribunto
text/plain
local export = {}
local decorations_module = "Module:decorations"
local load_module = "Module:load"
local parameters_module = "Module:parameters"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local insert = table.insert
local function deep_copy(...)
deep_copy = require(table_module).deepCopy
return deep_copy(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function format_decorations(...)
format_decorations = require(decorations_module).format_decorations
return format_decorations(...)
end
local function load_data(...)
load_data = require(load_module).load_data
return load_data(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local gender_and_number_data
local function get_gender_and_number_data()
gender_and_number_data, get_gender_and_number_data = load_data("Module:gender and number/data"), nil
return gender_and_number_data
end
--[==[ intro:
This module creates standardised displays for gender and number. It converts a gender specification into Wiki/HTML format.
A gender/number specification consists of one or more gender/number elements, separated by hyphens. Examples are:
{"n"} (neuter gender), {"f-p"} (feminine plural), {"m-an-p"} (masculine animate plural),
{"pf"} (perfective aspect). Each gender/number element has the following properties:
# A code, as used in the spec, e.g. {"f"} for feminine, {"p"} for plural.
# A type, e.g. `gender`, `number` or `animacy`. Each element in a given spec must be of a different type.
# A display form, which in turn consists of a display code and a tooltip gloss. The display code
may not be the same as the spec code, e.g. the spec code {"an"} has display code {"anim"} and tooltip
gloss ''animate''.
# A category into which lemmas of the right part of speech are placed if they have a gender/number
spec containing the given element. For example, a noun with gender/number spec {"m-an-p"} is placed
into the categories `<var>lang</var> masculine nouns`, `<var>lang</var> animate nouns` and `<var>lang</var> pluralia tantum`.
]==]
--[==[
Version of format_genders() that can be invoked from a template.
]==]
function export.show_list(frame)
local params = {
[1] = {list = true},
["lang"] = {type = "language"},
}
local iargs = process_params(frame.args, params)
local text, cats = export.format_genders(iargs[1], iargs.lang)
if not cats then
return text
end
return text .. format_categories(cats, iargs.lang)
end
function export.format_list(...)
-- FIXME: Added 2026-09-17. Remove after 2026-10-17.
error("Use format_genders instead of format_list")
end
local function autoadd_abbr(display)
if not display then
error("Internal error: '.display' for gender/number code is missing")
end
if display:find("<abbr", nil, true) then
return display
end
return ('%s'):format(display, display)
end
--[=[
Add decorations (qualifiers, labels and references) to a formatted gender/number spec. `spec` is the object
describing the gender/number spec, which should optionally contain:
* left qualifiers in `q`, an array of strings;
* right qualifiers in `qq`, an array of strings;
* left labels in `l`, an array of strings;
* right labels in `ll`, an array of strings;
* references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text`
(formatted reference text) and optionally `name` and/or `group`;
`formatted` is the formatted version of the term itself, and `lang` is the optional language object passed into
format_genders().
]=]
local function add_decorations(formatted, spec, lang)
local function field_non_empty(field)
local list = spec[field]
if not list then
return nil
end
if type(list) ~= "table" then
error(("Internal error: Wrong type for `spec.%s`=%s, should be \"table\""):format(
field, mw.dumpObject(list)))
end
return list[1]
end
if field_non_empty("q") or field_non_empty("qq") or field_non_empty("l") or field_non_empty("ll") or
field_non_empty("refs") then
formatted = format_decorations {
lang = lang,
text = formatted,
q = spec.q,
qq = spec.qq,
l = spec.l,
ll = spec.ll,
refs = spec.refs,
}
end
return formatted
end
--[==[
Format one or more gender/number specifications. Each spec is either a string, e.g. {"f-p"}, or a table of the form
{ {spec = "SPEC",
q = {"LEFT_QUALIFIER", "LEFT_QUALIFIER", ...},
qq = {"RIGHT_QUALIFIER", "RIGHT_QUALIFIER", ...},
l = {"LEFT_LABEL", "LEFT_LABEL", ...},
ll = {"RIGHT_LABEL", "RIGHT_LABEL", ...},
refs = {"FORMATTED_REFERENCE_TEXT", {text = "FORMATTED_REFERENCE_TEXT", name = "REFNAME", group = "REFGROUP"}, ...},
}} In the table format, only `spec` is required, and in the `refs` object format, only `text` is required.
If `lang` is not given, the funtion returns only one value, the formatted text.
If `lang` is given, the function returns two values:
# the formatted text;
# a list of the categories to add, which will be {nil} if there are no categories to add.
If `lang` (which should be a language object) and `pos_for_cat` (which should be a plural part of speech) are given,
gender categories such as `German masculine nouns` or `Russian imperfective verbs` are added to the categories, and
request categories such as `Requests for gender in <var>lang</var> entries` or
`Requests for animacy in <var>lang</var> entries` may also be added. Otherwise, if only `lang` is given, only request
categories may be returned. If both are omitted, only one return value is returned, as mentioned above.
]==]
function export.format_genders(specs, lang, pos_for_cat)
local formatted_specs, categories, seen_types = {}
local all_is_nounclass = nil
local full_langname = lang and lang:getFullName() or nil
local function do_gender_spec(spec, parts)
local types = {}
local codes = (gender_and_number_data or get_gender_and_number_data()).codes
for key, code in ipairs(parts) do
-- Is this code valid?
if not codes[code] then
error('The tag "' .. code .. '" in the gender specification "' .. spec.spec ..
'" is not valid. See [[Module:gender and number]] for a list of valid tags.')
end
-- Check for multiple genders/numbers/animacies in a single spec.
local typ = codes[code].type
if typ ~= "other" and types[typ] then
error('The gender specification "' .. spec.spec .. '" contains multiple tags of type "' .. typ .. '".')
end
types[typ] = true
parts[key] = autoadd_abbr(codes[code].display)
-- Generate categories if called for.
if lang and pos_for_cat then
local cat = codes[code].cat
if cat then
if not categories then
categories = {}
end
insert(categories, cat .. " bahasa " .. full_langname)
end
if not seen_types then
seen_types = {}
elseif seen_types[typ] and seen_types[typ] ~= code then
cat = (gender_and_number_data or get_gender_and_number_data()).multicode_cats[typ]
if cat then
if not categories then
categories = {}
end
insert(categories, cat .. " bahasa " .. full_langname)
end
end
seen_types[typ] = code
end
if lang and codes[code].req then
local type_for_req = typ
if code == "?" then
-- Keep in mind `pos_for_cat` may be nil here.
type_for_req = pos_for_cat == "kata kerja" and "aspek" or "genus"
end
if not categories then
categories = {}
end
insert(categories, "Permintaan untuk " .. type_for_req .. " dalam entri bahasa " .. full_langname)
end
end
-- Add the processed codes together with non-breaking spaces
if not parts[2] and parts[1] then
return parts[1]
else
return concat(parts, " ")
end
end
for _, spec in ipairs(specs) do
if type(spec) ~= "table" then
spec = {spec = spec}
end
local spec_spec, is_nounclass = spec.spec
-- If the specification starts with cX, then it is a noun class specification.
if spec_spec:match("^%d") or spec_spec:match("^c[^-]") then
is_nounclass = true
local code = spec_spec:gsub("^c", "")
local text
if code == "?" then
text = '<abbr class="noun-class" title="noun class missing">?</abbr>'
if lang then
if not categories then
categories = {}
end
insert(categories, "Permintaan kelas kata nama dalam entri bahasa " .. full_langname)
end
else
text = '<abbr class="noun-class" title="noun class ' .. code .. '">' .. code .. "</abbr>"
if lang and pos_for_cat then
if not categories then
categories = {}
end
insert(categories, "POS" .. code .. " kelas bahasa " .. full_langname)
end
end
local text_with_decorations = add_decorations(text, spec, lang)
insert(formatted_specs, text_with_decorations)
else
-- Split the parts and iterate over each part, converting it into its display form
local parts = split(spec.spec, "-", true, true)
local combined_codes = (gender_and_number_data or get_gender_and_number_data()).combinations
if lang then
-- Check if the specification is valid
--elseif langinfo.genders then
-- local valid_genders = {}
-- for _, g in ipairs(langinfo.genders) do valid_genders[g] = true end
--
-- if not valid_genders[spec.spec] then
-- local valid_string = {}
-- for i, g in ipairs(langinfo.genders) do valid_string[i] = g end
-- error('The gender specification "' .. spec.spec .. '" is not valid for ' .. langinfo.names[1] .. ". Valid are: " .. concat(valid_string, ", "))
-- end
--end
end
local has_combined = false
for _, code in ipairs(parts) do
if combined_codes[code] then
has_combined = true
break
end
end
if not has_combined then
if formatted_specs[1] then
insert(formatted_specs, "or")
end
insert(formatted_specs, add_decorations(do_gender_spec(spec, parts), spec, lang))
else
-- This logic is to handle combined gender specs like 'mf' and 'mfbysense'.
local all_parts = {{}}
local extra_displays
local this_formatted_specs = {}
for _, code in ipairs(parts) do
if combined_codes[code] then
local new_all_parts = {}
for _, one_parts in ipairs(all_parts) do
for _, one_code in ipairs(combined_codes[code].codes) do
local new_combined_parts = deep_copy(one_parts)
insert(new_combined_parts, one_code)
insert(new_all_parts, new_combined_parts)
end
end
all_parts = new_all_parts
if lang and pos_for_cat then
local extra_cat = combined_codes[code].cat
if extra_cat then
if not categories then
categories = {}
end
insert(categories, extra_cat .. " bahasa " .. full_langname)
end
end
local extra_display = combined_codes[code].display
if extra_display then
if not extra_displays then
extra_displays = {}
end
insert(extra_displays, autoadd_abbr(extra_display))
end
else
for _, one_parts in ipairs(all_parts) do
insert(one_parts, code)
end
end
end
for _, this_parts in ipairs(all_parts) do
if this_formatted_specs[1] then
insert(this_formatted_specs, "or")
end
insert(this_formatted_specs, do_gender_spec(spec, this_parts))
end
if extra_displays then
for _, display in ipairs(extra_displays) do
insert(this_formatted_specs, display)
end
end
insert(formatted_specs, add_decorations(
concat(this_formatted_specs, " "), spec, lang))
end
is_nounclass = false
end
-- Ensure that the specifications are either all noun classes, or none are.
if all_is_nounclass == nil then
all_is_nounclass = is_nounclass
elseif all_is_nounclass ~= is_nounclass then
error("Noun classes and genders cannot be mixed. Please use either one or the other.")
end
end
if categories and lang and pos_for_cat then
for i, cat in ipairs(categories) do
categories[i] = cat:gsub("POS", pos_for_cat)
end
end
local formatted_retval
if all_is_nounclass then
-- Add the processed codes together with slashes
formatted_retval = '<span class="gender">class ' .. concat(formatted_specs, "/") .. "</span>"
else
-- Add the processed codes together with spaces
formatted_retval = '<span class="gender">' .. concat(formatted_specs, " ") .. "</span>"
end
if lang then
return formatted_retval, categories
else
return formatted_retval
end
end
return export
2wkm2desvqy89qy6hbwj4b0amm5ltiu
Modul:links/templates
828
9792
375365
245095
2026-09-22T04:36:17Z
Hakimi97
2668
Kemas kini
375365
Scribunto
text/plain
-- Prevent substitution.
if mw.isSubsting() then
return require("Module:unsubst")
end
local export = {}
local links_module = "Module:links"
local process_params = require("Module:parameters").process
local remove = table.remove
local upper = require("Module:string utilities").upper
--[=[
Modules used:
[[Module:links]]
[[Module:languages]]
[[Module:scripts]]
[[Module:parameters]]
[[Module:debug]]
]=]
do
local function get_args(frame)
-- `compat` is a compatibility mode for {{term}}.
-- If given a nonempty value, the function uses lang= to specify the
-- language, and all the positional parameters shift one number lower.
local iargs = frame.args
iargs.compat = iargs.compat and iargs.compat ~= ""
iargs.langname = iargs.langname and iargs.langname ~= ""
iargs.notself = iargs.notself and iargs.notself ~= ""
local alias_of_4 = {alias_of = 4}
local boolean = {type = "boolean"}
local params = {
[1] = {required = true, type = "language", default = "und"},
[2] = true,
[3] = true,
[4] = true,
g = {list = true, type = "genders", flatten = true},
gloss = alias_of_4,
id = true,
lit = true,
ng = true,
pos = true,
sc = {type = "script"},
t = alias_of_4,
tr = true,
ts = true,
q = {type = "qualifier"},
qq = {type = "qualifier"},
l = {type = "labels"},
ll = {type = "labels"},
ref = {type = "references"},
["accel-form"] = true,
["accel-translit"] = true,
["accel-lemma"] = true,
["accel-lemma-translit"] = true,
["accel-gender"] = true,
["accel-nostore"] = boolean,
}
if iargs.compat then
params.lang = {type = "language", default = "und"}
remove(params, 1)
alias_of_4.alias_of = 3
end
if iargs.langname then
params.w = boolean
end
return process_params(frame:getParent().args, params), iargs
end
-- Used in [[Template:l]] and [[Template:m]].
function export.l_term_t(frame)
local args, iargs = get_args(frame)
local compat = iargs.compat
local lang = args[compat and "lang" or 1]
-- Tracking for und.
if not compat and lang:getCode() == "und" then
require("Module:debug").track("link/und")
end
local term = args[(compat and 1 or 2)]
local alt = args[(compat and 2 or 3)]
term = term ~= "" and term or nil
if not term and not alt and iargs.demo then
term = iargs.demo
end
local langname = iargs.langname and (
args.w and lang:makeWikipediaLink() or
lang:getCanonicalName()
) or nil
if langname and term == "-" then
return langname
end
-- Forward the information to full_link
return (langname and langname .. " " or "") .. require(links_module).full_link(
{
lang = lang,
sc = args.sc,
track_sc = true,
term = term,
alt = alt,
gloss = args[4],
id = args.id,
tr = args.tr,
ts = args.ts,
genders = args.g,
pos = args.pos,
ng = args.ng,
lit = args.lit,
q = args.q,
qq = args.qq,
l = args.l,
ll = args.ll,
refs = args.ref,
show_decorations = true,
accel = args["accel-form"] and {
form = args["accel-form"],
translit = args["accel-translit"],
lemma = args["accel-lemma"],
lemma_translit = args["accel-lemma-translit"],
gender = args["accel-gender"],
nostore = args["accel-nostore"],
} or nil
},
iargs.face,
not iargs.notself
)
end
-- Used in [[Template:link-annotations]].
function export.l_annotations_t(frame)
local args, iargs = get_args(frame)
-- Forward the information to format_link_annotations
return require(links_module).format_link_annotations(
{
lang = args[1],
tr = { args.tr },
ts = { args.ts },
genders = args.g,
pos = args.pos,
ng = args.ng,
lit = args.lit
},
iargs.face
)
end
end
-- Used in [[Template:ll]].
do
local function get_args(frame)
return process_params(frame:getParent().args, {
[1] = {required = true, type = "language", default = "und"},
[2] = {allow_empty = true},
[3] = true,
id = true,
sc = {type = "script"},
})
end
function export.ll(frame)
local args = get_args(frame)
local lang = args[1]
local sc = args.sc
local term = args[2]
term = term ~= "" and term or nil
return require(links_module).language_link{
lang = lang,
sc = sc,
term = term,
alt = args[3],
id = args.id
} or "<small>[Istilah?]</small>" ..
require("Module:utilities").format_categories(
{"Permintaan istilah bahasa " .. lang:getFullName()},
lang, "-", nil, nil, sc
)
end
end
function export.def_t(frame)
local args = process_params(frame:getParent().args, {
[1] = {required = true, default = ""},
})
local face = frame.args.face
local ret = require("Module:script utilities").tag_definition(require(links_module).embedded_language_links{
term = args[1],
lang = require("Module:languages").getByCode("ms"),
sc = require("Module:scripts").getByCode("Latn")
}, face)
if face == "non-gloss" then
return ret
end
return '<span class="mention-gloss-paren">(</span>' .. ret .. '<span class="mention-gloss-paren">)</span>'
end
function export.linkify_t(frame)
local args = process_params(frame:getParent().args, {
[1] = {required = true, default = ""},
})
args[1] = mw.text.trim(args[1])
if args[1] == "" or args[1]:find("[[", nil, true) then
return args[1]
end
return "[[" .. args[1] .. "]]"
end
function export.cap_t(frame)
local args = process_params(frame:getParent().args, {
[1] = {required = true},
[2] = true,
lang = {type = "language", default = "ms"},
})
local term = args[1]
return require(links_module).full_link{
lang = args.lang,
term = term,
alt = term:gsub("^.[\128-\191]*", upper) .. (args[2] or "")
}
end
function export.section_link_t(frame)
local args = process_params(frame:getParent().args, {
[1] = {},
})
return require(links_module).section_link(args[1])
end
return export
dhr763ljty2b15td9ft3hiqn36wxt5h
Modul:languages/data/3/k
828
9814
375377
373727
2026-09-22T05:35:12Z
Hakimi97
2668
Move "Dusun Tambunan" back to Module:languages/data/3/k
375377
Scribunto
text/plain
local m_langdata = require("Module:languages/data")
-- Loaded on demand, as it may not be needed (depending on the data).
local function u(...)
u = require("Module:string utilities").char
return u(...)
end
local c = m_langdata.chars
local p = m_langdata.puaChars
local s = m_langdata.shared
local m = {}
m["kaa"] = {
"Karakalpak",
33541,
"trk-kno",
"Latn, Cyrl, Arab",
dotted_dotless_i = true,
strip_diacritics = {
from = {"['’]"},
to = {"ʼ"}
},
sort_key = {
Latn = {
from = {
-- Sort the old orthography (using the apostrophe) after the new orthography (using the acute accent).
"í", "iʼ", "i", -- Ensure "i" comes after "í", "iʼ", "ı".
"sh", "ch",
"á", "aʼ", "ǵ", "gʼ", "x", p[4], p[5], "ı", "q", "ń", "nʼ", "ó", "oʼ", "ú", "uʼ", "c"
},
to = {
p[4], p[5], "i" .. p[3],
"z" .. p[1], "z" .. p[3],
"a" .. p[1], "a" .. p[2], "g" .. p[1], "g" .. p[2], "h" .. p[1], "i", "i" .. p[1], "i" .. p[2], "k" .. p[1], "n" .. p[1], "n" .. p[2], "o" .. p[1], "o" .. p[2], "u" .. p[1], "u" .. p[2], "z" .. p[2]
}
},
Cyrl = {
from = {"ә", "ғ", "ё", "қ", "ң", "ө", "ү", "ў", "ҳ"},
to = {"а" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1], "у" .. p[2], "х" .. p[1]}
},
},
}
m["kab"] = {
"Kabyle",
35853,
"ber",
"Latn, Arab, Tfng",
}
m["kac"] = {
"Jingpho",
33332,
"sit-jnp",
"Latn, Mymr",
}
m["kad"] = {
"Kadara",
3914011,
"nic-plc",
"Latn",
}
m["kae"] = {
"Ketangalan",
2779411,
"map",
}
m["kaf"] = {
"Katso",
246122,
"tbq-kzh",
}
m["kag"] = {
"Kajaman",
6348863,
"poz",
"Latn",
}
m["kah"] = {
"Fer",
5443742,
"csu-bgr",
"Latn",
}
m["kai"] = {
"Karekare",
3438770,
"cdc-wst",
"Latn",
}
m["kaj"] = {
"Jju",
35401,
"nic-plc",
"Latn",
}
m["kak"] = {
"Kayapa Kallahan",
3192220,
"phi",
"Latn",
}
m["kam"] = {
"Kamba",
2574767,
"bnt-kka",
"Latn",
}
m["kao"] = {
"Kassonke",
36905,
"dmn-wmn",
"Latn",
}
m["kap"] = {
"Bezhta",
33054,
"cau-ets",
"Cyrl",
translit = "cau-nec-translit",
override_translit = true,
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]},
}
m["kaq"] = {
"Capanahua",
2937196,
"sai-pan",
"Latn",
}
m["kaw"] = {
"Jawa Kuno",
49341,
"poz",
"Latn, Java, Kawi",
translit = "jv-translit", --same as jv
}
m["kax"] = {
"Kao",
3192799,
"paa-gto",
"Latn",
}
m["kay"] = {
"Kamayurá",
3192336,
"tup-gua",
"Latn",
}
m["kba"] = {
"Kalarko",
5517764,
"aus-pam",
"Latn",
}
m["kbb"] = {
"Kaxuyana",
12953626,
"sai-prk",
"Latn",
}
m["kbc"] = {
"Kadiwéu",
18168288,
"sai-guc",
"Latn",
}
m["kbd"] = {
"Kabardia",
33522,
"cau-cir",
"Cyrl, Latn, Arab",
translit = {
Cyrl = "cau-cir-translit",
Arab = "ar-translit",
},
override_translit = true,
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {
Cyrl = s["cau-Cyrl-stripdiacritics"],
Latn = s["cau-Latn-stripdiacritics"],
},
sort_key = {
Cyrl = {
from = {
"кхъу", "къӏу", -- 4 chars
"гъу", "джу", "дзу", "жъу", "къу", "кхъ", "къӏ", "кӏу", "кӏь", "лъу", "лӏу", "пӏу", "сӏу", "тӏу", "фӏу", "хъу", "цӏу", "чъу", "чӏу", "шъу", "шӏу", "щӏу", -- 3 chars
"гу", "гъ", "гь", "дж", "дз", "ё", "жъ", "жь", "ку", "къ", "кь", "кӏ", "лъ", "ль", "лӏ", "пӏ", "сӏ", "тӏ", "фӏ", "ху", "хъ", "хь", "цу", "цӏ", "чу", "чъ", "чӏ", "шъ", "шӏ", "щӏ", "ӏу", "ӏь", -- 2 chars
"э" -- 1 char
},
to = {
"к" .. p[5], "к" .. p[7],
"г" .. p[3], "д" .. p[2], "д" .. p[4], "ж" .. p[2], "к" .. p[3], "к" .. p[4], "к" .. p[6], "к" .. p[10], "к" .. p[11], "л" .. p[2], "л" .. p[5], "п" .. p[2], "с" .. p[2], "т" .. p[2], "ф" .. p[2], "х" .. p[3], "ц" .. p[3], "ч" .. p[3], "ч" .. p[5], "ш" .. p[2], "ш" .. p[4], "щ" .. p[2],
"г" .. p[1], "г" .. p[2], "г" .. p[4], "д" .. p[1], "д" .. p[3], "е" .. p[1], "ж" .. p[1], "ж" .. p[3], "к" .. p[1], "к" .. p[2], "к" .. p[8], "к" .. p[9], "л" .. p[1], "л" .. p[3], "л" .. p[4], "п" .. p[1], "с" .. p[1], "т" .. p[1], "ф" .. p[1], "х" .. p[1], "х" .. p[2], "х" .. p[4], "ц" .. p[1], "ц" .. p[2], "ч" .. p[1], "ч" .. p[2], "ч" .. p[4], "ш" .. p[1], "ш" .. p[3], "щ" .. p[1], "ӏ" .. p[1], "ӏ" .. p[2],
"а" .. p[1]
}
},
},
}
m["kbe"] = {
"Kanju",
10543322,
"aus-pam",
"Latn",
}
m["kbh"] = {
"Camsá",
2842667,
"qfa-iso",
"Latn",
}
m["kbi"] = {
"Kaptiau",
6367294,
"poz-oce",
"Latn",
}
m["kbj"] = {
"Kari",
6370438,
"bnt-boa",
"Latn",
}
m["kbk"] = {
"Grass Koiari",
12952642,
"ngf-koi",
"Latn",
}
m["kbm"] = {
"Iwal",
3156391,
"poz-ocw",
"Latn",
}
m["kbn"] = {
"Kare (Africa)",
35554,
"alv-mbm",
"Latn",
}
m["kbo"] = {
"Keliko",
11275553,
"csu-mma",
"Latn",
}
m["kbp"] = {
"Kabiyé",
35475,
"nic-gne",
"Latn",
}
m["kbq"] = {
"Kamano",
11732272,
"ngf-kya",
"Latn",
}
m["kbr"] = {
"Kafa",
35481,
"omv-gon",
"Ethi, Latn",
}
m["kbs"] = {
"Kande",
35556,
"bnt-tso",
"Latn",
}
m["kbt"] = {
"Gabadi",
3291159,
"poz-ocw",
"Latn",
}
m["kbu"] = {
"Kabutra",
10966761,
"raj",
}
m["kbv"] = {
"Kamberataro",
5261289,
"paa-sng",
"Latn",
}
m["kbw"] = {
"Kaiep",
6347632,
"poz-ocw",
"Latn",
}
m["kbx"] = {
"Ap Ma",
56298,
"paa-eke",
"Latn",
}
m["kbz"] = {
"Duhwa",
56295,
"cdc-wst",
"Latn",
}
m["kcb"] = {
"Kawacha",
11732302,
"ngf-woj",
"Latn",
}
m["kcc"] = {
"Lubila",
3914381,
"nic-uce",
"Latn",
}
m["kcd"] = {
"Ngkâlmpw Kanum",
12952566,
"paa-ngk",
"Latn",
}
m["kce"] = {
"Kaivi",
6348685,
"nic-kau",
"Latn",
}
m["kcf"] = {
"Ukaan",
36651,
"nic-bco",
"Latn",
}
m["kcg"] = {
"Tyap",
3912765,
"nic-plc",
"Latn",
}
m["kch"] = {
"Vono",
3913920,
"nic-kau",
"Latn",
}
m["kci"] = {
"Kamantan",
3914019,
"nic-plc",
}
m["kcj"] = {
"Kobiana",
35609,
"alv-nyn",
"Latn",
}
m["kck"] = {
"Kalanga",
33672,
"bnt-sho",
"Latn",
}
m["kcl"] = {
"Kala",
6349982,
"poz-ocw",
"Latn",
}
m["kcm"] = {
"Tar Gula",
277963,
"csu-bba",
}
m["kcn"] = {
"Nubi",
36388,
"crp",
"Latn, Arab",
ancestors = "apd",
strip_diacritics = {remove_diacritics = c.acute},
}
m["kco"] = {
"Kinalakna",
11732320,
"ngf-dal",
"Latn",
}
m["kcp"] = {
"Kanga",
6362384,
"qfa-kad",
"Latn",
}
m["kcq"] = {
"Kamo",
3914879,
"alv-wjk",
}
m["kcr"] = {
"Katla",
35688,
"nic-ktl",
"Latn",
}
m["kcs"] = {
"Koenoem",
3438755,
"cdc-wst",
}
m["kct"] = {
"Kaian",
6347538,
"paa-ott",
"Latn",
}
m["kcu"] = {
"Kikami",
3915212,
"bnt-ruv",
"Latn",
}
m["kcv"] = {
"Kete",
3195598,
"bnt-lub",
"Latn",
}
m["kcw"] = {
"Kabwari",
6344539,
"bnt-glb",
"Latn",
}
m["kcx"] = {
"Kachama-Ganjule",
12634070,
"omv-eom",
}
m["kcy"] = {
"Korandje",
33427,
"son",
}
m["kcz"] = {
"Konongo",
11732345,
"bnt-tkm",
"Latn",
}
m["kda"] = {
"Worimi",
3914062,
"aus-pam",
"Latn",
}
m["kdc"] = {
"Kutu",
6448634,
"bnt-ruv",
}
m["kdd"] = {
"Yankunytjatjara",
34207,
"aus-pam",
"Latn",
}
m["kde"] = {
"Makonde",
35172,
"bnt-rvm",
"Latn",
}
m["kdf"] = {
"Mamusi",
6746036,
"poz-ocw",
"Latn",
}
m["kdg"] = {
"Seba",
7442316,
"bnt-sbi",
"Latn",
}
m["kdh"] = {
"Tem",
36531,
"nic-gne",
"Latn",
}
m["kdi"] = {
"Kumam",
6443410,
"sdv-los",
}
m["kdj"] = {
"Karamojong",
56326,
"sdv-ttu",
"Latn",
}
m["kdk"] = {
"Numee",
3346774,
"poz-cln",
"Latn",
}
m["kdl"] = {
"Tsikimba",
3914404,
"nic-kam",
}
m["kdm"] = {
"Kagoma",
3914420,
"nic-plc",
}
m["kdn"] = {
"Kunda",
4121130,
"bnt-sna",
"Latn",
}
m["kdp"] = {
"Kaningdon-Nindem",
3914956,
"nic-nin",
}
m["kdq"] = {
"Koch",
56431,
"tbq-bdg",
}
m["kdr"] = {
"Karaim",
33725,
"trk-kcu",
"Cyrl, Latn, Hebr",
-- Hebr display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["kdt"] = {
"Kuy",
56310,
"mkh-kat",
"Thai, Khmr, Laoo",
}
m["kdu"] = {
"Kadaru",
35441,
"nub-hil",
"Latn",
}
m["kdv"] = {
"Kado",
7402721,
"sit-luu",
}
m["kdw"] = {
"Koneraw",
11732341,
"ngf-mom",
"Latn",
}
m["kdx"] = {
"Kam",
36753,
"alv-wjk",
}
m["kdy"] = {
"Keder",
6383641,
"paa-tor",
"Latn",
}
m["kdz"] = {
"Kwaja",
11128866,
"nic-nka",
"Latn",
}
m["kea"] = {
"Kabuverdianu",
35963,
"crp",
"Latn",
ancestors = "pt",
}
m["keb"] = {
"Kélé",
35559,
"bnt-kel",
}
m["kec"] = {
"Keiga",
3409311,
"qfa-kad",
"Latn",
}
m["ked"] = {
"Kerewe",
6393846,
"bnt-haj",
}
m["kee"] = {
"Keres Timur",
15649021,
"nai-ker",
"Latn",
}
m["kef"] = {
"Kpessi",
35748,
"alv-gbe",
}
m["keg"] = {
"Tese",
16887296,
"sdv",
}
m["keh"] = {
"Keak",
6382110,
"paa-nnd",
"Latn",
}
m["kei"] = {
"Kei",
2410352,
"poz-cet",
"Latn",
}
m["kej"] = {
"Kadar",
6345179,
"dra-mal",
}
m["kek"] = {
"Q'eqchi",
35536,
"myn",
"Latn",
}
m["kel"] = {
"Kela-Yela",
6385426,
"bnt-mon",
"Latn",
}
m["kem"] = {
"Kemak",
35549,
"poz-tim",
"Latn",
}
m["ken"] = {
"Kenyang",
35650,
"nic-mam",
"Latn",
}
m["keo"] = {
"Kakwa",
3033547,
"sdv-bri",
"Latn",
}
m["kep"] = {
"Kaikadi",
6347757,
"dra-tam",
}
m["keq"] = {
"Kamar",
14916877,
"inc-hal",
}
m["ker"] = {
"Kera",
56251,
"cdc-est",
"Latn",
}
m["kes"] = {
"Kugbo",
3813394,
"nic-cde",
"Latn",
}
m["ket"] = {
"Ket",
33485,
"qfa-yke",
"Cyrl",
translit = "ket-utils",
display_text = "ket-utils",
strip_diacritics = "ket-utils",
sort_key = {
from = {"ӷ", "ё", "ӄ", "ӈ", "ө", "ә", "ʼ"},
to = {"г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "ъ" .. p[1], "ь" .. p[1]}
},
}
m["keu"] = {
"Akebu",
35026,
"alv-ktg",
"Latn",
}
m["kev"] = {
"Kanikkaran",
6363201,
"dra-mal",
"Taml, Mlym",
-- Mlym translit in [[Module:scripts/data]] (NOTE: not present before, presumably an accidental omission)
}
m["kew"] = {
"Kewa",
12952619,
"ngf-ank",
"Latn",
}
m["kex"] = {
"Kukna",
5031131,
"inc-bhi",
wikipedia_article = "Dhodia–Kukna language",
}
m["key"] = {
"Kupia",
6445354,
"inc-eas",
}
m["kez"] = {
"Kukele",
3915391,
"nic-ucn",
"Latn",
}
m["kfa"] = {
"Kodava",
33531,
"dra-kod",
"Knda, Mlym",
-- Knda translit in [[Module:scripts/data]]
-- Mlym translit in [[Module:scripts/data]]
}
m["kfb"] = {
"Kolami",
33479,
"dra-knk",
"Deva, Telu",
translit = {
Telu = "te-translit",
},
}
m["kfc"] = {
"Konda-Dora",
35679,
"dra-kki",
"Orya, Telu",
translit = {
Orya = "gon-Orya-translit",
Telu = "te-translit",
},
}
m["kfd"] = {
"Korra Koraga",
12952655,
"dra-kor",
"Knda",
-- Knda translit in [[Module:scripts/data]]
}
m["kfe"] = {
"Kota (India)",
33483,
"dra-tkt",
"Taml",
translit = "ta-translit",
}
m["kff"] = {
"Koya",
33471,
"dra-gon",
"Telu, Orya, Deva, Latn",
}
m["kfg"] = {
"Kudiya",
12952667,
"dra-tlk",
}
m["kfh"] = {
"Kurichiya",
12952676,
"dra-mal",
"Mlym",
-- Mlym translit in [[Module:scripts/data]]
}
m["kfi"] = {
"Kannada Kurumba",
56589,
"dra-sdo",
}
m["kfj"] = {
"Kemiehua",
27144776,
"mkh-pal",
}
m["kfk"] = {
"Kinnauri",
2383208,
"sit-kin",
"Takr, Deva, Latn",
}
m["kfl"] = {
"Kung",
6444510,
"nic-rnc",
"Latn",
}
m["kfn"] = {
"Kuk",
6442398,
"nic-rnc",
"Latn",
}
m["kfo"] = {
"Koro (Afrika Barat)",
11160588,
"dmn-mnk",
"Latn, Nkoo",
}
m["kfp"] = {
"Korwa",
6432786,
"mun",
}
m["kfq"] = {
"Korku",
33715,
"mun",
"Deva",
}
m["kfr"] = {
"Kachchi",
56487,
"inc-snd",
"Gujr, Arab, Sind, Khoj",
translit = {
Gujr = "gu-translit",
Sind = "Sind-translit",
Arab = "sd-Arab-translit",
},
strip_diacritics = {
remove_diacritics = c.kashida .. c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.superalef,
from = {u(0x0671)},
to = {u(0x0627)}
},
}
m["kfs"] = {
"Bilaspuri",
12953397,
"him",
"Deva, Takr",
translit = "hi-translit",
}
m["kft"] = {
"Kanjari",
12953610,
"inc-pan",
ancestors = "pa",
}
m["kfu"] = {
"Katkari",
6377671,
"inc-sou",
}
m["kfv"] = {
"Kurmukar",
6446193,
"inc-eas",
}
m["kfw"] = {
"Kharam Naga",
12952906,
"tbq-kuk",
}
m["kfx"] = {
"Kullu Pahari",
6443148,
"him",
"Deva",
translit = "hi-translit",
}
m["kfy"] = {
"Kumaoni",
33529,
"inc-pac",
"Deva, Shrd, Takr",
-- Shrd translit in [[Module:scripts/data]] (NOTE: not present before, presumably an accidental omission)
}
m["kfz"] = {
"Koromfé",
35701,
"nic-gur",
"Latn",
}
m["kga"] = {
"Koyaga",
11155632,
"dmn-mnk",
}
m["kgb"] = {
"Kawe",
12952750,
"poz-hce",
"Latn",
}
m["kgd"] = {
"Kataang",
12953622,
"mkh",
}
m["kge"] = {
"Komering",
49224,
"poz-lgx",
"Latn, Arab",
}
m["kgf"] = {
"Kube",
11732359,
"ngf-kto",
"Latn",
}
m["kgg"] = {
"Kusunda",
33630,
"qfa-iso", -- central Nepal
"Latn",
}
m["kgi"] = {
"Selangor Sign Language",
33731,
"sgn-asl",
}
m["kgj"] = {
"Gamale Kham",
22236996,
"sit-kha",
"Deva",
}
m["kgk"] = {
"Kaiwá",
3111883,
"gn",
"Latn",
}
m["kgl"] = {
"Kunggari",
10550184,
"aus-pam",
"Latn",
}
m["kgn"] = {
"Karingani",
6371041,
"xme-ttc",
"Arab, Latn",
ancestors = "xme-ttc-nor",
}
m["kgo"] = {
"Krongo",
6438927,
"qfa-kad",
"Latn",
}
m["kgp"] = {
"Kaingang",
2665734,
"sai-sje",
"Latn",
}
m["kgq"] = {
"Kamoro",
6359001,
"ngf-ask",
"Latn",
}
m["kgr"] = {
"Abun",
56657,
"qfa-iso", -- Papuan; isolate in Ethnologue, Glottolog and Palmer (2018); grouped with West Papuan by Ross (2005)
"Latn",
}
m["kgs"] = {
"Kumbainggar",
3915412,
"aus-pam",
"Latn",
}
m["kgt"] = {
"Somyev",
3913354,
"nic-mmb",
"Latn",
}
m["kgu"] = {
"Kobol",
11732325,
"ngf-omo",
"Latn",
}
m["kgv"] = {
"Karas",
6368621,
"qfa-dis", -- Divergent Papuan language; grouped with Mbaham-Iha by Glottolog to form a (mainland) West Bomberai
-- family, but with Mbaham-Iha and Timor-Alor-Pantar by Wikipedia (following Usher and Schapper 2022)
-- into a (Greater) West Bomberai family.
"Latn",
}
m["kgw"] = {
"Karon Dori",
56817,
"paa-may",
"Latn",
}
m["kgx"] = {
"Kamaru",
12953604,
"poz-wot",
"Latn",
}
m["kgy"] = {
"Kyerung",
12952691,
"sit-kyk",
}
m["kha"] = {
"Khasi",
33584,
"aav-pkl",
"Latn, as-Beng",
}
m["khb"] = {
"Lü",
36948,
"tai-swe",
"Talu, Lana",
translit = {Talu = "Talu-translit"},
strip_diacritics = {remove_diacritics = c.ZWNJ},
sort_key = {
Talu = "Talu-sortkey",
Lana = "Lana-sortkey",
},
}
m["khc"] = {
"Tukang Besi Utara",
18611555,
"poz",
}
m["khd"] = {
"Bädi Kanum",
20888004,
"paa-ngk",
"Latn",
}
m["khe"] = {
"Korowai",
6432598,
"ngf-bda",
"Latn",
}
m["khf"] = {
"Khuen",
27144893,
"mkh",
}
m["khh"] = {
"Kehu",
10994953,
}
m["khj"] = {
"Kuturmi",
3914490,
"nic-plc",
"Latn",
}
m["khl"] = {
"Lusi",
3267788,
"poz-ocw",
"Latn",
}
m["kho"] = {
"Khotan",
6583551,
"xsc-sak",
"Brah, Khar",
-- Brah translit in [[Module:scripts/data]]
}
m["khp"] = {
"Kapauri",
3502575,
"qfa-dis", -- isolate per Glottolog, possibly Greater Kwerba per Wikipedia in Kapauri-Sause family
"Latn",
}
m["khq"] = {
"Koyra Chiini",
33600,
"son",
"Latn, Arab",
}
m["khr"] = {
"Kharia",
3915562,
"mun",
"Deva, Orya, Latn",
}
m["khs"] = {
"Kasua",
6374863,
"ngf-bos",
"Latn",
}
m["kht"] = {
"Khamti",
3915502,
"tai-swe",
"Mymr",
display_text = s["kht-displaytext"],
strip_diacritics = s["kht-stripdiacritics"],
}
m["khu"] = {
"Nkhumbi",
11019169,
"bnt-swb",
}
m["khv"] = {
"Khvarshi",
56425,
"cau-wts",
"Cyrl",
translit = "khv-translit",
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]},
}
m["khw"] = {
"Khowar",
938216,
"inc-chi",
"Aran",
strip_diacritics = {
-- character "ۂ" code U+06C2 to "ه" and "هٔ" (U+0647 + U+0654) to "ه"; hamzatu l-waṣli to a regular alif
from = {"هٔ", "ۂ", "ٱ"},
to = {"ہ", "ہ", "ا"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
}
m["khx"] = {
"Kanu",
12952571,
"bnt-lgb",
}
m["khy"] = {
"Ekele",
6385549,
"bnt-ske",
"Latn",
}
m["khz"] = {
"Keapara",
12952603,
"poz-ocw",
"Latn",
}
m["kia"] = {
"Kim",
35685,
"alv-kim",
}
m["kib"] = {
"Koalib",
35859,
"alv-hei",
}
m["kic"] = {
"Kickapoo",
20162127,
"alg-sfk",
"Latn",
}
m["kid"] = {
"Koshin",
35632,
"nic-beb",
"Latn",
}
m["kie"] = {
"Kibet",
56893,
}
m["kif"] = {
"Kham Parbate Timur",
12953022,
"sit-kha",
"Deva",
}
m["kig"] = {
"Kimaama",
11732321,
"paa-kol",
"Latn",
}
m["kih"] = {
"Kilmeri",
6408020,
"paa-bew",
"Latn",
}
m["kii"] = {
"Kitsai",
56627,
"cdd",
"Latn",
}
m["kij"] = {
"Kilivila",
3196601,
"poz-ocw",
"Latn",
}
m["kil"] = {
"Kariya",
3438708,
"cdc-wst",
}
m["kim"] = {
"Tofa",
36848,
"trk-ssb",
"Cyrl",
}
m["kio"] = {
"Kiowa",
56631,
"nai-kta",
"Latn",
}
m["kip"] = {
"Sheshi Kham",
12952622,
"sit-kha",
"Deva",
}
m["kiq"] = {
"Kosadle",
6432994,
"paa-kko",
"Latn",
}
m["kis"] = {
"Kis",
6416362,
"poz-ocw",
"Latn",
}
m["kit"] = {
"Agob",
3332143,
"paa-pah",
"Latn",
}
m["kiv"] = {
"Kimbu",
10997740,
"bnt-tkm",
}
m["kiw"] = {
"Kiwai Timur Laut",
11732324,
"paa-kiw",
"Latn",
}
m["kix"] = {
"Naga Khiamniungan",
6401546,
"sit-kch",
"Latn",
}
m["kiy"] = {
"Kirikiri",
6415159,
"paa-wlp",
"Latn",
}
m["kiz"] = {
"Kisi",
3912772,
"bnt-bki",
}
m["kja"] = {
"Mlap",
6885683,
"paa-nim",
"Latn",
}
m["kjb"] = {
"Q'anjob'al",
35551,
"myn",
"Latn",
}
m["kjc"] = {
"Konjo Pesisir",
3198689,
"poz",
"Latn",
}
m["kjd"] = {
"Kiwai Selatan",
11732322,
"paa-kiw",
"Latn",
}
m["kje"] = {
"Kisar",
3197441,
"poz",
"Latn",
}
m["kjg"] = {
"Khmu",
33335,
"mkh",
"Laoo",
ancestors = "mkh-khm-pro",
sort_key = "Laoo-sortkey",
}
m["kjh"] = {
"Khakas",
33575,
"trk-ssb",
"Cyrl",
translit = "kjh-translit",
override_translit = true,
}
m["kji"] = {
"Zabana",
379130,
"poz-ocw",
"Latn",
}
m["kjj"] = {
"Khinalug",
35278,
"cau-nec",
"Cyrl, Latn",
translit = "kjj-translit",
override_translit = true,
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {
Cyrl = s["cau-Cyrl-stripdiacritics"],
Latn = s["cau-Latn-stripdiacritics"],
},
}
m["kjk"] = {
"Highland Konjo",
3198688,
"poz",
}
m["kjl"] = {
"Kham",
22237017,
"sit-kha",
"Deva",
}
m["kjm"] = {
"Kháng",
6403501,
"mkh-pal",
}
m["kjn"] = {
"Kunjen",
3200468,
"aus-pmn",
"Latn",
}
m["kjo"] = {
"Harijan Kinnauri",
5657463,
"him",
"Takr, Deva",
}
m["kjp"] = {
"Pwo Timur",
5330390,
"kar",
"Mymr, Leke, Thai",
translit = "kjp-translit",
override_translit = true,
}
m["kjq"] = {
"Keres Barat",
12645568,
"nai-ker",
"Latn",
}
m["kjr"] = {
"Kurudu",
12952678,
"poz-hce",
"Latn",
}
m["kjs"] = {
"Kewa Timur",
20050949,
"ngf-ank",
"Latn",
}
m["kjt"] = {
"Phrae Pwo",
7187991,
"kar",
"Thai",
}
m["kju"] = {
"Kashaya",
3193689,
"nai-pom",
"Latn",
}
m["kjx"] = {
"Ramopa",
56830,
"paa-nbo",
"Latn",
}
m["kjy"] = {
"Erave",
12952416,
"ngf-ank",
"Latn",
}
m["kjz"] = {
"Bumthangkha",
2786408,
"sit-ebo",
"Tibt",
override_translit = true,
-- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["kka"] = {
"Kakanda",
3915342,
"alv-ngb",
}
m["kkb"] = {
"Kwerisa",
56881,
"paa-clp",
"Latn",
}
m["kkc"] = {
"Odoodee",
12952987,
"ngf-est",
"Latn",
}
m["kkd"] = {
"Kinuku",
6414422,
"nic-kau",
}
m["kke"] = {
"Kakabe",
3913966,
"dmn-mok",
"Latn",
}
m["kkf"] = {
"Monpa Kalaktang",
63257089,
"sit-tsk",
"Tibt, Latn, Deva",
override_translit = true,
-- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["kkg"] = {
"Kalinga Lembah Mabaka",
18753304,
"phi",
}
m["kkh"] = {
"Khün",
3545044,
"tai-swe",
"Lana, Thai",
sort_key = {
Lana = "Lana-sortkey",
Thai = "Thai-sortkey"
},
}
m["kki"] = {
"Kagulu",
12952537,
"bnt-ruv",
"Latn",
}
m["kkj"] = {
"Kako",
35755,
"bnt-kak",
}
m["kkk"] = {
"Kokota",
3198399,
"poz-ocw",
"Latn",
}
m["kkl"] = {
"Kosarek Yale",
6432995,
"ngf-mek",
"Latn",
}
m["kkm"] = {
"Kiong",
6414512,
"nic-ucr",
"Latn",
}
m["kkn"] = {
"Kon Keu",
6428686,
"mkh-pal",
}
m["kko"] = {
"Karko",
35529,
"nub-hil",
"Latn",
}
m["kkp"] = {
"Koko-Bera",
6426699,
"aus-pmn",
"Latn",
}
m["kkq"] = {
"Kaiku",
6347840,
"bnt-kbi",
"Latn",
}
m["kkr"] = {
"Kir-Balar",
3440527,
"cdc-wst",
"Latn",
}
m["kks"] = {
"Kirfi",
56242,
"cdc-wst",
"Latn",
}
m["kkt"] = {
"Koi",
6426194,
"sit-kiw",
}
m["kku"] = {
"Tumi",
3913934,
"nic-kau",
}
m["kkv"] = {
"Kangean",
2071325,
"poz-msa",
"Latn",
}
m["kkw"] = {
"Teke-Kukuya",
36560,
"bnt-tek",
}
m["kkx"] = {
"Kohin",
6425997,
"poz-brw",
}
m["kky"] = {
"Guugu Yimidhirr",
56543,
"aus-pam",
"Latn",
}
m["kkz"] = {
"Kaska",
20823,
"ath-nor",
"Latn",
}
m["kla"] = {
"Klamath-Modoc",
2669248,
"nai-plp",
"Latn",
}
m["klb"] = {
"Kiliwa",
3182593,
"nai-yuc",
"Latn",
}
m["klc"] = {
"Kolbila",
6427122,
"alv-lek",
}
m["kld"] = {
"Gamilaraay",
3111818,
"aus-cww",
"Latn",
}
m["kle"] = {
"Kulung",
6443304,
"sit-kic",
}
m["klf"] = {
"Kendeje",
56895,
}
m["klg"] = {
"Kalagan Tagakaulu",
18756514,
"phi",
"Latn",
}
m["klh"] = {
"Weliki",
7981017,
"ngf-uru",
"Latn",
}
m["kli"] = {
"Kalumpang",
13561407,
"poz",
}
m["klj"] = {
"Khalaj",
33455,
"trk",
"Arab, Latn",
ancestors = "klj-arg",
strip_diacritics = {
remove_diacritics = c.kashida .. c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun,
}
}
m["klk"] = {
"Kono (Nigeria)",
6429589,
"nic-kau",
"Latn",
}
m["kll"] = {
"Kalagan Kagan",
18748913,
"phi",
}
m["klm"] = {
"Kolom",
6844970,
"ngf-rai",
"Latn",
}
m["kln"] = {
"Kalenjin",
637228,
"sdv-nma",
"Latn",
}
m["klo"] = {
"Kapya",
6367410,
"nic-ykb",
}
m["klp"] = {
"Kamasa",
6356107,
"ngf-woj",
"Latn",
}
m["klq"] = {
"Rumu",
7379420,
"paa-tki",
"Latn",
}
m["klr"] = {
"Khaling",
56381,
"sit-kiw",
"Deva",
}
m["kls"] = {
"Kalasha",
33416,
"inc-chi",
"Latn, Aran",
}
m["klt"] = {
"Nukna",
7068874,
"ngf-uru",
"Latn",
}
m["klu"] = {
"Klao",
3914866,
"kro-wkr",
}
m["klv"] = {
"Maskelynes",
3297282,
"poz-vnc",
"Latn",
}
m["klw"] = {
"Lindu",
18390055,
"poz-kal",
"Latn",
}
m["klx"] = {
"Koluwawa",
6427954,
"poz-ocw",
"Latn",
}
m["kly"] = {
"Kalao",
6350643,
"poz-wot",
"Latn",
}
m["klz"] = {
"Kabola",
11732258,
"paa-alp",
"Latn",
}
m["kma"] = {
"Konni",
35680,
"nic-buk",
}
m["kmb"] = {
"Kimbundu",
35891,
"bnt-kmb",
"Latn",
}
m["kmc"] = {
"Kam Selatan",
35379,
"qfa-kms",
"Latn",
}
m["kmd"] = {
"Kalinga Madukayang",
18753305,
"phi",
}
m["kme"] = {
"Bakole",
35068,
"bnt-kpw",
"Latn",
}
m["kmf"] = {
"Kare (New Guinea)",
11732286,
"ngf-mab",
"Latn",
}
m["kmg"] = {
"Kâte",
3201059,
"ngf-kma",
"Latn",
}
m["kmh"] = {
"Kalam",
12952550,
"ngf-kak",
"Latn",
}
m["kmi"] = {
"Kami",
3915372,
"alv-ngb",
"Latn",
}
m["kmj"] = {
"Kumarbhag Paharia",
3130374,
"dra-mlo",
"Beng, Deva",
}
m["kmk"] = {
"Kalinga Limos",
18753303,
"phi",
"Latn",
}
m["kml"] = {
"Kalinga Tanudan",
18753307,
"phi",
"Latn",
}
m["kmm"] = {
"Kom (India)",
12952647,
"tbq-kuk",
}
m["kmn"] = {
"Awtuw",
3504217,
"paa-sep",
"Latn",
}
m["kmo"] = {
"Kwoma",
11732376,
"paa-sep",
"Latn",
}
m["kmp"] = {
"Gimme",
11152236,
"alv-dur",
}
m["kmq"] = {
"Kwama",
2591184,
"ssa-kom",
}
m["kmr"] = {
"Kurdi Utara",
36163,
"ku",
"Latn, Cyrl, Armn, Arab, Yezi",
translit = {
Cyrl = "kmr-translit",
-- Armn translit in [[Module:scripts/data]]
Arab = "ckb-translit",
},
strip_diacritics = {
Latn = {
remove_diacritics = "'’",
from = {"r̄", "R̄", "ẍ", "Ẍ"},
to = {"rr", "Rr", "x", "X"}
},
},
wikimedia_codes = "ku",
}
m["kms"] = {
"Kamasau",
6356117,
"paa-mar",
"Latn",
}
m["kmt"] = {
"Kemtuik",
6387179,
"paa-nim",
"Latn",
}
m["kmu"] = {
"Kanite",
12952567,
"ngf-kya",
"Latn",
}
m["kmv"] = {
"Perancis Kreol Karipúna",
2523999,
"crp",
"Latn",
ancestors = "fr",
sort_key = s["roa-oil-sortkey"],
}
m["kmw"] = {
"Kumu",
6428450,
"bnt-kbi",
"Latn",
}
m["kmx"] = {
"Waboda",
7958705,
"paa-kiw",
"Latn",
}
m["kmy"] = {
"Koma",
35634,
"alv-dur",
}
m["kmz"] = {
"Turki Khorasan",
35373,
"trk-ogz",
"Arab",
ancestors = "trk-oat",
}
m["kna"] = {
"Kanakuru",
56811,
"cdc-wst",
"Latn",
}
m["knb"] = {
"Kalinga Lubuagan",
12953602,
"phi",
"Latn",
}
m["knd"] = {
"Konda",
11732340,
"ngf-sbh",
"Latn",
}
m["kne"] = {
"Kankanaey",
18753329,
"phi",
"Latn",
strip_diacritics = {
Latn = {
remove_diacritics = c.grave .. c.acute .. c.circ .. c.diaer,
}
},
sort_key = {
Latn = "tl-sortkey",
},
standard_chars = {
Latn = "AaBbKkDdEeGgHhIiLlMmNnOoPpRrSsTtUuWwYy" .. c.punc,
},
}
m["knf"] = {
"Mankanya",
35789,
"alv-pap",
"Latn",
}
m["kni"] = {
"Kanufi",
3913297,
"nic-nin",
"Latn",
}
m["knj"] = {
"Akatek",
34923,
"myn",
"Latn",
}
m["knk"] = {
"Kuranko",
3198896,
"dmn-mok",
"Latn",
}
m["knl"] = {
"Keninjal",
6389309,
"poz-mly",
"Latn",
}
m["knm"] = { -- two unrelated lects have this name; this is the Katukinian one
"Kanamari",
3438373,
"sai-ktk",
"Latn",
}
m["kno"] = {
"Kono (Sierra Leone)",
35675,
"dmn-vak",
"Latn",
}
m["knp"] = {
"Kwanja",
35641,
"nic-mmb",
"Latn",
}
m["knq"] = {
"Kintaq",
6414335,
"mkh-asl",
"Latn",
}
m["knr"] = {
"Kaningra",
6363253,
"paa-sep",
"Latn",
}
m["kns"] = {
"Kensiu",
6391529,
"mkh-asl",
"Latn, Thai, Hani",
}
m["knt"] = {
"Katukina",
3194265,
"sai-pan",
"Latn",
}
m["knu"] = { -- a dialect of 'kpe'
"Kono (Guinea)",
3198703,
"dmn-msw",
"Latn, Kpel",
ancestors = "kpe",
}
m["knv"] = {
"Tabo",
7959888,
"aav",
}
m["knx"] = {
"Salako",
6388963,
"poz-mly",
"Latn",
}
m["kny"] = {
"Kanyok",
11110766,
"bnt-lub",
"Latn",
}
m["knz"] = {
"Kalamsé",
3914000,
"nic-gnn",
"Latn",
}
m["koa"] = {
"Konomala",
3198732,
"poz-ocw",
"Latn",
}
m["koc"] = {
"Kpati",
3913279,
"nic-nge",
"Latn",
}
m["kod"] = {
"Kodi",
4577633,
"poz-cet",
"Latn",
}
m["koe"] = {
"Kacipo-Balesi",
5364424,
"sdv",
"Latn",
}
m["kof"] = {
"Kubi",
3438718,
"cdc-wst",
"Latn",
}
m["kog"] = {
"Cogui",
3198286,
"cba",
"Latn",
}
m["koh"] = {
"Koyo",
35649,
"bnt-mbo",
"Latn",
}
m["koi"] = {
"Komi-Permyak",
56318,
"kv",
"Cyrl",
translit = "kv-translit",
strip_diacritics = {remove_diacritics = c.acute},
override_translit = true,
}
m["kok"] = {
"Konkani",
34239,
"inc-sou",
"Deva, Knda, Mlym, Arab, Latn",
translit = {
Deva = "mr-translit",
},
-- Knda translit in [[Module:scripts/data]]
-- Mlym translit in [[Module:scripts/data]]
strip_diacritics = {
-- FIXME: Separate out the scripts
from = {"च़", "ज़", "झ़", "ಚ಼", "ಜ಼", "ಝ಼"},
to = {"च", "ज", "झ", "ಚ", "ಜ", "ಝ"}
} ,
}
m["kol"] = {
"Kol (New Guinea)",
4227542,
}
m["koo"] = {
"Konzo",
2361829,
"bnt-glb",
"Latn",
}
m["kop"] = {
"Waube",
11732373,
"ngf-nur",
"Latn",
}
m["koq"] = {
"Kota (Gabon)",
35607,
"bnt-kel",
"Latn",
}
m["kos"] = {
"Kosrae",
33464,
"poz-mic",
"Latn",
}
m["kot"] = {
"Lagwan",
3502264,
"cdc-cbm",
"Latn",
}
m["kou"] = {
"Koke",
797249,
"alv-bua",
}
m["kov"] = {
"Kudu-Camo",
3915850,
"nic-jer",
}
m["kow"] = {
"Kugama",
3913307,
"alv-mye",
}
m["koy"] = {
"Koyukon",
28304,
"ath-nor",
"Latn",
}
m["koz"] = {
"Korak",
6431365,
"ngf-kow",
"Latn",
}
m["kpa"] = {
"Kutto",
3437656,
"cdc-wst",
}
m["kpb"] = {
"Mullu Kurumba",
19573111,
"dra-mal",
}
m["kpc"] = {
"Curripaco",
2882543,
"awd-nwk",
"Latn",
}
m["kpd"] = {
"Koba",
6424249,
"poz",
}
m["kpe"] = {
"Kpelle",
35673,
"dmn-msw",
"Latn, Kpel",
}
m["kpf"] = {
"Komba",
6428239,
"ngf-kab",
"Latn",
}
m["kpg"] = {
"Kapingamarangi",
35771,
"poz-pnp",
"Latn",
}
m["kph"] = {
"Kplang",
35628,
"alv-gng",
}
m["kpi"] = {
"Kofei",
6425665,
"paa-egb",
"Latn",
}
m["kpj"] = {
"Karajá",
10322066,
"sai-mje",
"Latn",
}
m["kpk"] = {
"Kpan",
3915380,
"nic-jkn",
"Latn",
}
m["kpl"] = {
"Kpala",
11154769,
"nic-nkk",
"Latn",
}
m["kpm"] = {
"Koho",
3511919,
"mkh-ban",
"Latn",
}
m["kpn"] = {
"Kepkiriwát",
3195366,
"tup",
"Latn",
}
m["kpo"] = {
"Ikposo",
35029,
"alv-ktg",
"Latn",
}
m["kpq"] = {
"Korupun-Sela",
6432769,
"ngf-mek",
"Latn",
}
m["kpr"] = {
"Korafe-Yegha",
11732347,
"ngf-gko",
"Latn",
}
m["kps"] = {
"Tehit",
7694851,
"paa-wbh",
"Latn",
}
m["kpt"] = {
"Karata",
56636,
"cau-and",
"Cyrl",
translit = "kpt-translit",
override_translit = true,
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]},
}
m["kpu"] = {
"Kafoa",
6346151,
"paa-alp",
"Latn",
}
m["kpv"] = {
"Komi-Zyrian",
34114,
"kv",
"Cyrl",
translit = "kv-translit",
override_translit = true,
wikimedia_codes = "kv",
}
m["kpw"] = {
"Kobon",
11732326,
"ngf-kak",
"Latn",
}
m["kpx"] = {
"Koiari Gunung",
6925030,
"ngf-koi",
"Latn",
}
m["kpy"] = {
"Koryak",
36199,
"qfa-ckn",
"Cyrl",
strip_diacritics = {
from = {"['’]"},
to = {"ʼ"}
},
sort_key = {
from = {"вʼ", "гʼ", "ё", "ӄ", "ӈ"},
to = {"в" .. p[1], "г" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1]}
},
translit = "kpy-translit",
}
m["kpz"] = {
"Kupsabiny",
56445,
"sdv-kln",
"Latn",
}
m["kqa"] = {
"Mum",
6935252,
"ngf-nso",
"Latn",
}
m["kqb"] = {
"Kovai",
6434822,
"ngf-ehu",
"Latn",
}
m["kqc"] = {
"Doromu-Koki",
5298175,
"paa-man",
"Latn",
}
m["kqd"] = {
"Koy Sanjaq Surat",
33463,
"sem-nna",
}
m["kqe"] = {
"Kalagan",
18748906,
"phi",
"Latn",
}
m["kqf"] = {
"Kakabai",
6349119,
"poz-ocw",
"Latn",
}
m["kqg"] = {
"Khe",
3914015,
"nic-gur",
"Latn",
}
m["kqh"] = {
"Kisankasa",
6416409,
"sdv",
}
m["kqi"] = {
"Koitabu",
6426363,
"ngf-koi",
"Latn",
}
m["kqj"] = {
"Koromira",
6432520,
"paa-sbo",
"Latn",
}
m["kqk"] = {
"Gbe Kotafon",
12952447,
"alv-pph",
}
m["kql"] = {
"Kyenele",
11732453,
"paa-yua",
"Latn",
}
m["kqm"] = {
"Khisa",
3913955,
"nic-gur",
"Latn",
}
m["kqn"] = {
"Kaonde",
33601,
"bnt-lub",
"Latn",
}
m["kqo"] = {
"Krahn Timur",
3915374,
"kro-wee",
"Latn",
}
m["kqp"] = {
"Kimré",
3441210,
"cdc-est",
"Latn",
}
m["kqq"] = {
"Krenak",
6436747,
"sai-cer",
"Latn",
}
m["kqr"] = {
"Kimaragang",
3196845,
"poz-san",
"Latn",
}
m["kqs"] = {
"Kissi Utara",
19921576,
"alv-kis",
}
m["kqt"] = {
"Kadazan Sungai Klias",
12953594,
"poz-san",
}
m["kqu"] = {
"Seroa",
33127766,
"khi-tuu",
"Latn",
}
m["kqv"] = {
"Okolod",
7082487,
"poz-san",
}
m["kqw"] = {
"Kandas",
3192590,
"poz-ocw",
"Latn",
}
m["kqx"] = {
"Mser",
3502347,
"cdc-cbm",
}
m["kqy"] = {
"Koorete",
6430753,
"omv-eom",
"Ethi, Latn",
}
m["kqz"] = {
"Korana",
2756709,
"khi-khk",
"Latn",
}
m["kra"] = {
"Kumhali",
13580783,
"inc-bih",
"Deva",
translit = "ne-translit",
}
m["krb"] = {
"Karkin",
3193345,
"nai-utn",
"Latn",
}
m["krc"] = {
"Karachay-Balkar",
33714,
"trk-kcu",
"Cyrl",
translit = "krc-translit",
sort_key = {
from = {"гъ", "дж", "ё", "къ", "нг"},
to = {"г" .. p[1], "д" .. p[1], "е" .. p[1], "к" .. p[1], "н" .. p[1]}
},
}
m["krd"] = {
"Kairui-Midiki",
12953277,
"poz-tim",
}
m["kre"] = {
"Panará",
3361895,
"sai-cer",
"Latn",
}
m["krf"] = {
"Koro (Vanuatu)",
3198995,
"poz-vnn",
"Latn",
}
m["krh"] = {
"Kurama",
35593,
"nic-kau",
}
m["kri"] = {
"Krio",
35744,
"crp",
"Latn",
ancestors = "en",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ},
sort_key = {
from = {"ɛ", "gb", "kp", "ɔ"},
to = {"e" .. p[1], "g" .. p[1], "k" .. p[1], "o" .. p[1]}
},
}
m["krj"] = {
"Kinaray-a",
33720,
"phi",
"Latn",
}
m["krk"] = {
"Kerek",
332792,
"qfa-ckn",
"Cyrl",
}
m["krl"] = {
"Karelia",
33557,
"urj-fin",
"Latn",
sort_key = {
from = {
"č", "š", "ž", "ü", "ä", "ö", -- 2 chars
"z", "'" -- 1 char
},
to = {
"c" .. p[1], "s" .. p[1], "s" .. p[3], "y" .. p[1], "y" .. p[2], "y" .. p[3],
"s" .. p[2], "y" .. p[4],
}
},
}
m["krm"] = {
"Krim",
35713,
"alv",
}
m["krn"] = {
"Sapo",
3915386,
"kro-wee",
}
m["krp"] = {
"Korop",
35626,
"nic-ucr",
"Latn",
}
m["krr"] = {
"Kru'ng",
12953650,
"mkh-ban",
}
m["krs"] = {
"Kresh",
56674,
"csu-bkr",
}
m["kru"] = {
"Kurukh",
33492,
"dra-kml",
"Deva, Tols",
translit = {
Deva = "hi-translit",
},
}
m["krv"] = {
"Kavet",
12953649,
"sai-ktk",
"Latn",
}
m["krw"] = {
"Krahn Barat",
10975611,
"kro-wee",
"Latn",
}
m["krx"] = {
"Karon",
35704,
"alv-jol",
}
m["kry"] = {
"Kryts",
35861,
"cau-ssm",
"Latn, Cyrl",
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {
Latn = s["cau-Latn-stripdiacritics"],
Cyrl = s["cau-Cyrl-stripdiacritics"],
},
}
m["krz"] = {
"Sota Kanum",
12952568,
"paa-kan",
"Latn",
}
m["ksa"] = {
"Shuwa-Zamani",
3913929,
"nic-kau",
}
m["ksb"] = {
"Shambala",
3788739,
"bnt-seu",
"Latn",
}
m["ksc"] = {
"Kalinga Selatan",
18753301,
"phi",
}
m["ksd"] = {
"Tolai",
35870,
"poz-ocw",
"Latn",
}
m["kse"] = {
"Kuni",
6444619,
"poz-ocw",
"Latn",
}
m["ksf"] = {
"Bafia",
34930,
"bnt-baf",
"Latn",
}
m["ksg"] = {
"Kusaghe",
3200638,
"poz-ocw",
"Latn",
}
m["ksi"] = {
"Krisa",
841704,
"paa-sko",
"Latn",
}
m["ksj"] = {
"Uare",
6450052,
"paa-kwa",
"Latn",
}
m["ksk"] = {
"Kansa",
3192772,
"sio-dhe",
"Latn",
}
m["ksl"] = {
"Kumalu",
17584381,
"poz-ocw",
"Latn",
}
m["ksm"] = {
"Kumba",
3913972,
"alv-mye",
}
m["ksn"] = {
"Kasiguranin",
6374525,
"phi",
}
m["kso"] = {
"Kofa",
56278,
"cdc-cbm",
}
m["ksp"] = {
"Kaba",
3915316,
"csu-sar",
}
m["ksq"] = {
"Kwaami",
3440525,
"cdc-wst",
}
m["ksr"] = {
"Borong",
4946263,
"ngf-kbm",
"Latn",
}
m["kss"] = {
"Kissi Selatan",
11028974,
"alv-kis",
}
m["kst"] = {
"Winyé",
3913360,
"nic-gnw",
}
m["ksu"] = {
"Khamyang",
6583541,
"tai-swe",
}
m["ksv"] = {
"Kusu",
6448199,
"bnt-tet",
}
m["ksw"] = {
"Karen S'gaw",
56410,
"kar",
"Mymr",
translit = "ksw-translit",
}
m["ksx"] = {
"Kedang",
6382520,
"poz",
"Latn",
}
m["ksy"] = {
"Kharia Thar",
6400661,
"inc-eas",
}
m["ksz"] = {
"Kodaku",
21179986,
"mun",
}
m["kta"] = {
"Katua",
6378404,
"mkh-ban",
}
m["ktb"] = {
"Kambaata",
35664,
"cus-hec",
"Latn",
}
m["ktc"] = {
"Kholok",
3440464,
"cdc-wst",
}
m["ktd"] = {
"Kokata",
10547021,
"aus-pam",
"Latn",
}
m["ktf"] = {
"Kwami",
12952687,
"bnt-lgb",
}
m["ktg"] = {
"Kalkatungu",
3914057,
"aus-pam",
"Latn",
}
m["kth"] = {
"Karanga",
713643,
}
m["kti"] = {
"Muyu Utara",
20857698,
"ngf-lok",
"Latn",
}
m["ktj"] = {
"Plapo Krumen",
10975356,
"kro-grb",
}
m["ktk"] = {
"Kaniet",
3399050,
"poz-aay",
"Latn",
}
m["ktl"] = {
"Koroshi",
3775265,
"ira-nwi",
ancestors = "bal",
}
m["ktm"] = {
"Kurti",
3200615,
"poz-aay",
"Latn",
}
m["ktn"] = {
"Karitiâna",
3112184,
"tup",
"Latn",
}
m["kto"] = {
"Kuot",
56537,
}
m["ktp"] = {
"Kaduo",
769809,
"tbq-bka",
}
m["ktq"] = {
"Katabaga",
3193895,
}
m["kts"] = {
"Muyu Selatan",
42308820,
"ngf-lok",
"Latn",
}
m["ktt"] = {
"Ketum",
12952616,
"ngf-dum",
"Latn",
}
m["ktu"] = {
"Kituba",
35746,
"crp",
"Latn",
ancestors = "kg",
}
m["ktv"] = {
"Katu Timur",
22808951,
"mkh-kat",
"Latn",
}
m["ktw"] = {
"Kato",
20831,
"ath-pco",
"Latn",
}
m["ktx"] = {
"Kaxararí",
6380124,
"sai-pan",
"Latn",
}
m["kty"] = {
"Kango",
6362818,
"bnt-bta",
"Latn",
}
m["ktz"] = {
"Juǀ'hoan",
1192295,
"khi-kxa",
"Latn",
}
m["kub"] = {
"Kutep",
35645,
"nic-jkn",
}
m["kuc"] = {
"Kwinsu",
6450460,
"paa-tor",
"Latn",
}
m["kud"] = {
"Auhelawa",
5166,
"poz-ocw",
"Latn",
}
m["kue"] = {
"Kuman",
137525,
"ngf-sim",
"Latn",
}
m["kuf"] = {
"Katu Barat",
6378400,
"mkh-kat",
"Laoo, Tale, Latn",
}
m["kug"] = {
"Kupa",
3915336,
"alv-ngb",
}
m["kuh"] = {
"Kushi",
3438747,
"cdc-wst",
}
m["kui"] = {
"Kuikúro",
3915522,
"sai-kui",
"Latn",
}
m["kuj"] = {
"Kuria",
6445968,
"bnt-lok",
"Latn",
}
m["kuk"] = {
"Kepo'",
6393217,
"poz",
}
m["kul"] = {
"Kulere",
3440506,
"cdc-wst",
}
m["kum"] = {
"Kumyk",
36209,
"trk-kcu",
"Cyrl",
translit = "kum-translit",
sort_key = {
from = {"гъ", "гь", "ё", "къ", "нг", "оь", "уь"},
to = {"г" .. p[1], "г" .. p[2], "е" .. p[1], "к" .. p[1], "н" .. p[1], "о" .. p[1], "у" .. p[1]}
},
}
m["kun"] = {
"Kunama",
36041,
}
m["kuo"] = {
"Kumukio",
11732362,
"ngf-dal",
"Latn",
}
m["kup"] = {
"Kunimaipa",
6444696,
"paa-kun",
"Latn",
}
m["kuq"] = {
"Karipuna",
6371071,
"tup-gua",
"Latn",
}
m["kus"] = {
"Kusaal",
35708,
"nic-dag",
"Latn",
}
m["kut"] = {
"Kutenai",
33434,
"qfa-iso",
"Latn",
}
m["kuu"] = {
"Upper Kuskokwim",
28062,
"ath-nor",
"Latn",
}
m["kuv"] = {
"Kur",
12635082,
"poz-cma",
"Latn",
}
m["kuw"] = {
"Kpagua",
11137573,
"bad-cnt",
}
m["kux"] = {
"Kukatja",
10549839,
"aus-pam",
"Latn",
}
m["kuy"] = {
"Kuuku-Ya'u",
10550697,
"aus-pmn",
"Latn",
}
m["kuz"] = {
"Kunza",
2669181,
"qfa-iso",
"Latn",
}
m["kva"] = {
"Bagvalal",
56638,
"cau-and",
"Cyrl",
translit = "cau-nec-translit",
override_translit = true,
display_text = {Cyrl = s["cau-Cyrl-displaytext"]},
strip_diacritics = {Cyrl = s["cau-Cyrl-stripdiacritics"]},
}
m["kvb"] = {
"Kubu",
6441341,
"poz-mly",
}
m["kvc"] = {
"Kove",
3199402,
"poz-ocw",
"Latn",
}
m["kvd"] = {
"Kui (Indonesia)",
6442230,
"paa-alp",
"Latn",
}
m["kve"] = {
"Kalabakan",
6350003,
"poz-san",
"Latn",
}
m["kvf"] = {
"Kabalai",
3440427,
"cdc-est",
}
m["kvg"] = {
"Kuni-Boazi",
2907551,
"paa-boa",
"Latn",
}
m["kvh"] = {
"Komodo",
3198565,
"poz-cet",
"Latn",
}
m["kvi"] = {
"Kwang",
3440398,
"cdc-est",
"Latn",
}
m["kvj"] = {
"Psikye",
56304,
"cdc-cbm",
}
m["kvk"] = {
"Bahasa Isyarat Korea",
3073428,
"sgn-jsl",
}
m["kvl"] = {
"Karen Brek",
12952577,
"kar",
}
m["kvm"] = {
"Kendem",
35751,
"nic-mam",
"Latn",
}
m["kvn"] = {
"Kuna Sempadan",
31777873,
"cba",
}
m["kvo"] = {
"Dobel",
5286559,
"poz",
"Latn",
}
m["kvp"] = {
"Kompane",
18343041,
"poz",
}
m["kvq"] = {
"Karen Geba",
12952581,
"kar",
"Latn, Mymr",
}
m["kvr"] = {
"Kerinci",
3195442,
"poz-mly",
"Latn, Arab", -- Also Incung, which we don't have
}
m["kvt"] = {
"Karen Lahta",
12952582,
"kar",
}
m["kvu"] = {
"Karen Yinbaw",
14426328,
"kar",
}
m["kvv"] = {
"Kola",
6426967,
"poz",
"Latn",
}
m["kvw"] = {
"Wersing",
7983599,
"paa-alp",
"Latn",
}
m["kvx"] = {
"Parkari Koli",
3244176,
"inc-wes",
}
m["kvy"] = {
"Karen Yintale",
14426329,
"kar",
}
m["kvz"] = {
"Tsakwambo",
7849438,
"ngf-kts",
"Latn",
}
m["kwa"] = {
"Dâw",
3042278,
"sai-nad",
"Latn",
}
m["kwb"] = {
"Baa",
34842,
"alv-ada",
}
m["kwc"] = {
"Likwala",
35597,
"bnt-mbo",
}
m["kwd"] = {
"Kwaio",
3200796,
"poz-sls",
"Latn",
}
m["kwe"] = {
"Kwerba",
6450328,
"paa-kwe",
"Latn",
}
m["kwf"] = {
"Kwara'ae",
3200829,
"poz-sls",
"Latn",
}
m["kwg"] = {
"Sara Kaba Deme",
3915384,
"csu-kab",
}
m["kwh"] = {
"Kowiai",
6435028,
"poz",
"Latn",
}
m["kwi"] = {
"Awa-Cuaiquer",
2603103,
"sai-bar",
"Latn",
}
m["kwj"] = {
"Kwanga",
3438383,
"paa-sep",
"Latn",
}
m["kwk"] = {
"Kwak'wala",
2640628,
"wak",
"Latn",
}
m["kwl"] = {
"Kofyar",
3441382,
"cdc-wst",
"Latn",
}
m["kwm"] = {
"Kwambi",
3487165,
"bnt-ova",
}
m["kwn"] = {
"Kwangali",
36334,
"bnt-kav",
"Latn",
}
m["kwo"] = {
"Kwomtari",
3508116,
"paa-kwo",
"Latn",
}
m["kwp"] = {
"Kodia",
3914867,
"kro-ekr",
}
m["kwq"] = {
"Kwak",
11014183,
"nic-nka",
ancestors = "yam",
}
m["kwr"] = {
"Kwer",
12635137,
"ngf-wok",
"Latn",
}
m["kws"] = {
"Kwese",
3200846,
"bnt-pen",
}
m["kwt"] = {
"Kwesten",
6450354,
"paa-tor",
"Latn",
}
m["kwu"] = {
"Kwakum",
35624,
"bnt-kak",
}
m["kwv"] = {
"Sara Kaba Náà",
3915361,
"csu-kab",
"Latn",
}
m["kww"] = {
"Kwinti",
721182,
"crp",
"Latn",
ancestors = "en"
}
m["kwx"] = {
"Khirwar",
12976968,
"dra",
}
m["kwz"] = {
"Kwadi",
2364661,
"khi-kkw",
"Latn",
}
m["kxa"] = {
"Kairiru",
3398785,
"poz-ocw",
"Latn",
}
m["kxb"] = {
"Krobu",
35586,
"alv-ptn",
"Latn",
}
m["kxc"] = {
"Khonso",
56624,
"cus-eas",
"Ethi, Latn",
}
m["kxd"] = {
"Melayu Brunei",
3182878,
"poz-mly",
"Latn, Arab",
}
m["kxe"] = {
"Kakihum",
3914433,
"nic-kam",
ancestors = "tvd",
}
m["kxf"] = {
"Karen Manumanaw",
12952592,
"kar",
"Mymr, Latn",
}
m["kxh"] = {
"Karo",
3447116,
"omv-aro",
}
m["kxi"] = {
"Murut Keningau",
6389308,
"poz-san",
"Latn",
}
m["kxj"] = {
"Kulfa",
713654,
"csu-kab",
}
m["kxk"] = {
"Karen Zayein",
14352960,
"kar",
}
-- Nepali Kurux [kxl] treated as part of Kurux [kru], consistent with ISO merger in 2020
m["kxm"] = {
"Khmer Utara",
3502234,
"mkh-kmr",
"Thai, Khmr",
ancestors = "xhm",
sort_key = {
from = {"[%pๆ]", "[็-๎]", "([เแโใไ])([ก-ฮ])"},
to = {"", "", "%2%1"}
},
}
m["kxn"] = {
"Kanowit",
6364300,
"poz-bnn",
"Latn",
}
m["kxo"] = {
"Kanoé",
4356223,
"qfa-iso",
"Latn",
}
m["kxp"] = {
"Wadiyara Koli",
12953645,
"inc-wes",
}
m["kxq"] = {
"Smärky Kanum",
12952569,
"paa-kan",
"Latn",
}
m["kxr"] = {
"Koro (New Guinea)",
3198994,
"poz-aay",
"Latn",
}
m["kxs"] = {
"Kangjia",
3182570,
"xgn-shr",
"Latn",
}
m["kxt"] = {
"Koiwat",
6426388,
"paa-nnd",
"Latn",
}
m["kxu"] = {
"Kui (India)",
33919,
"dra-kki",
"Orya",
translit = "kxv-translit",
strip_diacritics = {
remove_diacritics = "୕",
from = {"ଆଆ", "ଇଇ", "ଉଉ", "ଏଏ", "ଓଓ", "ିଇ", "ୁଉ", "େଏ", "ୋଓ"},
to = {"ଆ", "ଈ", "ଊ", "ଏ", "ଓ", "ୀ", "ୂ", "େ", "ୋ"},
},
}
m["kxv"] = {
"Kuvi",
3200721,
"dra-kki",
"Orya",
translit = "kxv-translit",
strip_diacritics = {
remove_diacritics = "୕",
from = {"ଆଆ", "ଇଇ", "ଉଉ", "ଏଏ", "ଓଓ", "([କ-ହ])ଆ", "ିଇ", "ୁଉ", "େଏ", "ୋଓ"},
to = {"ଆ", "ଈ", "ଊ", "ଏ", "ଓ", "%1ା", "ୀ", "ୂ", "େ", "ୋ"},
},
}
m["kxw"] = {
"Konai",
11732339,
"ngf-est",
"Latn",
}
m["kxx"] = {
"Likuba",
35646,
"bnt-bmo",
}
m["kxy"] = {
"Kayong",
6380673,
"mkh",
}
m["kxz"] = {
"Kerewo",
6393847,
"paa-kiw",
"Latn",
}
m["kya"] = {
"Kwaya",
6450276,
"bnt-haj",
"Latn",
}
m["kyb"] = {
"Kalinga Butbut",
18753300,
"phi",
"Latn",
}
m["kyc"] = {
"Kyaka",
12952690,
"ngf-enc",
"Latn",
}
m["kyd"] = {
"Karey",
6370196,
"poz",
}
m["kye"] = {
"Krache",
35658,
"alv-gng",
}
m["kyf"] = {
"Kouya",
35595,
"kro-bet",
}
m["kyg"] = {
"Keyagana",
6398208,
"ngf-kya",
"Latn",
}
m["kyh"] = {
"Karok",
1288440,
"qfa-iso", -- or Hokan?
"Latn",
}
m["kyi"] = {
"Kiput",
3038653,
"poz-swa",
"Latn",
}
m["kyj"] = {
"Karao",
3192950,
"phi",
"Latn",
}
m["kyk"] = {
"Kamayo",
3192339,
"phi",
"Latn",
}
m["kyl"] = {
"Kalapuya",
3192120,
"nai-klp",
}
m["kym"] = {
"Kpatili",
3913982,
"znd",
}
m["kyn"] = {
"Karolanos",
6373093,
"phi",
}
m["kyo"] = {
"Kelon",
6386414,
"paa-alp",
"Latn",
}
m["kyp"] = {
"Kang",
25559558,
"tai",
}
m["kyq"] = {
"Kenga",
35707,
"csu-bgr",
}
m["kyr"] = {
"Kuruáya",
3200633,
"tup",
"Latn",
}
m["kys"] = {
"Kayan Baram",
2883794,
"poz",
"Latn",
}
m["kyt"] = {
"Kayagar",
6380394,
"paa-kay",
"Latn",
}
m["kyu"] = {
"Kayah Barat",
12952596,
"kar",
"Kali, Mymr, Latn",
translit = {Kali = "Kali-translit"},
}
m["kyv"] = {
"Kayort",
6380675,
"inc-krd",
"Deva",
}
m["kyw"] = {
"Kudmali",
6446173,
"inc-sad",
"Deva, as-Beng, Orya, Chis",
}
m["kyx"] = {
"Rapoisi",
7294279,
"paa-nbo",
"Latn",
}
m["kyy"] = {
"Kambaira",
6356254,
"ngf-kai",
"Latn",
}
m["kyz"] = {
"Kayabí",
6380372,
"tup-gua",
"Latn",
}
m["kza"] = {
"Karaboro Barat",
36601,
"alv-krb",
}
m["kzb"] = {
"Kaibobo",
6347565,
"poz-cma",
"Latn",
}
m["kzc"] = {
"Bondoukou Kulango",
11031321,
"alv-kul",
"Latn",
}
m["kzd"] = {
"Kadai",
7679471,
"poz-cma",
"Latn",
}
--kze (Kosena) made an etym-only child of auy (Auyana) per [[Wiktionary:Language_treatment_requests#merge_Kosena_[kze]_into_Auyana_[auy]]]
m["kzf"] = {
"Da'a Kaili",
33103997,
"poz-kal",
"Latn",
}
m["kzg"] = {
"Kikai",
3196527,
"jpx-nry",
"Jpan",
translit = s["jpx-translit"],
display_text = s["jpx-displaytext"],
strip_diacritics = s["jpx-stripdiacritics"],
sort_key = s["jpx-sortkey"],
}
m["kzh"] = {
"Dongolawi",
5295991,
"nub",
"Latn",
}
m["kzi"] = {
"Kelabit",
6385445,
"poz-swa",
"Latn",
}
m["kzj"] = {
"Kadazan",
3307195,
"poz-san",
"Latn",
}
m["kzk"] = {
"Kazukuru",
1089069,
"poz-ocw",
}
m["kzl"] = {
"Kayeli",
4207444,
"poz-cma",
"Latn",
}
m["kzm"] = {
"Kais",
6348319,
"ngf-sbh",
"Latn",
}
m["kzn"] = {
"Kokola",
11128329,
"bnt-mak",
"Latn",
ancestors = "vmw",
}
m["kzo"] = {
"Kaningi",
35683,
"bnt-mbt",
}
m["kzp"] = {
"Kaidipang",
6347611,
"phi",
"Latn",
}
m["kzq"] = {
"Kaike",
10951226,
"sit-tam",
}
m["kzr"] = {
"Karang",
35681,
"alv-mbm",
"Latn",
}
m["kzs"] = {
"Dusun Sugut",
12953510,
"poz-san",
"Latn",
}
m["kzt"] = {
"Dusun Tambunan",
12953514,
"poz-san",
"Latn",
}
m["kzu"] = {
"Kayupulau",
6380723,
"poz-ocw",
}
m["kzv"] = {
"Komyandaret",
6428671,
"ngf-kts",
"Latn",
}
m["kzw"] = { -- contrast xoo, sai-kat, sai-xoc, the last of which the ISO conflated into this code
"Kariri",
12953620,
"sai-mje",
"Latn",
}
m["kzx"] = {
"Kamarian",
6356040,
"poz-cma",
"Latn",
}
m["kzy"] = {
"Kango-Sua",
11008360,
"bnt-kbi",
"Latn",
ancestors = "bip",
}
m["kzz"] = {
"Kalabra",
6350038,
"paa-wbh",
"Latn",
}
return require("Module:languages").finalizeData(m, "language")
3e9kmkh1hp86ecbfgihibcrgjsohcph
Modul:languages/data/3/t
828
9821
375375
373732
2026-09-22T05:33:45Z
Hakimi97
2668
Move "Temuan" back to Module:languages/data/3/t
375375
Scribunto
text/plain
local m_langdata = require("Module:languages/data")
-- Loaded on demand, as it may not be needed (depending on the data).
local function u(...)
u = require("Module:string utilities").char
return u(...)
end
local c = m_langdata.chars
local p = m_langdata.puaChars
local s = m_langdata.shared
local m = {}
m["taa"] = {
"Lower Tanana",
28565,
"ath-nor",
"Latn",
}
m["tab"] = {
"Tabasaran",
34079,
"cau-esm",
"Cyrl, Latn, Arab",
translit = {
Cyrl = "tab-translit",
},
override_translit = true,
display_text = {
Cyrl = s["cau-Cyrl-displaytext"]
},
strip_diacritics = {
Cyrl = s["cau-Cyrl-stripdiacritics"],
Latn = s["cau-Latn-stripdiacritics"],
},
sort_key = {
Cyrl = "tab-sortkey",
}
}
m["tac"] = {
"Lowland Tarahumara",
15616384,
"azc-trc",
"Latn",
}
m["tad"] = {
"Tause",
2356440,
"paa-wlp",
"Latn",
}
m["tae"] = {
"Tariana",
732726,
"awd-nwk",
"Latn",
}
m["taf"] = {
"Tapirapé",
7684673,
"tup-gua",
"Latn",
}
m["tag"] = {
"Tagoi",
36537,
"nic-ras",
"Latn",
}
m["taj"] = {
"Tamang Timur",
12953177,
"sit-tam",
"sit-tam-Tibt, Deva",
-- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
-- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake.
}
m["tak"] = {
"Tala",
3914494,
"cdc-wst",
"Latn",
}
m["tal"] = {
"Tal",
3440387,
"cdc-wst",
"Latn",
}
m["tan"] = {
"Tangale",
529921,
"cdc-wst",
"Latn",
}
m["tao"] = {
"Yami",
715760,
"phi",
"Latn",
}
m["tap"] = {
"Taabwa",
7673650,
"bnt-sbi",
"Latn",
}
m["tar"] = {
"Tarahumara Tengah",
20090009,
"azc-trc",
"Latn",
sort_key = {remove_diacritics = c.acute .. "ꞌ"},
}
m["tas"] = {
"Tây Bồi",
2233794,
"crp",
"Latn",
ancestors = "fr",
sort_key = s["roa-oil-sortkey"],
}
m["tau"] = {
"Upper Tanana",
28281,
"ath-nor",
"Latn",
}
m["tav"] = {
"Tatuyo",
2524007,
"sai-tuc",
"Latn",
}
m["taw"] = {
"Tai",
7675861,
"ngf-kak",
"Latn",
}
m["tax"] = {
"Tamki",
3449082,
"cdc-est",
"Latn",
}
m["tay"] = {
"Atayal",
715766,
"map-ata",
"Latn",
}
m["taz"] = {
"Tocho",
36680,
"alv-tal",
"Latn",
}
m["tba"] = {
"Aikanã",
3409307,
"qfa-iso",
"Latn",
}
-- Tapeba [tbb] is spurious
m["tbc"] = {
"Takia",
3514336,
"poz-oce",
"Latn",
}
m["tbd"] = {
"Kaki Ae",
6349417,
"qfa-iso", -- isolate in Glottolog and Pawley and Hammarström (2018); tentatively in Eleman family by Ross (2005)
-- (and Usher?), but they don't address counterarguments of Clifton 1997
"Latn",
}
m["tbe"] = {
"Tanimbili",
3515188,
"poz-tem",
"Latn",
}
m["tbf"] = {
"Mandara",
3285424,
"poz-ocw",
"Latn",
}
m["tbg"] = {
"Tairora Utara",
20210398,
"ngf-tai",
"Latn",
}
m["tbh"] = {
"Thurawal",
3537135,
"aus-yuk",
"Latn",
}
m["tbi"] = {
"Gaam",
35338,
"sdv-eje",
"Latn",
}
m["tbj"] = {
"Tiang",
3528020,
"poz-ocw",
"Latn",
}
m["tbk"] = {
"Calamian Tagbanwa",
3915487,
"phi-kal",
"Tagb, Latn",
}
m["tbl"] = {
"Tboli",
7690594,
"phi",
"Latn",
}
m["tbm"] = {
"Tagbu",
7675188,
"nic-ser",
}
m["tbn"] = {
"Barro Negro Tunebo",
12953943,
"cba",
}
m["tbo"] = {
"Tawala",
7689206,
"poz-ocw",
"Latn",
}
m["tbp"] = {
"Taworta",
7689337,
"paa-elp",
"Latn",
}
m["tbr"] = {
"Tumtum",
3407029,
"qfa-kad",
}
m["tbs"] = {
"Tanguat",
7683166,
"paa-ata",
"Latn",
}
m["tbt"] = {
"Kitembo",
13123561,
"bnt-shh",
"Latn",
}
m["tbu"] = {
"Tubar",
56730,
"azc-trc",
"Latn",
}
m["tbv"] = {
-- considered a dialect of Kulungtfu-Yuanggeng-Tobo [kgf] by Glottolog
"Tobo",
7811712,
"ngf-kto",
"Latn",
}
m["tbw"] = {
"Tagbanwa",
3915475,
"phi",
"Latn",
}
m["tbx"] = {
"Kapin",
6366665,
"poz-ocw",
"Latn",
}
m["tby"] = {
"Tabaru",
11732670,
"paa-gto",
"Latn",
}
m["tbz"] = {
"Ditammari",
35186,
"nic-eov",
"Latn",
}
m["tca"] = {
"Ticuna",
1815205,
"sai-tyu",
"Latn",
}
m["tcb"] = {
"Tanacross",
28268,
"ath-nor",
"Latn",
}
m["tcc"] = {
"Datooga",
35327,
"sdv-nis",
"Latn",
}
m["tcd"] = {
"Tafi",
36545,
"alv-ktg",
}
m["tce"] = {
"Tutchone Selatan",
31091048,
"ath-nor",
"Latn",
}
m["tcf"] = {
"Tlapanec Malinaltepec",
25559732,
"omq",
"Latn",
}
m["tcg"] = {
"Tamagario",
7680531,
"paa-kay",
"Latn",
}
m["tch"] = {
"Inggeris Kreol Turks dan Caicos",
7855478,
"crp",
"Latn",
ancestors = "en",
}
m["tci"] = {
"Wára",
20825638,
"paa-wko",
"Latn",
}
m["tck"] = {
"Tchitchege",
36595,
"bnt-tek",
}
m["tcl"] = {
"Taman (Myanmar)",
15616518,
"sit-jnp",
"Latn",
}
m["tcm"] = {
"Tanahmerah",
3514927,
"qfa-dis", -- Papuan; isolate per Glottolog and Palmer (2018), considered an independent branch of TNG by Usher
-- (2020); seems based only on some pronoun correspondences
"Latn",
}
m["tco"] = {
"Taungyo",
12953186,
"tbq-brm",
ancestors = "obr",
}
m["tcp"] = {
"Chin Tawr",
7689338,
"tbq-kuk",
}
m["tcq"] = {
"Kaiy",
6348709,
"paa-clp",
"Latn",
}
m["tcs"] = {
"Kreol Selat Torres",
36648,
"crp",
"Latn",
ancestors = "en",
}
m["tct"] = {
"T'en",
3442330,
"qfa-kms",
}
m["tcu"] = {
"Tarahumara Tenggara",
36807,
"azc-trc",
"Latn",
}
m["tcw"] = {
"Tecpatlán Totonac",
7692795,
"nai-ttn",
"Latn",
}
m["tcx"] = {
"Toda",
34042,
"dra-tkt",
"Taml",
--translit = {Taml = "Taml-translit"},
}
m["tcy"] = {
"Tulu",
34251,
"dra-tlk",
"Tutg, Mlym, Knda", -- Mlym is nearer than Knda but both lack ɛ/ɛː.
translit = {
Tutg = "tcy-Tutg-translit",
},
-- Knda translit in [[Module:scripts/data]]
-- Mlym translit in [[Module:scripts/data]]
}
m["tcz"] = {
"Chin Thado",
6583558,
"tbq-kuk",
}
m["tda"] = {
"Tagdal",
36570,
"son",
}
m["tdb"] = {
"Panchpargania",
21946879,
"inc-sad",
"Deva, as-Beng, Orya, Chis",
}
m["tdc"] = {
"Emberá-Tadó",
3052041,
"sai-chc",
"Latn",
}
m["tdd"] = {
"Tai Nüa",
36556,
"tai-swe",
"Tale",
translit = "Tale-translit",
strip_diacritics = {remove_diacritics = c.ZWNJ .. c.ZWJ},
}
m["tde"] = {
"Tiranige Diga Dogon",
5313387,
"nic-dgw",
}
m["tdf"] = {
"Talieng",
37525108,
"mkh-ban",
}
m["tdg"] = {
"Tamang Barat",
12953178,
"sit-tam",
"sit-tam-Tibt, Deva",
-- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
-- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake.
}
m["tdh"] = {
"Thulung",
56553,
"sit-kiw",
}
m["tdi"] = {
"Tomadino",
7818197,
"poz-btk",
"Latn",
}
m["tdj"] = {
"Tajio",
7676870,
"poz",
"Latn",
}
m["tdk"] = {
"Tambas",
3440392,
"cdc-wst",
}
m["tdl"] = {
"Sur",
3914453,
"nic-tar",
}
m["tdm"] = {
"Taruma",
5559094,
}
m["tdn"] = {
"Tondano",
3531514,
"phi",
"Latn",
}
m["tdo"] = {
"Teme",
3913994,
"alv-mye",
}
m["tdq"] = {
"Tita",
3914899,
"nic-bco",
}
m["tdr"] = {
"Todrah",
7812881,
"mkh",
}
m["tds"] = {
"Doutai",
5302331,
"paa-clp",
"Latn",
}
m["tdt"] = {
"Tetun Dili",
12643484,
"poz-tim",
"Latn",
}
m["tdv"] = {
"Toro",
3438367,
"nic-alu",
"Latn",
}
m["tdy"] = {
"Tadyawan",
7674700,
"phi",
"Latn",
}
m["tea"] = {
"Temiar",
3914693,
"mkh-asl",
"Latn",
}
m["teb"] = {
"Tetete",
7706087,
"sai-tuc",
"Latn",
}
m["tec"] = {
"Terik",
3518379,
"sdv-nma",
}
m["ted"] = {
"Tepo Krumen",
11152243,
"kro-grb",
}
m["tee"] = {
"Tepehua Huehuetla",
56455,
"nai-ttn",
"Latn",
}
m["tef"] = {
"Teressa",
3518362,
"aav-nic",
}
m["teg"] = {
"Teke-Tege",
36478,
"bnt-tek",
}
m["teh"] = {
"Tehuelche",
33930,
"sai-cho",
"Latn",
}
m["tei"] = {
"Torricelli",
3450788,
"paa-kom",
"Latn",
}
m["tek"] = {
"Ibali Teke",
2802914,
"bnt-tek",
}
m["tem"] = {
"Temne",
36613,
"alv-mel",
"Latn",
}
m["ten"] = {
"Tama (Colombia)",
3832969,
"sai-tuc",
"Latn",
}
m["teo"] = {
"Ateso",
29474,
"sdv-ttu",
"Latn",
}
m["tep"] = {
"Tepecano",
3915525,
"azc-pim",
"Latn",
}
m["teq"] = {
"Temein",
7698064,
"sdv",
}
m["ter"] = {
"Tereno",
3314742,
"awd",
"Latn",
}
m["tes"] = {
"Tengger",
12473479,
"poz",
"Latn, Java",
}
m["tet"] = {
"Tetum",
34125,
"poz-tim",
"Latn",
}
m["teu"] = {
"Soo",
3437607,
"ssa-klk",
}
m["tev"] = {
"Teor",
12953198,
"poz-cma",
"Latn",
}
m["tew"] = {
"Tewa",
56492,
"nai-kta",
"Latn",
}
m["tex"] = {
"Tennet",
56346,
"sdv",
}
m["tey"] = {
"Tulishi",
12911106,
"qfa-kad",
"Latn",
}
m["tez"] = {
"Tetserret",
7706841,
"ber",
"Latn",
}
m["tfi"] = {
"Tofin Gbe",
3530330,
"alv-pph",
}
m["tfn"] = {
"Dena'ina",
27785,
"ath-nor",
"Latn",
}
m["tfo"] = {
"Tefaro",
7694618,
"paa-egb",
"Latn",
}
m["tfr"] = {
"Teribe",
36533,
"cba",
"Latn",
}
m["tft"] = {
"Ternate",
3518492,
"paa-tti",
"Latn, Arab",
}
m["tga"] = {
"Sagalla",
12953082,
"bnt-cht",
}
m["tgb"] = {
"Tobilung",
12953913,
"poz-san",
"Latn",
}
m["tgc"] = {
"Tigak",
3528276,
"poz-ocw",
"Latn",
}
m["tgd"] = {
"Ciwogai",
3438799,
"cdc-wst",
"Latn",
}
m["tge"] = {
"Tamang Gorkha Timur",
12953175,
"sit-tam",
"sit-tam-Tibt, Deva",
-- sit-tam-Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
-- NOTE: Formerly there was no sort_key or translit specified; I assume that's a mistake.
}
m["tgf"] = {
"Chali",
3695197,
"sit-ebo",
"Tibt, Latn",
override_translit = true,
-- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["tgh"] = {
"Inggeris Kreol Tobago",
7811541,
"crp",
ancestors = "en",
}
m["tgi"] = {
"Lawunuia",
3219937,
"poz-ocw",
}
m["tgn"] = {
"Tandaganon",
63311769,
"phi",
"Latn",
}
m["tgo"] = {
"Sudest",
7675351,
"poz-ocw",
}
m["tgp"] = {
"Tangoa",
2410276,
"poz-vnn",
"Latn",
}
m["tgq"] = {
"Tring",
7842360,
"poz-swa",
}
m["tgr"] = {
"Tareng",
25559541,
"mkh",
}
m["tgs"] = {
"Nume",
3346290,
"poz-vnn",
"Latn",
}
m["tgt"] = {
"Tagbanwa Tengah",
3915515,
"phi",
"Tagb",
}
m["tgu"] = {
"Tanggu",
7682930,
"paa-ata",
"Latn",
}
m["tgv"] = {
"Tingui-Boto",
7808195,
"sai-mje",
"Latn",
}
m["tgw"] = {
"Tagwana Senoufo",
36514,
"alv-tdj",
}
m["tgx"] = {
"Tagish",
28064,
"ath-nor",
"Latn",
}
m["tgy"] = {
"Togoyo",
36825,
"nic-ser",
}
m["thc"] = {
"Tai Hang Tong",
7675753,
"tai-nor",
}
m["thd"] = {
"Kuuk Thaayorre",
6448718,
"aus-pmn",
"Latn",
}
m["the"] = {
"Chitwania Tharu",
22083804,
"inc-tha",
"Deva",
}
m["thf"] = {
"Thangmi",
7710314,
"sit-new",
"Deva",
}
m["thh"] = {
"Tarahumara Utara",
15616395,
"azc-trc",
"Latn",
}
m["thi"] = {
"Tai Long",
25559562,
"tai-swe",
}
m["thk"] = {
"Tharaka",
15407179,
"bnt-kka",
}
m["thl"] = {
"Dangaura Tharu",
22083815,
"inc-tha",
"Deva",
}
m["thm"] = {
"Thavung",
34780,
"mkh-vie",
"Thai", --Laoo is feasible but no evidence yet.
sort_key = "Thai-sortkey",
}
m["thn"] = {
"Thachanadan",
7708880,
"dra-mal",
}
m["thp"] = {
"Thompson",
1755054,
"sal",
"Latn, Dupl",
}
m["thq"] = {
"Kochila Tharu",
22083826,
"inc-tha",
}
m["thr"] = {
"Rana Tharu",
12953920,
"inc-tha",
"Deva",
}
m["ths"] = {
"Thakali",
7709348,
"sit-tam",
}
m["tht"] = {
"Tahltan",
30125,
"ath-nor",
"Latn",
}
m["thu"] = {
"Thuri",
7799291,
"sdv-lon",
}
m["thy"] = {
"Tha",
3915849,
"alv-bwj",
}
m["tic"] = {
"Tira",
36677,
"alv-hei",
}
m["tif"] = {
"Tifal",
11732691,
"ngf-mok",
"Latn",
}
m["tig"] = {
"Tigre",
34129,
"sem-eth",
"Ethi",
translit = "Ethi-translit",
}
m["tih"] = {
"Murut Timugon",
7807680,
"poz-san",
"Latn",
}
m["tii"] = {
"Tiene",
36469,
"bnt-tek",
}
m["tij"] = {
"Tilung",
7803037,
"sit-kiw",
}
m["tik"] = {
"Tikar",
36483,
"nic-bdn",
"Latn",
}
m["til"] = {
"Tillamook",
2109432,
"sal",
"Latn",
}
m["tim"] = {
"Timbe",
7804599,
"ngf-kab",
"Latn",
}
m["tin"] = {
"Tindi",
36860,
"cau-and",
"Cyrl",
display_text = s["cau-Cyrl-displaytext"],
strip_diacritics = s["cau-Cyrl-stripdiacritics"],
}
m["tio"] = {
"Teop",
3518239,
"poz-ocw",
"Latn",
}
m["tip"] = {
"Trimuris",
7842270,
"paa-kwe",
"Latn",
}
m["tiq"] = {
"Tiéfo",
3914874,
"alv-sav",
}
m["tis"] = {
"Masadiit Itneg",
18748769,
"phi",
}
m["tit"] = {
"Tinigua",
3029805,
"sai-tin",
"Latn",
}
m["tiu"] = {
"Adasen",
11214797,
"phi",
"Latn",
}
m["tiv"] = {
"Tiv",
34131,
"nic-tvc",
"Latn",
}
m["tiw"] = {
"Tiwi",
1656014,
"qfa-iso",
"Latn",
}
m["tix"] = {
"Tiwa Selatan",
7570552,
"nai-kta",
"Latn",
}
m["tiy"] = {
"Tiruray",
7809425,
"phi",
"Latn",
}
m["tiz"] = {
"Tai Hongjin",
3915716,
"tai-swe",
}
m["tja"] = {
"Tajuasohn",
3915326,
"kro-wkr",
}
m["tjg"] = {
"Tunjung",
3542117,
"poz",
"Latn",
}
m["tji"] = {
"Tujia Utara",
12953229,
"sit-tja",
"Latn",
}
m["tjl"] = {
"Tai Laing",
7675773,
"tai-swe",
"Mymr",
}
m["tjm"] = {
"Timucua",
638300,
"qfa-iso",
"Latn",
}
m["tjn"] = {
"Tonjon",
3913372,
"dmn-jje",
}
m["tjs"] = {
"Tujia Selatan",
12633994,
"sit-tja",
"Latn",
}
m["tju"] = {
"Tjurruru",
3913834,
"aus-nga",
"Latn",
}
m["tjw"] = {
"Chaap Wuurong",
5285187,
"aus-pam",
"Latn",
}
m["tka"] = {
"Truká",
7847648,
}
m["tkb"] = {
"Buksa",
20983638,
"inc-eas",
"Deva",
}
m["tkd"] = {
"Tukudede",
36863,
"poz-tim",
"Latn",
}
m["tke"] = {
"Takwane",
11030092,
"bnt-mak",
"Latn",
ancestors = "vmw",
}
m["tkf"] = {
"Tukumanféd",
42330115,
"tup-gua",
"Latn",
}
m["tkl"] = {
"Tokelau",
34097,
"poz-pnp",
"Latn",
}
m["tkm"] = {
"Takelma",
56710,
}
m["tkn"] = {
"Toku-No-Shima",
3530484,
"jpx-nry",
"Jpan",
translit = s["jpx-translit"],
display_text = s["jpx-displaytext"],
strip_diacritics = s["jpx-stripdiacritics"],
sort_key = s["jpx-sortkey"],
}
m["tkp"] = {
"Tikopia",
36682,
"poz-pnp",
"Latn",
}
m["tkq"] = {
"Tee",
3075144,
"nic-ogo",
"Latn",
}
m["tkr"] = {
"Tsakhur",
36853,
"cau-wsm",
"Cyrl, Latn, Arab",
translit = "tkr-translit",
override_translit = true,
display_text = {
Cyrl = s["cau-Cyrl-displaytext"]
},
strip_diacritics = {
Cyrl = s["cau-Cyrl-stripdiacritics"],
Latn = s["cau-Latn-stripdiacritics"],
},
}
m["tks"] = {
"Ramandi",
25261947,
"xme-ttc",
"Arab",
ancestors = "xme-ttc-sou",
}
m["tkt"] = {
"Kathoriya Tharu",
22083822,
"inc-tha",
}
m["tku"] = {
"Upper Necaxa Totonac",
56343,
"nai-ttn",
"Latn",
}
m["tkv"] = {
"Mur Pano",
16939373,
"poz-ocw",
"Latn",
}
m["tkw"] = {
"Teanu",
3516731,
"poz-tem",
"Latn",
}
m["tkx"] = {
"Tangko",
7682993,
"ngf-tna",
"Latn",
}
m["tkz"] = {
"Takua",
7678544,
"mkh",
}
m["tla"] = {
"Tepehuan Barat Daya",
3518245,
"azc-pim",
"Latn",
}
m["tlb"] = {
"Tobelo",
1142333,
"paa-gto",
"Latn",
}
m["tlc"] = {
"Misantla Totonac",
56460,
"nai-ttn",
"Latn",
}
m["tld"] = {
"Talaud",
7678964,
"phi",
"Latn",
}
m["tlf"] = {
"Telefol",
7696150,
"ngf-mok",
"Latn",
}
m["tlg"] = {
"Tofanma",
4461493,
"paa-nto",
"Latn",
}
m["tlh"] = {
"Klingon",
10134,
"art",
"Latn",
type = "appendix-constructed",
}
m["tli"] = {
"Tlingit",
27792,
"xnd",
"Latn, Cyrl",
}
m["tlj"] = {
"Talinga-Bwisi",
7679530,
"bnt-haj",
}
m["tlk"] = {
"Taloki",
3514563,
"poz-btk",
}
m["tll"] = {
"Tetela",
2613465,
"bnt-tet",
}
m["tlm"] = {
"Tolomako",
3130514,
"poz-vnn",
"Latn",
}
m["tln"] = {
"Talondo'",
7680293,
"poz-ssw",
}
m["tlo"] = {
"Talodi",
36525,
"alv-tal",
}
m["tlp"] = {
"Filomena Mata-Coahuitlán Totonac",
5449202,
"nai-ttn",
"Latn",
}
m["tlq"] = {
"Tai Loi",
7675784,
"mkh-pal",
}
m["tlr"] = {
"Talise",
3514510,
"poz-sls",
"Latn",
}
m["tls"] = {
"Tambotalo",
7681065,
"poz-vnn",
"Latn",
}
m["tlt"] = {
"Teluti",
12953194,
"poz-cma",
}
m["tlu"] = {
"Tulehu",
7852006,
"poz-cma",
}
m["tlv"] = {
"Taliabu",
3514498,
"poz-cma",
"Latn",
}
m["tlx"] = {
"Khehek",
3196124,
"poz-aay",
}
m["tly"] = {
"Talysh",
34318,
"xme-ttc",
"Latn, Cyrl, Arab",
}
m["tma"] = {
"Tama (Chad)",
57001,
"sdv-tmn",
}
m["tmb"] = {
"Avava",
2157461,
"poz-vnc",
"Latn",
}
m["tmc"] = {
"Tumak",
3121045,
"cdc-est",
}
m["tmd"] = {
"Haruai",
12632146,
"paa-pia",
"Latn",
}
m["tme"] = {
"Tremembé",
5246937,
}
m["tmf"] = {
"Toba-Maskoy",
3033544,
"sai-mas",
"Latn",
}
m["tmg"] = {
"Ternateño",
7232597,
}
m["tmh"] = {
"Tuareg",
34065,
"ber",
"Latn, Tfng, Arab",
strip_diacritics = {
Latn = {remove_diacritics = c.grave .. c.acute .. c.circ},
},
}
m["tmi"] = {
"Tutuba",
7857052,
"poz-vnn",
"Latn",
}
m["tmj"] = {
"Samarokena",
7408865,
"paa-saa",
"Latn",
}
m["tml"] = {
"Tamnim Citak",
12643315,
"ngf-asm",
"Latn",
}
m["tmm"] = {
"Tai Thanh",
7675842,
"tai-swe",
}
m["tmn"] = {
"Taman (Indonesia)",
7680671,
"poz",
"Latn",
}
m["tmo"] = {
"Temoq",
7698205,
"mkh-asl",
}
m["tmq"] = {
"Tumleo",
7852641,
"poz-ocw",
}
m["tms"] = {
"Tima",
36684,
"nic-ktl",
}
m["tmt"] = {
"Tasmate",
7687571,
"poz-vnn",
"Latn",
}
m["tmu"] = {
"Iau",
56867,
"paa-lpl",
"Latn",
}
m["tmv"] = {
"Motembo",
11013108,
"bnt-bun",
}
m["tmw"] = {
"Temuan",
3025610,
"poz-mly",
"Latn",
}
m["tmy"] = {
"Tami",
3514812,
"poz-oce",
}
m["tmz"] = {
"Tamanaku",
3441435,
"sai-ven",
"Latn",
}
m["tna"] = {
"Tacana",
3182551,
"sai-tac",
"Latn",
}
m["tnb"] = {
"Tunebo Barat",
3181238,
"cba",
}
m["tnc"] = {
"Tanimuca-Retuarã",
36535,
"sai-tuc",
"Latn",
}
m["tnd"] = {
"Angosturas Tunebo",
25559604,
"cba",
}
m["tne"] = {
"Tinoc Kallahan",
3192219,
}
m["tng"] = {
"Tobanga",
3440501,
"cdc-est",
}
m["tnh"] = {
"Maiani",
6735243,
"ngf-kau",
"Latn",
}
m["tni"] = {
"Tandia",
7682454,
"poz-hce",
"Latn",
}
m["tnk"] = {
"Kwamera",
3200806,
"poz-vns",
"Latn",
}
m["tnl"] = {
"Lenakel",
3229429,
"poz-vns",
"Latn",
}
m["tnm"] = {
"Tabla",
7673105,
"paa-sen",
"Latn",
}
m["tnn"] = {
"Tanna Utara",
957945,
"poz-vns",
"Latn",
}
m["tno"] = {
"Toromono",
510544,
"sai-tac",
"Latn",
}
m["tnp"] = {
"Whitesands",
3063761,
"poz-vns",
"Latn",
}
m["tnq"] = {
"Taíno",
5232952,
"awd-taa",
"Latn",
}
m["tnr"] = {
"Bedik",
35096,
"alv-ten",
"Latn",
}
m["tns"] = {
"Tenis",
7699870,
"poz-stm",
"Latn",
}
m["tnt"] = {
"Tontemboan",
3531666,
"phi",
"Latn",
}
m["tnu"] = {
"Tay Khang",
6362363,
"tai",
}
m["tnv"] = {
"Tanchangya",
7682361,
"inc-bas",
"Cakm",
ancestors = "inc-obn",
}
m["tnw"] = {
"Tonsawang",
3531660,
"phi",
"Latn",
}
m["tnx"] = {
"Tanema",
2106984,
"poz-tem",
"Latn",
}
m["tny"] = {
"Tongwe",
7821200,
"bnt",
}
m["tnz"] = {
"Ten'edn",
3073453,
"mkh-asl",
"Latn",
}
m["tob"] = {
"Toba",
3113756,
"sai-guc",
"Latn",
}
m["toc"] = {
"Coyutla Totonac",
15615591,
"nai-ttn",
"Latn",
}
m["tod"] = {
"Toma",
11055484,
"dmn-msw",
"Latn, Loma"
}
m["tof"] = {
"Gizrra",
5565941,
"paa-etf",
"Latn",
}
m["tog"] = {
"Tonga (Malawi)",
3847648,
"bnt-nys",
"Latn",
}
m["toh"] = {
"Tonga (Mozambique)",
7820988,
"bnt-bso",
}
m["toi"] = {
"Tonga (Zambia)",
34101,
"bnt-bot",
"Latn",
}
m["toj"] = {
"Tojolabal",
36762,
"myn",
"Latn",
}
m["tok"] = {
"Toki Pona",
36846,
"art",
"Latn",
type = "appendix-constructed",
}
m["tol"] = {
"Tolowa",
20827,
"ath-pco",
"Latn",
}
m["tom"] = {
"Tombulu",
3531199,
"phi",
"Latn",
}
m["too"] = {
"Xicotepec de Juárez Totonac",
8044353,
"nai-ttn",
"Latn",
}
m["top"] = {
"Papantla Totonac",
56329,
"nai-ttn",
"Latn",
}
m["toq"] = {
"Toposa",
3033588,
"sdv-ttu",
}
m["tor"] = {
"Togbo-Vara Banda",
11002922,
"bad-cnt",
}
m["tos"] = {
"Highland Totonac",
13154149,
"nai-ttn",
"Latn",
}
m["tou"] = {
"Tho",
22694631,
"mkh-vie",
"Latn",
}
m["tov"] = {
"Upper Taromi",
12953183,
"xme-ttc",
ancestors = "xme-ttc-cen",
}
m["tow"] = {
"Jemez",
3912876,
"nai-kta",
"Latn",
}
m["tox"] = {
"Tobian",
34022,
"poz-mic",
}
m["toy"] = {
"Topoiyo",
7824977,
"poz-kal",
}
m["toz"] = {
"To",
7811216,
"alv-mbm",
}
m["tpa"] = {
"Taupota",
7688832,
"poz-ocw",
}
m["tpc"] = {
"Azoyú Me'phaa",
25559730,
"omq",
"Latn",
}
m["tpe"] = {
"Tippera",
16115423,
"tbq-bdg",
}
m["tpf"] = {
"Tarpia",
12953185,
"poz-ocw",
"Latn",
}
m["tpg"] = {
"Kula",
6442714,
"paa-alp",
"Latn",
}
m["tpi"] = {
"Tok Pisin",
34159,
"crp",
"Latn",
ancestors = "en",
}
m["tpj"] = {
"Tapieté",
3121063,
"gn",
"Latn",
}
m["tpk"] = {
"Tupinikin",
33924,
"tup-gua",
}
m["tpl"] = {
"Tlacoapa Me'phaa",
16115511,
"omq",
}
m["tpm"] = {
"Tampulma",
36590,
"nic-gnw",
}
m["tpn"] = {
"Tupinambá",
31528147,
"tup-gua",
"Latn",
}
m["tpo"] = {
"Tai Pao",
7675795,
"tai-nor",
}
m["tpp"] = {
"Pisaflores Tepehua",
56349,
"nai-ttn",
}
m["tpq"] = {
"Tukpa",
12953230,
"sit-las",
}
m["tpr"] = {
"Tuparí",
3542217,
"tup",
"Latn",
}
m["tpt"] = {
"Tlachichilco Tepehua",
56330,
"nai-ttn",
}
m["tpu"] = {
"Tampuan",
3514882,
"mkh-ban",
"Khmr",
}
m["tpv"] = {
"Tanapag",
3397371,
"poz-mic",
}
m["tpw"] = {
"Tupi Kuno",
56944,
"tup-gua",
"Latn",
}
m["tpx"] = {
"Acatepec Me'phaa",
31157882,
"omq",
"Latn",
}
m["tpy"] = {
"Trumai",
12294279,
"qfa-iso",
}
m["tpz"] = {
"Tinputz",
3529205,
"poz-ocw",
}
m["tqb"] = {
"Tembé",
10322157,
"tup-gua",
"Latn",
}
m["tql"] = {
"Lehali",
3229119,
"poz-vnn",
"Latn",
}
m["tqm"] = {
"Turumsa",
7856508,
"paa-dtu",
"Latn",
}
m["tqn"] = {
"Tenino",
15699255,
"nai-shp",
"Latn",
ancestors = "nai-spt",
}
m["tqo"] = {
"Toaripi",
7811403,
"paa-eel",
"Latn",
}
m["tqp"] = {
"Tomoip",
3531388,
"poz-ocw",
}
m["tqq"] = {
"Tunni",
3514343,
"cus-som",
}
m["tqr"] = {
"Torona",
36679,
"alv-tal",
}
m["tqt"] = {
"Totonac Barat",
7116691,
"nai-ttn",
"Latn",
}
m["tqu"] = {
"Touo",
56750,
}
m["tqw"] = {
"Tonkawa",
2454881,
"qfa-iso",
"Latn",
}
m["tra"] = {
"Tirahi",
3812406,
"inc-koh",
}
m["trb"] = {
"Terebu",
7701797,
"poz-ocw",
}
m["trc"] = {
"Copala Triqui",
12953935,
"omq-tri",
"Latn",
}
m["trd"] = {
"Turi",
7854914,
"mun",
}
m["tre"] = {
"Tarangan Timur",
18609750,
"poz",
}
m["trf"] = {
"Inggeris Kreol Trinidad",
7842493,
"crp",
"Latn",
ancestors = "en",
}
m["trg"] = {
"Lishán Didán",
56473,
"sem-nna",
"Hebr",
}
m["trh"] = {
"Turaka",
12953237,
"ngf-dag",
"Latn",
}
m["tri"] = {
"Trió",
56885,
"sai-tar",
"Latn",
}
m["trj"] = {
"Toram",
3441225,
"cdc-est",
}
m["trl"] = {
"Traveller Scottish",
3915671,
"qfa-mix",
"Latn",
ancestors = "rom, sco",
}
m["trm"] = {
"Tregami",
34081,
"nur-sou",
}
m["trn"] = {
"Trinitario",
3539279,
"awd",
}
m["tro"] = {
"Tarao",
3515603,
"tbq-kuk",
"Latn",
}
m["trp"] = {
"Kokborok",
35947,
"tbq-bdg",
"Beng, Latn" -- WP lists 2 more
}
m["trq"] = {
"San Martín Itunyoso Triqui",
12953934,
"omq-tri",
"Latn",
}
m["trr"] = {
"Taushiro",
1957508,
nil,
"Latn",
}
m["trs"] = {
"Chicahuaxtla Triqui",
3539587,
"omq-tri",
"Latn",
}
m["trt"] = {
"Tunggare",
615071,
"paa-egb",
"Latn",
}
m["tru"] = {
"Turoyo",
34040,
"sem-cna",
"Syrc, Latn",
translit = {
Syrc = "tru-translit",
},
strip_diacritics = {
Syrc = "Syrc-stripdiacritics",
},
}
m["trv"] = {
"Taroko",
716686,
"map-ata",
"Latn",
}
m["trw"] = {
"Torwali",
2665246,
"inc-koh",
"Aran",
}
m["trx"] = {
"Tringgus",
7842365,
"day",
}
m["try"] = {
"Turung",
7856514,
"tai-swe",
"as-Beng",
}
m["trz"] = {
"Torá",
7827518,
"sai-cpc",
}
m["tsa"] = {
"Tsaangi",
36675,
"bnt-nze",
}
m["tsb"] = {
"Tsamai",
2371358,
"cus-eas",
}
m["tsc"] = {
"Tswa",
2085051,
"bnt-tsr",
}
m["tsd"] = {
"Tsakonian",
220607,
"grk",
"Grek",
ancestors = "grc-dor",
translit = "el-translit",
-- Grek display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["tse"] = {
"Bahasa Isyarat Tunisia",
7853191,
"sgn",
}
m["tsg"] = {
"Suluk",
34142,
"phi",
"Latn, Arab",
}
m["tsh"] = {
"Tsuvan",
3502326,
"cdc-cbm",
"Latn",
}
m["tsi"] = {
"Tsimshian",
20085721,
"nai-tsi",
"Latn",
}
m["tsj"] = {
"Tshangla",
36840,
"sit-tsk",
"Tibt, Latn, Deva",
override_translit = true,
-- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["tsl"] = {
"Ts'ün-Lao",
3446675,
"tai",
}
m["tsm"] = {
"Bahasa Isyarat Turki",
36885,
"sgn",
}
m["tsp"] = {
"Toussiana Utara",
11155635,
"alv-sav",
}
m["tsq"] = {
"Bahasa Isyarat Thai",
7709156,
"sgn-asl",
"Sgnw",
}
m["tsr"] = {
"Akei",
2828964,
"poz-vnn",
"Latn",
}
m["tss"] = {
"Bahasa Isyarat Taiwan",
34019,
"sgn-jsl",
}
m["tsu"] = {
"Tsou",
716681,
"map",
"Latn",
}
m["tsv"] = {
"Tsogo",
36674,
"bnt-tso",
}
m["tsw"] = {
"Tsishingini",
13123571,
"nic-kam",
}
m["tsx"] = {
"Mubami",
6930815,
"paa-wig",
"Latn",
}
m["tsy"] = {
"Bahasa Isyarat Tebul",
7692090,
"sgn",
}
m["tta"] = {
"Tutelo",
2311602,
"sio-ohv",
"Latn",
}
m["ttb"] = {
"Gaa",
3438361,
"nic-dak",
"Latn",
}
m["ttc"] = {
"Tektiteko",
36686,
"myn",
"Latn",
}
m["ttd"] = {
"Tauade",
7688634,
"qfa-dis", -- Papuan; isolate per Glottolog; Glottolog says "A Goilalan family uniting Kunimaipan, Tauade and Fuyug
-- is often posited based on the lexicostatistical figures reported in Tom E. Dutton 1975: 631-632" but
-- goes on to say the data "is clearly insufficient, as the lexical links so far proposed are few and
-- show irregular one-consonant correspondences".
"Latn",
}
m["tte"] = {
"Bwanabwana",
5003667,
"poz-ocw",
"Latn",
}
m["ttf"] = {
"Tuotomb",
7853459,
"nic-mbw",
"Latn",
}
m["ttg"] = {
"Tutong",
3507990,
"poz-swa",
"Latn",
}
m["tth"] = {
"Upper Ta'oih",
3512660,
"mkh-kat",
}
m["tti"] = {
"Tobati",
7811556,
"poz-ocw",
"Latn",
}
m["ttj"] = {
"Tooro",
7824218,
"bnt-nyg",
"Latn",
}
m["ttk"] = {
"Totoro",
3532756,
"sai-bar",
"Latn",
}
m["ttl"] = {
"Totela",
10962316,
"bnt-bot",
"Latn",
}
m["ttm"] = {
"Tutchone Utara",
20822,
"ath-nor",
"Latn",
}
m["ttn"] = {
"Towei",
7829606,
"paa-wpw",
"Latn",
}
m["tto"] = {
"Lower Ta'oih",
25559539,
"mkh-kat",
}
m["ttp"] = {
"Tombelala",
6799663,
"poz-kal",
}
m["ttr"] = {
"Tera",
56267,
"cdc-cbm",
}
m["tts"] = {
"Isan",
33417,
"tai-swe",
"Thai", -- also Tai Noi/Lao Buhan script
sort_key = "Thai-sortkey",
}
m["ttt"] = {
"Tat",
56489,
"ira-swi",
"Cyrl, Latn, Armn, Arab",
-- Armn translit in [[Module:scripts/data]] (NOTE: formerly not present, probably an accidental omission)
ancestors = "fa",
}
m["ttu"] = {
"Torau",
3532208,
"poz-ocw",
}
m["ttv"] = {
"Titan",
3445811,
"poz-aay",
"Latn",
}
m["ttw"] = {
"Long Wat",
7856961,
"poz-swa",
}
m["tty"] = {
"Sikaritai",
7513600,
"paa-clp",
"Latn",
}
m["ttz"] = {
"Tsum",
12953223,
"sit-kyk",
}
m["tua"] = {
"Wiarumus",
7998045,
"paa-mmu",
"Latn",
}
m["tub"] = {
"Tübatulabal",
56704,
"azc",
"Latn",
}
m["tuc"] = {
"Mutu",
3331003,
"poz-ocw",
"Latn",
}
m["tud"] = {
"Tuxá",
7857217,
}
m["tue"] = {
"Tuyuca",
2520538,
"sai-tuc",
"Latn",
}
m["tuf"] = {
"Tunebo Tengah",
12953942,
"cba",
"Latn",
}
m["tug"] = {
"Tunia",
863721,
"alv-bua",
}
m["tuh"] = {
"Taulil",
3516141,
}
m["tui"] = {
"Tupuri",
36646,
"alv-mbm",
"Latn",
}
m["tuj"] = {
"Tugutil",
12953228,
"paa-gto",
"Latn",
}
m["tul"] = {
"Tula",
3914907,
"alv-wjk",
}
m["tum"] = {
"Tumbuka",
34138,
"bnt-nys",
"Latn",
}
m["tun"] = {
"Tunica",
56619,
"qfa-iso",
"Latn",
}
m["tuo"] = {
"Tucano",
3541834,
"sai-tuc",
"Latn",
}
m["tuq"] = {
"Tedaga",
36639,
"ssa-sah",
"Latn",
}
m["tus"] = {
"Tuscarora",
36944,
"iro-nor",
"Latn",
}
m["tuu"] = {
"Tututni",
20627,
"ath-pco",
"Latn",
}
m["tuv"] = {
"Turkana",
36958,
"sdv-ttu",
"Latn",
}
m["tux"] = {
"Tuxináwa",
7857204,
"sai-pan",
"Latn",
}
m["tuy"] = {
"Tugen",
3541935,
"sdv-nma",
}
m["tuz"] = {
"Turka",
36643,
"nic-gur",
"Latn",
}
m["tva"] = {
"Vaghua",
3553248,
"poz-ocw",
"Latn",
}
m["tvd"] = {
"Tsuvadi",
3914936,
"nic-kam",
}
m["tve"] = {
"Te'un",
7690709,
"poz-cet",
"Latn",
}
m["tvk"] = {
"Ambrym Tenggara",
252411,
"poz-vnc",
"Latn",
}
m["tvl"] = {
"Tuvalu",
34055,
"poz-pnp",
"Latn",
}
m["tvm"] = {
"Tela-Masbuar",
7695666,
"poz-tim",
}
m["tvn"] = {
"Tavoyan",
7689158,
"tbq-brm",
"Mymr",
ancestors = "obr",
}
m["tvo"] = {
"Tidore",
3528199,
"paa-tti",
"Latn, Arab",
}
m["tvs"] = {
"Taveta",
15632387,
"bnt-par",
"Latn",
}
m["tvt"] = {
"Tutsa Naga",
7856987,
"sit-tno",
}
m["tvu"] = {
"Tunen",
36632,
"nic-mbw",
}
m["tvw"] = {
"Sedoa",
7445362,
"poz-kal",
}
m["tvx"] = {
"Taivoan",
1975271,
"map",
"Latn",
}
m["tvy"] = {
"Timor Pidgin",
4904029,
"crp",
ancestors = "pt",
}
m["twa"] = {
"Twana",
7857412,
"sal",
"Latn",
}
m["twb"] = {
"Tawbuid Barat",
12953912,
"phi",
}
m["twc"] = {
"Teshenawa",
3436597,
"cdc-wst",
"Latn",
}
m["twe"] = {
"Teiwa",
3519302,
"paa-alp",
"Latn",
}
m["twf"] = {
"Taos",
7684320,
"nai-kta",
"Latn",
}
m["twg"] = {
"Tereweng",
12953200,
"paa-alp",
"Latn",
}
m["twh"] = {
"Tai Dón",
7675751,
"tai-swe",
"Tavt",
--translit = "Tavt-translit",
sort_key = {
from = {"[꪿ꫀ꫁ꫂ]", "([ꪵꪶꪹꪻꪼ])([ꪀ-ꪯ])"},
to = {"", "%2%1"}
},
}
m["twm"] = {
"Tawang Monpa",
36586,
"sit-ebo",
"Tibt",
override_translit = true,
-- Tibt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["twn"] = {
"Twendi",
7857682,
"nic-mmb",
}
m["two"] = {
"Tswapong",
3446241,
"bnt-sts",
}
m["twp"] = {
"Ere",
3056045,
"poz-aay",
"Latn",
}
m["twq"] = {
"Tasawaq",
36564,
"son",
}
m["twr"] = {
"Tarahumara Barat Daya",
12953909,
"azc-trc",
"Latn",
}
m["twt"] = {
"Turiwára",
3542307,
"tup-gua",
"Latn",
}
m["twu"] = {
"Termanu",
7702572,
"poz-tim",
"Latn",
}
m["tww"] = {
"Tuwari",
7857159,
"paa-wal",
"Latn",
}
m["twy"] = {
"Tawoyan",
3513542,
"poz-bre",
"Latn",
}
m["txa"] = {
"Tombonuo",
7818692,
"poz-san",
"Latn",
}
m["txb"] = {
"Tocharia B",
3199353,
"ine-toc",
"Latn",
standard_chars = "AaÄäĀāCcEeIiKkLlMmṂṃNnṄṅÑñOoPpRrSsŚśṢṣTtUuWwYy" .. c.punc,
}
m["txc"] = {
"Tsetsaut",
20829,
"ath-nor",
"Latn",
}
m["txe"] = {
"Totoli",
7828387,
"poz-tot",
"Latn",
}
m["txg"] = {
"Tangut",
2727930,
"ero",
"Tang",
-- Tang translit in [[Module:scripts/data]]
}
m["txh"] = {
"Thracian",
36793,
"ine",
"Latn, Polyt",
-- Polyt translit, display_text, strip_diacritics, sort_key in [[Module:scripts/data]]
}
m["txi"] = {
"Ikpeng",
9344891,
"sai-pek",
"Latn",
}
m["txj"] = {
"Tarjumo",
24906088,
"ssa-sah",
"Latn, Arab",
}
m["txm"] = {
"Tomini",
7818911,
"poz",
"Latn",
}
m["txn"] = {
"Tarangan Barat",
3515594,
"poz",
"Latn",
}
m["txo"] = {
"Toto",
36709,
"sit-dhi",
"Beng, Toto"
}
m["txq"] = {
"Tii",
7801784,
"poz-tim",
}
m["txr"] = {
"Tartessian",
36795,
"qfa-unc", -- extinct, no consensus on classification
}
m["txs"] = {
"Tonsea",
3531659,
"phi",
"Latn",
}
m["txt"] = {
"Citak",
3447279,
"ngf-asm",
"Latn",
}
m["txu"] = {
"Kayapó",
3101212,
"sai-nje",
"Latn",
}
m["txx"] = {
"Tatana",
18643518,
"poz-san",
"Latn",
}
m["tya"] = {
"Tauya",
7688978,
"ngf-rai",
"Latn",
}
m["tye"] = {
"Kyenga",
3913304,
"dmn-bbu",
"Latn",
}
m["tyh"] = {
"O'du",
3347428,
"mkh",
}
m["tyi"] = {
"Teke-Tsaayi",
33123613,
"bnt-nze",
}
m["tyj"] = {
"Tai Do",
7675746,
"tai-nor", -- Chamberlain (1991), but Pittayaporn (2009) suggests tai-swe
"Latn, Tayo", -- Vietnam
}
m["tyl"] = {
"Thu Lao",
12953921,
"tai-cen",
}
m["tyn"] = {
"Kombai",
6428241,
"ngf-nde",
"Latn",
}
m["typ"] = {
"Kuku-Thaypan",
3915693,
"aus-pmn",
"Latn",
}
m["tyr"] = {
"Tai Daeng",
3915207,
"tai-swe",
"Tavt",
}
m["tys"] = {
"Sapa",
3446668,
"tai-sap",
"Latn",
}
m["tyt"] = {
"Tày Tac",
7862029,
"tai-swe",
}
m["tyu"] = {
"Kua",
3832933,
"khi-kal",
"Latn",
}
m["tyv"] = {
"Tuva",
34119,
"trk-ssb",
"Cyrl",
translit = "tyv-translit",
override_translit = true,
sort_key = "tyv-sortkey",
}
m["tyx"] = {
"Teke-Tyee",
36634,
"bnt-nze",
"Latn",
}
m["tyz"] = {
"Tày", -- This does not mean its family "Tai" languages.
2511476,
"tai-tay",
"Latn, Hani",
sort_key = {
Hani = "Hani-sortkey"
},
}
m["tza"] = {
"Bahasa Isyarat Tanzania",
7684177,
"sgn",
}
m["tzh"] = {
"Tzeltal",
36808,
"myn",
"Latn",
}
m["tzj"] = {
"Tz'utujil",
36941,
"myn",
"Latn",
}
m["tzl"] = {
"Talossan",
1063911,
"art",
"Latn",
type = "appendix-constructed",
sort_key = "tzl-sortkey",
}
m["tzm"] = {
"Tamazight Atlas Tengah",
49741,
"ber",
"Tfng, Arab, Latn",
translit = {
Tfng = "Tfng-translit",
},
}
m["tzn"] = {
"Tugun",
12953225,
"poz-tim",
"Latn",
}
m["tzo"] = {
"Tzotzil",
36809,
"myn",
"Latn",
}
m["tzx"] = {
"Tabriak",
56872,
"paa-lse",
"Latn",
}
return require("Module:languages").finalizeData(m, "language")
seo5fecw6ixaf75o4763th5in98arnd
Modul:translations
828
9941
375372
344653
2026-09-22T05:12:48Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708794|92708794]])
375372
Scribunto
text/plain
local export = {}
local anchors_module = "Module:anchors"
local debug_track_module = "Module:debug/track"
local languages_module = "Module:languages"
local links_module = "Module:links"
local pages_module = "Module:pages"
local parameters_module = "Module:parameters"
local string_utilities_module = "Module:string utilities"
local templatestyles_module = "Module:TemplateStyles"
local utilities_module = "Module:utilities"
local wikimedia_languages_module = "Module:wikimedia languages"
local concat = table.concat
local html_create = mw.html.create
local insert = table.insert
local load_data = mw.loadData
local new_title = mw.title.new
local require = require
--[==[
Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==]
local function decode_uri(...)
decode_uri = require(string_utilities_module).decode_uri
return decode_uri(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function get_wikimedia_lang(...)
get_wikimedia_lang = require(wikimedia_languages_module).getByCode
return get_wikimedia_lang(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function normalize_anchor(...)
normalize_anchor = require(anchors_module).normalize_anchor
return normalize_anchor(...)
end
local function plain_link(...)
plain_link = require(links_module).plain_link
return plain_link(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function remove_links(...)
remove_links = require(links_module).remove_links
return remove_links(...)
end
local function split_on_slashes(...)
split_on_slashes = require(links_module).split_on_slashes
return split_on_slashes(...)
end
local function templatestyles(...)
templatestyles = require(templatestyles_module)
return templatestyles(...)
end
local function track(...)
track = require(debug_track_module)
return track(...)
end
--[==[
Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==]
local en
local function get_en()
en, get_en = require(languages_module).getByCode("ms"), nil
return en
end
local headword_data
local function get_headword_data()
headword_data, get_headword_data = load_data("Module:headword/data"), nil
return headword_data
end
local parameters_data
local function get_parameters_data()
parameters_data, get_parameters_data = load_data("Module:parameters/data"), nil
return parameters_data
end
local translations_data
local function get_translations_data()
translations_data, get_translations_data = load_data("Module:translations/data"), nil
return translations_data
end
local function is_translation_subpage(pagename)
if (headword_data or get_headword_data()).page.namespace ~= "" then
return false
elseif not pagename then
pagename = (headword_data or get_headword_data()).encoded_pagename
end
return pagename:match("./translations$") and true or false
end
local function canonical_pagename()
local pagename = (headword_data or get_headword_data()).encoded_pagename
return is_translation_subpage(pagename) and pagename:sub(1, -14) or pagename
end
local function interwiki(terminfo, term, lang, langcode)
-- No interwiki link if term is empty/missing
if not term or #term < 1 then
terminfo.interwiki = false
return
end
-- Percent-decode the term.
term = decode_uri(terminfo.term, "PATH")
-- Don't show an interwiki link if it's an invalid title.
if not new_title(term) then
terminfo.interwiki = false
return
end
local interwiki_langcode = (translations_data or get_translations_data()).interwiki_langs[langcode]
local wmlangs = interwiki_langcode and {get_wikimedia_lang(interwiki_langcode)} or lang:getWikimediaLanguages()
-- Don't show the interwiki link if the language is not recognised by Wikimedia.
if #wmlangs == 0 then
terminfo.interwiki = false
return
end
local sc = terminfo.sc
local target_page = get_link_page(term, lang, sc)
local split = split_on_slashes(target_page)
if not split[1] then
terminfo.interwiki = false
return
end
target_page = split[1]
local wmlangcode = wmlangs[1]:getCode()
local interwiki_link = language_link{
lang = lang,
sc = sc,
term = wmlangcode .. ":" .. target_page,
alt = "(" .. wmlangcode .. ")",
tr = "-"
}
terminfo.interwiki = tostring(html_create("span")
:addClass("tpos")
:wikitext(" " .. interwiki_link)
)
end
function export.show_terminfo(terminfo, check)
local lang = terminfo.lang
local langcode, langname = lang:getCode(), lang:getCanonicalName()
-- Translations must be for mainspace languages.
if not lang:hasType("regular") then
error("Translations must be for attested and approved main-namespace languages.")
else
local disallowed = (translations_data or get_translations_data()).disallowed
local err_msg = disallowed[langcode]
if err_msg then
error("Translations not allowed in " .. langname .. " (" .. langcode .. "). " .. langname .. " translations should " .. err_msg)
end
local fullcode = lang:getFullCode()
if fullcode ~= langcode then
err_msg = disallowed[fullcode]
if err_msg then
langname = lang:getFullName()
error("Translations not allowed in " .. langname .. " (" .. fullcode .. "). " .. langname .. " translations should " .. err_msg)
end
end
end
if langcode == "ms" then
if terminfo.interwiki then
error("Interwiki translations not allowed for English; they should always link to a different Wiktionary")
end
local current_L2 = require(pages_module).get_current_L2()
if current_L2 ~= "Rentas bahasa" and mw.title.getCurrentTitle().nsText ~= "Wikikamus" then
if current_L2 then
error("English translations only allowed in Translingual section, not in " .. current_L2)
else
error("English translations only allowed in Translingual section, not outside of any L2")
end
end
end
local term = terminfo.term
-- Check if there is a term. Don't show the interwiki link if there is nothing to link to.
if not term then
-- Track entries that don't provide a term.
-- FIXME: This should be a category.
track("translations/no term")
track("translations/no term/" .. langcode)
end
if terminfo.interwiki then
interwiki(terminfo, term, lang, langcode)
end
langcode = lang:getFullCode()
if (translations_data or get_translations_data()).need_super[langcode] then
local tr = terminfo.tr
if tr ~= nil then
terminfo.tr = tr:gsub("%d[%d%*%-]*%f[^%d%*]", "<sup>%0</sup>")
end
end
terminfo.show_decorations = true
local link = full_link(terminfo, "translation")
local canonical_name = lang:getCanonicalName()
local full_name = lang:getFullName()
local categories = {"Perkataan dengan terjemahan bahasa " .. canonical_name}
if canonical_name ~= full_name then
insert(categories, "Perkataan dengan terjemahan bahasa " .. full_name)
end
if check then
link = tostring(html_create("span")
:addClass("ttbc")
:tag("sup")
:addClass("ttbc")
:wikitext("(sila [[WT:Terjemahan#Terjemahan untuk disemak|sahkan]])")
:done()
:wikitext(" " .. link)
)
insert(categories, "Permintaan pengesahan terjemahan bahasa " .. langname )
end
return link .. format_categories(categories, en or get_en(), nil, canonical_pagename())
end
-- Implements {{t}}, {{t+}}, {{t-check}} and {{t+check}}.
function export.show(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation"])
local check = frame.args.check
return export.show_terminfo({
lang = args[1],
sc = args.sc,
track_sc = true,
term = args[2],
alt = args.alt,
id = args.id,
genders = args[3],
tr = args.tr,
ts = args.ts,
lit = args.lit,
q = args.q,
qq = args.qq,
l = args.l,
ll = args.ll,
refs = args.ref,
interwiki = frame.args.interwiki,
}, check and check ~= "")
end
local function add_id(div, id)
return id and div:attr("id", normalize_anchor("Translations-" .. id)) or div
end
-- Implements {{ter-atas}} and part of {{ter-atas-juga}}.
local function top(args, title, id, navhead)
local column_width = (args["column-width"] == "wide" or args["column-width"] == "narrow") and "-" .. args["column-width"] or ""
local div = html_create("div")
:addClass("NavFrame")
:node(navhead)
:tag("div")
:addClass("NavContent")
:tag("table")
:addClass("translations")
:attr("role", "presentation")
:attr("data-gloss", title or "")
:tag("tr")
:tag("td")
:addClass("translations-cell")
:addClass("multicolumn-list" .. column_width)
:attr("colspan", "3")
:allDone()
div = add_id(div, id)
local categories = {}
if not title then
insert(categories, "Pengepala jadual terjemahan kekurangan padanan kata")
end
local pagename = canonical_pagename()
if is_translation_subpage() then
insert(categories, "Sublaman terjemahan")
end
return (tostring(div):gsub("</td></tr></table></div></div>$", "")) ..
(#categories > 0 and format_categories(categories, en or get_en(), nil, pagename) or "") ..
-- Category to trigger [[MediaWiki:Gadget-TranslationAdder.js]]; we want this even on
-- user pages and such.
format_categories("Perkataan dengan kotak terjemahan", nil, nil, nil, true) ..
templatestyles("Modul:translations/styles.css")
end
-- Entry point for {{ter-atas}}.
function export.top(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas"])
local title = args[1]
local id = args.id or title
title = title and remove_links(title)
return top(args, title, id, html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(title or "Terjemahan")
)
end
-- Entry point for {{checktrans-top}}.
function export.check_top(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["checktrans-top"])
local text = "\n:''The translations below need to be checked and inserted above into the appropriate translation tables. See instructions at " ..
frame:expandTemplate{
title = "section link",
args = {"Wikikamus:Susun atur entri#Terjemahan"}
} ..
".''\n"
local header = html_create("div")
:addClass("checktrans")
:wikitext(text)
local subtitle = args[1]
local title = "Terjemahan untuk disemak"
if subtitle then
title = title .. "‌: \"" .. subtitle .. "\""
end
-- No ID, since these should always accompany proper translation tables, and can't be trusted anyway (i.e. there's no use-case for links).
return tostring(header) .. "\n" .. top(args, title, nil, html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(title or "Terjemahan")
)
end
-- Implements {{ter-bawah}}.
function export.bottom(frame)
-- Check nothing is being passed as a parameter.
process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-bawah"])
return "</table></div></div>"
end
-- Implements {{ter-lihat}} and part of {{ter-atas-juga}}.
local function see(args, see_text)
local navhead = html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(args[1] .. " ")
:tag("span")
:css("font-weight", "normal")
:wikitext("— ")
:tag("i")
:wikitext(see_text)
:allDone()
local terms, id = args[2], args.id
if #terms == 0 then
terms[1] = args[1]
end
for i = 1, #terms do
local term_id = id[i] or id.default
local data = {
term = terms[i],
id = term_id and "Translations-" .. term_id or "Translations",
}
terms[i] = plain_link(data)
end
return navhead:wikitext(concat(terms, ",‎ "))
end
-- Entry point for {{ter-lihat}}.
function export.see(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-lihat"])
local div = html_create("div")
:addClass("pseudo")
:addClass("NavFrame")
:node(see(args, "see "))
return tostring(add_id(div, args.id.default or args[1]))
end
-- Entry point for {{ter-atas-juga}}.
function export.top_also(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas-juga"])
local navhead = see(args, "lihat juga ")
local title = args[1]
local id = args.id.default or title
title = remove_links(title)
return top(args, title, id, navhead)
end
-- Implements {{translation subpage}}.
function export.subpage(frame)
process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation subpage"])
if not is_translation_subpage() then
error("This template should only be used on translation subpages, which have titles that end with '/translations'.")
end
-- "Translation subpages" category is handled by {{trans-top}}.
return ("''This page contains translations for ''%s''. See the main entry for more information.''"):format(full_link{
lang = en or get_en(),
term = canonical_pagename(),
})
end
-- Implements {{t-needed}}.
function export.needed(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["t-needed"])
local lang, category = args[1], ""
local span = html_create("span")
:addClass("trreq")
:attr("data-lang", lang:getCode())
:tag("i")
:wikitext("sila tambah terjemahan ini jika boleh")
:done()
if not args.nocat then
local type, sort = args[2], args.sort
if type == "quote" then
category = "Permintaan terjemahan petikan bahasa " .. lang:getCanonicalName()
elseif type == "usex" then
category = "Permintaan terjemahan contoh penggunaan bahasa " .. lang:getCanonicalName()
else
category = "Permintaan terjemahan ke dalam bahasa " .. lang:getCanonicalName()
lang = en or get_en()
end
category = format_categories(category, lang, sort, not sort and canonical_pagename() or nil)
end
return tostring(span) .. category
end
-- Implements {{no equivalent translation}}.
function export.no_equivalent(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no equivalent translation"])
local text = "tiada padanan kata dalam bahasa " .. args[1]:getCanonicalName()
if not args.noend then
text = text .. ", tetapi lihat"
end
return tostring(html_create("i"):wikitext(text))
end
-- Implements {{no attested translation}}.
function export.no_attested(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no attested translation"])
local langname = args[1]:getCanonicalName()
local text = "tiada perkataan [[WT:ATTEST|disahkan]] dalam bahasa " .. langname
local category = ""
if not args.noend then
text = text .. ", but see"
local sort = args.sort
category = format_categories("Terjemahan bahasa " .. langname .. " yang tidak disahkan", en or get_en(), sort, not sort and canonical_pagename() or nil)
end
return tostring(html_create("i"):wikitext(text)) .. category
end
-- Implements {{not used}}.
function export.not_used(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["not used"])
return tostring(html_create("i"):wikitext((args[2] or "tidak digunakan") .. " dalam " .. args[1]:getCanonicalName()))
end
return export
hj4av85y0jqga4vvfg9ql3zt8i96adp
375388
375372
2026-09-22T07:30:23Z
Hakimi97
2668
Betulkan nama English ke Malay
375388
Scribunto
text/plain
local export = {}
local anchors_module = "Module:anchors"
local debug_track_module = "Module:debug/track"
local languages_module = "Module:languages"
local links_module = "Module:links"
local pages_module = "Module:pages"
local parameters_module = "Module:parameters"
local string_utilities_module = "Module:string utilities"
local templatestyles_module = "Module:TemplateStyles"
local utilities_module = "Module:utilities"
local wikimedia_languages_module = "Module:wikimedia languages"
local concat = table.concat
local html_create = mw.html.create
local insert = table.insert
local load_data = mw.loadData
local new_title = mw.title.new
local require = require
--[==[
Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no overhead after the first call, since the target functions are called directly in any subsequent calls.]==]
local function decode_uri(...)
decode_uri = require(string_utilities_module).decode_uri
return decode_uri(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function get_wikimedia_lang(...)
get_wikimedia_lang = require(wikimedia_languages_module).getByCode
return get_wikimedia_lang(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function normalize_anchor(...)
normalize_anchor = require(anchors_module).normalize_anchor
return normalize_anchor(...)
end
local function plain_link(...)
plain_link = require(links_module).plain_link
return plain_link(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function remove_links(...)
remove_links = require(links_module).remove_links
return remove_links(...)
end
local function split_on_slashes(...)
split_on_slashes = require(links_module).split_on_slashes
return split_on_slashes(...)
end
local function templatestyles(...)
templatestyles = require(templatestyles_module)
return templatestyles(...)
end
local function track(...)
track = require(debug_track_module)
return track(...)
end
--[==[
Loaders for objects, which load data (or some other object) into some variable, which can then be accessed as "foo or get_foo()", where the function get_foo sets the object to "foo" and then returns it. This ensures they are only loaded when needed, and avoids the need to check for the existence of the object each time, since once "foo" has been set, "get_foo" will not be called again.]==]
local en
local function get_en()
en, get_en = require(languages_module).getByCode("ms"), nil
return en
end
local headword_data
local function get_headword_data()
headword_data, get_headword_data = load_data("Module:headword/data"), nil
return headword_data
end
local parameters_data
local function get_parameters_data()
parameters_data, get_parameters_data = load_data("Module:parameters/data"), nil
return parameters_data
end
local translations_data
local function get_translations_data()
translations_data, get_translations_data = load_data("Module:translations/data"), nil
return translations_data
end
local function is_translation_subpage(pagename)
if (headword_data or get_headword_data()).page.namespace ~= "" then
return false
elseif not pagename then
pagename = (headword_data or get_headword_data()).encoded_pagename
end
return pagename:match("./translations$") and true or false
end
local function canonical_pagename()
local pagename = (headword_data or get_headword_data()).encoded_pagename
return is_translation_subpage(pagename) and pagename:sub(1, -14) or pagename
end
local function interwiki(terminfo, term, lang, langcode)
-- No interwiki link if term is empty/missing
if not term or #term < 1 then
terminfo.interwiki = false
return
end
-- Percent-decode the term.
term = decode_uri(terminfo.term, "PATH")
-- Don't show an interwiki link if it's an invalid title.
if not new_title(term) then
terminfo.interwiki = false
return
end
local interwiki_langcode = (translations_data or get_translations_data()).interwiki_langs[langcode]
local wmlangs = interwiki_langcode and {get_wikimedia_lang(interwiki_langcode)} or lang:getWikimediaLanguages()
-- Don't show the interwiki link if the language is not recognised by Wikimedia.
if #wmlangs == 0 then
terminfo.interwiki = false
return
end
local sc = terminfo.sc
local target_page = get_link_page(term, lang, sc)
local split = split_on_slashes(target_page)
if not split[1] then
terminfo.interwiki = false
return
end
target_page = split[1]
local wmlangcode = wmlangs[1]:getCode()
local interwiki_link = language_link{
lang = lang,
sc = sc,
term = wmlangcode .. ":" .. target_page,
alt = "(" .. wmlangcode .. ")",
tr = "-"
}
terminfo.interwiki = tostring(html_create("span")
:addClass("tpos")
:wikitext(" " .. interwiki_link)
)
end
function export.show_terminfo(terminfo, check)
local lang = terminfo.lang
local langcode, langname = lang:getCode(), lang:getCanonicalName()
-- Translations must be for mainspace languages.
if not lang:hasType("regular") then
error("Translations must be for attested and approved main-namespace languages.")
else
local disallowed = (translations_data or get_translations_data()).disallowed
local err_msg = disallowed[langcode]
if err_msg then
error("Translations not allowed in " .. langname .. " (" .. langcode .. "). " .. langname .. " translations should " .. err_msg)
end
local fullcode = lang:getFullCode()
if fullcode ~= langcode then
err_msg = disallowed[fullcode]
if err_msg then
langname = lang:getFullName()
error("Translations not allowed in " .. langname .. " (" .. fullcode .. "). " .. langname .. " translations should " .. err_msg)
end
end
end
if langcode == "ms" then
if terminfo.interwiki then
error("Interwiki translations not allowed for Malay; they should always link to a different Wiktionary")
end
local current_L2 = require(pages_module).get_current_L2()
if current_L2 ~= "Rentas bahasa" and mw.title.getCurrentTitle().nsText ~= "Wikikamus" then
if current_L2 then
error("Malay translations only allowed in Translingual section, not in " .. current_L2)
else
error("Malay translations only allowed in Translingual section, not outside of any L2")
end
end
end
local term = terminfo.term
-- Check if there is a term. Don't show the interwiki link if there is nothing to link to.
if not term then
-- Track entries that don't provide a term.
-- FIXME: This should be a category.
track("translations/no term")
track("translations/no term/" .. langcode)
end
if terminfo.interwiki then
interwiki(terminfo, term, lang, langcode)
end
langcode = lang:getFullCode()
if (translations_data or get_translations_data()).need_super[langcode] then
local tr = terminfo.tr
if tr ~= nil then
terminfo.tr = tr:gsub("%d[%d%*%-]*%f[^%d%*]", "<sup>%0</sup>")
end
end
terminfo.show_decorations = true
local link = full_link(terminfo, "translation")
local canonical_name = lang:getCanonicalName()
local full_name = lang:getFullName()
local categories = {"Perkataan dengan terjemahan bahasa " .. canonical_name}
if canonical_name ~= full_name then
insert(categories, "Perkataan dengan terjemahan bahasa " .. full_name)
end
if check then
link = tostring(html_create("span")
:addClass("ttbc")
:tag("sup")
:addClass("ttbc")
:wikitext("(sila [[WT:Terjemahan#Terjemahan untuk disemak|sahkan]])")
:done()
:wikitext(" " .. link)
)
insert(categories, "Permintaan pengesahan terjemahan bahasa " .. langname )
end
return link .. format_categories(categories, en or get_en(), nil, canonical_pagename())
end
-- Implements {{t}}, {{t+}}, {{t-check}} and {{t+check}}.
function export.show(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation"])
local check = frame.args.check
return export.show_terminfo({
lang = args[1],
sc = args.sc,
track_sc = true,
term = args[2],
alt = args.alt,
id = args.id,
genders = args[3],
tr = args.tr,
ts = args.ts,
lit = args.lit,
q = args.q,
qq = args.qq,
l = args.l,
ll = args.ll,
refs = args.ref,
interwiki = frame.args.interwiki,
}, check and check ~= "")
end
local function add_id(div, id)
return id and div:attr("id", normalize_anchor("Translations-" .. id)) or div
end
-- Implements {{ter-atas}} and part of {{ter-atas-juga}}.
local function top(args, title, id, navhead)
local column_width = (args["column-width"] == "wide" or args["column-width"] == "narrow") and "-" .. args["column-width"] or ""
local div = html_create("div")
:addClass("NavFrame")
:node(navhead)
:tag("div")
:addClass("NavContent")
:tag("table")
:addClass("translations")
:attr("role", "presentation")
:attr("data-gloss", title or "")
:tag("tr")
:tag("td")
:addClass("translations-cell")
:addClass("multicolumn-list" .. column_width)
:attr("colspan", "3")
:allDone()
div = add_id(div, id)
local categories = {}
if not title then
insert(categories, "Pengepala jadual terjemahan kekurangan padanan kata")
end
local pagename = canonical_pagename()
if is_translation_subpage() then
insert(categories, "Sublaman terjemahan")
end
return (tostring(div):gsub("</td></tr></table></div></div>$", "")) ..
(#categories > 0 and format_categories(categories, en or get_en(), nil, pagename) or "") ..
-- Category to trigger [[MediaWiki:Gadget-TranslationAdder.js]]; we want this even on
-- user pages and such.
format_categories("Perkataan dengan kotak terjemahan", nil, nil, nil, true) ..
templatestyles("Modul:translations/styles.css")
end
-- Entry point for {{ter-atas}}.
function export.top(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas"])
local title = args[1]
local id = args.id or title
title = title and remove_links(title)
return top(args, title, id, html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(title or "Terjemahan")
)
end
-- Entry point for {{checktrans-top}}.
function export.check_top(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["checktrans-top"])
local text = "\n:''The translations below need to be checked and inserted above into the appropriate translation tables. See instructions at " ..
frame:expandTemplate{
title = "section link",
args = {"Wikikamus:Susun atur entri#Terjemahan"}
} ..
".''\n"
local header = html_create("div")
:addClass("checktrans")
:wikitext(text)
local subtitle = args[1]
local title = "Terjemahan untuk disemak"
if subtitle then
title = title .. "‌: \"" .. subtitle .. "\""
end
-- No ID, since these should always accompany proper translation tables, and can't be trusted anyway (i.e. there's no use-case for links).
return tostring(header) .. "\n" .. top(args, title, nil, html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(title or "Terjemahan")
)
end
-- Implements {{ter-bawah}}.
function export.bottom(frame)
-- Check nothing is being passed as a parameter.
process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-bawah"])
return "</table></div></div>"
end
-- Implements {{ter-lihat}} and part of {{ter-atas-juga}}.
local function see(args, see_text)
local navhead = html_create("div")
:addClass("NavHead")
:css("text-align", "left")
:wikitext(args[1] .. " ")
:tag("span")
:css("font-weight", "normal")
:wikitext("— ")
:tag("i")
:wikitext(see_text)
:allDone()
local terms, id = args[2], args.id
if #terms == 0 then
terms[1] = args[1]
end
for i = 1, #terms do
local term_id = id[i] or id.default
local data = {
term = terms[i],
id = term_id and "Translations-" .. term_id or "Translations",
}
terms[i] = plain_link(data)
end
return navhead:wikitext(concat(terms, ",‎ "))
end
-- Entry point for {{ter-lihat}}.
function export.see(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-lihat"])
local div = html_create("div")
:addClass("pseudo")
:addClass("NavFrame")
:node(see(args, "see "))
return tostring(add_id(div, args.id.default or args[1]))
end
-- Entry point for {{ter-atas-juga}}.
function export.top_also(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["ter-atas-juga"])
local navhead = see(args, "lihat juga ")
local title = args[1]
local id = args.id.default or title
title = remove_links(title)
return top(args, title, id, navhead)
end
-- Implements {{translation subpage}}.
function export.subpage(frame)
process_params(frame:getParent().args, (parameters_data or get_parameters_data())["translation subpage"])
if not is_translation_subpage() then
error("This template should only be used on translation subpages, which have titles that end with '/translations'.")
end
-- "Translation subpages" category is handled by {{trans-top}}.
return ("''This page contains translations for ''%s''. See the main entry for more information.''"):format(full_link{
lang = en or get_en(),
term = canonical_pagename(),
})
end
-- Implements {{t-needed}}.
function export.needed(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["t-needed"])
local lang, category = args[1], ""
local span = html_create("span")
:addClass("trreq")
:attr("data-lang", lang:getCode())
:tag("i")
:wikitext("sila tambah terjemahan ini jika boleh")
:done()
if not args.nocat then
local type, sort = args[2], args.sort
if type == "quote" then
category = "Permintaan terjemahan petikan bahasa " .. lang:getCanonicalName()
elseif type == "usex" then
category = "Permintaan terjemahan contoh penggunaan bahasa " .. lang:getCanonicalName()
else
category = "Permintaan terjemahan ke dalam bahasa " .. lang:getCanonicalName()
lang = en or get_en()
end
category = format_categories(category, lang, sort, not sort and canonical_pagename() or nil)
end
return tostring(span) .. category
end
-- Implements {{no equivalent translation}}.
function export.no_equivalent(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no equivalent translation"])
local text = "tiada padanan kata dalam bahasa " .. args[1]:getCanonicalName()
if not args.noend then
text = text .. ", tetapi lihat"
end
return tostring(html_create("i"):wikitext(text))
end
-- Implements {{no attested translation}}.
function export.no_attested(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["no attested translation"])
local langname = args[1]:getCanonicalName()
local text = "tiada perkataan [[WT:ATTEST|disahkan]] dalam bahasa " .. langname
local category = ""
if not args.noend then
text = text .. ", but see"
local sort = args.sort
category = format_categories("Terjemahan bahasa " .. langname .. " yang tidak disahkan", en or get_en(), sort, not sort and canonical_pagename() or nil)
end
return tostring(html_create("i"):wikitext(text)) .. category
end
-- Implements {{not used}}.
function export.not_used(frame)
local args = process_params(frame:getParent().args, (parameters_data or get_parameters_data())["not used"])
return tostring(html_create("i"):wikitext((args[2] or "tidak digunakan") .. " dalam " .. args[1]:getCanonicalName()))
end
return export
hg552qndw0fl79qt8gujt8rt8an4ytr
Modul:IPA
828
9946
375359
227150
2026-09-22T03:14:30Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92723663|92723663]])
375359
Scribunto
text/plain
local export = {}
local force_cat = false -- for testing
local decorations_module = "Module:decorations"
local pages_module = "Module:pages"
local qualifier_module = "Module:qualifier"
local string_utilities_module = "Module:string utilities"
local syllables_module = "Module:syllables"
local utilities_module = "Module:utilities"
local m_data = mw.loadData("Module:IPA/data")
local m_str_utils = require(string_utilities_module)
local m_syllables -- [[Module:syllables]]; loaded below if needed
local m_symbols = mw.loadData("Module:IPA/data/symbols")
local concat = table.concat
local decode_entities = m_str_utils.decode_entities
local find = string.find
local gcodepoint = m_str_utils.gcodepoint
local gmatch = m_str_utils.gmatch
local gsub = string.gsub
local insert = table.insert
local is_preview = require(pages_module).is_preview
local len = m_str_utils.len
local listToText = mw.text.listToText
local match = string.match
local pattern_escape = m_str_utils.pattern_escape
local sub = string.sub
local u = m_str_utils.char
local ugsub = m_str_utils.gsub
local umatch = m_str_utils.match
local usub = m_str_utils.sub
local function with_codepoints(s)
if find(s, "%S%s") then
local parts = {}
for ch in gmatch(s, "%S") do
parts[#parts + 1] = with_codepoints(ch)
end
return concat(parts, ", ")
end
local cps = {}
for cp in gcodepoint(s) do
cps[#cps + 1] = ("U+%04X"):format(cp)
end
return s .. " [" .. concat(cps, " ") .. "]"
end
local namespace = mw.title.getCurrentTitle().nsText
local function is_content_page(lang, namespace)
return namespace == "" or namespace == "Rekonstruksi" or
lang and lang:hasType("appendix-constructed") and namespace == "Lampiran"
end
-- Etymology-only languages are not L2 entry languages; IPA should use the parent full language.
local function assert_not_etymology_only_lang(lang)
if lang and lang.hasType and lang:hasType("language", "etymology-only") then
local parent_code = lang.getParentCode and lang:getParentCode() or nil
error(("Cannot use IPA with the etymology-only language %q; use the parent full language %q instead."):format(lang:getCode(), parent_code))
end
end
local function track(page)
require("Module:debug/track")("IPA/" .. page)
return true
end
local function process_maybe_split_categories(split_output, categories, prontext, lang, errtext)
if split_output ~= "raw" then
if categories[1] then
categories = require(utilities_module).format_categories(categories, lang, nil, nil, force_cat)
else
categories = ""
end
end
if split_output then -- for use of IPA in links, etc.
if errtext then
return prontext, categories, errtext
else
return prontext, categories
end
else
return prontext .. (errtext or "") .. categories
end
end
--[==[
Format a line of one or more IPA pronunciations as {{tl|IPA}} would do it, i.e. with a preceding {"IPA:"} followed by
the word {"key"} linking to an Appendix page describing the language's phonology, and with an added category
` ``lang`` terms with IPA pronunciation`. Other than the extra preceding text and category, this is identical
to {format_IPA_multiple()}, and the considerations described there in the documentation apply here as well. There is a
single parameter `data`, an object with the following fields:
* `lang`: Object representing the language of the pronunciations, which is used when adding cleanup categories for
pronunciations with invalid phonemes; for determining how many syllables the pronunciations have in them, in order to
add a category such as [[:Category:Italian 2-syllable words]] (for certain languages only); for adding a category
` ``lang`` terms with IPA pronunciation`; and for determining the proper sort keys for categories. Unlike
for {format_IPA_multiple()}, `lang` may not be {nil}.
* `items`: List of pronunciations, in exactly the same format as for {format_IPA_multiple()}.
* `err`: If not {nil}, a string containing an error message to use in place of the link to the language's phonology.
* `separator`: The default separator to use when separating formatted items. Defaults to {", "}. Does not apply to the
first item, where the default separator is always the empty string. Overridden by the per-item `separator` field in
`items`.
* `sort_key`: Explicit sort key used for categories.
* `no_count`: Suppress adding a {#-syllable words} category such as [[:Category:Italian 2-syllable words]]. Note that
only certain languages add such categories to begin with, because it depends on knowing how to count syllables in a
given language, which depends on the phonology of the language. Also, this does not suppress the addition of cleanup
or other categories. If you need them suppressed, use `split_output` to return the categories separately and ignore
them.
* `split_output`: If not given, the return value is a concatenation of the formatted pronunciation and formatted
categories. Otherwise, two values are returned: the formatted pronunciation and the categories. If `split_output` is
the value {"raw"}, the categories are returned in list form, where the list elements are a combination of category
strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]]. If
`split_output` is any other value besides {nil}, the categories are returned as a pre-formatted concatenated string.
* `include_langname`: If specified, prefix the result with the language name, followed by a colon.
* `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display at the beginning, before the formatted
pronunciations and preceding {"IPA:"}.
* `qq`: {nil} or a list of right qualifiers to display after all formatted pronunciations.
* `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display at the beginning, before the formatted
pronunciations and preceding {"IPA:"}.
* `aa`: {nil} or a list of right accent qualifiers to display after all formatted pronunciations.
]==]
function export.format_IPA_full(data)
if type(data) ~= "table" or data.getCode then
error("Must now supply a table of arguments to format_IPA_full(); first argument should be that table, not a language object")
end
local lang = data.lang
local items = data.items
local err = data.err
local separator = data.separator
local sort_key = data.sort_key
local no_count = data.no_count
local split_output = data.split_output
local q = data.q
local qq = data.qq
local a = data.a
local aa = data.aa
local include_langname = data.include_langname
if data.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("overall `.qualifiers` is no longer supported; change the code to use `.q` or `.qq`")
end
local hasKey = m_data.langs_with_infopages
if not lang or not lang.getCode then
error("Must specify language to format_IPA_full()")
end
assert_not_etymology_only_lang(lang)
local langname = lang:getCanonicalName()
local prefix_text
if err then
prefix_text = '<span class="error">' .. err .. '</span>'
else
if hasKey[lang:getCode()] then
prefix_text = "Lampiran:Sebutan bahasa " .. langname
else
prefix_text = "wikipedia:Fonologi bahasa " .. langname
end
prefix_text = "[[" .. prefix_text .. "|kekunci]]"
end
local prefix = "[[Wikikamus:Abjad Fonetik Antarabangsa|AFA]]<sup>(" .. prefix_text .. ")</sup>: "
local IPAs, categories = export.format_IPA_multiple(lang, items, separator, no_count, "raw")
if is_content_page(lang, namespace) then
insert(categories, {
cat = "Perkataan dengan sebutan AFA bahasa " .. langname,
sort_key = sort_key
})
end
local prontext = prefix .. IPAs
if q and q[1] or qq and qq[1] or a and a[1] or aa and aa[1] then
prontext = require(decorations_module).format_decorations {
lang = lang,
text = prontext,
q = q,
qq = qq,
a = a,
aa = aa,
}
end
if include_langname then
prontext = langname .. ": " .. prontext
end
return process_maybe_split_categories(split_output, categories, prontext, lang)
end
local function split_phonemic_phonetic(pron)
local reconstructed, phonemic, phonetic = match(pron, "^(%*?)(/.-/)%s+(%[.-%])$")
if reconstructed then
return reconstructed .. phonemic, reconstructed .. phonetic
else
return pron, nil
end
end
local function determine_repr(pron)
local reconstructed
-- Temporarily remove any initial asterisk before representation marks,
-- which avoids having to account for it in the data, but set the
-- `reconstructed` flag.
if sub(pron, 1, 1) == "*" then
reconstructed = true
pron = sub(pron, 2)
end
-- Some representation types have aliases for convenience (e.g. "// //" is
-- an alias for "⫽ ⫽"). and these need to be substituted in before checking
-- for other data.
local opening, n = match(pron, "^.[\128-\191]*")
local subs_data = m_data.representation_subs[opening]
if subs_data then
pron, n = ugsub(pron, subs_data[1], subs_data[2])
-- If the substitution was made, `opening` needs to be changed to the
-- new opening character.
if n ~= 0 then
opening = subs_data[3]
end
end
-- Get the type data based on the opening character (if any), and set the
-- representation type if the closing character matches.
local type_data, repr, closing = m_data.representation_types[opening]
if type_data then
closing = type_data[2]
if type_data and match(pron, pattern_escape(closing) .. "$", #opening + 1) then
repr = type_data[1]
end
end
-- Default to the empty string.
if not repr then
opening, closing = "", ""
end
-- Reattach the asterisk if reconstructed.
if reconstructed then
pron = "*" .. pron
end
return pron, repr, opening, closing, reconstructed
end
local function hasInvalidSeparators(transcription)
-- Escape certain characters as well as pauses, which have the format "(...)" (with any number of dots), to avoid false-positives.
transcription = transcription:gsub(".[\128-\191]*", m_symbols.separator_escapes)
:gsub("%(%.+%)", "\3")
:gsub("[()]+", "")
return (
transcription:find("..", nil, true) or
transcription:match("%.%f[%z \1\2\3,:;]") or
transcription:match("\1%f[%z \2\3,:;]") or
transcription:match("\2%f[%z \1\3,:;]") or
transcription:match("\3[:;]") or
transcription:match("%f[^%z \1\2\3,]%.")
) and true or false
end
--[==[
Format a line of one or more bare IPA pronunciations (i.e. without any preceding {"IPA:"} and without adding to a
category ` ``lang`` terms with IPA pronunciation`). Individual pronunciations are formatted using
{format_IPA()} and are combined with separators, decorations, pre-text, post-text, etc. to form a line of pronunciations.
Parameters accepted are:
* `lang` is an object representing the language of the pronunciations, which is used when adding cleanup categories for
pronunciations with invalid phonemes; for determining how many syllables the pronunciations have in them, in order to
add a category such as [[:Category:Italian 2-syllable words]] (for certain languages only); and for computing the
proper sort keys for categories. `lang` may be {nil}.
* `items` is a list of pronunciations, each of which is an object with the following properties:
** `pron`: the pronunciation, in the same format as is accepted by {format_IPA()}, i.e. it should be either phonemic
(surrounded by {/.../}), phonetic (surrounded by {[...]}), orthographic (surrounded by {⟨...⟩}) or a rhyme
(beginning with a hyphen);
** `pretext`: text to display directly before the formatted pronunciation, inside of any qualifiers or accent
qualifiers;
** `posttext`: text to display directly after the formatted pronunciation, inside of any qualifiers or accent
qualifiers;
** `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display before the formatted pronunciation;
** `qq`: {nil} or a list of right qualifiers to display after the formatted pronunciation;
** `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display before the formatted pronunciation;
** `aa`: {nil} or a list of right accent qualifiers to after before the formatted pronunciation;
** `refs`: {nil} or a list of references or reference specs to add after the pronunciation and any posttext and
qualifiers; the value of a list item is either a string containing the reference text (typically a call to a
citation template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text`
(the reference text), `name` (the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}}
or {{cd|<nowiki><ref name="foo" /></nowiki>}}) and/or `group` (the group of the reference, as in
{{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or
{{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference
appropriately and insert a footnote number that hyperlinks to the actual reference, located in the
{{cd|<nowiki><references /></nowiki>}} section;
** `gloss`: {nil} or a gloss (definition) for this item, if different definitions have different pronunciations;
** `pos`: {nil} or a part of speech for this item, if different parts of speech have different pronunciations;
** `separator`: the separator text to insert directly before the formatted pronunciation and all decorations
and pre-text; defaults to the outer `separator` parameter.
* `separator`: The default separator to use when separating formatted items. Defaults to {", "}. Does not apply to the
first item, where the default separator is always the empty string. Overridden by the per-item `separator` field in
`items`.
* `no_count`: Suppress adding a {#-syllable words} category such as [[:Category:Italian 2-syllable words]]. Note that
only certain languages add such categories to begin with, because it depends on knowing how to count syllables in a
given language, which depends on the phonology of the language. Also, this does not suppress the addition of cleanup
categories. If you need them suppressed, use `split_output` to return the categories separately and ignore them.
* `split_output`: If not given, the return value is a concatenation of the formatted pronunciation and formatted
categories. Otherwise, two values are returned: the formatted pronunciation and the categories. If `split_output` is
the value {"raw"}, the categories are returned in list form, where the list elements are a combination of category
strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]]. If
`split_output` is any other value besides {nil}, the categories are returned as a pre-formatted concatenated string.
]==]
function export.format_IPA_multiple(lang, items, separator, no_count, split_output)
local categories = {}
separator = separator or ", "
if not lang then
track("format-multiple-nolang")
else
assert_not_etymology_only_lang(lang)
end
-- Format
if not items[1] then
if namespace == "Templat" then
insert(items, {pron = "/aɪ piː ˈeɪ/"})
else
insert(categories, "Templat sebutan tanpa sebutan")
end
end
local bits = {}
for i, item in ipairs(items) do
local bit
-- If the pronunciation is entirely empty, allow this and don't do anything, so that e.g. the pretext and/or
-- posttext can be specified to force something like ''unknown'' to appear in place of the pronunciation
-- (as happens e.g. when ? is used as a respelling in [[Module:ca-IPA]]; see [[guèiser]] for an example).
if item.pron == "" then
bit = ""
else
local item_categories, errtext
bit, item_categories, errtext = export.format_IPA(lang, item.pron, "raw")
bit = bit .. errtext
for _, cat in ipairs(item_categories) do
insert(categories, cat)
end
end
if item.pretext then
bit = item.pretext .. bit
end
if item.posttext then
bit = bit .. item.posttext
end
if item.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("`.qualifiers` is no longer supported; change the code to use `.q` or `.qq`")
end
local has_decorations = item.q and item.q[1] or item.qq and item.qq[1] or item.a and item.a[1] or
item.aa and item.aa[1] or item.refs and item.refs[1]
local has_gloss_or_pos = item.gloss or item.pos
if has_decorations or has_gloss_or_pos then
-- FIXME: Currently we tack the gloss and POS (in that order) onto the end of the regular left qualifiers.
-- Should we do something different?
local q = item.q
if has_gloss_or_pos then
q = mw.clone(item.q) or {}
if item.gloss then
local m_qualifier = require(qualifier_module)
insert(q, m_qualifier.wrap_qualifier_css("“", "quote") .. item.gloss ..
m_qualifier.wrap_qualifier_css("”", "quote"))
end
if item.pos then
-- FIXME: Consider expanding aliases as found in [[Module:headword/data]] or similar.
insert(q, item.pos)
end
end
bit = require(decorations_module).format_decorations {
lang = lang,
text = bit,
q = q,
qq = item.qq,
a = item.a,
aa = item.aa,
refs = item.refs,
}
end
bit = (item.separator or (i == 1 and "" or separator)) .. bit
insert(bits, bit)
--[=[ [[Special:WhatLinksHere/Wiktionary:Tracking/IPA/syntax-error]]
The length or gemination symbol should not appear after a syllable break or stress symbol. ]=]
-- The nature of the following pattern match is such that we don't have to split a combined '/.../ [...]' spec
-- into its parts in order to process.
if match(item.pron, "[.\203][\136\140]?\203[\144\145]") then -- [.ˈˌ][ːˑ]
track("syntax-error")
end
if lang then
-- Add syllable count if the language's diphthongs are listed in [[Module:syllables]].
-- Don't do this if the term has spaces, a liaison mark (‿) or isn't in mainspace.
if not no_count and namespace == "" then
m_syllables = m_syllables or require(syllables_module)
local langcode = lang:getCode()
if m_data.langs_to_generate_syllable_count_categories[langcode] then
local raw_phonemic, phonetic, use_it = split_phonemic_phonetic(item.pron)
local phonemic, repr = determine_repr(raw_phonemic)
if not phonetic then -- not a '/.../ [...]' combined pronunciation
if m_data.langs_to_use_phonetic_or_phonemic_notation[langcode] then
use_it = phonemic
elseif m_data.langs_to_use_phonetic_notation[langcode] then
use_it = repr == "phonetic" and phonemic or nil
else
use_it = repr == "phonemic" and phonemic or nil
end
elseif repr == "phonetic" then
use_it = phonetic
elseif repr == "phonemic" then
use_it = phonemic
end
-- Note: two uses of find with plain patterns is much faster than umatch with [ ‿].
if use_it and not (find(use_it, " ") or find(use_it, "‿")) then
local syllable_count = m_syllables.getVowels(use_it, lang)
if syllable_count then
insert(categories, "Perkataan " .. syllable_count .. " suku kata bahasa " .. lang:getCanonicalName())
end
end
end
end
end
end
return process_maybe_split_categories(split_output, categories, concat(bits), lang)
end
--[=[
Format a single IPA pronunciation, which cannot be a combined spec (such as {/.../ [...]}). This has been extracted from
{format_IPA()} to allow the latter to handle such combined specs. This works like {format_IPA()} but requires that
pre-created {err} (for error messages) and {categories} lists be passed in, and adds any generated error messages and
categories to those lists. A single value is returned, the pronunciation, which is usually the same as passed in, but
may have HTML added surrounding invalid characters so they appear in red.
]=]
local function format_one_IPA(lang, raw_pron, err, categories)
-- Disallow wikilinks.
if match(raw_pron, "%[%[.-%]%]") then
error("IPA input must not contain wikilinks.")
end
raw_pron = decode_entities(raw_pron)
-- Detect the type of transcription.
local pron, repr, opening, closing, reconstructed = determine_repr(raw_pron)
-- Strip any reconstruction asterisk and representation marks.
pron = sub(pron, #opening + 1 + (reconstructed and 1 or 0), -#closing - 1)
if not repr then
insert(categories, "Sebutan AFA dengan tanda perwakilan tidak sah")
-- insert(err, "tanda perwakilan tidak sah")
-- Removed because it's annoying when previewing pronunciation pages.
end
if repr ~= "orthographic" and lang and lang:getCode() == "en" and hasInvalidSeparators(pron) then
insert(categories, "Sebutan AFA bahasa Inggeris dengan pemisah tidak sah")
end
if pron == "" then
insert(categories, "Sebutan AFA dengan tiada sebutan")
end
-- Check for obsolete and nonstandard symbols
for _, symbol in ipairs(m_data.nonstandard) do
local result
for nonstandard in gmatch(pron, symbol) do
if not result then
result = {}
end
insert(result, nonstandard)
insert(categories,
{cat = "Sebutan AFA dengan aksara usang atau tidak standard", sort_key = nonstandard}
)
end
if result then
insert(err, "aksara usang atau tidak standard (" .. concat(result) .. ")")
break
end
end
--[[ Check for invalid symbols after removing the following:
1. wikilinks (handled above)
2. paired HTML tags
3. bolding
4. italics
5. asterisk at beginning of transcription
6. comma followed by spacing characters
7. superscripts enclosed in superscript parentheses ]]
local found_HTML
local result = gsub(pron, "<(%a+)[^>]*>([^<]+)</%1>",
function(tagName, content)
found_HTML = true
return content
end)
result = gsub(result, "'''([^']*)'''", "%1")
result = gsub(result, "''([^']*)''", "%1")
result = gsub(result, "^%*", "")
result = ugsub(result, ",%s+", "")
-- VS15
local vs15_class = "[" .. m_symbols.add_vs15 .. "]"
if umatch(pron, vs15_class) then
local vs15 = u(0xFE0E)
if find(result, vs15) then
result = gsub(result, vs15, "")
pron = gsub(pron, vs15, "")
end
pron = ugsub(pron, vs15_class, "%0" .. vs15)
end
if result ~= "" then
local content_page = is_content_page(lang, namespace)
if lang then
-- Get the per_lang_valid data, and convert any per-language valid sequences to spaces.
local per_lang_valid = m_symbols.per_lang_valid[lang:getCode()]
if per_lang_valid then
if type(per_lang_valid) == "table" then
for _, pattern in pairs(per_lang_valid) do
result = ugsub(result, pattern, " ")
end
else -- Should be a string.
result = ugsub(result, per_lang_valid, " ")
end
end
end
local suggestions = {}
-- Check for any invalid sequences, excluding anything in the per-language lookup table.
for k, v in pairs(m_symbols.invalid) do
if find(result, k, nil, true) then
insert(suggestions, with_codepoints(k) .. " dengan " .. with_codepoints(v))
end
end
if suggestions[1] then
local replacements = "menggantikan " .. listToText(suggestions)
if content_page then
error("Invalid IPA: " .. replacements)
end
insert(err, replacements)
end
-- Convert any valid character sequences to spaces
for _, pattern in pairs(m_symbols.valid) do
result = ugsub(result, pattern, " ")
end
if not match(result, "^ *$") then
local category = "Sebutan AFA dengan aksara AFA tidak sah"
if not content_page then
category = category .. "/non_mainspace"
end
insert(categories, category)
insert(err, "aksara AFA tidak sah: " .. with_codepoints(result))
end
end
if found_HTML then
insert(categories, "Sebutan AFA dengan tag HTML berpasangan")
end
if (repr == "phonemic" or repr == "rhyme") and lang and m_data.phonemes[lang:getCode()] then
local valid_phonemes = m_data.phonemes[lang:getCode()]
local rest = pron
local phonemes = {}
while #rest > 0 do
local longestmatch, longestmatch_len = "", 0
local rest_init = sub(rest, 1, 1)
if rest_init == "(" or rest_init == ")" then
longestmatch = rest_init
longestmatch_len = 1
else
for _, phoneme in ipairs(valid_phonemes) do
local phoneme_len = len(phoneme)
if phoneme_len > longestmatch_len and usub(rest, 1, phoneme_len) == phoneme then
longestmatch = phoneme
longestmatch_len = len(longestmatch)
end
end
end
if longestmatch_len > 0 then
insert(phonemes, longestmatch)
rest = usub(rest, longestmatch_len + 1)
else
local phoneme = usub(rest, 1, 1)
insert(phonemes, "<span style=\"color: var(--wikt-palette-red,red)\">" .. phoneme .. "</span>")
rest = usub(rest, 2)
insert(categories, "Sebutan AFA dengan fonem tidak sah/" .. lang:getCode())
track("fonem tidak sah/" .. phoneme)
end
end
pron = concat(phonemes)
end
return (reconstructed and "*" or "") .. opening .. pron .. closing
end
--[==[
Format an IPA pronunciation. This wraps the pronunciation in appropriate CSS classes and adds cleanup categories and
error messages as needed. The pronunciation `pron` should be either phonemic (surrounded by {/.../}), phonetic
(surrounded by {[...]}), orthographic (surrounded by {⟨...⟩}), a rhyme (beginning with a hyphen) or a combined
phonemic/phonetic spec (of the form {/.../ [...]}). `lang` indicates the language of the pronunciation and can be {nil}.
If not {nil}, and the specified language has data in [[Module:IPA/data]] indicating the allowed phonemes, then the page
will be added to a cleanup category and an error message displayed next to the outputted pronunciation. Note that {lang}
also determines sort key processing in the added cleanup categories. If `split_output` is not given, the return value is
a concatenation of the formatted pronunciation, error messages and formatted cleanup categories. Otherwise, three values
are returned: the formatted pronunciation, the cleanup categories and the concatenated error messages. If `split_output`
is the value {"raw"}, the cleanup categories are returned in list form, where the list elements are a combination of
category strings and category objects of the form suitable for passing to {format_categories()} in [[Module:utilities]].
If `split_output` is any other value besides {nil}, the cleanup categories are returned as a pre-formatted concatenated
string.
]==]
function export.format_IPA(lang, pron, split_output)
local err = {}
local categories = {}
-- `pron` shouldn't contain ref tags.
if match(pron, "\127'\"`UNIQ%-%-ref%-[%dA-F]+%-QINU`\"'\127") then
error("<ref> tags found inside pronunciation parameter.")
end
if not lang then
track("format-nolang")
else
assert_not_etymology_only_lang(lang)
end
local phonemic, phonetic = split_phonemic_phonetic(pron)
pron = format_one_IPA(lang, phonemic, err, categories)
if phonetic then
track("phonemic-phonetic") -- There's no benefit to supporting the "/.../ [...]" format within one parameter.
phonetic = format_one_IPA(lang, phonetic, err, categories)
pron = pron .. " " .. phonetic
end
if err[1] and is_preview() then
err = '<span class="error" style="font-size: small;> ' .. concat(err, ", ") .. "</span>"
else
err = ""
end
return process_maybe_split_categories(split_output, categories, '<span class="IPA nowrap">' .. pron .. "</span>", lang,
err)
end
--[==[
Format a line of one or more enPR pronunciations as {{tl|enPR}} would do it, i.e. with a preceding {"enPR:"} (linked to
[[Appendix:English pronunciation]]) followed by one or more formatted, comma-separated enPR pronunciations. The
pronunciations are formatted by wrapping them in the `AHD` and `enPR` CSS classes and adding any decorations
(qualifiers, accent qualifiers and references). In addition, the overall result is wrapped in any overall decorations.
There is a single parameter `data`, an object with the following fields:
* `items` is a list of enPR pronunciations, each of which is an object with the following properties:
** `pron`: the enPR pronunciation;
** `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display before the formatted pronunciation;
** `qq`: {nil} or a list of right qualifiers to display after the formatted pronunciation;
** `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display before the formatted pronunciation;
** `aa`: {nil} or a list of right accent qualifiers to after before the formatted pronunciation.
* `q`: {nil} or a list of left qualifiers (as in {{tl|q}}) to display at the beginning, before the formatted
pronunciations and preceding {"enPR:"}.
* `qq`: {nil} or a list of right qualifiers to display after all formatted pronunciations.
* `a`: {nil} or a list of left accent qualifiers (as in {{tl|a}}) to display at the beginning, before the formatted
pronunciations and preceding {"enPR:"}.
* `aa`: {nil} or a list of right accent qualifiers to display after all formatted pronunciations.
]==]
function export.format_enPR_full(data)
local prefix = "[[Appendix:English pronunciation|enPR]]: "
local lang = require("Module:languages").getByCode("en")
local parts = {}
for _, item in ipairs(data.items) do
local part = '<span class="AHD enPR">' .. item.pron .. "</span>"
if item.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("`.qualifiers` is no longer supported; change the code to use `.q` or `.qq`")
end
if item.q and item.q[1] or item.qq and item.qq[1] or item.a and item.a[1] or item.aa and item.aa[1] then
part = require(decorations_module).format_decorations {
lang = lang,
text = part,
q = item.q,
qq = item.qq,
a = item.a,
aa = item.aa,
}
end
insert(parts, part)
end
local prontext = prefix .. concat(parts, ", ")
if data.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("overall `.qualifiers` is no longer supported; change the code to use `.q` or `.qq`")
end
if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then
prontext = require(decorations_module).format_decorations {
lang = lang,
text = prontext,
q = data.q,
qq = data.qq,
a = data.a,
aa = data.aa,
}
end
return prontext
end
return export
1qhkiy8wyqpetrqk0p5yanxnrkmiezk
Modul:IPA/data
828
10199
375367
218692
2026-09-22T04:48:13Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91333763|91333763]])
375367
Scribunto
text/plain
local list_to_set = require("Module:table").listToSet
local data = {}
--[=[
A list of representation types (e.g. /foo/ for phonemic and [bar] for phonetic),
given as a table. The key is the opening character, the first value the
representation type, and the second value the closing symbol.]=]
data.representation_types = {
["/"] = {"phonemic", "/"},
["["] = {"phonetic", "]"},
["⫽"] = {"morphophonemic", "⫽"},
["⟨"] = {"orthographic", "⟩"},
["-"] = {"rhyme", ""},
}
--[=[
A list of convenience inputs for certain representation types. The key is the
opening character, and the table is a three-item array consisting of (1) an
mw.ustring.gsub pattern which is anchored to the start and end of the string,
with a single capture group that excludes the characters to be substituted,
(2) a corresponding replacement pattern to be used with the pattern, and (3) the
replacement opening character.]=]
data.representation_subs = {
["<"] = {"^<(.*)>$", "⟨%1⟩", "⟨"},
["/"] = {"^//(.*)//$", "⫽%1⫽", "⫽"},
}
--[=[
This should list the language codes of all languages that have a pronunciation
page in the appendix of the form ''Appendix:LANG pronunciation'', e.g.
[[Appendix:Russian pronunciation]]. For these languages, the text "key" next to
the generated pronunciation links to such pages; for other languages, it links
to the "LANG phonology" page in Wikipedia (which may or may not exist).
[[Module:IPA]] is responsible for this linking; see format_IPA_full().]=]
data.langs_with_infopages = list_to_set{
"acw",
"ady",
"ang",
"arc",
"ba",
"bg",
"bo",
"ca",
"cho",
"cmn",
"cs",
"cv",
"cy",
"da",
"de",
"dsb",
"dz",
"egl",
"egy",
"el",
"en",
"enm",
"eo",
"es",
"fa",
"fi",
"fo",
"fr",
"fy",
"ga",
"gd",
"ghc",
"gmh",
"gmw-msc",
"got",
"he",
"hi",
"hrx",
"hu",
"hy",
"id",
"ii",
"is",
"it",
"iu",
"ja",
"jbo",
"ka",
"kls",
"ko",
"kw",
"la",
"lb",
"liv",
"lt",
"lv",
"mdf",
"mfe",
"mic",
"mk",
"mns-nor",
"ms",
"mt",
"mul",
"my",
"nan",
"nci",
"nl",
"nn",
"no",
"nov",
"nv",
"pjt",
"pl",
"ps",
"pt",
"ro",
"ru",
"scn",
"sco",
"sga",
"sh",
"sl",
"sq",
"sv",
"sw",
"syc",
"szl",
"tg",
"th",
"tl",
"tpw",
"tr",
"tyv",
"ug",
"uk",
"vi",
"vo",
"wlm",
"yi",
"yrl",
"yue",
"zlw-mas"
}
--[=[
This should list the diphthongs of a language (in the form of Lua patterns),
provided they do *NOT* contain semivowel symbols such as /j w ɰ ɥ/ or vowels
with nonsyllabic diacritics such as /i̯ u̯/. For example, list /au/ or /aʊ/,
but do not list /aw/ or /au̯/. The data in this table is used to count the
number of syllables in a word. [[Module:syllables]] automatically knows how
to correctly handle semivowel symbols and nonsyllabic diacritics.
Any language listed here will automatically have categories of the form
"LANG #-syllable words" generated. In addition, any language listed below under
`langs_to_generate_syllable_count_categories` will also have such categories
generated.
NOTE: There are some additional languages that have these categories.
For example:
* Thai words have these categories added by [[Module:th-pron]].]=]
data.diphthongs = {
["cs"] = { -- [[w:Czech phonology#Diphthongs]]
"[aeo]u",
},
["de"] = {
"a[ɪʊ]",
"ɔ[ʏɪ]",
},
["en"] = { -- from [[Appendix:English pronunciation]] mostly, but /ʌɪ/ is from the OED
"[aɑæeɛoɔʌ][ɪi]",
"[ɑɒæo]e",
"[əɐ]ʉ",
"[aɒəoɔæ]ʊ",
"æo",
"[ɛeɪiɔʊʉ]ə", -- /iə/ is a diphthong in NZE, but a disyllabic sequence in GA.
-- /ɪə/ is both a disyllabic sequence and a diphthong in old-fashioned RP.
"[aʌ][ʊɪ]ə", -- May be a disyllabic sequence in some or all dialects?
},
["grc"] = {
"[aeyo]i",
"[ae]u",
"[ɛɔa]ː[iu]",
},
["hrx"] = {
"aɪ̯",
"aʊ̯",
"oɪ̯",
"eʊ̯",
},
["is"] = { -- [[w:Icelandic phonology#Vowels]]
"[aeɔœʏ]i", -- diphthongs as the module generates them
"[ao]u", -- diphthongs as the module generates them
"ø[iɪy]", -- additional forms that may occur; Wikipedia is oddly specific about the second element: ei and ai, but øɪ.
},
["it"] = {
"[aeɛoɔu]i",
"[aeɛioɔ]u",
},
["lb"] = {
"[iu]ə",
"[ɜoæɑ]ɪ",
"[əæɑ]ʊ",
},
["lt"] = {
"ɐɪ", "ɒʊ", "ɛɪ", "ɛʊ", "ʊɪ", "ɔɪ", "ɔʊ", -- Simple diphthongs (unstressed forms)
"iɛ", "uɔ", -- Complex diphthongs
"ɑˑɪ", "ɑˑʊ", "æˑɪ", "æˑʊ", "oˑɪ", -- Falling tone (acute)
"ɐɪˑ", "ɒʊˑ", "ɛɪˑ", "ɛʊˑ", "ʊɪˑ", -- Rising tone (tilde) - lengthened second element
-- Note: Mixed diphthongs (e.g., ɐlˑ, æˑn, ʊl, etc.) are omitted since they are inherently monosyllabic
},
}
--[=[
This should list any languages for which categories of the form
"LANG #-syllable words", e.g. [[:Category:Russian 3-syllable words]], should be
generated. Do not list languages here if they have an entry above under
`data.diphthongs`; such languages are automatically added to this list.]=]
local langs_to_generate_syllable_count_categories = list_to_set{
"ar", -- Arabic has diphthongs, but they are transcribed
-- with semivowel symbols.
"ary", -- Moroccan Arabic has diphthongs, but they are transcribed
-- with semivowel symbols.
"bg", -- Bulgarian has diphthongs with /j/ and marginally with /w/,
-- but these are semivowels.
"ca", -- Catalan has diphthongs, but they are generally transcribed using
-- /w/ and /j/, so do not need to be listed (see [[w:Catalan language#Diphthongs and triphthongs]].
"eo",
"es", -- Spanish has diphthongs, but they are transcribed with i̯ etc.
"eu", -- Basque has dipthongs, but they are transcribed with i̯ and u̯.
"fi", -- Finnish has diphthongs, but they are now automatically transcribed with
-- the nonsyllabic diacritic
"fr", -- French has diphthongs, but they are transcribed
-- with semivowel symbols: [[w:French phonology#Glides and diphthongs]].
"hnn",
"id", -- Indonesian has diphthongs, but they are transcribed with i̯ or /j/ etc.
"ka",
"kne",
"kmr",
"ku",
"la", -- All diphthongs transcribed with e̯ or /j/ etc.
"mk",
"ms", -- Malay has diphthongs, but they are transcribed with i̯ or /j/ etc.
"mt", -- Maltese has diphthongs, but they are transcribed
-- with semivowel symbols.
"pl", -- No diphthongs, properly speaking; sequences of a vowel and /w/ or /j/ though.
"pt", -- Portuguese has diphthongs, but they are transcribed with i̯ or /j/ etc.
"rsk", -- No diphthongs but there are sequences of vowel and /j/ or /w/.
"ru", -- No diphthongs, properly speaking; sequences of a vowel and /j/ though.
"sk", -- Slovak has rising diphthongs, /i̯e, i̯a, i̯u, u̯o/, which are probably always spelled with the nonsyllabic diacritic, so do not need to be listed.
"sl", -- No diphthongs, properly speaking; sequences of a vowel, /j/ and /w/ though
"sq", -- [[w:Albanian language#Vowels]] doesn't mention anything about diphthongs.
"szy", -- All diphthongs are transcribed with /j/ or /w/
"tl", -- Tagalog has diphthongs, but they are transcribed with i̯ or /j/ etc
"tsg",
"ug", -- No diphthongs.
}
-- Also add languages listed under `data.diphthongs`.
for langcode, _ in pairs(data.diphthongs) do
langs_to_generate_syllable_count_categories[langcode] = true
end
data.langs_to_generate_syllable_count_categories = langs_to_generate_syllable_count_categories
-- Languages to use the phonetic not phonemic notation to compute syllable counts.
data.langs_to_use_phonetic_notation = list_to_set{
"bg",
"es",
"id",
"la",
"lt",
"mk",
"ms",
"rsk",
"ru",
}
-- Languages to use the phonetic or phonemic notation to compute syllable counts, whichever is available.
data.langs_to_use_phonetic_or_phonemic_notation = list_to_set{
-- [[Module:is-IPA]] generates [...] but many manual pronuns use /.../.
"is",
}
-- Non-standard or obsolete IPA symbols.
data.nonstandard = {
--[[ The following symbols consist of more than one character,
so we can't put them in the line below. ]]
"ɑ̢", "ɔ̗", "ɔ̖",
"[?ƍσƺƪƞƛłščžǰǧǯẋⱻʚω∅ØȣᴀᴇⱻQKPT]"
}
-- See valid IPA characters at [[Module:IPA/data/symbols]].
data.phonemes = {}
data.phonemes["dz"] = {
"m", "n", "ŋ",
"p", "t", "ʈ", "k",
"pʰ", "tʰ", "ʈʰ", "kʰ",
"t͡s", "t͡ɕ",
"t͡sʰ", "t͡ɕʰ",
"w", "s", "z", "ɬ", "l", "r", "ɕ", "ʑ", "j", "h",
"ɑ", "e", "i", "o", "u",
"ɑː", "eː", "ɛː", "iː", "oː", "øː", "uː", "yː",
"ɑ˥", "e˥", "i˥", "o˥", "u˥",
"ɑː˥", "eː˥", "ɛː˥", "iː˥", "oː˥", "øː˥", "uː˥", "yː˥",
"m˥", "n˥", "ŋ˥", "p˥", "k˥", "k̚˥", "w˥", "l˥", "r˥", "ɕ˥", "j˥", ")˥",
"ɑ˩", "e˩", "i˩", "o˩", "u˩",
"ɑː˩", "eː˩", "ɛː˩", "iː˩", "oː˩", "øː˩", "uː˩", "yː˩",
"m˩", "n˩", "ŋ˩", "p˩", "k˩", "k̚˩", "w˩", "l˩", "r˩", "ɕ˩", "j˩", ")˩",
".", ",", "-",
}
data.phonemes["eo"] = {
"a", "b", "d", "d͡ʒ", "d͡z", "e", "f", "h", "i", "j", "k",
"l", "m", "n", "o", "p", "r", "s", "t", "t͡s", "t͡ʃ",
"u", "u̯", "v", "w", "x", "z", "ɡ", "ʃ", "ʒ",
"ˈ", ".", " ", "-", "u̯", "i̯"
}
data.phonemes["hy"] = {
"ɑ", "b", "ɡ", "d", "e", "z", "ə", "tʰ", "ʒ", "i", "l", "χ", "t͡s",
"k", "h", "d͡z", "ʁ", "t͡ʃ", "m", "j", "n", "ʃ", "ɔ", "t͡ʃʰ", "p", "d͡ʒ",
"r", "s", "v", "t", "ɾ", "t͡sʰ", "v", "pʰ", "kʰ", "o", "f", "ŋɡ", "ŋk",
"ŋχ", "u", "œ", "ʏ", "ˈ", "ˌ", ".", " ", "ː",
}
data.phonemes["nl"] = {
"m", "n", "ŋ",
"p", "b", "t", "d", "k", "ɡ",
"f", "v", "s", "z", "ʃ", "ʒ", "x", "ɣ", "ɦ",
"ʋ", "l", "j", "r",
"ɪ", "ʏ", "ɛ", "ə", "ɔ", "ɑ",
"i", "iː", "y", "yː", "u", "uː", "eː", "øː", "oː", "ɛː", "œː", "ɔː", "aː",
"ɛi̯", "œy̯", "ɔi̯", "ɑu̯", "ɑi̯",
"iu̯", "yu̯", "ui̯", "eːu̯", "oːi̯", "aːi̯",
"ˈ", "ˌ", ".", " ", "-",
}
data.phonemes["mt"] = {
"m", "n",
"p", "t", "k", "ʔ",
"b", "d", "ɡ",
"t͡s", "t͡ʃ",
"d͡z", "d͡ʒ",
"f", "s", "ʃ", "ħ",
"v", "z", "ʒ", "ɣ",
"l", "j", "w",
"r",
"ɪ", "ɛ", "ɔ", "a", "u",
"ɛˤ", "ɔˤ", "aˤ", "əˤ",
"ɛˤː", "ɔˤː", "aˤː", "əˤː", "ɪˤː",
"iː", "ɪː", "ɛː", "ɔː", "aː", "uː",
"ˈ", "ˌ", ".", " ", "‿", "-"
}
return data
9j3yfr5dzmr70htnh6pfy97jzb11aph
Modul:columns
828
10232
375356
229188
2026-09-22T03:13:54Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708759|92708759]])
375356
Scribunto
text/plain
local export = {}
local collation_module = "Module:collation"
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local headword_data_module = "Module:headword/data"
local JSON_module = "Module:JSON"
local languages_module = "Module:languages"
local links_module = "Module:links"
local pages_module = "Module:pages"
local parameter_utilities_module = "Module:parameter utilities"
local parameters_module = "Module:parameters"
local parse_utilities_module = "Module:parse utilities"
local qualifier_module = "Module:qualifier"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local yesno_module = "Module:yesno"
local m_str_utils = require(string_utilities_module)
local concat = table.concat
local html = mw.html.create
local is_substing = mw.isSubsting
local insert = table.insert
local rmatch = m_str_utils.match
local remove = table.remove
local sub = string.sub
local trim = m_str_utils.trim
local u = m_str_utils.char
local dump = mw.dumpObject
local function track(page)
require(debug_track_module)("columns/" .. page)
return true
end
local function deepEquals(...)
deepEquals = require(table_module).deepEquals
return deepEquals(...)
end
local function term_already_linked(term)
return term == "?" or -- signals an unknown term
-- optimization to avoid unnecessarily loading [[Module:parse utilities]]
(term:find("[<{]") and require(parse_utilities_module).term_already_linked(term))
end
local function convert_delimiter_to_separator(item, itemind, args)
if itemind == 1 then
item.separator = nil
elseif item.delimiter == " " then
item.separator = args.space_delim
elseif item.delimiter == "~" then
item.separator = args.tilde_delim
else
item.separator = args.comma_delim
end
end
local function get_horizontal_separator(args_horiz, embedded_comma)
return args_horiz == "bullet" and " · " or embedded_comma and "; " or ", "
end
-- Suppress false positives in categories like [[Category:English links with redundant wikilinks]] so people won't
-- be tempted to "correct" them; terms like embedded ~ like [[Micros~1]] or embedded comma not followed by a space
-- such as [[1,6-Cleves acid]] need to have a link around them to avoid the tilde or comma being interpreted as a
-- delimiter.
local function suppress_redundant_wikilink_cat(term, _alt)
return term:find("~") or term:find(",%S")
end
local function full_link_and_track_self_links(item, face, nochecktr)
if item.term then
local pagename = mw.loadData(headword_data_module).pagename
local term_is_pagename = item.term == pagename
local term_contains_pagename = item.term:find("%[%[" .. m_str_utils.pattern_escape(pagename) .. "[|%]]")
if term_is_pagename or term_contains_pagename then
local current_L2 = require(pages_module).get_current_L2()
if current_L2 then
local current_L2_lang = require(languages_module).getByCanonicalName(current_L2)
if current_L2_lang and current_L2_lang:getCode() == item.lang:getCode() then
if term_is_pagename then
track("term-is-pagename")
else
track("term-contains-pagename")
end
end
end
end
end
item.suppress_redundant_wikilink_cat = suppress_redundant_wikilink_cat
item.never_call_transliteration_module = nochecktr
return require(links_module).full_link(item, face)
end
local function format_subitem(subitem, lang, face, compute_embedded_comma, nochecktr)
local embedded_comma = false
local text
if subitem.term and term_already_linked(subitem.term) then
text = subitem.term
if compute_embedded_comma then
embedded_comma = not not require(utilities_module).get_plaintext(text):find(",")
end
else
text = full_link_and_track_self_links(subitem, face, nochecktr)
if compute_embedded_comma then
-- We don't check decoration text for commas as it's inside parens or displayed elsewhere.
local subitem_plaintext = subitem.alt or subitem.term
if subitem_plaintext then
embedded_comma = not not subitem_plaintext:find(",")
end
end
end
-- We could use the `show_decorations` field to full_link() but not when term_already_linked().
if subitem.q and subitem.q[1] or subitem.qq and subitem.qq[1] or subitem.l and subitem.l[1] or
subitem.ll and subitem.ll[1] or subitem.refs and subitem.refs[1] then
text = require(decorations_module).format_decorations {
lang = subitem.lang or lang,
text = text,
q = subitem.q,
qq = subitem.qq,
l = subitem.l,
ll = subitem.ll,
refs = subitem.refs,
}
end
return text, embedded_comma
end
function export.format_item(item, args, face)
local compute_embedded_comma = args.horiz == "comma"
local embedded_comma = false
local nochecktr = args.noautotr
if type(item) == "table" then
if item.terms then
local parts = {}
local is_first = true
for _, subitem in ipairs(item.terms) do
if subitem == false then
-- omitted subitem; do nothing
else
local separator = subitem.separator or not is_first and (args.subitem_separator or ", ")
if separator then
if compute_embedded_comma then
embedded_comma = embedded_comma or not not separator:find(",")
end
insert(parts, separator)
end
local formatted, this_embedded_comma = format_subitem(subitem, args.lang, face,
compute_embedded_comma, nochecktr)
embedded_comma = embedded_comma or this_embedded_comma
insert(parts, formatted)
is_first = false
end
end
return concat(parts), embedded_comma
else
return format_subitem(item, args.lang, face, compute_embedded_comma, nochecktr)
end
else
if compute_embedded_comma then
embedded_comma = not not require(utilities_module).get_plaintext(item):find(",")
end
if args.lang and not term_already_linked(item) then
return full_link_and_track_self_links({lang = args.lang, term = item, sc = args.sc}, face, nochecktr), embedded_comma
else
return item, embedded_comma
end
end
end
function export.construct_old_style_header(header, horiz)
local old_style_header
local function ib_colon()
return tostring(html("span"):addClass("ib-colon"):addClass("ib-content"):wikitext(":"))
end
if horiz then
old_style_header = require(qualifier_module).format_qualifiers {
qualifiers = header,
open = false,
close = false,
} .. ib_colon() .. " "
else
old_style_header = require(qualifier_module).format_qualifiers {
qualifiers = header
} .. ib_colon()
old_style_header = tostring(html("div"):wikitext(old_style_header))
end
return old_style_header
end
-- Construct the sort base of a single term. As a hack, sort appendices after mainspace items.
local function term_sortbase(val)
if not val then
-- This should not normally happen.
return u(0x10FFFF)
elseif val:find("^%[*Appendix:") then
return u(0x10FFFE) .. val
else
return val
end
end
-- Construct the sort base of a single item, using the display form preferentially, otherwise the term itself.
-- As a hack, sort appendices after mainspace items.
local function item_sortbase(item)
return term_sortbase(item.alt or item.term)
end
local function make_sortbase(item)
if item == false then
return "*" -- doesn't matter, will be omitted in create_list()
elseif type(item) == "table" then
if item.terms then
-- Optimize for the common case of only a single term
if item.terms[2] then
local parts = {}
-- multiple terms
local first = true
for _, subitem in ipairs(item.terms) do
if subitem ~= false then
if not first then
insert(parts, ", ")
end
insert(parts, item_sortbase(subitem))
first = false
end
end
if parts[1] then
return concat(parts)
end
else
local subitem = item.terms[1]
if subitem ~= false then
return item_sortbase(subitem)
end
end
return "*" -- doesn't matter, entire group will be omitted in create_list()
else
return item_sortbase(item)
end
else
return item
end
end
local function make_node_sortbase(node)
return make_sortbase(node.item)
end
-- Sort a sublist of `list` in place, keeping the first `keepfirst` and last `keeplast` items fixed.
-- `lang` is the language of the items and `make_sortbase` creates the appropriate sort base.
local function sort_sublist(list, lang, make_sortbase_fn, keepfirst, keeplast)
if keepfirst == 0 and keeplast == 0 then
require(collation_module).sort(list, lang, make_sortbase_fn)
else
local sublist = {}
for i = keepfirst + 1, #list - keeplast do
sublist[i - keepfirst] = list[i]
end
require(collation_module).sort(sublist, lang, make_sortbase_fn)
for i = keepfirst + 1, #list - keeplast do
list[i] = sublist[i - keepfirst]
end
end
end
--[=[ Unused but could be useful in the future
-- URL-encode only the characters that serve as template delimiters (left and right brace, vertical bar, equal sign
-- and percent sign since it's the escape character).
local function bot_url_encode(txt)
return (txt:gsub("[%%|{}=&]",
{["%"] = "%25", ["|"] = "%7C", ["{"] = "%7B", ["}"] = "%7D", ["="] = "%3D", ["&"] = "%26"}))
end
]=]
-- Reverse the action of bot_url_encode().
local function bot_url_decode(txt)
return (txt:gsub("%%7([BCD])", {B = "{", C = "|", D = "}"}):gsub("%%3D", "="):gsub("%%26", "&"):gsub("%%25", "%%"))
end
--[==[
Bot-callable function to generate a number of sortkeys simultaneously. {{para|1}} contains the langcode, and remaining
numeric parameters contain "bot-URL-encoded" strings whose sort keys will be computed and returned as a JSON array.
Here, "bot-URL-encoded" means that the six characters `{ | } = & %` should be converted to
their URL-encoded representation (respectively `%7B %7C %7D %3D %26 %25`), and will be decoded appropriately
before computing the sortkey.
]==]
function export.make_sortkey(frame)
local iparams = {
[1] = {type = "language"},
[2] = {list = true},
}
local iargs = require(parameters_module).process(frame.args, iparams)
local make_sortkey = require(collation_module).make_lang_sortkey_function(iargs[1], term_sortbase)
local retval = {}
for _, arg in ipairs(iargs[2]) do
arg = bot_url_decode(arg)
insert(retval, make_sortkey(arg))
end
return require(JSON_module).toJSON(retval)
end
local large_text_scripts = {
["Arab"] = true,
["Beng"] = true,
["Deva"] = true,
["Gujr"] = true,
["Guru"] = true,
["Hebr"] = true,
["Khmr"] = true,
["Knda"] = true,
["Laoo"] = true,
["Mlym"] = true,
["Mong"] = true,
["Mymr"] = true,
["Orya"] = true,
["Sinh"] = true,
["Syrc"] = true,
["Taml"] = true,
["Telu"] = true,
["Tfng"] = true,
["Thai"] = true,
["Tibt"] = true,
}
--[==[
Format a list of items using HTML. `args` is an object specifying the items to add and related properties, with the
following fields:
* `content`: A list of the items to format. See below for the format of the items.
* `lang`: The language object of the items to format, if the items in `content` are strings.
* `sc`: The script object of the items to format, if the items in `content` are strings.
* `raw`: If true, return the list raw, without any collapsing or columns.
* `class`: The CSS class of the surrounding {<div>}.
* `column_count`: Number of columns to format the list into.
* `alphabetize`: If true, sort the items in the table.
* `collapse`: If true, make the table partially collapsed by default, with a "Show more" button at the bottom.
* `toggle_category`: Value of `data-toggle-category` property grouping collapsible elements.
* `header`: If specified, Wikicode to prepend to the output.
* `title_new_style`: If true, the header is treated as a title and displayed in a new style. This is ignored if `horiz`
is non-nil.
* `subitem_separator`: Separator used between subitems when multiple subitems occur on a line, if not specified in the
subitem itself (using the `separator` field). Defaults to {", "}.
* `keepfirst`: If > 0, keep this many rows unsorted at the beginning of the top level.
* `keeplast`: If > 0, keep this many rows unsorted at the end of the top level.
* `horiz`: If non-nil, format the items horizontally. If the value is "bullet", put a center dot/bullet (·) between
items. If the value is "comma", put a comma between items (but if there is an embedded comma in any item,
put a semicolon between all items).
Each item in `content` is in one of the following formats:
* A string. This is for compatibility and should not be used by new callers.
* An object describing an item to format, in the format expected by full_link() in [[Module:links]], including
decorations (left or right qualifiers, left or right labels, or references).
* An object describing a list of subitems to format, displayed side-by-side, separated by a comma or other separator.
This format is identified by the presence of a key `terms` specifying the list of subitems. Each subitem is in
the same format as for a single top-level item, except that it should also have a `separator` field specifying the
separator to display before each item (which will typically be a blank string before the first item).
]==]
function export.create_list(args)
if type(args) ~= "table" then
error("expected table, got " .. type(args))
end
local column_count = args.column_count or 1
local toggle_category = args.toggle_category or "kata terbitan"
local keepfirst = args.keepfirst or 0
local keeplast = args.keeplast or 0
if keepfirst > 0 then
track("keepfirst")
end
if keeplast > 0 then
track("keeplast")
end
-- maybe construct old-style header
local old_style_header = nil
if args.header and (args.horiz or not args.title_new_style) then
old_style_header = export.construct_old_style_header(args.header, args.horiz)
end
if args.horiz then
old_style_header = "* " .. (old_style_header or "")
end
local list
local any_extra_indented_item = false
for _, item in ipairs(args.content) do
if item == false then
-- do nothing
elseif type(item) == "table" and item.extra_indent and item.extra_indent > 0 then
any_extra_indented_item = true
break
end
end
-- If any extra indented item, convert the items to a nested structure, which is necessary both for sorting and
-- for converting to HTML.
if any_extra_indented_item then
local function make_node(item)
return {
item = item
}
end
local root_node = make_node(nil)
local node_stack = {root_node}
local last_indent = 0
local function append_subnode(node, subnode)
if not node.subnodes then
node.subnodes = {}
end
insert(node.subnodes, subnode)
end
for i, item in ipairs(args.content) do
if item == false then
-- do nothing
else
local this_indent
if type(item) ~= "table" then
this_indent = 1
else
this_indent = (item.extra_indent or 0) + 1
end
local node = make_node(item)
if this_indent == last_indent then
append_subnode(node_stack[#node_stack], node)
elseif this_indent > last_indent + 1 then
error(("Element #%s (%s) has indent %s, which is more than one greater than the previous item with indent %s"):format(
i, make_sortbase(item), this_indent, last_indent))
elseif this_indent > last_indent then
-- Start a new sublist attached to the last item of the sublist one level up; but we need special
-- handling for the root node (last_indent == 0).
if last_indent > 0 then
local subnodes = node_stack[#node_stack].subnodes
if not subnodes then
error(("Internal error: Not first item and no subnodes at preceding level %s: %s"):format(
#node_stack, dump(node_stack)))
end
insert(node_stack, subnodes[#subnodes])
end
append_subnode(node_stack[#node_stack], node)
last_indent = this_indent
else
while last_indent > this_indent do
local finished_node = table.remove(node_stack)
if args.alphabetize then
require(collation_module).sort(finished_node.subnodes, args.lang, make_node_sortbase)
end
last_indent = last_indent - 1
end
append_subnode(node_stack[#node_stack], node)
end
end
end
if args.alphabetize then
while node_stack[1] do
local finished_node = table.remove(node_stack)
if node_stack[1] then
-- We're sorting something other than the root node.
require(collation_module).sort(finished_node.subnodes, args.lang, make_node_sortbase)
else
-- We're sorting the root node; honor `keepfirst` and `keeplast`.
sort_sublist(finished_node.subnodes, args.lang, make_node_sortbase, keepfirst, keeplast)
end
end
end
local function format_node(node, depth)
local sublist
local embedded_comma = false
if node.subnodes then
if args.horiz then
sublist = {}
else
sublist = html("ul")
end
local prevnode = nil
for _, subnode in ipairs(node.subnodes) do
local thisnode, this_embedded_comma = format_node(subnode, depth + 1)
embedded_comma = embedded_comma or this_embedded_comma
if not prevnode or not args.alphabetize or not deepEquals(prevnode, thisnode) then
if args.horiz then
table.insert(sublist, thisnode)
else
sublist = sublist:node(thisnode)
end
prevnode = thisnode
end
end
if args.horiz then
sublist = table.concat(sublist, get_horizontal_separator(args.horiz, embedded_comma))
end
end
if not node.item then
-- At the root.
return sublist, embedded_comma
end
local formatted, listitem
-- Ignore embedded commas in subitems inside of parens or square brackets.
formatted, embedded_comma = export.format_item(node.item, args)
if args.horiz then
listitem = formatted
if sublist then
-- Use parens for the first, third, fifth, etc. sublists and square brackets for the remainder.
if depth % 2 == 1 then
listitem = ("%s (%s)"):format(listitem, sublist)
else
listitem = ("%s [%s]"):format(listitem, sublist)
end
end
else
listitem = html("li"):wikitext(formatted)
if sublist then
listitem = listitem:node(sublist)
end
end
return listitem, embedded_comma
end
list = format_node(root_node, 0)
else
if args.alphabetize then
sort_sublist(args.content, args.lang, make_sortbase, keepfirst, keeplast)
end
if args.horiz then
list = {}
else
list = html("ul")
end
local previtem = nil
local embedded_comma = false
for _, item in ipairs(args.content) do
if item == false then
-- omitted item; do nothing
else
local thisitem, this_embedded_comma = export.format_item(item, args)
embedded_comma = embedded_comma or this_embedded_comma
if not previtem or not args.alphabetize or previtem ~= thisitem then
if args.horiz then
table.insert(list, thisitem)
else
list = list:node(html("li"):wikitext(thisitem))
end
previtem = thisitem
end
end
end
if args.horiz then
list = table.concat(list, get_horizontal_separator(args.horiz, embedded_comma))
end
end
local output
if args.horiz then
output = list
else
output = html("div"):addClass("term-list"):node(list)
if args.class then
output:addClass(args.class)
end
if not args.raw then
output:addClass("ul-column-count")
:attr("data-column-count", column_count)
if args.collapse then
output = html("div")
:node(output)
:addClass("list-switcher")
:attr("data-toggle-category", toggle_category)
-- identify commonly used scripts that use large text and
-- provide a special CSS class to make the template bigger
local sc = args.sc
if sc == nil then
local scripts = args.lang:getScripts()
if #scripts > 0 then
sc = scripts[1]
end
end
if sc ~= nil then
local scriptcode = sc:getParentCode()
if scriptcode == "top" then
scriptcode = sc:getCode()
end
if large_text_scripts[scriptcode] then
output:addClass("list-switcher-large-text")
end
end
end
end
if args.collapse or args.title_new_style then
-- wrap in wrapper to prevent interference from floating elements
local list_switcher_wrapper = html("div")
:addClass("list-switcher-wrapper")
if args.title_new_style then
list_switcher_wrapper
:node(
html("div")
:addClass("list-switcher-header")
:wikitext(args.header)
)
end
list_switcher_wrapper:node(output)
output = list_switcher_wrapper
end
output = tostring(output)
end
return (old_style_header or "") .. output
end
-- This function is for compatibility with earlier version of [[Module:columns]]
-- (now found in [[Module:columns/old]]).
function export.create_table(...)
-- Earlier arguments to create_table:
-- n_columns, content, alphabetize, bg, collapse, class, title, column_width, line_start, lang
local args = {}
args.column_count, args.content, args.alphabetize,
args.collapse, args.class, args.header, args.column_width,
args.line_start, args.lang = ...
return export.create_list(args)
end
function export.display_from(frame_args, parent_args, frame)
local boolean = {type = "boolean"}
local iparams = {
["class"] = true,
-- Default for auto-collapse. Overridable by template |collapse= param.
["collapse"] = boolean,
-- If specified, this specifies the number of columns, and no columns parameter is available on the template.
-- Otherwise, the columns parameter is named |n=.
["columns"] = {type = "number"},
-- If specified, this specifies the default language code, which can be overridden using |lang= in the template.
-- Otherwise, the language-code parameter is required and normally found in |1=, but for compatibility can be
-- specified as |lang= (which leads to deprecation handling).
["lang"] = {type = "language"},
-- Default for auto-sort. Overridable by template |sort= param.
["sort"] = boolean,
["toggle_category"] = true,
-- Minimum number of rows required to format into a multicolumn list. If below this, the list is displayed "raw"
-- (no columns, no collapsbility).
["minrows"] = {type = "number", default = 5},
-- Disables automatic transliteration; entries without a manual transliteration will have none at all.
-- Used on large pages, especially Chinese ones, because zh-translit works by fetching and parsing
-- the target of the page, which is a performance killer on large pages with potentially thousands
-- of link targets.
-- Note: noautotr also disables redundant transliteration checks.
["noautotr"] = boolean,
}
local iargs = require(parameters_module).process(frame_args, iparams)
local langcode_in_lang = iargs.lang or parent_args.lang
local lang_param = langcode_in_lang and "lang" or 1
local deprecated = not iargs.lang and langcode_in_lang
local ret = export.handle_display_from_or_topic_list(iargs, parent_args, nil)
return deprecated and frame:expandTemplate{title = "check deprecated lang param usage",
-- FIXME: Accessing undefined global var
args = {ret, lang = args[lang_param]}} or ret
end
--[==[
Implement `display_from()` [the internal entry point for {{tl|col}} and variants, which enter originally through
`display()`] as well as regular (column-oriented) topic lists, invoked through [[Module:topic list]].
`iargs` are the invocation args of {{tl|col}}, and `raw_item_args` are the arguments specifying the values of
each row as well as other properties, corresponding to the user-specified template arguments of {{tl|col}}. Note that
`show()` in [[Module:topic list]] is normally invoked directly by a topic list template, whose invocation
arguments are passed in using `raw_item_args` and are similar to the template arguments of {{tl|col}}. `iargs` for
topic-list invocations is hard-coded, and template arguments to a topic-list template are processed in
[[Module:topic list]] itself. Note that the handling of topic lists is currently implemented almost entirely
through callbacks in `topic_list_data` (which is nil if we're processing {{tl|col}} rather than a topic list) in an
attempt to reduce the coupling and keep the topic-list-specific code in [[Module:topic list]], but IMO the coupling
is still too tight. Probably the control structure should be reversed and the following function split up into
subfunctions, which are invoked as needed by {{tl|col}} and/or [[Module:topic list]].
]==]
function export.handle_display_from_or_topic_list(iargs, raw_item_args, topic_list_data)
local boolean = {type = "boolean"}
local langcode_in_lang = iargs.lang or raw_item_args.lang
local lang_param = langcode_in_lang and "lang" or 1
local first_content_param = langcode_in_lang and 1 or 2
local params = {
[lang_param] = {required = not iargs.lang, type = "language",
template_default = not iargs.lang and "und" or nil},
["n"] = not iargs.columns and {type = "number"} or nil,
[first_content_param] = {list = true, allow_holes = true},
["title"] = {},
["collapse"] = boolean,
["sort"] = boolean,
["sc"] = {type = "script"},
-- used when calling from [[Module:saurus]] so the page displaying the synonyms/antonyms doesn't occur in the
-- list
["omit"] = {list = true},
["keepfirst"] = {type = "number", default = 0},
["keeplast"] = {type = "number", default = 0},
["horiz"] = {},
["notr"] = boolean,
["noautotr"] = boolean,
["allow_space_delim"] = boolean,
["tilde_delim"] = {},
["space_delim"] = {},
["comma_delim"] = {},
}
if topic_list_data then
topic_list_data.add_topic_list_params(params)
end
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{default = true, require_index = true},
{group = "link"}, -- sc has separate_no_index = true; that's the only one
-- It makes no sense to have overall l=, ll=, q= or qq= params for columnar display.
{group = {"ref", "l", "q"}, require_index = true},
}
m_param_utils.augment_params_with_modifiers(params, param_mods)
local processed_args = require(parameters_module).process(raw_item_args, params)
local horiz = processed_args.horiz
if horiz and horiz ~= "comma" and horiz ~= "bullet" then
horiz = require(yesno_module)(horiz)
if horiz == nil then
error(("Unrecognized value |horiz=%s; should be 'comma', 'bullet' or a recognized Boolean value such " ..
"as 'yes' or '1' (same as 'bullet') or 'no' or '0'"):format(processed_args.horiz))
end
if horiz == true then
horiz = "bullet"
end
processed_args.horiz = horiz
end
-- If default argument values specified, set them after parsing the caller-specified arguments in `raw_item_args`.
if topic_list_data then
topic_list_data.set_default_arguments(processed_args)
end
-- Now set defaults for the various delimiters, depending in some cases on whether horiz was set.
-- We can't set these defaults (even regardless of their dependency on horiz=) in `local params` above
-- because we want any defaults specified in `default_props` to override these.
if not processed_args.tilde_delim then
local tilde_with_abbr = '<abbr title="near equivalent">~</abbr>'
processed_args.tilde_delim = processed_args.horiz and tilde_with_abbr or " " .. tilde_with_abbr .. " "
end
if not processed_args.space_delim then
processed_args.space_delim = " "
end
if not processed_args.comma_delim then
processed_args.comma_delim = processed_args.horiz and "/" or ", "
end
-- Check for extra term indent. Do this before calling parse_list_with_inline_modifiers_and_separate_params()
-- because sometimes space is a delimiter and the space in the indent will confuse things and get interpreted as a
-- delimiter.
local extra_indent_by_termno = {}
local termargs = processed_args[first_content_param]
for i = 1, termargs.maxindex do
local term = termargs[i]
if term then
local extra_indent, actual_term = rmatch(term, "^(%*+)%s+(.-)$")
if extra_indent then
termargs[i] = actual_term
extra_indent_by_termno[i] = #extra_indent
end
end
end
local groups, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
param_mods = param_mods,
processed_args = processed_args,
termarg = first_content_param,
parse_lang_prefix = true,
allow_multiple_lang_prefixes = true,
disallow_custom_separators = true,
track_module = "columns",
lang = iargs.lang or lang_param,
sc = "sc.default",
splitchar = processed_args.allow_space_delim and "[,~ ]" or "[,~]",
no_show_decorations = true, -- since we handle them ourselves in format_subitem()
}
local lang = iargs.lang or args[lang_param]
local langcode = lang:getCode()
local sc = args.sc.default
local sort = iargs.sort
if args.sort ~= nil then
if not args.sort then
track("nosort")
end
sort = args.sort
else
-- HACK! For Japanese-script languages (Japanese, Okinawan, Miyako, etc.), sorting doesn't yet work properly, so
-- disable it.
for _, langsc in ipairs(lang:getScriptCodes()) do
if langsc == "Jpan" then
sort = false
break
end
end
end
local collapse = iargs.collapse
if args.collapse ~= nil then
if not args.collapse then
track("nocollapse")
end
collapse = args.collapse
end
local title = args.title
local formatted_cats
if topic_list_data then
title, formatted_cats = topic_list_data.get_title_and_formatted_cats(args, lang, sc, topic_list_data)
end
local number_of_groups = 0
for i, group in ipairs(groups) do
local number_of_items = 0
group.extra_indent = extra_indent_by_termno[group.orig_index]
for j, item in ipairs(group.terms) do
convert_delimiter_to_separator(item, j, args)
if args.notr then
item.tr = "-"
elseif args.noautotr then
item.tr = item.tr or "-"
end
-- If a separate language code was given for the term, display the language name as a right qualifier.
-- (Briefly we made them labels but this leads to non-obvious behavior e.g. "French" becoming "France" under
-- some circumstances.) Otherwise it may not be obvious that the term is in a separate language (e.g. if the
-- main language is 'zh' and the term language is a Chinese lect such as Min Nan). But don't do this for
-- Translingual terms, which are often added to the list of English and other-language terms.
if item.termlangs then
local qqs = {}
for _, termlang in ipairs(item.termlangs) do
local termlangcode = termlang:getCode()
if termlangcode ~= langcode and termlangcode ~= "mul" then
insert(qqs, termlang:getCanonicalName())
end
end
if item.qq then
for _, qq in ipairs(item.qq) do
insert(qqs, qq)
end
end
item.qq = qqs
end
local omitted = false
for _, omitted_item in ipairs(args.omit) do
if omitted_item == item.term then
omitted = true
break
end
end
if omitted then
-- signal create_list() to omit this item
group.terms[j] = false
else
number_of_items = number_of_items + 1
end
end
if number_of_items == 0 then
-- omit the whole group
groups[i] = false
else
number_of_groups = number_of_groups + 1
end
end
local column_count = iargs.columns or args.n
-- FIXME: This needs a total rewrite.
if column_count == nil then
column_count = number_of_groups <= 3 and 1 or
number_of_groups <= 9 and 2 or
number_of_groups <= 27 and 3 or
number_of_groups <= 81 and 4 or
5
end
local raw = number_of_groups < iargs.minrows
local horiz_edit_button
if topic_list_data and args.horiz then
-- append edit button to title
horiz_edit_button = topic_list_data.make_horiz_edit_button(topic_list_data.topic_list_template)
end
return export.create_list {
column_count = column_count,
raw = raw,
content = groups,
alphabetize = sort,
header = title,
title_new_style = (title ~= nil and title ~= ''),
collapse = collapse,
toggle_category = iargs.toggle_category,
-- columns-bg (in [[MediaWiki:Gadget-Site.css]]) provides the background color
class = (iargs.class and iargs.class .. " columns-bg" or "columns-bg"),
lang = lang,
sc = sc,
subitem_separator = ", ",
keepfirst = args.keepfirst,
keeplast = args.keeplast,
horiz = args.horiz,
noautotr = args.noautotr,
} .. (horiz_edit_button or "") .. (formatted_cats or "")
end
function export.display(frame)
if not is_substing() then
return export.display_from(frame.args, frame:getParent().args, frame, false)
end
-- If substed, unsubst template with newlines between each term, redundant wikilinks removed, and remove duplicates + sort terms if sort is enabled.
local m_table = require("Module:table")
local m_template_parser = require("Module:template parser")
local parent = frame:getParent()
local elems = m_table.shallowCopy(parent.args)
local code = remove(elems, 1)
code = code and trim(code)
local lang = require("Module:languages").getByCode(code, 1)
local i = 1
while true do
local elem = elems[i]
while elem do
elem = trim(elem, "%s")
if elem ~= "" then
break
end
remove(elems, i)
elem = elems[i]
end
if not elem then
break
elseif not ( -- Strip redundant wikilinks.
not elem:match("^()%[%[") or
elem:find("[[", 3, true) or
elem:find("]]", 3, true) ~= #elem - 1 or
elem:find("|", 3, true)
) then
elem = sub(elem, 3, -3)
elem = trim(elem, "%s")
end
elems[i] = elem .. "\n"
i = i + 1
end
-- If sort is enabled, remove duplicates then sort elements.
if require("Module:yesno")(frame.args.sort) then
elems = m_table.removeDuplicates(elems)
require("Module:collation").sort(elems, lang)
end
-- Readd the langcode.
insert(elems, 1, code .. "\n")
-- TODO: Place non-numbered parameters after 1 and before 2.
local template = m_template_parser.getTemplateInvocationName(mw.title.new(parent:getTitle()))
return "{{" .. concat(m_template_parser.buildTemplate(template, elems), "|") .. "}}"
end
return export
ayrhhchd4gj67psdazy9iw6oj4i10ai
Modul:ca-headword
828
10361
375387
223898
2026-09-22T07:28:03Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714447|92714447]])
375387
Scribunto
text/plain
local export = {}
local pos_functions = {}
local force_cat = false -- for testing; if true, categories appear in non-mainspace pages
local require_when_needed = require("Module:utilities/require when needed")
local m_table = require("Module:table")
local com = require("Module:ca-common")
local ca_IPA_module = "Module:ca-IPA"
local ca_verb_module = "Module:ca-verb"
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local headword_utilities_module = "Module:headword utilities"
local inflection_utilities_module = "Module:inflection utilities"
local parse_utilities_module = "Module:parse utilities"
local romut_module = "Module:romance utilities"
local m_en_utilities = require_when_needed(en_utilities_module)
local m_headword_utilities = require_when_needed(headword_utilities_module)
local m_string_utilities = require_when_needed("Module:string utilities")
local glossary_link = require_when_needed(headword_utilities_module, "glossary_link")
local lang = require("Module:languages").getByCode("ca")
local langname = lang:getCanonicalName()
local list_to_text = mw.text.listToText
local insert = table.insert
local concat = table.concat
local rfind = m_string_utilities.find
local rmatch = m_string_utilities.match
local rsplit = m_string_utilities.split
local usub = m_string_utilities.sub
local rsub = com.rsub
local function track(page)
require("Module:debug/track")("ca-headword/" .. page)
return true
end
local list_param = {list = true, disallow_holes = true}
local boolean_param = {type = "boolean"}
-----------------------------------------------------------------------------------------
-- Main entry point --
-----------------------------------------------------------------------------------------
-- The main entry point.
-- This is the only function that can be invoked from a template.
function export.show(frame)
local poscat = frame.args[1] or error("Part of speech has not been specified. Please pass parameter 1 to the module invocation.")
local params = {
["head"] = list_param,
["id"] = true,
["splithyph"] = boolean_param,
["nolinkhead"] = boolean_param,
["json"] = boolean_param,
["pagename"] = true, -- for testing
}
if pos_functions[poscat] then
for key, val in pairs(pos_functions[poscat].params) do
params[key] = val
end
end
local args = require("Module:parameters").process(frame:getParent().args, params)
local pagename = args.pagename or mw.loadData("Module:headword/data").pagename
local user_specified_heads = args.head
local heads = user_specified_heads
if args.nolinkhead then
if #heads == 0 then
heads = {pagename}
end
else
local romut = require(romut_module)
local auto_linked_head = romut.add_links_to_multiword_term(pagename, args.splithyph)
if #heads == 0 then
heads = {auto_linked_head}
else
for i, head in ipairs(heads) do
if head:find("^~") then
head = romut.apply_link_modifiers(auto_linked_head, usub(head, 2))
heads[i] = head
end
if head == auto_linked_head then
track("redundant-head")
end
end
end
end
local data = {
lang = lang,
pos_category = pos_functions[poscat] and pos_functions[poscat].pos_category or poscat,
categories = {},
heads = heads,
user_specified_heads = user_specified_heads,
no_redundant_head_cat = #user_specified_heads == 0,
genders = {},
inflections = {},
pagename = pagename,
id = args.id,
force_cat_output = force_cat,
checkredlinks = pos_functions[poscat] and pos_functions[poscat].redlink_pos or true,
}
if pagename:find("^%-") and poscat ~= "bentuk akhiran" then
data.is_suffix = true
data.pos_category = "suffixes"
data.checkredlinks = true
local singular_poscat = require(en_utilities_module).singularize(poscat)
insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes")
insert(data.inflections, {label = singular_poscat .. "-forming suffix"})
end
if pos_functions[poscat] then
pos_functions[poscat].func(args, data)
end
if args.json then
return require("Module:JSON").toJSON(data)
end
local post_note = data.post_note and "; " .. data.post_note or ""
return require("Module:headword").full_headword(data) .. post_note
end
-----------------------------------------------------------------------------------------
-- Utility functions --
-----------------------------------------------------------------------------------------
local function replace_hash_with_lemma(term, lemma)
-- If there is a % sign in the lemma, we have to replace it with %% so it doesn't get interpreted as a capture replace
-- expression.
lemma = lemma:gsub("%%", "%%%%")
-- Assign to a variable to discard second return value.
term = term:gsub("#", lemma)
return term
end
-- Parse and insert an inflection not requiring additional processing into `data.inflections`. The raw arguments come
-- from `args[field]`, which is parsed for inline modifiers. `label` is the label that the inflections are given;
-- `accel` is the accelerator form, or nil.
local function parse_and_insert_inflection(data, args, field, label, accel)
m_headword_utilities.parse_and_insert_inflection {
headdata = data,
forms = args[field],
paramname = field,
splitchar = ",",
label = label,
accel = accel and {form = accel} or nil,
}
end
-- Insert default plurals generated when a given plural had the value of + and default plurals were fetched as a result.
-- `plobj` is the parsed object whose `term` field is "+". `defpls` is the list of default plurals. `dest` is the list
-- into which the plurals are inserted (which inherit their decorations from `plobj`).
local function insert_defpls(defpls, plobj, dest)
if not defpls then
-- Happens e.g. with [[S.A.]] where the default plural algorithm returns nothing.
return
end
if #defpls == 1 then
plobj.term = defpls[1]
insert(dest, plobj)
else
for _, defpl in ipairs(defpls) do
local newplobj = m_table.shallowCopy(plobj)
newplobj.term = defpl
insert(dest, newplobj)
end
end
end
-----------------------------------------------------------------------------------------
-- Adjectives --
-----------------------------------------------------------------------------------------
local function do_adjective(args, data, is_superlative)
local feminines = {}
local masculine_plurals = {}
local feminine_plurals = {}
-- Use "participle" not "past participle" for categories such as 'invariable paticiples'
local category_plpos = data.checkredlinks
if category_plpos == true then
category_plpos = data.pos_category
end
local category_pos = m_en_utilities.singularize(category_plpos)
if args.sp then
local romut = require(romut_module)
if not romut.allowed_special_indicators[args.sp] then
local indicators = {}
for indic, _ in pairs(romut.allowed_special_indicators) do
insert(indicators, "'" .. indic .. "'")
end
table.sort(indicators)
error("Special inflection indicator beginning can only be " ..
list_to_text(indicators) .. ": " .. args.sp)
end
end
local lemma = data.pagename
local function fetch_inflections(field)
local retval = m_headword_utilities.parse_term_list_with_modifiers {
paramname = field,
forms = args[field],
splitchar = ",",
}
if not retval[1] then
return {{term = "+"}}
end
return retval
end
local function insert_inflection(terms, label, accel)
m_headword_utilities.insert_inflection {
headdata = data,
terms = terms,
label = label,
accel = accel and {form = accel} or nil,
}
end
if args.f[1] == "ind" or args.f[1] == "inv" then
-- invariable adjective
insert(data.inflections, {label = glossary_link("invariable")})
insert(data.categories, langname .. " indeclinable " .. category_plpos)
if args.sp or args.f[2] or args.pl[1] or args.mpl[1] or args.fpl[1] then
error("Can't specify inflections with an invariable " .. category_pos)
end
elseif args.fonly then
-- feminine-only
if args.f[1] then
error("Can't specify explicit feminines with feminine-only " .. category_pos)
end
if args.pl[1] then
error("Can't specify explicit plurals with feminine-only " .. category_pos .. ", use fpl=")
end
if args.mpl[1] then
error("Can't specify explicit masculine plurals with feminine-only " .. category_pos)
end
local argsfpl = fetch_inflections("fpl")
for _, fpl in ipairs(argsfpl) do
if fpl.term == "+" then
-- Generate default feminine plural.
local defpls = com.make_plural(lemma, "f", args.sp)
if not defpls then
error("Unable to generate default plural of '" .. lemma .. "'")
end
insert_defpls(defpls, fpl, feminine_plurals)
else
fpl.term = replace_hash_with_lemma(fpl.term, lemma)
insert(feminine_plurals, fpl)
end
end
insert(data.inflections, {label = "feminine-only"})
insert_inflection(feminine_plurals, "feminine plural", "f|p")
else
-- Gather feminines.
for _, f in ipairs(fetch_inflections("f")) do
if f.term == "mf" then
f.term = lemma
elseif f.term == "+" then
-- Generate default feminine.
f.term = com.make_feminine(lemma, args.sp)
else
f.term = replace_hash_with_lemma(f.term, lemma)
end
insert(feminines, f)
end
local fem_like_lemma = #feminines == 1 and feminines[1].term == lemma and
not m_headword_utilities.termobj_has_decorations(feminines[1])
if fem_like_lemma then
insert(data.categories, langname .. " epicene " .. category_plpos)
end
local mpl_field = "mpl"
local fpl_field = "fpl"
if args.pl[1] then
if args.mpl[1] or args.fpl[1] then
error("Can't specify both pl= and mpl=/fpl=")
end
mpl_field = "pl"
fpl_field = "pl"
end
local argsmpl = fetch_inflections(mpl_field)
local argsfpl = fetch_inflections(fpl_field)
for _, mpl in ipairs(argsmpl) do
if mpl.term == "+" then
-- Generate default masculine plural.
local defpls
-- First, some special hacks based on the feminine singular.
if not fem_like_lemma and not args.sp and not lemma:find(" ") then
for _, f in ipairs(feminines) do
if f.term:find("ssa$") then
-- If the feminine ends in -ssa, assume that the -ss- is also in the
-- masculine plural form
defpls = {rsub(f.term, "a$", "os")}
break
elseif f.term == lemma .. "na" then
defpls = {lemma .. "ns"}
break
elseif lemma:find("ig$") and f.term:find("ja$") then
-- Adjectives in -ig have two masculine plural forms, one derived from
-- the m.sg. and the other derived from the f.sg.
defpls = {lemma .. "s", rsub(f.term, "ja$", "jos")}
break
end
end
end
defpls = defpls or com.make_plural(lemma, "m", args.sp)
if not defpls then
error("Unable to generate default plural of '" .. lemma .. "'")
end
insert_defpls(defpls, mpl, masculine_plurals)
else
mpl.term = replace_hash_with_lemma(mpl.term, lemma)
insert(masculine_plurals, mpl)
end
end
for _, fpl in ipairs(argsfpl) do
if fpl.term == "+" then
-- First, some special hacks based on the feminine singular.
if fem_like_lemma and not args.sp and not lemma:find(" ") and lemma:find("[çx]$") then
-- Adjectives ending in -ç or -x behave as mf-type in the singular, but
-- regular type in the plural.
local defpls = com.make_plural(lemma .. "a", "f")
if not defpls then
error("Unable to generate default plural of '" .. lemma .. "a'")
end
insert_defpls(defpls, fpl, feminine_plurals)
else
for _, f in ipairs(feminines) do
-- Generate default feminine plural; f is a table.
local defpls = com.make_plural(f.term, "f", args.sp)
if not defpls then
error("Unable to generate default plural of '" .. f.term .. "'")
end
for _, defpl in ipairs(defpls) do
local fplobj = m_table.shallowCopy(fpl)
fplobj.term = defpl
m_headword_utilities.combine_termobj_decorations(fplobj, f)
insert(feminine_plurals, fplobj)
end
end
end
else
fpl.term = replace_hash_with_lemma(fpl.term, lemma)
insert(feminine_plurals, fpl)
end
end
local fem_pl_like_masc_pl = masculine_plurals[1] and feminine_plurals[1] and
m_table.deepEquals(masculine_plurals, feminine_plurals)
local masc_pl_like_lemma = #masculine_plurals == 1 and masculine_plurals[1].term == lemma and
not m_headword_utilities.termobj_has_decorations(masculine_plurals[1])
if fem_like_lemma and fem_pl_like_masc_pl and masc_pl_like_lemma then
-- actually invariable
insert(data.inflections, {label = glossary_link("invariable")})
insert(data.categories, langname .. " indeclinable " .. category_plpos)
else
-- Make sure there are feminines given and not same as lemma.
if not fem_like_lemma then
insert_inflection(feminines, "feminine", "f|s")
elseif args.gneut then
data.genders = {"gneut"}
else
data.genders = {"mf"}
end
if fem_pl_like_masc_pl then
if args.gneut then
insert_inflection(masculine_plurals, "plural", "p")
else
insert_inflection(masculine_plurals, "masculine and feminine plural", "p")
end
else
insert_inflection(masculine_plurals, "masculine plural", "m|p")
insert_inflection(feminine_plurals, "feminine plural", "f|p")
end
end
end
parse_and_insert_inflection(data, args, "comp", "comparative")
parse_and_insert_inflection(data, args, "sup", "superlative")
parse_and_insert_inflection(data, args, "dim", "diminutive")
parse_and_insert_inflection(data, args, "aug", "augmentative")
if args.irreg and is_superlative then
insert(data.categories, langname .. " irregular superlative " .. category_plpos)
end
end
local function get_adjective_params(adjtype)
local params = {
["sp"] = true, -- special indicator: "first", "first-last", etc.
["f"] = list_param, --feminine form(s)
[1] = {alias_of = "f", list = false},
["pl"] = list_param, --plural override(s)
["mpl"] = list_param, --masculine plural override(s)
["fpl"] = list_param, --feminine plural override(s)
}
if adjtype == "base" then
params["comp"] = list_param --comparative(s)
params["sup"] = list_param --superlative(s)
params["dim"] = list_param --diminutive(s)
params["aug"] = list_param --augmentative(s)
params["fonly"] = boolean_param -- feminine only
params["hascomp"] = {} -- has comparative
end
if adjtype == "sup" then
params["irreg"] = boolean_param
end
return params
end
-- Display additional inflection information for an adjective
pos_functions["adjectives"] = {
params = get_adjective_params("base"),
func = do_adjective,
}
pos_functions["past participles"] = {
params = get_adjective_params("part"),
func = do_adjective,
redlink_pos = "participles",
}
pos_functions["determiners"] = {
params = get_adjective_params("det"),
func = do_adjective,
}
pos_functions["pronouns"] = {
params = get_adjective_params("pron"),
func = do_adjective,
}
-----------------------------------------------------------------------------------------
-- Nouns --
-----------------------------------------------------------------------------------------
local allowed_genders = m_table.listToSet(
{"m", "f", "mf", "mfbysense", "mfequiv", "gneut", "n", "m-p", "f-p", "mf-p", "mfbysense-p", "mfequiv-p", "gneut-p", "n-p", "?", "?-p"}
)
local function validate_genders(genders)
for _, g in ipairs(genders) do
if type(g) == "table" then
g = g.spec
end
if not allowed_genders[g] then
error("Unrecognized gender: " .. g)
end
end
end
local function do_noun(args, data, is_proper)
local is_plurale_tantum = false
local has_singular = false
local category_plpos = data.checkredlinks
if category_plpos == true then
category_plpos = data.pos_category
end
local category_pos = m_en_utilities.singularize(category_plpos)
validate_genders(args[1])
data.genders = args[1]
local saw_m = false
local saw_f = false
local saw_gneut = false
local gender_for_irreg_ending, gender_for_default_plural
-- Check for specific genders and pluralia tantum.
for _, g in ipairs(args[1]) do
if type(g) == "table" then
g = g.spec
end
if g:find("-p$") then
is_plurale_tantum = true
else
has_singular = true
if g == "m" or g == "mf" or g == "mfbysense" then
saw_m = true
end
if g == "f" or g == "mf" or g == "mfbysense" then
saw_f = true
end
if g == "gneut" then
saw_gneut = true
end
end
end
if saw_m and saw_f then
gender_for_irreg_ending = "mf"
elseif saw_f then
gender_for_irreg_ending = "f"
else
gender_for_irreg_ending = "m"
end
gender_for_default_plural =
saw_gneut and "gneut" or gender_for_irreg_ending == "mf" and "m" or gender_for_irreg_ending
local lemma = data.pagename
-- Plural
local plurals = {}
local function insert_noun_inflection(terms, label, accel)
m_headword_utilities.insert_inflection {
headdata = data,
terms = terms,
label = label,
accel = accel and {form = accel} or nil,
}
end
if is_plurale_tantum and not has_singular then
if args[2][1] then
error("Can't specify plurals of plurale tantum " .. category_pos)
end
insert(data.inflections, {label = glossary_link("plural only")})
else
plurals = m_headword_utilities.parse_term_list_with_modifiers {
paramname = {2, "pl"},
forms = args[2],
splitchar = ",",
}
-- Check for special plural signals
local mode = nil
local pl1 = plurals[1]
if pl1 and #pl1.term == 1 then
mode = pl1.term
if mode == "?" or mode == "!" or mode == "-" or mode == "~" then
pl1.term = nil
if next(pl1) then
error(("Can't specify inline modifiers with plural code '%s'"):format(mode))
end
table.remove(plurals, 1) -- Remove the mode parameter
elseif mode ~= "+" and mode ~= "#" then
error(("Unexpected plural code '%s'"):format(mode))
end
end
if is_plurale_tantum then
-- both singular and plural
insert(data.inflections, {label = "sometimes " .. glossary_link("plural only") .. ", in variation"})
end
if mode == "?" then
-- Plural is unknown
insert(data.categories, langname .. " " .. category_plpos .. " with unknown or uncertain plurals")
elseif mode == "!" then
-- Plural is not attested
insert(data.inflections, {label = "plural not attested"})
insert(data.categories, langname .. " " .. category_plpos .. " with unattested plurals")
if plurals[1] then
error("Can't specify any plurals along with unattested plural code '!'")
end
elseif mode == "-" then
-- Uncountable noun; may occasionally have a plural
insert(data.categories, langname .. " uncountable " .. category_plpos)
-- If plural forms were given explicitly, then show "usually"
if plurals[1] then
insert(data.inflections, {label = "usually " .. glossary_link("uncountable")})
insert(data.categories, langname .. " countable " .. category_plpos)
else
insert(data.inflections, {label = glossary_link("uncountable")})
end
else
-- Countable or mixed countable/uncountable
if not plurals[1] and not is_proper then
plurals[1] = {term = "+"}
end
if mode == "~" then
-- Mixed countable/uncountable noun, always has a plural
insert(data.inflections, {label = glossary_link("countable") .. " and " .. glossary_link("uncountable")})
insert(data.categories, langname .. " uncountable " .. category_plpos)
insert(data.categories, langname .. " countable " .. category_plpos)
elseif plurals[1] then
-- Countable nouns
insert(data.categories, langname .. " countable " .. category_plpos)
else
-- Uncountable nouns
insert(data.categories, langname .. " uncountable " .. category_plpos)
end
end
-- Gather plurals, handling requests for default plurals.
local has_default_or_hash = false
for _, pl in ipairs(plurals) do
if pl.term:find("^%+") or pl.term:find("#") then
has_default_or_hash = true
break
end
end
if has_default_or_hash then
local newpls = {}
for _, pl in ipairs(plurals) do
if pl.term == "+" then
local default_pls = com.make_plural(lemma, gender_for_default_plural)
insert_defpls(default_pls, pl, newpls)
elseif pl.term:find("^%+") then
pl.term = require(romut_module).get_special_indicator(pl.term)
local default_pls = com.make_plural(lemma, gender_for_default_plural, pl.term)
insert_defpls(default_pls, pl, newpls)
else
pl.term = replace_hash_with_lemma(pl.term, lemma)
insert(newpls, pl)
end
end
plurals = newpls
end
local pl1 = plurals[1]
if pl1 and not plurals[2] and pl1.term == lemma then
insert(data.inflections, {label = glossary_link("invariable"),
q = pl1.q, qq = pl1.qq, l = pl1.l, ll = pl1.ll, refs = pl1.refs
})
insert(data.categories, langname .. " indeclinable " .. category_plpos)
else
insert_noun_inflection(plurals, "plural", "p")
end
if plurals[2] then
insert(data.categories, langname .. " " .. category_plpos .. " with multiple plurals")
end
end
-- Gather masculines/feminines. For each one, generate the corresponding plural. `field` is the name of the field
-- containing the masculine or feminine forms (normally "m" or "f"); `inflect` is a function of one or two arguments
-- to generate the default masculine or feminine from the lemma (the arguments are the lemma and optionally a
-- "special" flag to indicate how to handle multiword lemmas, and the function is normally make_feminine or
-- make_masculine from [[Module:ca-common]]); and `default_plurals` is a list into which the corresponding default
-- plurals of the gathered or generated masculine or feminine forms are stored.
local function handle_mf(field, inflect, default_plurals)
local function call_inflect(special)
if inflect then
-- Generate default feminine.
return inflect(lemma, special)
else
-- FIXME
error("Can't generate default masculine currently")
end
end
local mfs = m_headword_utilities.parse_term_list_with_modifiers {
paramname = field,
forms = args[field],
splitchar = ",",
frob = function(term)
if term == "+" then
-- Generate default masculine/feminine.
term = call_inflect()
else
term = replace_hash_with_lemma(term, lemma)
end
local special = require(romut_module).get_special_indicator(term)
if special then
term = call_inflect(special)
end
return term
end
}
for _, mf in ipairs(mfs) do
local mfpls = com.make_plural(mf.term, gender, special)
if mfpls then
for _, mfpl in ipairs(mfpls) do
local plobj = m_table.shallowCopy(mf)
plobj.term = mfpl
-- Add an accelerator for each masculine/feminine plural whose lemma
-- is the corresponding singular, so that the accelerated entry
-- that is generated has a definition that looks like
-- # {{plural of|ca|MFSING}}
plobj.accel = {form = "p", lemma = mf.term}
table.insert(default_plurals, plobj)
end
end
end
return mfs
end
local feminine_plurals = {}
local feminines = handle_mf("f", com.make_feminine, feminine_plurals)
local masculine_plurals = {}
local masculines = handle_mf("m", com.make_masculine, masculine_plurals)
local function handle_mf_plural(mfplfield, default_plurals, singulars)
local mfpl = m_headword_utilities.parse_term_list_with_modifiers {
paramname = mfplfield,
forms = args[mfplfield],
splitchar = ",",
}
local new_mfpls = {}
local saw_plus
for i, mfpl in ipairs(mfpl) do
local accel
if #mfpl == #singulars then
-- If same number of overriding masculine/feminine plurals as singulars, assume each plural goes with
-- the corresponding singular and use each corresponding singular as the lemma in the accelerator. The
-- generated entry will have
-- # {{plural of|ca|SINGULAR}}
-- as the definition.
accel = {form = "p", lemma = singulars[i].term}
else
accel = nil
end
if mfpl.term == "+" then
-- We should never see + twice. If we do, it will lead to problems since we overwrite the values of
-- default_plurals the first time around.
if saw_plus then
error(("Saw + twice when handling %s="):format(mfplfield))
end
saw_plus = true
if not default_plurals[1] then
-- FIXME: Can this happen? Not in corresponding Spanish code and the old Portuguese code tried to
-- handle this condition by generating the default plural from the lemma.
error("Internal error: Something wrong, no generated default m/f plurals at this stage")
end
for _, defpl in ipairs(default_plurals) do
-- defpl is already a table and has an accel field
m_headword_utilities.combine_termobj_decorations(defpl, mfpl)
insert(new_mfpls, defpl)
end
elseif mfpl.term:find("^%+") then
mfpl.term = require(romut_module).get_special_indicator(mfpl.term)
for _, mf in ipairs(singulars) do
local default_mfpls = com.make_plural(mf.term, gender, mfpl.term)
for _, defp in ipairs(default_mfpls) do
local mfplobj = m_table.shallowCopy(mfpl)
mfplobj.term = defp
mfplobj.accel = accel
m_headword_utilities.combine_termobj_decorations(mfplobj, mf)
insert(new_mfpls, mfplobj)
end
end
else
mfpl.accel = accel
mfpl.term = replace_hash_with_lemma(mfpl.term, lemma)
insert(new_mfpls, mfpl)
end
end
return new_mfpls
end
if args.fpl[1] then
-- Override any existing feminine plurals.
feminine_plurals = handle_mf_plural("fpl", feminine_plurals, feminines)
end
if args.mpl[1] then
-- Override any existing masculine plurals.
masculine_plurals = handle_mf_plural("mpl", masculine_plurals, masculines)
end
local function parse_and_insert_noun_inflection(field, label, accel)
parse_and_insert_inflection(data, args, field, label, accel)
end
insert_noun_inflection(feminines, "feminine", "f")
insert_noun_inflection(feminine_plurals, "feminine plural")
insert_noun_inflection(masculines, "masculine")
insert_noun_inflection(masculine_plurals, "masculine plural")
parse_and_insert_noun_inflection("dim", "diminutive")
parse_and_insert_noun_inflection("aug", "augmentative")
parse_and_insert_noun_inflection("pej", "pejorative")
parse_and_insert_noun_inflection("dem", "demonym")
parse_and_insert_noun_inflection("fdem", "female demonym")
-- Is this a noun with an unexpected ending (for its gender)?
-- Only check if the term is one word (there are no spaces in the term).
local irreg_gender_lemma = rsub(lemma, " .*", "") -- only look at first word
if (gender_for_irreg_ending == "m" or gender_for_irreg_ending == "mf") and irreg_gender_lemma:find("a$") then
insert(data.categories, langname .. " masculine " .. category_plpos .. " ending in -a")
elseif (gender_for_irreg_ending == "f" or gender_for_irreg_ending == "mf") and not (
irreg_gender_lemma:find("a$") or irreg_gender_lemma:find("ió$") or irreg_gender_lemma:find("tat$") or
irreg_gender_lemma:find("tud$") or irreg_gender_lemma:find("[dt]riu$")) then
insert(data.categories, langname .. " feminine " .. category_plpos .. " with no feminine ending")
end
end
local function get_noun_params(is_proper)
return {
[1] = {list = "g", disallow_holes = true, required = not is_proper, default = "?", type = "genders",
flatten = true}, -- gender(s)
[2] = {list = "pl", disallow_holes = true}, --plural override(s)
["f"] = list_param, --feminine form(s)
["m"] = list_param, --masculine form(s)
["fpl"] = list_param, --feminine plural override(s)
["mpl"] = list_param, --masculine plural override(s)
["dim"] = list_param, --diminutive(s)
["aug"] = list_param, --diminutive(s)
["pej"] = list_param, --pejorative(s)
["dem"] = list_param, --demonym(s)
["fdem"] = list_param, --female demonym(s)
}
end
pos_functions["Kata nama"] = {
params = get_noun_params(),
func = do_noun,
}
pos_functions["Kata nama khas"] = {
params = get_noun_params("is proper"),
func = function(args, data)
do_noun(args, data, "is proper")
end,
}
-----------------------------------------------------------------------------------------
-- Verbs --
-----------------------------------------------------------------------------------------
pos_functions["Kata kerja"] = {
params = {
[1] = true,
["pres"] = list_param, --present
["pres_qual"] = {list = "pres\1_qual", allow_holes = true},
["pres3s"] = list_param, --third-singular present
["pres3s_qual"] = {list = "pres3s\1_qual", allow_holes = true},
["pret"] = list_param, --preterite
["pret_qual"] = {list = "pret\1_qual", allow_holes = true},
["part"] = list_param, --participle
["part_qual"] = {list = "part\1_qual", allow_holes = true},
["short_part"] = list_param, --short participle
["short_part_qual"] = {list = "short_part\1_qual", allow_holes = true},
["noautolinktext"] = boolean_param,
["noautolinkverb"] = boolean_param,
["attn"] = boolean_param,
["pres_1_sg"] = true, -- accept any ignore old-style param
["past_part"] = true, -- accept any ignore old-style param
["root"] = true, -- FIXME: Implement root-stressed vowel quality
},
func = function(args, data, tracking_categories, frame)
local preses, preses_3s, prets, parts, short_parts
if args.attn then
insert(tracking_categories, "Requests for attention concerning " .. langname)
return
end
local ca_verb = require(ca_verb_module)
local alternant_multiword_spec = ca_verb.do_generate_forms(args, "ca-verb", data.heads[1])
local specforms = alternant_multiword_spec.forms
local function slot_exists(slot)
return specforms[slot] and #specforms[slot] > 0
end
local function do_finite(slot_tense, label_tense)
-- Use pres_3s if it exists and pres_1s doesn't exist (e.g. impersonal verbs); similarly for pres_3p (only3p verbs);
-- but fall back to pres_1s if neither pres_1s nor pres_3s nor pres_3p exist (e.g. [[empedernir]]).
local has_1s = slot_exists(slot_tense .. "_1s")
local has_3s = slot_exists(slot_tense .. "_3s")
local has_3p = slot_exists(slot_tense .. "_3p")
if has_1s or (not has_3s and not has_3p) then
return {
slot = slot_tense .. "_1s",
label = ("first-person singular %s"):format(label_tense),
}, true
elseif has_3s then
return {
slot = slot_tense .. "_3s",
label = ("third-person singular %s"):format(label_tense),
}, false
else
return {
slot = slot_tense .. "_3p",
label = ("third-person plural %s"):format(label_tense),
}, false
end
end
local did_pres_1s
preses, did_pres_1s = do_finite("pres", "present")
preses_3s = {
slot = "pres_3s",
label = "third-person singular present",
}
prets = do_finite("pret", "preterite")
parts = {
slot = "pp_ms",
label = "past participle",
}
short_parts = {
slot = "short_pp_ms",
label = "short past participle",
}
if args.pres[1] or args.pres3s[1] or args.pret[1] or args.part[1] or args.short_part[1] then
track("verb-old-multiarg")
end
local function strip_brackets(qualifiers)
if not qualifiers then
return nil
end
local stripped_qualifiers = {}
for _, qualifier in ipairs(qualifiers) do
local stripped_qualifier = qualifier:match("^%[(.*)%]$")
if not stripped_qualifier then
error("Internal error: Qualifier should be surrounded by brackets at this stage: " .. qualifier)
end
insert(stripped_qualifiers, stripped_qualifier)
end
return stripped_qualifiers
end
local function do_verb_form(args, qualifiers, slot_desc, skip_if_empty)
local forms
local to_insert
if #args == 0 then
forms = specforms[slot_desc.slot]
if not forms or #forms == 0 then
if skip_if_empty then
return
end
forms = {{form = "-"}}
end
elseif #args == 1 and args[1] == "-" then
forms = {{form = "-"}}
else
forms = {}
for i, arg in ipairs(args) do
local qual = qualifiers[i]
if qual then
-- FIXME: It's annoying we have to add brackets and strip them out later. The inflection
-- code adds all footnotes with brackets around them; we should change this.
qual = {"[" .. qual .. "]"}
end
local form = arg
if not args.noautolinkverb then
-- [[Module:inflection utilities]] already loaded by [[Module:ca-verb]]
form = require(inflection_utilities_module).add_links(form)
end
insert(forms, {form = form, footnotes = qual})
end
end
if forms[1].form == "-" then
to_insert = {label = "no " .. slot_desc.label}
else
local into_table = {label = slot_desc.label}
for _, form in ipairs(forms) do
local qualifiers = strip_brackets(form.footnotes)
-- Strip redundant brackets surrounding entire form. These may get generated e.g.
-- if we use the angle bracket notation with a single word.
local stripped_form = rmatch(form.form, "^%[%[([^%[%]]*)%]%]$") or form.form
-- Don't include accelerators if brackets remain in form, as the result will be wrong.
-- FIXME: For now, don't include accelerators. We should use the new {{ca-verb form of}}.
-- local this_accel = not stripped_form:find("%[%[") and accel or nil
local this_accel = nil
insert(into_table, {term = stripped_form, q = qualifiers, accel = this_accel})
end
to_insert = into_table
end
insert(data.inflections, to_insert)
end
local skip_pres_if_empty
if alternant_multiword_spec.no_pres1_and_sub then
insert(data.inflections, {label = "no first-person singular present"})
insert(data.inflections, {label = "no present subjunctive"})
end
if alternant_multiword_spec.no_pres_stressed then
insert(data.inflections, {label = "no stressed present indicative or subjunctive"})
skip_pres_if_empty = true
end
if alternant_multiword_spec.only3s then
insert(data.inflections, {label = glossary_link("impersonal")})
elseif alternant_multiword_spec.only3sp then
insert(data.inflections, {label = "third-person only"})
elseif alternant_multiword_spec.only3p then
insert(data.inflections, {label = "third-person plural only"})
end
local has_vowel_alt
if alternant_multiword_spec.vowel_alt then
for _, vowel_alt in ipairs(alternant_multiword_spec.vowel_alt) do
if vowel_alt ~= "+" and vowel_alt ~= "í" and vowel_alt ~= "ú" then
has_vowel_alt = true
break
end
end
end
do_verb_form(args.pres, args.pres_qual, preses, skip_pres_if_empty)
-- We want to include both the pres_1s and pres_3s if there is a vowel alternation in the present singular. But we
-- don't want to redundantly include the pres_3s if we already included it.
if did_pres_1s and has_vowel_alt then
do_verb_form(args.pres3s, args.pres3s_qual, preses_3s, skip_pres_if_empty)
end
do_verb_form(args.pret, args.pret_qual, prets)
do_verb_form(args.part, args.part_qual, parts)
do_verb_form(args.short_part, args.short_part_qual, short_parts, "skip if empty")
-- Add categories.
for _, cat in ipairs(alternant_multiword_spec.categories) do
insert(data.categories, cat)
end
-- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to
-- incorporate any links in that head into the 1= specification, use the infinitive generated by
-- [[Module:ca-verb]] in place of the user-specified or auto-generated head. This was copied from
-- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on
-- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the
-- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian
-- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Catalan equivalent).
if #data.user_specified_heads == 0 or (
#data.user_specified_heads == 1 and alternant_multiword_spec.incorporated_headword_head_into_lemma
) then
data.heads = {}
for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do
local quals, refs = require(inflection_utilities_module).
convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes)
insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs})
end
end
if args.root then
local m_ca_IPA = require(ca_IPA_module)
local parsed_respellings = {}
local function set_parsed_respelling(dialect, parsed)
-- Validate the individual root vowel specs.
for _, termobj in ipairs(parsed.terms) do
if not rfind(termobj.words[1].term, "^" .. m_ca_IPA.mid_vowel_hint_c .. "$") then
error(("Root vowel spec '%s' should be one of the vowels %s"):format(
termobj.words[1].term, m_ca_IPA.mid_vowel_hints))
end
end
if not dialect then
for _, dial in ipairs(m_ca_IPA.dialects) do
-- Need to clone as we destructively modify each one later with the pronun.
parsed_respellings[dial] = m_table.deepCopy(parsed)
end
elseif m_ca_IPA.dialect_groups[dialect] then
for _, dial in ipairs(m_ca_IPA.dialect_groups[dialect]) do
-- Need to clone as we destructively modify each one later with the pronun.
parsed_respellings[dial] = m_table.deepCopy(parsed)
end
else
parsed_respellings[dialect] = parsed
end
end
local function check_dialect_or_dialect_group(dialect)
if not m_table.contains(m_ca_IPA.dialects, dialect) and not
m_ca_IPA.dialect_groups[dialect] then
local dialect_list = {}
for _, dial in ipairs(m_ca_IPA.dialects) do
insert(dialect_list, "'" .. dial .. "'")
end
dialect_list = list_to_text(dialect_list, nil, " or ")
local dialect_group_list = {}
for dialect_group, _ in pairs(m_ca_IPA.dialect_groups) do
insert(dialect_group_list, "'" .. dialect_group .. "'")
end
dialect_group_list = list_to_text(dialect_group_list, nil, " or ")
error(("Unrecognized dialect '%s': Should be a dialect %s or a dialect group %s"):format(
dialect, dialect_list, dialect_group_list))
end
end
-- Parse the root vowel specs.
if args.root:find("[<%[]") then
local put = require(parse_utilities_module)
-- Parse balanced segment runs involving either [...] (substitution notation) or <...> (inline
-- modifiers). We do this because we don't want commas or semicolons inside of square or angle brackets
-- to count as respelling delimiters. However, we need to rejoin square-bracketed segments with nearby
-- ones after splitting alternating runs on comma and semicolon.
local segments = put.parse_multi_delimiter_balanced_segment_run(args.root, {{"<", ">"}, {"[", "]"}})
local semicolon_separated_groups = put.split_alternating_runs(segments, "%s*;%s*")
for _, group in ipairs(semicolon_separated_groups) do
local first_element = group[1]
local dialect
if first_element:find("^[a-z]+:") then
-- a dialect-specific spec
local rest
dialect, rest = first_element:match("^([a-z]+):(.*)$")
check_dialect_or_dialect_group(dialect)
group[1] = rest
end
local comma_separated_groups = put.split_alternating_runs_on_comma(group)
-- Process each value.
local outer_container = m_ca_IPA.parse_comma_separated_groups(comma_separated_groups, true, args.root,
"root")
set_parsed_respelling(dialect, outer_container)
end
else
for _, dialect_spec in ipairs(rsplit(args.root, "%s*;%s*")) do
local dialect
if dialect_spec:find("^[a-z]+:") then
-- a dialect-specific spec
local rest
dialect, rest = dialect_spec:match("^([a-z]+):(.*)$")
check_dialect_or_dialect_group(dialect)
dialect_spec = rest
end
local termobjs = {}
for _, word in ipairs(rsplit(dialect_spec, ",")) do
insert(termobjs, {words = {{term = word}}})
end
set_parsed_respelling(dialect, {
terms = termobjs,
})
end
end
-- Convert each canonicalized respelling to phonemic/phonetic IPA.
m_ca_IPA.generate_phonemic_phonetic(parsed_respellings)
-- Group the results.
local grouped_pronuns = m_ca_IPA.group_pronuns_by_dialect(parsed_respellings)
-- Format for display.
for _, grouped_pronun_spec in pairs(grouped_pronuns) do
local pronunciations = {}
local function ins(text)
insert(pronunciations, text)
end
-- Loop through each pronunciation. For each one, format the phonetic version "raw".
for j, pronun in ipairs(grouped_pronun_spec.pronuns) do
-- Add dialect tags to left accent qualifiers if first one
local as = pronun.a
if j == 1 then
if as then
as = m_table.deepCopy(as)
else
as = {}
end
for _, dialect in ipairs(grouped_pronun_spec.dialects) do
insert(as, m_ca_IPA.dialects_to_names[dialect])
end
else
ins(", ")
end
local slash_pron = "/" .. pronun.phonetic:gsub("ˈ", "") .. "/"
if as or pronun.q or pronun.qq or pronun.aa then
ins(require(decorations_module).format_decorations {
lang = lang,
text = slash_pron,
q = pronun.q,
a = as,
qq = pronun.qq,
aa = pronun.aa
})
else
ins(slash_pron)
end
if pronun.refs then
-- FIXME: Copied from [[Module:IPA]]. Should be in a module.
local refs = {}
if #pronun.refs > 0 then
for _, refspec in ipairs(pronun.refs) do
if type(refspec) ~= "table" then
refspec = {text = refspec}
end
local refargs
if refspec.name or refspec.group then
refargs = {name = refspec.name, group = refspec.group}
end
insert(refs, mw.getCurrentFrame():extensionTag("ref", refspec.text, refargs))
end
ins(concat(refs))
end
end
end
grouped_pronun_spec.formatted = concat(pronunciations)
end
-- Concatenate formatted results.
local formatted = {}
for _, grouped_pronun_spec in ipairs(grouped_pronuns) do
insert(formatted, grouped_pronun_spec.formatted)
end
data.post_note = "''root stress'': " .. concat(formatted, "; ")
end
end
}
-----------------------------------------------------------------------------------------
-- Numerals --
-----------------------------------------------------------------------------------------
-- Display additional inflection information for a numeral
pos_functions["Kata bilangan"] = {
params = {
[1] = true,
[2] = true,
},
func = function(args, data)
if args[1] then
insert(data.genders, "m")
parse_and_insert_inflection(data, args, 1, "feminine")
parse_and_insert_inflection(data, args, 2, "noun form")
else
insert(data.genders, "m")
insert(data.genders, "f")
end
end
}
-----------------------------------------------------------------------------------------
-- Phrases --
-----------------------------------------------------------------------------------------
pos_functions["frasa"] = {
params = {
["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true},
["m"] = list_param,
["f"] = list_param,
},
func = function(args, data)
validate_genders(args.g)
data.genders = args.g
parse_and_insert_inflection(data, args, "m", "masculine")
parse_and_insert_inflection(data, args, "f", "feminine")
end,
}
-----------------------------------------------------------------------------------------
-- Suffix forms --
-----------------------------------------------------------------------------------------
pos_functions["bentuk akhiran"] = {
params = {
[1] = {required = true, list = true, disallow_holes = true},
["g"] = {list = true, disallow_holes = true, type = "genders", flatten = true},
},
func = function(args, data)
validate_genders(args.g)
data.genders = args.g
local suffix_type = {}
for _, typ in ipairs(args[1]) do
insert(suffix_type, typ .. "-forming suffix")
end
insert(data.inflections, {label = "non-lemma form of " .. m_table.serialCommaJoin(suffix_type, {conj = "or"})})
end,
}
return export
oqhw4n4h2xa5p5vi46bv3kr1mv8txpb
Modul:affix
828
10384
375358
373580
2026-09-22T03:14:21Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708528|92708528]])
375358
Scribunto
text/plain
local export = {}
local debug_force_cat = false -- if set to true, always display categories even on userspace pages
local m_links = require("Module:links")
local m_str_utils = require("Module:string utilities")
local m_table = require("Module:table")
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local etymology_module = "Module:etymology"
local scripts_module = "Module:scripts"
local utilities_module = "Module:utilities"
-- Export this so the category code in [[Module:category tree/etymology]] can access it.
export.affix_lang_data_module_prefix = "Module:affix/lang-data/"
local ulen = m_str_utils.len
local rfind = m_str_utils.find
local rmatch = m_str_utils.match
local pluralize = require(en_utilities_module).pluralize
local singularize = require(en_utilities_module).singularize
local u = m_str_utils.char
local ucfirst = m_str_utils.ucfirst
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
function export.affix_variants(canonical, variants)
local mappings = {}
for _, variant in ipairs(variants) do
mappings[variant] = canonical
end
return mappings
end
function export.id_mapping(default, ids)
local mapping = { default = default }
if ids then
for id, target in pairs(ids) do
mapping[id] = target
end
end
return mapping
end
function export.id_mapping_with_affix_variants(base, id_variants)
local mappings = {}
for id, variants in pairs(id_variants) do
for _, variant in ipairs(variants) do
mappings[variant] = export.id_mapping(base, {[id] = base})
end
end
return mappings
end
function export.merge_tables(...)
local result = {}
for i = 1, select('#', ...) do
local t = select(i, ...)
if t then
for k, v in pairs(t) do
result[k] = v
end
end
end
return result
end
-- Export this so the category code in [[Module:category tree/etymology]] can access it.
export.langs_with_lang_specific_data = {
["az"] = true,
["fi"] = true,
["fr"] = true,
["izh"] = true,
["la"] = true,
["sah"] = true,
["tr"] = true,
["trk-pro"] = true,
}
local default_pos = "perkataan"
local function pluralize_pos(pos)
return pluralize(singularize(pos))
end
--[==[ intro:
===About different types of hyphens ("template", "display" and "lookup"):===
* The "template hyphen" is the per-script hyphen character that is used in template calls to indicate that a term is an
affix. This is always a single Unicode char, but there may be multiple possible hyphens for a given script. Normally
this is just the regular hyphen character "-", but for some non-Latin-script languages (currently only right-to-left
languages), it is different.
* The "display hyphen" is the string (which might be an empty string) that is added onto a term as displayed and linked,
to indicate that a term is an affix. Currently this is always either the same as the template hyphen or an empty
string, but the code below is written generally enough to handle arbitrary display hyphens. Specifically:
*# For East Asian languages, the display hyphen is always blank.
*# For Arabic-script languages, either tatweel (ـ) or ZWNJ (zero-width non-joiner) are allowed as template hyphens,
where ZWNJ is supported primarily for Farsi, because some suffixes have non-joining behavior. The display hyphen
corresponding to tatweel is also tatweel, but the display hyphen corresponding to ZWNJ is blank (tatweel is also
the default display hyphen, for calls to {{tl|prefix}}/{{tl|suffix}}/etc. that don't include an explicit hyphen).
* The "lookup hyphen" is the hyphen that is used when looking up language-specific affix mappings. (These mappings are
discussed in more detail below when discussing link affixes.) It depends only on the script of the affix in question.
Most scripts (including East Asian scripts) use a regular hyphen "-" as the lookup hyphen, but Hebrew and Arabic
have their own lookup hyphens (respectively maqqef and tatweel). Note that for Arabic in particular, there are
three possible template hyphens that are recognized (tatweel, ZWNJ and regular hyphen), but mappings must use tatweel.
===About different types of affixes ("template", "display", "link", "lookup" and "category"):===
* A "template affix" is an affix in its source form as it appears in a template call. Generally, a template affix has an
attached template hyphen (see above) to indicate that it is an affix and indicate what type of affix it is (prefix,
suffix, interfix or circumfix), but some of the older-style templates such as {{tl|suffix}}, {{tl|prefix}},
{{tl|confix}}, etc. have "positional" affixes where the presence of the affix in a certain position (e.g. the second
or third parameter) indicates that it is a certain type of affix, whether or not it has an attached template hyphen.
* A "display affix" is the corresponding affix as it is actually displayed to the user. The display affix may differ
from the template affix for various reasons:
*# The display affix may be specified explicitly using the {{para|alt<var>N</var>}} parameter, the `<alt:...>` inline
modifier or a piped link of the form e.g. `<nowiki>[[-kas|-käs]]</nowiki>` (here indicating that the affix should
display as `-käs` but be linked as `-kas`). Here, the template affix is arguably the entire piped link, while the
display affix is `-käs`.
*# Even in the absence of {{para|alt<var>N</var>}} parameters, `<alt:...>` inline modifiers and piped links, certain
languages have differences between the "template hyphen" specified in the template (which always needs to be
specified somehow or other in templates like {{tl|affix}}, to indicate that the term is an affix and what type of
affix it is) and the display hyphen (see above), with corresponding differences between template and display
affixes.
* A (regular) "link affix" is the affix that is linked to when the affix is shown to the user. The link affix is usually
the same as the display affix, but will differ in one of three circumstances:
*# The display and link affixes are explicitly made different using {{para|alt<var>N</var>}} parameters, `<alt:...>`
inline modifiers or piped links, as described above under "display affix".
*# For certain languages, certain affixes are mapped to canonical form using language-specific mappings. For example,
in Finnish, the adjective-forming suffix {{m|fi|-kas}} appears as {{m|fi|-käs}} after front vowels, but logically
both forms are the same suffix and should be linked and categorized the same. Similarly, in Latin, the negative and
intensive prefixes spelled {{m|la|in-}} (etymologically two distinct prefixes) appear variously as {{m|la|il-}},
{{m|la|im-}} or {{m|la|ir-}} before certain consonants. Mappings are supplied in [[Module:affix/lang-data/LANGCODE]]
to convert Finnish {{m|fi|-käs}} to {{m|fi|-kas}} for linking and categorization purposes. Note that the affixes in
the mappings use "lookup hyphens" to indicate the different types of affixes, which is usually the same as the
template hyphen but differs for Arabic scripts, because there are multiple possible template hyphens recognized but
only one lookup hyphen (tatweel). The form of the affix as used to look up in the mapping tables is called the
"lookup affix"; see below.
* A "stripped link affix" is a link affix that has been passed through the language's `stripDiacritics()` function, which
may strip certain diacritics: e.g. macrons in Latin and Old English (indicating length); acute and grave accents in
Russian and various other Slavic languages (indicating stress); vowel diacritics in most Arabic-script languages; and
also tatweel in some Arabic-script languages (currently, for example, Persian, Arabic and Urdu strip tatweel, but
Ottoman Turkish does not). Stripped link affixes are currently what are used in category names.
* A "lookup affix" is the form of the affix as it is looked up in the language-specific lookup mappings described above
under link affixes. There are actually two lookup stages:
*# First, the affix is looked up in a modified display form (specifically, the same as the display affix but using
lookup hyphens). Note that this lookup does not occur if an explicit display form is given using
{{para|alt<var>N</var>}} or an `<alt:...>` inline modifier, or if the template affix contains a piped or embedded
link.
*# If no entry is found, the affix is then looked up in a modified link form (specifically, the modified display
form passed through the language's `stripDiacritics()` function, which strips out certain diacritics, but with the
lookup hyphen re-added if it was stripped out, as in the case of tatweel in many Arabic-script languages).
The reason for this double lookup procedure is to allow for mappings that are sensitive to the extra diacritics, but
also allow for mappings that are not sensitive in this fashion (e.g. Russian {{m|ru|-ливый}} occurs both stressed and
unstressed, but is the same prefix either way).
* A "category affix" is the affix as it appears in categories such as [[:Category:Finnish terms suffixed with -kas|
Category:Finnish terms suffixed with ''-kas'']]. The category affix is currently always the same as the stripped link
affix. This means that for Arabic-script languages, it may or may not have a tatweel, even if the correponding display
affix and regular link affix have a tatweel. As mentioned above, stripDiacritics() strips tatweel for Arabic, Persian
and Urdu, but not for Ottoman Turkish. Hence affix categories for Arabic, Persian and Urdu will be missing the
tatweel, but affix categories for Ottoman Turkish will have it. An additional complication is that if the template
affix contains a ZWNJ, the display (and hence the link and category affixes) will have no hyphen attached in any case.
]==]
-----------------------------------------------------------------------------------------
-- Template and display hyphens --
-----------------------------------------------------------------------------------------
--[=[
Per-script template hyphens. The template hyphen is what appears in the {{affix}}/{{prefix}}/{{suffix}}/etc. template
(in the wikicode). See above.
They key below is a script code, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong'
and 'xwo-Mong' will match 'Mong'.
The value below is a string consisting of one or more hyphen characters. If there is more than one character, the
default hyphen must come last and a non-default function must be specified for the script in display_hyphens[] so
the correct display hyphen will be specified when no template hyphen is given (in {{suffix}}/{{prefix}}/etc.).
Script detection is normally done when linking, but we need to do it earlier. However, under most circumstances we
don't need to do script detection. Specifically, we only need to do script detection for a given language if
(a) the language has multiple scripts; and
(b) at least one of those scripts is listed below or in display_hyphens.
]=]
local ZWNJ = u(0x200C) -- zero-width non-joiner
local template_hyphens = {
-- This covers all Arabic scripts. See above.
["Arab"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen
["Aran"] = "ـ" .. ZWNJ .. "-", -- tatweel + zero-width non-joiner + regular hyphen
["Hebr"] = "־", -- Hebrew-specific hyphen termed "maqqef"
["Mong"] = "᠊",
-- FIXME! What about the following right-to-left scripts?
-- Adlm (Adlam)
-- Armi (Imperial Aramaic)
-- Avst (Avestan)
-- Cprt (Cypriot)
-- Khar (Kharoshthi)
-- Mand (Mandaic/Mandaean)
-- Mani (Manichaean)
-- Mend (Mende/Mende Kikakui)
-- Narb (Old North Arabian)
-- Nbat (Nabataean/Nabatean)
-- Nkoo (N'Ko)
-- Orkh (Orkhon runes)
-- Phli (Inscriptional Pahlavi)
-- Phlp (Psalter Pahlavi)
-- Phlv (Book Pahlavi)
-- Phnx (Phoenician)
-- Prti (Inscriptional Parthian)
-- Rohg (Hanifi Rohingya)
-- Samr (Samaritan)
-- Sarb (Old South Arabian)
-- Sogd (Sogdian)
-- Sogo (Old Sogdian)
-- Syrc (Syriac)
-- Thaa (Thaana)
}
-- Hyphens used when looking up an affix in a lang-specific affix mapping. Defaults to regular hyphen (-). The keys
-- are script codes, after removing a hyphen and anything preceding. Hence, script codes like 'mnc-Mong' and 'xwo-Mong'
-- will match 'Mong'. The value should be a single character.
local lookup_hyphens = {
["Hebr"] = "־",
-- This covers all Arabic scripts. See above.
["Arab"] = "ـ",
["Aran"] = "ـ",
}
-- Default display-hyphen function.
local function default_display_hyphen(script, hyph)
if not hyph then
return template_hyphens[script] or "-"
end
return hyph
end
local function arab_get_display_hyphen(_script, hyph)
if not hyph then
return "ـ" -- tatweel
elseif hyph == ZWNJ then
return ""
else
return hyph
end
end
local function no_display_hyphen(_script, _hyph)
return ""
end
-- Per-script function to return the correct display hyphen given the script and template hyphen. The function should
-- also handle the case where the passed-in template hyphen is nil, corresponding to the situation in
-- {{prefix}}/{{suffix}}/etc. where no template hyphen is specified. The key is the script code after removing a hyphen
-- and anything preceding, so 'mnc-Mong', 'xwo-Mong' etc. will match 'Mong'.
local display_hyphens = {
-- This covers all Arabic scripts. See above.
["Arab"] = arab_get_display_hyphen,
["Aran"] = arab_get_display_hyphen,
["Bopo"] = no_display_hyphen,
["Hani"] = no_display_hyphen,
["Hans"] = no_display_hyphen,
["Hant"] = no_display_hyphen,
-- The following is a mixture of several scripts. Hopefully the specs here are correct!
["Jpan"] = no_display_hyphen,
["Jurc"] = no_display_hyphen,
["Kitl"] = no_display_hyphen,
["Kits"] = no_display_hyphen,
["Laoo"] = no_display_hyphen,
["Nshu"] = no_display_hyphen,
["Shui"] = no_display_hyphen,
["Tang"] = no_display_hyphen,
["Thaa"] = no_display_hyphen,
["Thai"] = no_display_hyphen,
["Tibt"] = no_display_hyphen,
}
-----------------------------------------------------------------------------------------
-- Basic Utility functions --
-----------------------------------------------------------------------------------------
local function glossary_link(entry, text)
text = text or entry
return "[[Lampiran:Glosari#" .. entry .. "|" .. text .. "]]"
end
local function track(page)
if type(page) == "table" then
for i, pg in ipairs(page) do
page[i] = "affix/" .. pg
end
else
page = "affix/" .. page
end
require("Module:debug/track")(page)
end
local function ine(val)
return val ~= "" and val or nil
end
-----------------------------------------------------------------------------------------
-- Compound types --
-----------------------------------------------------------------------------------------
local function make_compound_type(typ, alttext)
return {
text = "kata majmuk " .. glossary_link(typ, alttext),
cat = "Kata majmuk " .. typ,
}
end
-- Make a compound type entry with a simple rather than glossary link.
-- These should be replaced with a glossary link when the entry in the glossary
-- is created.
local function make_non_glossary_compound_type(typ, alttext)
local link = alttext and "[[" .. typ .. "|" .. alttext .. "]]" or "[[" .. typ .. "]]"
return {
text = "kata majmuk " .. link,
cat = "Kata majmuk " .. typ,
}
end
local function make_raw_compound_type(typ, alttext)
return {
text = glossary_link(typ, alttext),
cat = pluralize(typ),
}
end
local function make_borrowing_type(typ, alttext)
return {
text = glossary_link(typ, alttext),
borrowing_type = pluralize(typ),
}
end
export.etymology_types = {
["adapted borrowing"] = make_borrowing_type("pinjaman tersuai"),
["adap"] = "pinjaman tersuai",
["abor"] = "pinjaman tersuai",
["alliterative"] = make_non_glossary_compound_type("alliterative"),
["allit"] = "aliteratif",
["antonymous"] = make_non_glossary_compound_type("antonim"),
["ant"] = "antonim",
["bahuvrihi"] = make_compound_type("bahuvrihi", "bahuvrīhi"),
["bahu"] = "bahuvrihi",
["bv"] = "bahuvrihi",
["coordinative"] = make_compound_type("koordinatif"),
["coord"] = "coordinative",
["descriptive"] = make_compound_type("deskriptif"),
["desc"] = "deskriptif",
["determinative"] = make_compound_type("determinative"),
["det"] = "determinative",
["dvandva"] = make_compound_type("dvandva"),
["dva"] = "dvandva",
["dvigu"] = make_compound_type("dvigu"),
["dvi"] = "dvigu",
["endocentric"] = make_compound_type("endosentrik"),
["endo"] = "endocentric",
["exocentric"] = make_compound_type("eksosentrik"),
["exo"] = "exocentric",
["izafet I"] = make_compound_type("izafet I"),
["iz1"] = "izafet I",
["izafet II"] = make_compound_type("izafet II"),
["iz2"] = "izafet II",
["izafet III"] = make_compound_type("izafet III"),
["iz3"] = "izafet III",
["karmadharaya"] = make_compound_type("karmadharaya", "karmadhāraya"),
["karma"] = "karmadharaya",
["kd"] = "karmadharaya",
["kenning"] = make_raw_compound_type("kenning"),
["ken"] = "kenning",
["rhyming"] = make_non_glossary_compound_type("berima"),
["rhy"] = "berima",
["synonymous"] = make_non_glossary_compound_type("sinonim"),
["syn"] = "sinonim",
["tatpurusa"] = make_compound_type("tatpurusa", "tatpuruṣa"),
["tat"] = "tatpurusa",
["tp"] = "tatpurusa",
}
local function process_etymology_type(typ, nocap, notext, has_parts)
local text_sections = {}
local categories = {}
local borrowing_type
if typ then
local typdata = export.etymology_types[typ]
if type(typdata) == "string" then
typdata = export.etymology_types[typdata]
end
if not typdata then
error("Internal error: Unrecognized type '" .. typ .. "'")
end
local text = typdata.text
if not nocap then
text = ucfirst(text)
end
local cat = typdata.cat
borrowing_type = typdata.borrowing_type
local oftext = typdata.oftext or " bagi"
if not notext then
table.insert(text_sections, text)
if has_parts then
table.insert(text_sections, oftext)
table.insert(text_sections, " ")
end
end
if cat then
table.insert(categories, cat)
end
end
return text_sections, categories, borrowing_type
end
-----------------------------------------------------------------------------------------
-- Utility functions --
-----------------------------------------------------------------------------------------
-- Iterate an array up to the greatest integer index found.
local function ipairs_with_gaps(t)
local indices = m_table.numKeys(t)
local max_index = #indices > 0 and math.max(unpack(indices)) or 0
local i = 0
return function()
if i < max_index then
i = i + 1
return i, t[i]
end
end
end
export.ipairs_with_gaps = ipairs_with_gaps
--[==[
Join formatted parts (in `parts_formatted`) together with any overall {{para|lit}} spec (in `lit`) plus categories,
which are formatted by prepending the language name as found in `lang`. The value of an entry in `categories` can be
either a string (which is formatted using `sort_key`) or a table of the form `{ {cat=<var>category</var>,
sort_key=<var>sort_key</var>, sort_base=<var>sort_base</var>}`, specifying the sort key and sort base to use when
formatting the category. If `nocat` is given, no categories are added; otherwise, `force_cat` causes categories to be
added even on userspace pages.
]==]
function export.join_formatted_parts(data)
local cattext
local lang = data.data.lang
local force_cat = data.data.force_cat or debug_force_cat
if data.data.nocat then
cattext = ""
else
for i, cat in ipairs(data.categories) do
if type(cat) == "table" then
data.categories[i] = require(utilities_module).format_categories(cat.cat .. " bahasa " .. lang:getFullName(),
lang, cat.sort_key, cat.sort_base, force_cat)
else
data.categories[i] = require(utilities_module).format_categories(cat .. " bahasa " .. lang:getFullName(), lang,
data.data.sort_key, nil, force_cat)
end
end
cattext = table.concat(data.categories)
end
local result = table.concat(data.parts_formatted, not data.separator_already_added and " +‎ " or nil) ..
(data.data.lit and ", secara harfiah " .. m_links.mark(data.data.lit, "gloss") or "")
local q = data.data.q
local qq = data.data.qq
local l = data.data.l
local ll = data.data.ll
if q and q[1] or qq and qq[1] or l and l[1] or ll and ll[1] then
result = require(decorations_module).format_decorations {
lang = lang,
text = result,
q = q,
qq = qq,
l = l,
ll = ll,
}
end
return result .. cattext
end
-- Remove links and call lang:stripDiacritics(term).
local function strip_diacritics_no_links(lang, term)
return lang:stripDiacritics(m_links.remove_links(term))
end
--[=[
Convert a raw part as passed into an entry point into a part ready for linking. `lang` and `sc` are the overall
language and script objects. This uses the overall language and script objects as defaults for the part and parses off
any fragment from the term. We need to do the latter so that fragments don't end up in categories and so that we
correctly do affix mapping even in the presence of fragments.
]=]
local function canonicalize_part(part, lang, sc)
if not part then
return
end
-- Save the original (user-specified, part-specific) value of `lang`. If such a value is specified, we don't insert
-- a '*fixed with' category, and we format the part using format_derived() in [[Module:etymology]] rather than
-- full_link() in [[Module:links]].
part.part_lang = part.lang
part.lang = part.lang or lang
part.sc = part.sc or sc
local term = part.term
if not term then
return
elseif not part.fragment then
part.term, part.fragment = m_links.get_fragment(term)
else
part.term = m_links.get_fragment(term)
end
end
--[==[
Construct a single linked part based on the information in `part`, for use by `show_affix()` and other entry points.
This should be called after `canonicalize_part()` is called on the part. This is a thin wrapper around `full_link()` in
[[Module:links]] unless `part.part_lang` is specified (indicating that a part-specific language was given), in which
case `format_derived()` in [[Module:etymology]] is called to display a term in a language other than the language of
the overall term (specified in `data.lang`). `data` contains the entire object passed into the entry point and is used
to access information for constructing the categories added by `format_derived()`.
]==]
function export.link_term(part, data, include_separator)
local result
if part.part_lang then
result = require(etymology_module).format_derived {
lang = data.lang,
terms = {part},
sources = {part.lang},
sort_key = data.sort_key,
nocat = data.nocat,
template_name = "affix",
decorations_on_outside = true,
borrowing_type = data.borrowing_type,
force_cat = data.force_cat or debug_force_cat,
}
else
result = m_links.full_link(part, "term")
end
if include_separator and part.separator then
return part.separator .. result
else
return result
end
end
local function canonicalize_script_code(scode)
-- Convert 'mnc-Mong', 'xwo-Mong' etc. to 'Mong'.
return (scode:gsub("^.*%-", ""))
end
-----------------------------------------------------------------------------------------
-- Affix-handling functions --
-----------------------------------------------------------------------------------------
-- Figure out the appropriate script for the given affix and language (unless the script is explicitly passed in), and
-- return the values of template_hyphens[], display_hyphens[] and lookup_hyphens[] for that script, substituting
-- default values as appropriate. Four values are returned:
-- DETECTED_SCRIPT, TEMPLATE_HYPHEN, DISPLAY_HYPHEN, LOOKUP_HYPHEN
local function detect_script_and_hyphens(text, lang, sc)
local scode
-- 1. If the script is explicitly passed in, use it.
if sc then
scode = sc:getCode()
else
local possible_script_codes = lang:getScriptCodes()
-- YUCK! `possible_script_codes` comes from loadData() so #possible_scripts doesn't work (always returns 0).
local num_possible_script_codes = m_table.length(possible_script_codes)
if num_possible_script_codes == 0 then
-- This shouldn't happen; if the language has no script codes,
-- the list {"None"} should be returned.
error("Something is majorly wrong! Language " .. lang:getCanonicalName() .. " has no script codes.")
end
if num_possible_script_codes == 1 then
-- 2. If the language has only one possible script, use it.
scode = possible_script_codes[1]
else
-- 3. Check if any of the possible scripts for the language have non-default values for template_hyphens[]
-- or display_hyphens[]. If so, we need to do script detection on the text. If not, just use "Latn",
-- which may not be technically correct but produces the right results because Latn has all default
-- values for template_hyphens[] and display_hyphens[].
local may_have_nondefault_hyphen = false
for _, script_code in ipairs(possible_script_codes) do
script_code = canonicalize_script_code(script_code)
if template_hyphens[script_code] or display_hyphens[script_code] then
may_have_nondefault_hyphen = true
break
end
end
if not may_have_nondefault_hyphen then
scode = "Latn"
else
scode = lang:findBestScript(text):getCode()
end
end
end
scode = canonicalize_script_code(scode)
local template_hyphen = template_hyphens[scode] or "-"
local lookup_hyphen = lookup_hyphens[scode] or "-"
local display_hyphen = display_hyphens[scode] or default_display_hyphen
return scode, template_hyphen, display_hyphen, lookup_hyphen
end
--[=[
Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to
the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string,
specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen,
or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that
matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only
the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix).
]=]
local function reconstruct_term_per_hyphens(term, affix_type, scode, thyph_re, new_hyphen)
local function get_hyphen(hyph)
if type(new_hyphen) == "string" then
return new_hyphen
end
return new_hyphen(scode, hyph)
end
if affix_type == "non-affix" then
return term
elseif affix_type == "apitan" then
local before, before_hyphen, after_hyphen, after = rmatch(term, "^(.*)" .. thyph_re .. " " .. thyph_re
.. "(.*)$")
if not before or ulen(term) <= 3 then
-- Unlike with other types of affixes, don't try to add hyphens in the middle of the term to convert it to
-- a circumfix. Also, if the term is just hyphen + space + hyphen, return it.
return term
end
return before .. get_hyphen(before_hyphen) .. " " .. get_hyphen(after_hyphen) .. after
elseif affix_type == "sisipan" or affix_type == "jalinan" then
local before_hyphen, middle, after_hyphen = rmatch(term, "^" .. thyph_re .. "(.*)" .. thyph_re .. "$")
if before_hyphen and ulen(term) <= 1 then
-- If the term is just a hyphen, return it.
return term
end
return get_hyphen(before_hyphen) .. (middle or term) .. get_hyphen(after_hyphen)
elseif affix_type == "awalan" then
local middle, after_hyphen = rmatch(term, "^(.*)" .. thyph_re .. "$")
if middle and ulen(term) <= 1 then
-- If the term is just a hyphen, return it.
return term
end
return (middle or term) .. get_hyphen(after_hyphen)
elseif affix_type == "akhiran" then
local before_hyphen, middle = rmatch(term, "^" .. thyph_re .. "(.*)$")
if before_hyphen and ulen(term) <= 1 then
-- If the term is just a hyphen, return it.
return term
end
return get_hyphen(before_hyphen) .. (middle or term)
else
error(("Internal error: Unrecognized affix type '%s'"):format(affix_type))
end
end
--[=[
Look up a mapping from a given affix variant to the canonical form used in categories and links. The lookup tables are
language-specific according to `lang`, and may be ID-specific according to `affix_id`. The affixes as they appear in the
lookup tables (both the variant and the canonical form) are in "lookup affix" format (approximately speaking, they use a
regular hyphen for most scripts, but a tatweel for Arabic-script entries and a maqqef for Hebrew-script entries), but
the passed-in `affix` param is in "template affix" format (which differs from the lookup affix for Arabic-script
entries, because more types of hyphens are allowed in template affixes; see the comments at the top of the file). The
remaining parameters to this function are used to convert from template affixes to lookup affixes; see the
reconstruct_term_per_hyphens() function above.
If the affix contains brackets, no lookup is done. Otherwise, a two-stage process is used, first looking up the affix
directly and then stripping diacritics and looking it up again. The reason for this is documented above in the comments
at the top of the file (specifically, the comments describing lookup affixes).
The value of a mapping can either be a string (do the mapping regardless of affix ID) or a table indexed by affix ID
(where the special value `false` indicates no affix ID). The values of entries in this table can also be strings, or
tables with keys `affix` and `id` (again, use `false` to indicate no ID). This allows an affix mapping to map from one
ID to another (for example, this is used in English to map the [[an-]] prefix with no ID to the [[a-]] prefix with the
ID 'not').
The Given a template affix `term` and an affix type `affix_type`, change the relevant template hyphen(s) in the affix to
the display or lookup hyphen specified in `new_hyphen`, or add them if they are missing. `new_hyphen` can be a string,
specifying a fixed hyphen, or a function of two arguments (the script code `scode` and the discovered template hyphen,
or nil of no relevant template hyphen is present). `thyph_re` is a Lua pattern (which must be enclosed in parens) that
matches the possible template hyphens. Note that not all template hyphens present in the affix are changed, but only
the "relevant" ones (e.g. for a prefix, a relevant template hyphen is one coming at the end of the affix).
]=]
local function lookup_affix_mapping(affix, affix_type, lang, scode, thyph_re, lookup_hyph, affix_id)
local function do_lookup(afx)
-- Ensure that the affix uses lookup hyphens regardless of whether it used a different type of hyphens before
-- or no hyphens.
local lookup_affix = reconstruct_term_per_hyphens(afx, affix_type, scode, thyph_re, lookup_hyph)
local function do_lookup_for_langcode(langcode)
if export.langs_with_lang_specific_data[langcode] then
local langdata = mw.loadData(export.affix_lang_data_module_prefix .. langcode)
if langdata.affix_mappings then
local mapping = langdata.affix_mappings[lookup_affix]
if mapping then
if type(mapping) == "table" then
mapping = mapping[affix_id] or mapping.default or mapping[affix_id or false]
if mapping then
return mapping
end
else
return mapping
end
end
end
end
end
-- If `lang` is an etymology-only language, look for a mapping both for it and its full parent.
local langcode = lang:getCode()
local mapping = do_lookup_for_langcode(langcode)
if mapping then
return mapping
end
local full_langcode = lang:getFullCode()
if full_langcode ~= langcode then
mapping = do_lookup_for_langcode(full_langcode)
if mapping then
return mapping
end
end
return nil
end
if affix:find("%[%[") then
return nil
end
return do_lookup(affix) or do_lookup(lang:stripDiacritics(affix)) or nil
end
--[==[
For a given template term in a given language (see the definition of "template affix" near the top of the file),
possibly in an explicitly specified script `sc` (but usually nil), return the term's affix type ({"prefix"},
{"interfix"}, {"suffix"}, {"circumfix"} or {"non-affix"}) along with the corresponding link and display affixes
(see definitions near the top of the file); also the corresponding lookup affix (if `return_lookup_affix` is specified).
The term passed in should already have any fragment (after the # sign) parsed off of it. Four values are returned:
`affix_type`, `link_term`, `display_term` and `lookup_term`. The affix type can be passed in instead of autodetected; in
this case, the template term need not have any attached hyphens, and the appropriate hyphens will be added in the
appropriate places. If `do_affix_mapping` is specified, look up the affix in the lang-specific affix mappings, as
described in the comment at the top of the file; otherwise, the link and display terms will always be the same. (They
will be the same in any case if the template term has a bracketed link in it or is not an affix.) If
`return_lookup_affix` is given, the fourth return value contains the term with appropriate lookup hyphens in the
appropriate places; otherwise, it is the same as the display term. (This functionality is used in
[[Module:category tree/affixes and compounds]] to convert link affixes into lookup affixes so that they can be looked up
in the affix mapping tables.)
Exported because used by [[Module:headword utilities]] to determine the affix type of a given pagename.
]==]
function export.parse_term_for_affixes(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id)
if not term then
return "non-affix", nil, nil, nil
end
if term == "^" then
-- Indicates a null term to emulate the behavior of {{suffix|foo||bar}}.
term = ""
return "non-affix", term, term, term
end
if term:find("^%^") then
-- HACK! ^ at the beginning of Korean languages has a special meaning, triggering capitalization of the
-- transliteration. Don't interpret it as "force non-affix" for those languages.
local langcode = lang:getCode()
if langcode ~= "ko" and langcode ~= "okm" and langcode ~= "jje" then
-- Formerly we allowed ^ to force non-affix type; this is now handled using an inline modifier
-- <naf>, <root>, etc. Throw an error for the moment when the old way is encountered.
error("Use of ^ to force non-affix status is no longer supported; use an inline modifier <naf> or <root> " ..
"after the component")
end
end
-- Remove an asterisk if the morpheme is reconstructed and add it back at the end.
local reconstructed = ""
if term:find("^%*") then
reconstructed = "*"
term = term:gsub("^%*", "")
end
local scode, thyph, dhyph, lhyph = detect_script_and_hyphens(term, lang, sc)
thyph = "([" .. thyph .. "])"
if not affix_type then
if rfind(term, thyph .. " " .. thyph) then
affix_type = "apitan"
else
local has_beginning_hyphen = rfind(term, "^" .. thyph)
local has_ending_hyphen = rfind(term, thyph .. "$")
if has_beginning_hyphen and has_ending_hyphen then
affix_type = "jalinan"
elseif has_ending_hyphen then
affix_type = "awalan"
elseif has_beginning_hyphen then
affix_type = "akhiran"
else
affix_type = "non-affix"
end
end
end
local link_term, display_term, lookup_term
if affix_type == "non-affix" then
link_term = term
display_term = term
lookup_term = term
else
display_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, dhyph)
if do_affix_mapping then
link_term = lookup_affix_mapping(term, affix_type, lang, scode, thyph, lhyph, affix_id)
-- The return value of lookup_affix_mapping() may be an affix mapping with lookup hyphens if a mapping
-- was found, otherwise nil if a mapping was not found. We need to convert to display hyphens in
-- either case, but in the latter case we can reuse the display term, which has already been converted.
if link_term then
link_term = reconstruct_term_per_hyphens(link_term, affix_type, scode, thyph, dhyph)
else
link_term = display_term
end
else
link_term = display_term
end
if return_lookup_affix then
lookup_term = reconstruct_term_per_hyphens(term, affix_type, scode, thyph, lhyph)
else
lookup_term = display_term
end
end
link_term = reconstructed .. link_term
display_term = reconstructed .. display_term
lookup_term = reconstructed .. lookup_term
return affix_type, link_term, display_term, lookup_term
end
--[==[
Add a hyphen to a term in the appropriate place, based on the specified affix type, stripping off any existing hyphens
in that place. For example, if `affix_type` == {"prefix"}, we'll add a hyphen onto the end if it's not already there (or
is of the wrong type). Three values are returned: the link term, display term and lookup term. This function is a thin
wrapper around `parse_term_for_affixes`; see the comments above that function for more information. Note that this
function is exposed externally because it is called by [[Module:category tree/affixes and compounds]]; see the comment
in `parse_term_for_affixes` for more information.
]==]
function export.make_affix(term, lang, sc, affix_type, do_affix_mapping, return_lookup_affix, affix_id)
if not (affix_type == "awalan" or affix_type == "akhiran" or affix_type == "apitan" or affix_type == "sisipan" or
affix_type == "jalinan" or affix_type == "non-affix") then
error("Internal error: Invalid affix type " .. (affix_type or "(nil)"))
end
local _, link_term, display_term, lookup_term = export.parse_term_for_affixes(term, lang, sc, affix_type,
do_affix_mapping, return_lookup_affix, affix_id)
return link_term, display_term, lookup_term
end
-----------------------------------------------------------------------------------------
-- Main entry points --
-----------------------------------------------------------------------------------------
--[==[
Core categorization logic for affixes. This is shared between show_affix(), show_compound_like() and
get_affix_categories_only(). Returns the categories array and other metadata needed for formatting.
]==]
local function generate_affix_categories(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
local text_sections, categories, borrowing_type =
process_etymology_type(data.type, data.surface_analysis or data.nocap, data.notext, #data.parts > 0)
data.borrowing_type = borrowing_type
-- Process each part
local whole_words = 0
local is_affix_or_compound = false
-- Canonicalize and generate links for all the parts first; then do categorization in a separate step, because when
-- processing the first part for categorization, we may access the second part and need it already canonicalized.
for i, part in ipairs_with_gaps(data.parts) do
part = part or {}
data.parts[i] = part
canonicalize_part(part, data.lang, data.sc)
-- Determine affix type and get link and display terms (see text at top of file). Store them in the part
-- (in fields that won't clash with fields used by full_link() in [[Module:links]] or link_term()), so they
-- can be used in the loop below when categorizing.
part.affix_type, part.affix_link_term, part.affix_display_term = export.parse_term_for_affixes(part.term,
part.lang, part.sc, part.type, not part.alt, nil, part.id)
-- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with inline
-- modifiers. The intention in either case is not to link the term.
part.term = ine(part.affix_link_term)
-- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being
-- redundant alt text.
part.alt = part.alt or (part.affix_display_term ~= part.affix_link_term and part.affix_display_term) or nil
end
if not data.noaffixcat then
-- Now do categorization.
for i, part in ipairs_with_gaps(data.parts) do
local affix_type = part.affix_type
if affix_type ~= "non-affix" then
is_affix_or_compound = true
-- Make a sort key. For the first part, use the second part as the sort key; the intention is that if the
-- term has a prefix, sorting by the prefix won't be very useful so we sort by what follows, which is
-- presumably the root.
local part_sort_base = nil
local part_sort = part.sort or data.sort_key
if i == 1 and data.parts[2] and data.parts[2].term then
local part2 = data.parts[2]
-- If the second-part link term is empty, the user requested an unlinked term; avoid a wikitext error
-- by using the alt value if available.
part_sort_base = ine(part2.affix_link_term) or ine(part2.alt)
if part_sort_base then
part_sort_base = strip_diacritics_no_links(part2.lang, part_sort_base)
end
end
if part.pos and rfind(part.pos, "patronym") then
table.insert(categories, {cat = "Patronim", sort_key = part_sort, sort_base = part_sort_base})
end
if data.pos ~= "perkataan" and part.pos and rfind(part.pos, "diminutive") then
table.insert(categories, {cat = ucfirst(data.pos) .. " diminutif", sort_key = part_sort,
sort_base = part_sort_base})
end
-- Don't add a '*fixed with' category if the link term is empty or is in a different language.
if ine(part.affix_link_term) and not part.part_lang then
table.insert(categories, {cat = ucfirst(data.pos) .. " dengan " .. affix_type .. " " ..
strip_diacritics_no_links(part.lang, part.affix_link_term) ..
(part.id and " (" .. part.id .. ")" or ""),
sort_key = part_sort, sort_base = part_sort_base})
end
else
whole_words = whole_words + 1
if whole_words == 2 then
is_affix_or_compound = true
table.insert(categories, ucfirst(data.pos) .. " majmuk")
end
end
end
-- Make sure there was either an affix or a compound (two or more non-affix terms).
if not is_affix_or_compound and not data.allow_no_affixes_or_compounds then
error("The parameters did not include any affixes, and the term is not a compound. Please provide at least one affix.")
end
end
return text_sections, categories, borrowing_type
end
--[==[
Implementation of {{tl|affix}} and {{tl|surface analysis}}. `data` contains all the information describing the affixes to
be displayed, and contains the following:
* `.lang` ('''required'''): Overall language object. Different from term-specific language objects (see `.parts` below).
* `.sc`: Overall script object (usually omitted). Different from term-specific script objects.
* `.parts` ('''required'''): List of objects describing the affixes to show. The general format of each object is as would
be passed to `full_link()`, except that the `.lang` field should be missing unless the term is of a language
different from the overall `.lang` value (in such a case, the language name is shown along with the term and
an additional "derived from" category is added). '''WARNING''': The data in `.parts` will be destructively
modified.
* `.pos`: Overall part of speech (used in categories, defaults to {"terms"}). Different from term-specific part of speech.
* `.sort_key`: Overall sort key. Normally omitted except e.g. in Japanese.
* `.type`: Type of compound, if the parts in `.parts` describe a compound. Strictly optional, and if supplied, the
compound type is displayed before the parts (normally capitalized, unless `.nocap` is given).
* `.nocap`: Don't capitalize the first letter of text displayed before the parts (relevant only if `.type` or
`.surface_analysis` is given).
* `.notext`: Don't display any text before the parts (relevant only if `.type` or `.surface_analysis` is given).
* `.nocat`: Disable all categorization.
* `.noaffixcat`: Disable affix (and compound) categorization. Relevant for e.g. blends, which may otherwise
be incorrectly categorized as compound terms.
* `.lit`: Overall literal definition. Different from term-specific literal definitions.
* `.force_cat`: Always display categories, even on userspace pages.
* `.surface_analysis`: Implement {{surface analysis}}; adds `By surface analysis, ` before the parts.
'''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`.
]==]
function export.show_affix(data)
local text_sections, categories, _ = generate_affix_categories(data)
-- Process each part for display
local parts_formatted = {}
for i, part in ipairs_with_gaps(data.parts) do
-- Make a link for the part
table.insert(parts_formatted, export.link_term(part, data, "include_separator"))
end
if data.surface_analysis then
local text = "mengikut " .. glossary_link("analisis permukaan") .. ", "
if not data.nocap then
text = ucfirst(text)
end
table.insert(text_sections, 1, text)
end
table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted,
categories = categories, separator_already_added = true })
return table.concat(text_sections)
end
--[==[
Get only the categories that would be generated by show_affix(), without any text output or formatting.
This is used by Module:etymon to get affix categorization.
Returns an array of category objects, where
each entry is either a string (simple category name) or a table with keys `cat`, `sort_key`,
and `sort_base` for more complex categorization.
`data` should have the same structure as passed to show_affix():
* `.lang` (required): Overall language object
* `.parts` (required): Array of affix part objects with `.term`, `.lang`, `.id`, etc.
* `.pos`: Part of speech (defaults to "terms")
* `.sort_key`: Overall sort key for categories
'''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`.
]==]
function export.get_affix_categories_only(data)
local _, categories, _ = generate_affix_categories(data)
return categories
end
function export.show_surface_analysis(data)
data.surface_analysis = true
data.allow_no_affixes_or_compounds = true
return export.show_affix(data)
end
--[==[
Implementation of {{tl|compound}}.
'''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`.
]==]
function export.show_compound(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
local text_sections, categories, borrowing_type =
process_etymology_type(data.type, data.nocap, data.notext, #data.parts > 0)
data.borrowing_type = borrowing_type
local parts_formatted = {}
table.insert(categories, ucfirst(data.pos) .. " majmuk")
-- Make links out of all the parts
local whole_words = 0
for i, part in ipairs(data.parts) do
canonicalize_part(part, data.lang, data.sc)
-- Determine affix type and get link and display terms (see text at top of file).
local affix_type, link_term, display_term = export.parse_term_for_affixes(part.term, part.lang, part.sc,
part.type, not part.alt, nil, part.id)
-- If the term is an interfix or the type was explicitly given, recognize it as such (which means e.g. that we
-- will display the term without hyphens for East Asian languages). Otherwise, ignore the fact that it looks
-- like an affix and display as specified in the template (but pay attention to the detected affix type for
-- certain tracking purposes).
if affix_type == "jalinan" or (part.type and part.type ~= "non-affix") then
-- If link_term is an empty string, either a bare ^ was specified or an empty term was used along with
-- inline modifiers. The intention in either case is not to link the term. Don't add a '*fixed with'
-- category in this case, or if the term is in a different language.
-- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being
-- redundant alt text.
if link_term and link_term ~= "" and not part.part_lang then
table.insert(categories, {cat = ucfirst(data.pos) .. " dengan " .. affix_type .. " " ..
strip_diacritics_no_links(part.lang, link_term), sort_key = part.sort or data.sort_key})
end
part.term = link_term ~= "" and link_term or nil
part.alt = part.alt or (display_term ~= link_term and display_term) or nil
else
if affix_type ~= "non-affix" then
local langcode = data.lang:getCode()
-- If `data.lang` is an etymology-only language, track both using its code and its full parent's code.
track { affix_type, affix_type .. "/lang/" .. langcode }
local full_langcode = data.lang:getFullCode()
if langcode ~= full_langcode then
track(affix_type .. "/lang/" .. full_langcode)
end
else
whole_words = whole_words + 1
end
end
table.insert(parts_formatted, export.link_term(part, data, "include_separator"))
end
if whole_words == 1 then
track("one whole word")
elseif whole_words == 0 then
track("looks like confix")
end
table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted,
categories = categories, separator_already_added = true })
return table.concat(text_sections)
end
--[==[
Implementation of {{tl|blend}}, {{tl|univerbation}} and similar "compound-like" templates.
'''WARNING''': This destructively modifies both `data` and the individual structures within `.parts`.
]==]
function export.show_compound_like(data)
data.allow_no_affixes_or_compounds = true
local text_sections, categories, _ = generate_affix_categories(data)
if data.cat then
table.insert(categories, data.cat)
end
-- Process each part for display
local parts_formatted = {}
for i, part in ipairs_with_gaps(data.parts) do
-- Make a link for the part
table.insert(parts_formatted, export.link_term(part, data, "include_separator"))
end
if #data.parts > 0 and data.oftext then
table.insert(text_sections, 1, " " .. data.oftext .. " ")
end
if data.text then
table.insert(text_sections, 1, data.text)
end
table.insert(text_sections, export.join_formatted_parts { data = data, parts_formatted = parts_formatted,
categories = categories, separator_already_added = true })
return table.concat(text_sections)
end
--[==[
Make `part` (a structure holding information on an affix part) into an affix of type `affix_type`, and apply any
relevant affix mappings. For example, if the desired affix type is "suffix", this will (in general) add a hyphen onto
the beginning of the term, alt, tr and ts components of the part if not already present. The hyphen that's added is the
"display hyphen" (see above) and may be script-specific. (In the case of East Asian scripts, the display hyphen is an
empty string whereas the template hyphen is the regular hyphen, meaning that any regular hyphen at the beginning of the
part will be effectively removed.) `lang` and `sc` hold overall language and script objects.
Note that this also applies any language-specific affix mappings, so that e.g. if the language is Finnish and the user
specified [[-käs]] in the affix and didn't specify an `.alt` value, `part.term` will contain [[-kas]] and `part.alt` will
contain [[-käs]].
This function is used by the "legacy" templates ({{tl|prefix}}, {{tl|suffix}}, {{tl|confix}}, etc.) where the nature of
the affix is specified by the template itself rather than auto-determined from the affix, as is the case with
{{tl|affix}}.
'''WARNING''': This destructively modifies `part`.
]==]
local function make_part_into_affix(part, lang, sc, affix_type)
canonicalize_part(part, lang, sc)
local link_term, display_term = export.make_affix(part.term, part.lang, part.sc, affix_type, not part.alt, nil, part.id)
part.term = link_term
-- When we don't specify `do_affix_mapping` to make_affix(), link and display terms (first and second retvals of
-- make_affix()) are the same.
-- If part.alt would be the same as part.term, make it nil, so that it isn't erroneously tracked as being
-- redundant alt text.
part.alt = part.alt and export.make_affix(part.alt, part.lang, part.sc, affix_type) or (display_term ~= link_term and display_term) or nil
local Latn = require(scripts_module).getByCode("Latn")
part.tr = export.make_affix(part.tr, part.lang, Latn, affix_type)
part.ts = export.make_affix(part.ts, part.lang, Latn, affix_type)
end
local function track_wrong_affix_type(template, part, expected_affix_type)
if part and not part.type then
local affix_type = export.parse_term_for_affixes(part.term, part.lang, part.sc)
if affix_type ~= expected_affix_type then
local part_name = expected_affix_type or "base"
local langcode = part.lang:getCode()
local full_langcode = part.lang:getFullCode()
require("Module:debug/track") {
template,
template .. "/" .. part_name,
template .. "/" .. part_name .. "/" .. (affix_type or "none"),
template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. langcode
}
-- If `part.lang` is an etymology-only language, track both using its code and its full parent's code.
if full_langcode ~= langcode then
require("Module:debug/track")(
template .. "/" .. part_name .. "/" .. (affix_type or "none") .. "/lang/" .. full_langcode
)
end
end
end
end
local function insert_affix_category(categories, pos, affix_type, part, sort_key, sort_base)
-- Don't add a '*fixed with' category if the link term is empty or is in a different language.
if part.term and not part.part_lang then
local cat = ucfirst(pos) .. " dengan " .. affix_type .. " " .. strip_diacritics_no_links(part.lang, part.term) ..
(part.id and " (" .. part.id .. ")" or "")
if sort_key or sort_base then
table.insert(categories, {cat = cat, sort_key = sort_key, sort_base = sort_base})
else
table.insert(categories, cat)
end
end
end
--[==[
Implementation of {{tl|circumfix}}.
'''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`.
]==]
function export.show_circumfix(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
canonicalize_part(data.base, data.lang, data.sc)
-- Hyphenate the affixes and apply any affix mappings.
make_part_into_affix(data.prefix, data.lang, data.sc, "awalan")
make_part_into_affix(data.suffix, data.lang, data.sc, "akhiran")
track_wrong_affix_type("apitan", data.prefix, "awalan")
track_wrong_affix_type("apitan", data.base, nil)
track_wrong_affix_type("apitan", data.suffix, "akhiran")
-- Create circumfix term.
local circumfix = nil
if data.prefix.term and data.suffix.term then
circumfix = data.prefix.term .. " " .. data.suffix.term
data.prefix.alt = data.prefix.alt or data.prefix.term
data.suffix.alt = data.suffix.alt or data.suffix.term
data.prefix.term = circumfix
data.suffix.term = circumfix
end
-- Make links out of all the parts.
local parts_formatted = {}
local categories = {}
local sort_base
if data.base.term then
sort_base = strip_diacritics_no_links(data.base.lang, data.base.term)
end
table.insert(parts_formatted, export.link_term(data.prefix, data))
table.insert(parts_formatted, export.link_term(data.base, data))
table.insert(parts_formatted, export.link_term(data.suffix, data))
-- Insert the categories, but don't add a '*fixed with' category if the link term is in a different language.
if not data.prefix.part_lang then
table.insert(categories, {cat=ucfirst(data.pos) .. " dengan apitan " .. strip_diacritics_no_links(data.prefix.lang,
circumfix), sort_key=data.sort_key, sort_base=sort_base})
end
return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories }
end
--[==[
Implementation of {{tl|confix}}.
'''WARNING''': This destructively modifies both `data` and `.prefix`, `.base` and `.suffix`.
]==]
function export.show_confix(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
canonicalize_part(data.base, data.lang, data.sc)
-- Hyphenate the affixes and apply any affix mappings.
make_part_into_affix(data.prefix, data.lang, data.sc, "awalan")
make_part_into_affix(data.suffix, data.lang, data.sc, "akhiran")
track_wrong_affix_type("confix", data.prefix, "awalan")
track_wrong_affix_type("confix", data.base, nil)
track_wrong_affix_type("confix", data.suffix, "akhiran")
-- Make links out of all the parts.
local parts_formatted = {}
local prefix_sort_base
if data.base and data.base.term then
prefix_sort_base = strip_diacritics_no_links(data.base.lang, data.base.term)
elseif data.suffix.term then
prefix_sort_base = strip_diacritics_no_links(data.suffix.lang, data.suffix.term)
end
-- Insert the categories and parts.
local categories = {}
table.insert(parts_formatted, export.link_term(data.prefix, data))
insert_affix_category(categories, data.pos, "awalan", data.prefix, data.sort_key, prefix_sort_base)
if data.base then
table.insert(parts_formatted, export.link_term(data.base, data))
end
table.insert(parts_formatted, export.link_term(data.suffix, data))
-- FIXME, should we be specifying a sort base here?
insert_affix_category(categories, data.pos, "akhiran", data.suffix)
return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories }
end
--[==[
Implementation of {{tl|infix}}.
'''WARNING''': This destructively modifies both `data` and `.base` and `.infix`.
]==]
function export.show_infix(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
canonicalize_part(data.base, data.lang, data.sc)
-- Hyphenate the affixes and apply any affix mappings.
make_part_into_affix(data.infix, data.lang, data.sc, "sisipan")
track_wrong_affix_type("sisipan", data.base, nil)
track_wrong_affix_type("sisipan", data.infix, "sisipan")
-- Make links out of all the parts.
local parts_formatted = {}
local categories = {}
table.insert(parts_formatted, export.link_term(data.base, data))
table.insert(parts_formatted, export.link_term(data.infix, data))
-- Insert the categories.
-- FIXME, should we be specifying a sort base here?
insert_affix_category(categories, data.pos, "sisipan", data.infix)
return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories }
end
--[==[
Implementation of {{tl|prefix}}.
'''WARNING''': This destructively modifies both `data` and the structures within `.prefixes`, as well as `.base`.
]==]
function export.show_prefix(data)
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
canonicalize_part(data.base, data.lang, data.sc)
-- Hyphenate the affixes and apply any affix mappings.
for i, prefix in ipairs(data.prefixes) do
make_part_into_affix(prefix, data.lang, data.sc, "awalan")
end
for i, prefix in ipairs(data.prefixes) do
track_wrong_affix_type("awalan", prefix, "awalan")
end
track_wrong_affix_type("awalan", data.base, nil)
-- Make links out of all the parts.
local parts_formatted = {}
local first_sort_base = nil
local categories = {}
if data.prefixes[2] then
first_sort_base = ine(data.prefixes[2].term) or ine(data.prefixes[2].alt)
if first_sort_base then
first_sort_base = strip_diacritics_no_links(data.prefixes[2].lang, first_sort_base)
end
elseif data.base then
first_sort_base = ine(data.base.term) or ine(data.base.alt)
if first_sort_base then
first_sort_base = strip_diacritics_no_links(data.base.lang, first_sort_base)
end
end
for i, prefix in ipairs(data.prefixes) do
table.insert(parts_formatted, export.link_term(prefix, data))
insert_affix_category(categories, data.pos, "awalan", prefix, data.sort_key, i == 1 and first_sort_base or nil)
end
if data.base then
table.insert(parts_formatted, export.link_term(data.base, data))
else
table.insert(parts_formatted, "")
end
return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories }
end
--[==[
Implementation of {{tl|suffix}}.
'''WARNING''': This destructively modifies both `data` and the structures within `.suffixes`, as well as `.base`.
]==]
function export.show_suffix(data)
local categories = {}
data.pos = data.pos or default_pos
data.pos = pluralize_pos(data.pos)
canonicalize_part(data.base, data.lang, data.sc)
-- Hyphenate the affixes and apply any affix mappings.
for i, suffix in ipairs(data.suffixes) do
make_part_into_affix(suffix, data.lang, data.sc, "akhiran")
end
track_wrong_affix_type("akhiran", data.base, nil)
for i, suffix in ipairs(data.suffixes) do
track_wrong_affix_type("akhiran", suffix, "akhiran")
end
-- Make links out of all the parts.
local parts_formatted = {}
if data.base then
table.insert(parts_formatted, export.link_term(data.base, data))
else
table.insert(parts_formatted, "")
end
for i, suffix in ipairs(data.suffixes) do
table.insert(parts_formatted, export.link_term(suffix, data))
end
-- Insert the categories.
for i, suffix in ipairs(data.suffixes) do
-- FIXME, should we be specifying a sort base here?
insert_affix_category(categories, data.pos, "akhiran", suffix)
if suffix.pos and rfind(suffix.pos, "patronym") then
table.insert(categories, "Patronim")
end
end
return export.join_formatted_parts { data = data, parts_formatted = parts_formatted, categories = categories }
end
return export
00ctt3fkh5mg0bo9nx60necwwv3han4
Modul:homophones
828
10482
375351
227296
2026-09-22T03:12:54Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708691|92708691]])
375351
Scribunto
text/plain
local export = {}
local decorations_module = "Module:decorations"
local links_module = "Module:links"
local parameter_utilities_module = "Module:parameter utilities"
local function track(page)
return require("Module:debug/track")("homophones/" .. page)
end
--[==[
Meant to be called from a module. `data` is a table containing the following fields:
* `lang`: language object for the homophones;
* `homophones`: a list of homophones, each described by an object which can contain all the fields in the object
passed to {full_link()} in [[Module:links]] except for `lang` and `sc` (which are copied from the outer level),
including decoration fields;
** `term`: the homophone itself;
** `separator`: {nil} or the string used to separate this homophone from the preceding one when displayed; defaults to
the top-level `separator`;
** `alt`: display text for the homophone, as in {{tl|l}};
** `gloss`: gloss for the homophone, as in {{tl|l}};
** `tr`: transliteration for the homophone, as in {{tl|l}};
** `ts`: transcription for the homophone, as in {{tl|l}};
** `genders`: list of genders for the homophone, as in {{tl|l}};
** `pos`: part of speech of the homophone, as in {{tl|l}};
** `ng`: non-gloss text for the homophone, as in {{tl|l}};
** `lit`: literal meaning of the homophone, as in {{tl|l}};
** `id`: sense ID for the homophone, as in {{tl|l}};
** `lang`: optional lang code, overriding the lang code in the top-level `lang` field;
** `sc`: optional script code, overriding the script code in the top-level `sc` field;
** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the homophone in question;
** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the homophone in question;
** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]) and displayed directly
before the homophone in question;
** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the homophone in question;
** `refs`: {nil} or a list of references or reference specs to add after the pronunciation and any posttext and
qualifiers; the value of a list item is either a string containing the reference text (typically a call to a citation
template such as {{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the
reference text), `name` (the name of the reference, as in `<nowiki><ref name="foo">...</ref></nowiki>` or
`<nowiki><ref name="foo" /></nowiki>`) and/or `group` (the group of the reference, as in
`<nowiki><ref name="foo" group="bar">...</ref></nowiki>` or `<nowiki><ref name="foo" group="bar"/></nowiki>`);
this uses a parser function to format the reference appropriately and insert a footnote number that hyperlinks to the
actual reference, located in the `<nowiki><references /></nowiki>` section;
* `separator`: {nil} or a string, specifying the separator displayed before all homophones but the first; by default,
{", "}; overridable at the individual homophone level;
* `q`: {nil} or a list of left regular qualifier strings, displayed before the initial caption;
* `qq`: {nil} or a list of right regular qualifier strings, displayed after all homophones;
* `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the initial
caption;
* `aa`: {nil} or a list of right accent qualifier strings, displayed after all homophones;
* `sc`: {nil} or script object for the homophones;
* `sort`: {nil} or sort key;
* `caption`: {nil} or string specifying the caption to use, in place of {"Homophone"} (if there is a single homophone),
or {"Homophones"} (otherwise); a colon and space is automatically added after the caption;
* `nocaption`: If true, suppress the caption display.
* `nocat`: If true, suppress categorization.
If both regular and accent qualifiers on the same side and at the same level are specified, the accent qualifiers
precede the regular qualifiers on both left and right.
'''WARNING''': Destructively modifies the objects inside the `homophones` field.
]==]
function export.format_homophones(data)
local hmptexts = {}
local hmpcats = {}
local m_links = require(links_module)
local overall_sep = data.separator or ", "
for i, hmp in ipairs(data.homophones) do
hmp.lang = hmp.lang or data.lang
hmp.sc = hmp.sc or data.sc
hmp.show_decorations = true -- make full_link() display decorations (`.q`, `.qq`, `.a`, `.aa` and `.refs`)
if hmp.qualifiers then
error("`.qualifiers` is no longer supported; change the code to use `.qq` or `.q`")
end
local text = m_links.full_link(hmp)
table.insert(hmptexts, hmp.separator or i > 1 and overall_sep or "")
table.insert(hmptexts, text)
end
table.insert(hmpcats, "Perkataan dengan homofon bahasa " .. data.lang:getFullName())
local text = table.concat(hmptexts)
local caption = data.nocaption and "" or (
data.caption or "[[Lampiran:Glosari#homofon|Homofon" .. (#data.homophones > 1 and "" or "") .. "]]"
) .. ": "
text = caption .. text
if data.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("overall `.qualifiers` is no longer supported; change the code to use `.qq` or `.q`")
end
if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then
text = require(decorations_module).format_decorations {
lang = data.lang,
text = text,
q = data.q,
qq = data.qq,
a = data.a,
aa = data.aa,
}
end
text = "<span class=\"homophones\">" .. text .. "</span>"
if not data.nocat then
local categories = require("Module:utilities").format_categories(hmpcats, data.lang, data.sort)
text = text .. categories
end
return text
end
--[==[
Entry point for {{tl|homophones}} template (also written {{tl|homophone}} and {{tl|hmp}}).
]==]
function export.show(frame)
local parent_args = frame:getParent().args
local compat = parent_args.lang
local offset = compat and 0 or 1
local lang_arg = compat and "lang" or 1
local params = {
[lang_arg] = {required = true, type = "language", default = "ms"},
[1 + offset] = {list = true, required = true, allow_holes = true, default = "perkataan"},
["caption"] = {},
["nocaption"] = {type = "boolean"},
["nocat"] = {type = "boolean"},
["sort"] = {},
}
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{group = {"link", "ref", "a", "q"}},
}
local homophones, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 1 + offset,
parse_lang_prefix = true,
track_module = "homophones",
lang = lang_arg,
sc = "sc.default",
}
local data = {
lang = args[lang_arg],
homophones = homophones,
caption = args.caption,
nocaption = args.nocaption,
nocat = args.nocat,
sc = args.sc.default,
sort = args.sort,
q = args.q.default,
qq = args.qq.default,
a = args.a.default,
aa = args.aa.default,
}
return export.format_homophones(data)
end
return export
1wlerteggfy10doldx9e5ckzpge6vq4
Modul:ar-headword
828
11149
375386
266015
2026-09-22T07:16:37Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92714469|92714469]])
375386
Scribunto
text/plain
-- Author: primarily Benwing2; some work by Fenakhay, Erutuon; early version by Rua
local export = {}
local pos_functions = {}
local force_cat = false -- for testing; if true, categories appear in non-mainspace pages
local ar_translit = require("Module:ar-translit")
local ar_verb_module = "Module:ar-verb"
local ar_utilities_module = "Module:ar-utilities"
local ar = require(ar_utilities_module)
local en_utilities_module = "Module:en-utilities"
local headword_module = "Module:headword"
local headword_utilities_module = "Module:headword utilities"
local links_module = "Module:links"
local inflection_utilities_module = "Module:inflection utilities"
local parse_utilities_module = "Module:parse utilities"
local require_when_needed = require("Module:utilities/require when needed")
local remove_links = require_when_needed(links_module, "remove_links")
local m_table = require("Module:table")
local m_str_utils = require("Module:string utilities")
local m_en_utilities = require_when_needed(en_utilities_module)
local m_headword_utilities = require_when_needed(headword_utilities_module)
local glossary_link = require_when_needed(headword_utilities_module, "glossary_link")
local boolean_param = {type = "boolean"}
local list_to_set = m_table.listToSet
local rfind = m_str_utils.find
local rmatch = m_str_utils.match
local rsubn = m_str_utils.gsub
local u = m_str_utils.char
local rsplit = m_str_utils.split
local insert = table.insert
local concat = table.concat
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local langcode = "ar"
local lang = require("Module:languages").getByCode(langcode)
local langname = lang:getCanonicalName()
local TEMPCOMMA = u(0xFFF0)
local TEMPARCOMMA = u(0xFFF1)
local misc_pos_with_gender = list_to_set {
"suffixes",
"adjective forms",
"noun forms",
"proper noun forms",
"pronoun forms",
"determiner forms",
}
-----------------------------------------------------------------------------------------
-- Utility functions --
-----------------------------------------------------------------------------------------
local dump = mw.dumpObject
-- version of mw.ustring.gsub() that discards all but the first return value
local function rsub(term, foo, bar)
local retval = rsubn(term, foo, bar)
return retval
end
local function ine(val)
if val == "" then return nil else return val end
end
-- Replace comma with a temporary char in comma + whitespace.
local function escape_comma_whitespace(run)
local escaped = false
if run:find("\\,") then
run = run:gsub("\\,", "\\" .. TEMPCOMMA)
escaped = true
end
if run:find("\\،") then
run = run:gsub("\\،", "\\" .. TEMPARCOMMA)
escaped = true
end
if run:find(",%s") then
run = run:gsub(",(%s)", TEMPCOMMA .. "%1")
escaped = true
end
if run:find("،%s") then
run = run:gsub("،(%s)", TEMPARCOMMA .. "%1")
escaped = true
end
return run, escaped
end
-- Undo replacement of comma with a temporary char in comma + whitespace.
local function unescape_comma_whitespace(run)
return (run:gsub(TEMPCOMMA, ","):gsub(TEMPARCOMMA, "،"))
end
-- Split an argument on comma or Arabic comma, but not either type of comma followed by whitespace.
local function split_on_comma(val)
if rfind(val, "[,،]%s") or val:find("\\") then
return export.split_escaping(val, "[,،]", false, escape_comma_whitespace, unescape_comma_whitespace)
else
return rsplit(val, "[,،]")
end
end
local function replace_tr_ending(tr, from, to)
if not tr then
return nil
end
local pref = tr:match("^(.*)" .. from .. "$")
if not pref then
error(("Translit '%s' does not end in -%s, as expected"):format(tr, from))
end
return pref .. to
end
-----------------------------------------------------------------------------------------
-- Tracking functions --
-----------------------------------------------------------------------------------------
local trackfn = require("Module:debug/track")
local function track(page)
trackfn(langcode .. "-headword/" .. page)
return true
end
--[==[
Examples of what you can find by looking at what links to the given
pages:
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized]]
all unvocalized pages
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/pl]]
all unvocalized pages where the plural is unvocalized,
whether specified using pl=, pl2=, etc.
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head]]
all unvocalized pages where the head is unvocalized
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/nouns]]
all nouns excluding proper nouns, collective nouns,
singulative nouns where the head is unvocalized
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/proper]]
nouns all proper nouns where the head is unvocalized
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/head/not]]
proper nouns all words that are not proper nouns
where the head is unvocalized
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized/adjectives]]
all adjectives where any parameter is unvocalized;
currently only works for heads,
so equivalent to .../unvocalized/head/adjectives
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-empty-head]]
all pages with an empty head
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-manual-translit]]
all unvocalized pages with manual translit
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-manual-translit/head/nouns]]
all nouns where the head is unvocalized but has manual translit
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/unvocalized-no-translit]]
all unvocalized pages without manual translit
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab]]
all pages with any parameter containing i3rab
of either -un, -u, -a or -i
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-un]]
all pages with any parameter containing an -un i3rab ending
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-un/pl]]
all pages where a form specified using pl=, pl2=, etc.
contains an -un i3rab ending
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab-u/head]]
all pages with a head containing an -u i3rab ending
[[Special:WhatLinksHere/Wiktionary:Tracking/ar-headword/i3rab/head/proper]]
nouns (all proper nouns with a head containing i3rab
of either -un, -u, -a or -i)
In general, the format is one of the following:
Wiktionary:Tracking/ar-headword/FIRSTLEVEL
Wiktionary:Tracking/ar-headword/FIRSTLEVEL/ARGNAME
Wiktionary:Tracking/ar-headword/FIRSTLEVEL/POS
Wiktionary:Tracking/ar-headword/FIRSTLEVEL/ARGNAME/POS
FIRSTLEVEL can be one of "unvocalized", "unvocalized-empty-head" or its
opposite "unvocalized-specified", "unvocalized-manual-translit" or its
opposite "unvocalized-no-translit", "i3rab", "i3rab-un", "i3rab-u",
"i3rab-a", or "i3rab-i".
ARGNAME is either "head" or an argument such as "pl", "f", "cons", etc.
This automatically includes arguments specified as head2=, pl3=, etc.
POS is a part of speech, lowercase and singular, e.g. "kata nama",
"Kata sifat", "Kata nama khas", "collective nouns", etc. or
"not proper noun", which includes all parts of speech but proper nouns.
]==]
local function track_form(argname, form, translit, pos)
form = ar.reorder_shadda(remove_links(form))
function dotrack(page)
track(page)
track(page .. "/" .. argname)
if pos then
track(page .. "/" .. pos)
track(page .. "/" .. argname .. "/" .. pos)
if pos ~= "Kata nama khas" then
track(page .. "/not proper noun")
track(page .. "/" .. argname .. "/not proper noun")
end
end
end
function track_i3rab(arabic, tr)
if rfind(form, arabic .. "$") then
dotrack("i3rab")
dotrack("i3rab-" .. tr)
end
end
track_i3rab(ar.UN, "un")
track_i3rab(ar.U, "u")
track_i3rab(ar.A, "a")
track_i3rab(ar.I, "i")
if form == "" or not (lang:transliterate(form)) then
dotrack("unvocalized")
if form == "" then
dotrack("unvocalized-empty-head")
else
dotrack("unvocalized-specified")
end
if translit then
dotrack("unvocalized-manual-translit")
else
dotrack("unvocalized-no-translit")
end
end
end
-----------------------------------------------------------------------------------------
-- Inflection-parsing functions --
-----------------------------------------------------------------------------------------
-- Construct the default construct state or informal form of a term in lemma format. Usually this is the same as the
-- lemma but is different for final-weak nouns and adjectives ending in -n in their lemma. NOTE: Input must be
-- shadda-reordered for this to work properly.
local function default_construct_state_or_informal(term, tr)
local pref = term:match("^(.*)" .. ar.HAMZA .. ar.IN .."$")
-- Hamza on the line with -in changes to hamza-on-yā with -ī.
if pref then
return pref .. ar.HAMZA_ON_YA .. ar.II, replace_tr_ending(tr, "in", "ī")
end
-- Otherwise just change -in to -ī.
pref = term:match("^(.*)" .. ar.IN .. "$")
if pref then
return pref .. ar.II, replace_tr_ending(tr, "in", "ī")
end
-- Change -an with alif maqṣūra to -ā with alif maqṣūra.
pref = term:match("^(.*)" .. ar.AN .. ar.AMAQ .. "$")
if pref then
return pref .. ar.AAMAQ, replace_tr_ending(tr, "an", "ā")
end
-- Change -an with tall alif (e.g. عَصًا) to -ā with tall alif.
pref = term:match("^(.*)" .. ar.AN .. ar.ALIF .. "$")
if pref then
return pref .. ar.AA, replace_tr_ending(tr, "an", "ā")
end
return term, tr
end
local function generate_construct_state_or_informal_default(data, args)
local heads = data.heads
local consobjs = {}
local different_cons = false
for _, headobj in ipairs(data.heads) do
local consterm, constr = default_construct_state_or_informal(headobj.term, headobj.tr)
different_cons = different_cons or consterm ~= headobj.term or constr ~= headobj.tr
local consobj = m_table.shallowCopy(headobj)
consobj.term = consterm
consobj.tr = constr
insert(consobjs, consobj)
end
if different_cons then
return consobjs
else
return {}
end
end
local noun_field_cons = {
field = "cons", label = "<<construct state>>", generate_default = generate_construct_state_or_informal_default,
default_when_not_explicit = function(args, data) return true end,
}
local noun_field_inf = {field = "inf", label = "informal"}
local noun_field_obl = {field = "obl", label = "<<oblique>>"}
local noun_field_def = {field = "def", label = "<<definite>> state"}
local noun_inflections = {
noun_field_cons,
noun_field_inf,
noun_field_obl,
noun_field_def,
}
local adj_field_inf = {
field = "inf", label = "informal", generate_default = generate_construct_state_or_informal_default,
default_when_not_explicit = function(args, data) return true end,
}
local adj_field_obl = noun_field_obl
local adj_field_def = noun_field_def
local adjective_inflections = {
adj_field_inf,
adj_field_obl,
adj_field_def,
}
local function has_construct_state(data)
return data.pos_category ~= "Kata sifat"
end
local function parse_nominal_inflection(paramname, val, parse_err)
return m_headword_utilities.parse_term_with_modifiers {
val = val,
paramname = paramname,
splitchar = ",",
include_mods = {"tr", "g"},
}
end
local function make_nominal_inflection_param_mod_spec(paramname)
return {convert = function(val, parse_err)
return parse_nominal_inflection(paramname, val, parse_err)
end}
end
-- Parse an inflection. The raw arguments come from `args[field]`, which is parsed for inline modifiers. Multiple
-- comma-separated values are allowed.
local function parse_inflection(data, args, field, is_head)
local argfield = field
local argpref = field
if type(argfield) == "table" then
argpref = argfield[2]
argfield = argfield[1]
end
local include_mods
if is_head then
include_mods = {"tr"}
else
include_mods = {"tr", "g"}
for _, spec in ipairs(has_construct_state(data) and noun_inflections or adjective_inflections) do
insert(include_mods, {spec.field, make_nominal_inflection_param_mod_spec(argpref .. "." .. spec.field)})
end
end
if is_head then
local retval
if args[argfield] then
retval = m_headword_utilities.parse_term_with_modifiers {
val = args[argfield],
paramname = field,
splitchar = ",",
is_head = is_head,
include_mods = include_mods,
}
end
return retval or {}
else
return m_headword_utilities.parse_term_list_with_modifiers {
forms = args[argfield],
paramname = field,
splitchar = ",",
is_head = is_head,
include_mods = include_mods,
}
end
end
local function insert_inflection(data, terms, label, accel, defgender, track_field, no_label, usually_no_label)
local track_pos = m_en_utilities.singularize(data.pos_category)
for _, termobj in ipairs(terms) do
-- If the user supplied a construct state or informal form for the term with a value of "+", substitute the
-- default value for the term. If the user supplied a value of "--", they want no value displayed. Otherwise,
-- if the user didn't supply any value, we check to see if the default construct state or informal form is
-- different from the lemma and display it if so; this applies particularly to terms in '-in' and '-an', where
-- the default construct state or informal form is almost always correct.
local field = has_construct_state(data) and "cons" or "inf"
if not termobj[field] then
local defcons, defconstr = default_construct_state_or_informal(termobj.term, termobj.tr)
if termobj.term ~= defcons or termobj.tr ~= defconstr then
-- We don't want to copy decorations from the term object because we're a subinflection of the term
-- object.
termobj[field] = {{term = defcons, tr = defconstr}}
end
elseif termobj[field][1].term == "--" then
if termobj[field][2] then
error("Can't specify more than one value for <" .. field .. ":...> if first value is '--', meaning \"don't insert anything\"")
end
termobj[field] = nil
else
for i, consobj in ipairs(termobj[field]) do
if consobj.term == "+" then
if consobj.tr then
error("Can't specify translit for default value '+'")
end
consobj.term, consobj.tr = default_construct_state_or_informal(termobj.term, termobj.tr)
elseif consobj.term == "~" then
if consobj.tr then
error("Can't specify translit for term-requesting value '~'")
end
consobj.term, consobj.tr = termobj.term, termobj.tr
end
end
end
if defgender and not termobj.genders then
termobj.genders = {{spec = defgender}}
end
local function insert_nested_inflection(field, label)
if termobj[field] then
m_headword_utilities.insert_inflection {
headdata = data,
inflobj = termobj,
terms = termobj[field],
label = label
}
end
end
for _, spec in ipairs(has_construct_state(data) and noun_inflections or adjective_inflections) do
insert_nested_inflection(spec.field, spec.label)
end
track_form(track_field, termobj.term, termobj.tr, track_pos)
end
m_headword_utilities.insert_inflection {
headdata = data,
terms = terms,
label = label,
accel = accel and {form = accel} or nil,
no_label = no_label,
usually_no_label = usually_no_label,
}
end
-----------------------------------------------------------------------------------------
-- Main entry point --
-----------------------------------------------------------------------------------------
function export.show(frame)
local iparams = {
[1] = true,
}
local iargs = require("Module:parameters").process(frame.args, iparams)
local parargs = frame:getParent().args
local poscat = iargs[1]
local pos_in_1 = not poscat
if pos_in_1 then
poscat = ine(parargs[1]) or
mw.title.getCurrentTitle().fullText == "Template:" .. langcode .. "-head" and "interjection" or
error("Part of speech must be specified in 1=")
poscat = require(headword_module).canonicalize_pos(poscat)
end
local indexing_poscat = pos_in_1 and (misc_pos_with_gender[poscat] and "head_with_gender" or "head") or poscat
local params = {
["suffix"] = boolean_param,
["nosuffix"] = boolean_param,
["id"] = true,
["json"] = boolean_param,
["pagename"] = {}, -- for testing
}
if pos_in_1 then
params[1] = {required = true} -- required but ignored as already processed above
end
local head_is_head = pos_functions[indexing_poscat] and pos_functions[indexing_poscat].head_is_not_1
local headfield = head_is_head and "head" or pos_in_1 and 2 or 1
params[headfield] = head_is_head and true or {default = "+"}
params.head2 = {replaced_by = false, instead = "use multiple comma-separated values in |" .. headfield .. "="}
local tr_replaced_by = {replaced_by = false, instead = "use <tr:...> inline modifier on |" .. headfield .. "="}
params.tr = tr_replaced_by
params.tr2 = tr_replaced_by
if pos_functions[indexing_poscat] then
for key, val in pairs(pos_functions[indexing_poscat].params()) do
params[key] = val
end
end
local parargs = frame:getParent().args
local args = require("Module:parameters").process(parargs, params)
local pagename = args.pagename or mw.loadData("Module:headword/data").pagename
local data = {
lang = lang,
pos_category = poscat,
orig_pos_category = poscat,
categories = {},
heads = {},
genders = {},
inflections = {enable_auto_translit = true},
pagename = pagename,
id = args.id,
sort_key = args.sort,
force_cat_output = force_cat,
-- We expect a head always so the redundant head cat will be inaccurate.
no_redundant_head_cat = true,
}
data.heads = parse_inflection(data, args, headfield, "is_head")
for _, headobj in ipairs(data.heads) do
if headobj.term == "+" then
headobj.term = pagename
end
end
data.is_suffix = false
if args.suffix or (
not args.nosuffix and pagename:find("^%-") and poscat ~= "suffixes" and poscat ~= "suffix forms"
) then
data.is_suffix = true
data.pos_category = "suffixes"
local singular_poscat = m_en_utilities.singularize(poscat)
insert(data.categories, langname .. " " .. singular_poscat .. "-forming suffixes")
insert(data.inflections, {label = singular_poscat .. "-forming suffix"})
end
if pos_functions[indexing_poscat] then
pos_functions[indexing_poscat].func(data, args)
end
-- Do this after calling pos_functions[poscat].func() as it may modify data.heads (as verbs do).
local irreg_translit = false
for _, head in ipairs(data.heads) do
if ar_translit.irregular_translit(head.term, head.tr) then
irreg_translit = true
break
end
end
if irreg_translit then
insert(data.categories, langname .. " terms with irregular pronunciations")
end
if args.json then
return require("Module:JSON").toJSON(data)
end
return require(headword_module).full_headword(data)
end
-----------------------------------------------------------------------------------------
-- Gender handling --
-----------------------------------------------------------------------------------------
local valid_bare_genders = {false, "m", "f", "mf", "mfbysense", "mfequiv"}
local valid_bare_numbers = {false, "d", "p"}
local valid_bare_animacies = {false, "pr", "np"}
local valid_genders = {}
for _, gender in ipairs(valid_bare_genders) do
for _, number in ipairs(valid_bare_numbers) do
for _, animacy in ipairs(valid_bare_animacies) do
local parts = {}
local function ins_part(part)
if part then
insert(parts, part)
end
end
ins_part(gender)
ins_part(number)
ins_part(animacy)
local full_gender = concat(parts, "-")
valid_genders[full_gender == "" and "?" or full_gender] = true
end
end
end
local function is_masc_sg(g)
return g == "m" or g == "m-pr" or g == "m-np"
end
local function is_fem_sg(g)
return g == "f" or g == "f-pr" or g == "f-np"
end
local function is_masc_fem_sg(g)
g = g:gsub("%-pr", ""):gsub("%-np", "")
return g == "mf" or g == "mfequiv" or g == "mfbysense"
end
local function add_gender_params(params, default)
params[2] = {type = "genders", default = default or "?", template_default = "m"}
params["g2"] = {replaced_by = false, instead = "use comma-separated values in |g="}
end
-- Handle gender in params 2=, inserting into `data.genders`. Also, if a lemma, insert categories into `data.categories`
-- if the gender is unexpected for the form of the noun. (Note: If there are multiple genders,
-- [[Module:gender and number]] will automatically insert 'Arabic POS with multiple genders'.)
local function handle_gender(data, args, nonlemma, field)
if not args[field or 2] then
return
end
for _, gspec in ipairs(args[field or 2]) do
if not valid_genders[gspec.spec] then
error("Unrecognized gender: " .. gspec.spec)
end
end
data.genders = args[field or 2]
if nonlemma then
return
end
for _, gspec in ipairs(data.genders) do
local g = gspec.spec
if is_masc_sg(g) or is_fem_sg(g) or is_masc_fem_sg(g) then
local head = data.heads[1]
if head then
head = rsub(ar.reorder_shadda(remove_links(head.term)), ar.UNUOPT .. "$", "")
local ends_with_tam = rfind(head, "^[^ ]*" .. ar.TAM .. "$") or
rfind(head, "^[^ ]*" .. ar.TAM .. " ")
if (is_masc_sg(g) or is_masc_fem_sg(g)) and ends_with_tam then
insert(data.categories, langname .. " masculine terms with feminine ending")
elseif (is_fem_sg(g) or is_masc_fem_sg(g)) and not ends_with_tam and
not rfind(head, "[" .. ar.ALIF .. ar.AMAQ .. "]$") and
not rfind(head, ar.ALIF .. ar.HAMZA .. "$") then
insert(data.categories, langname .. " feminine terms lacking feminine ending")
end
end
end
end
end
-----------------------------------------------------------------------------------------
-- Inflection handlers --
-----------------------------------------------------------------------------------------
-- Add list parameters to `params` (a structure as passed to [[Module:parameters]]) for a parameter named `argpref`.
-- If `argpref` is "*", add the nominal inflection parameters for construct state, definite state, etc. Related
-- transliteration and gender parameters are no longer supported in favor of inline modifiers, and error messages are
-- output if these parameters are used.
local function add_infl_params(params, argpref)
params[argpref] = {list = true, disallow_holes = true}
params[argpref .. "tr"] = {replaced_by = false, instead = "use <tr:...> inline modifier on |" .. argpref .. "="}
params[argpref .. "g"] = {replaced_by = false, instead = "use <g:...> inline modifier on |" .. argpref .. "="}
end
--[=[
Fetch a list of inflections from the arguments in `args` based on argument `field` (e.g. "pl"). Label with `label`
(e.g. "plural"), which will appear in the headword. Insert into `data.inflections`, where `data` is the structure
passed to [[Module:headword]]. If `generate_default` is specified, it should be a function of two arguments
(`data`, `args`), which should generate the default value if no values are specified or if "+" is explicitly given.
If `generate_default` isn't specified and the user gave no values, no inflection will be inserted.
]=]
local function handle_infl(data, args, spec)
local newinfls = parse_inflection(data, args, spec.field, false)
if not newinfls[1] and spec.default_when_not_explicit and spec.default_when_not_explicit(data, args) then
newinfls = {{term = "+"}}
end
if spec.handle then
spec.handle(data, args, newinfls)
end
local default_specs = spec.allowed_defspecs
if not default_specs then
default_specs = spec.generate_default and {["+"] = true} or {}
end
local saw_defspec = false
for _, newinfl in ipairs(newinfls) do
if default_specs[newinfl.term] or newinfl.term == "~" then
saw_defspec = true
break
end
end
if saw_defspec then
local newnewinfls = {}
for _, newinfl in ipairs(newinfls) do
if default_specs[newinfl.term] then
if newinfl.tr then
error("Can't specify translit for default value '" .. newinfl.term .. "'")
end
local definfls = spec.generate_default(data, args, newinfl.term)
for _, definfl in ipairs(definfls) do
m_headword_utilities.combine_termobj_decorations(definfl, newinfl)
insert(newnewinfls, definfl)
end
elseif newinfl.term == "~" then
if newinfl.tr then
error("Can't specify translit for head-requesting value '~'")
end
for _, headobj in ipairs(data.heads) do
headobj = m_table.shallowCopy(headobj)
m_headword_utilities.combine_termobj_decorations(headobj, newinfl)
insert(newnewinfls, headobj)
end
else
insert(newnewinfls, newinfl)
end
end
newinfls = newnewinfls
end
if newinfls[1] then
if newinfls[1].term == "--" then
if newinfls[2] then
error("Can't specify more than one term if first term is '--', meaning \"don't insert anything\"")
end
else
insert_inflection(data, newinfls, spec.label, nil, spec.defgender, spec.field, spec.no_label,
spec.usually_no_label)
end
end
end
local function add_infl_list_params(params, infl_list)
for _, infl in ipairs(infl_list) do
add_infl_params(params, infl.field)
end
end
local function handle_infl_list_args(data, args, infl_list)
for _, infl in ipairs(infl_list) do
handle_infl(data, args, infl)
end
end
-----------------------------------------------------------------------------------------
-- Default ending generators --
-----------------------------------------------------------------------------------------
local function make_conditional_default(specs)
return function(data, args)
local heads = data.heads
if not heads[1] then
heads = {{term = data.pagename}}
end
local newobjs = {}
for _, headobj in ipairs(heads) do
local term = ar.reorder_shadda(headobj.term)
local tr = headobj.tr
local matched = false
for _, spec in ipairs(specs) do
local from, fromtr, to, totr = unpack(spec)
if from:find("^%^") then
pref = rmatch(term, from .. "$")
else
pref = rmatch(term, "^(.*)" .. from .. "$")
end
if pref then
term = pref .. to
tr = replace_tr_ending(tr, fromtr, totr)
matched = true
headobj = m_table.shallowCopy(headobj)
headobj.term = ar.undo_reorder_shadda(term)
headobj.tr = tr
insert(newobjs, headobj)
break
end
end
if not matched then
error(("Internal error: No matching spec: head=%s"):format(dump(headobj)))
end
end
return newobjs
end
end
local default_feminine = make_conditional_default {
{ar.AN .. ar.AMAQ, "an", ar.AAH, "āh"},
{ar.AN .. ar.ALIF, "an", ar.AAH, "āh"}, -- e.g. مُحْيًا
{ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAH, "iya"},
{ar.IN, "in", ar.IYAH, "iya"},
{"", "", ar.AH, "a"},
}
local default_masculine = make_conditional_default {
-- tall alif substitutes for alif maqṣūra after a yāʔ
{ar.Y .. ar.AAH, "āh", ar.AN .. ar.ALIF, "an"},
{ar.AAH, "āh", ar.AN .. ar.AMAQ, "an"},
-- handle the common case of final-weak feminine active participle with preceding hamza;
-- the hamza-on-yāʔ always converts back to hamza on the line when preceded by ā (alif) but
-- may not otherwise, so we just leave it alone in that case
{ar.ALIF .. ar.HAMZA_ON_YA .. ar.IYAH, "iya", ar.HAMZA .. ar.IN, "in"},
{ar.IYAH, "iya", ar.IN, "in"},
{ar.AH, "a", "", ""},
{"", "", "", ""},
}
local default_masculine_plural = make_conditional_default {
{ar.AN .. ar.AMAQ, "an", ar.AWN, "awn"},
{ar.AN .. ar.ALIF, "an", ar.AWN, "awn"}, -- e.g. مُحْيًا
{ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_WAW .. ar.UUN, "ūn"},
{ar.IN, "in", ar.UUN, "ūn"},
{"", "", ar.UUN, "ūn"},
}
local default_feminine_plural = make_conditional_default {
-- صَلَاة pl. صَلَوَات and أَدَاة pl. أَدَوَات and similar; but نَوَاة and وَفَاة with a و in them become نَوَيَات and وَفَيَات;
-- and longer terms like مُبَارَاة and كُمَّثْرَاة invariably form their plural in -يَات.
{"^([^و]" .. ar.A .. "[^و])" .. ar.AAH, "āh", ar.A .. ar.W .. ar.AAT, "awāt"},
{ar.AAH, "āh", ar.AYAAT, "ayāt"},
{ar.AN .. ar.AMAQ, "an", ar.AYAAT, "ayāt"},
{ar.AN .. ar.ALIF, "an", ar.AYAAT, "ayāt"}, -- e.g. مُحْيًا
{ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAAT, "iyāt"},
{ar.IN, "in", ar.IYAAT, "iyāt"},
{ar.AH, "a", ar.AAT, "āt"},
{"", "", ar.AAT, "āt"},
}
local default_masculine_dual = make_conditional_default {
{ar.AN .. ar.AMAQ, "an", ar.AYAAN, "ayān"},
{ar.AN .. ar.ALIF, "an", ar.AYAAN, "ayān"}, -- e.g. مُحْيًا
{ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IYAAN, "iyān"},
{ar.IN, "in", ar.IYAAN, "iyān"},
{"", "", ar.AAN, "ān"},
}
local default_feminine_dual = make_conditional_default {
{ar.AN .. ar.AMAQ, "an", ar.AATAAN, "ātān"},
{ar.AN .. ar.ALIF, "an", ar.AATAAN, "ātān"}, -- e.g. مُحْيًا
{ar.HAMZA .. ar.IN, "in", ar.HAMZA_ON_YA .. ar.IY .. ar.ATAAN, "iyatān"},
{ar.IN, "in", ar.IY .. ar.ATAAN, "iyatān"},
{"", "", ar.ATAAN, "atān"},
}
-- Return whether `term` is a nisba noun or adjective, ending in -iyy or -iyyah. `nisba_val` is the value of
-- args.nisba; if non-nil, it overrides any auto-determination based on the shape of the term.
local function term_is_nisba(term, nisba_val)
if nisba_val ~= nil then
return nisba_val
end
term = ar.reorder_shadda(term) -- necessary to avoid issues with e.g. أُورُوبِّيّ.
local pref = rmatch(term, "^(.*)" .. ar.IYY .. ar.UN .. "?$")
if not pref then
pref = rmatch(term, "^(.*)" .. ar.IYYAH .. ar.UN .. "?$")
end
-- Avoid false positives for words like قَوِيّ "strong" and صَبِيّ "boy". There may be other false positives
-- but this should catch most of them and will avoid very many false negatives.
return pref and not rfind(pref, "^[^ا]" .. ar.A .. ".$")
end
-----------------------------------------------------------------------------------------
-- Adjectives --
-----------------------------------------------------------------------------------------
local function is_defaulting_adjective(data, args)
return data.orig_pos_category == "defaulting adjectives"
end
local adj_field_elative = {field = "el", label = "<<elative>>"}
local adj_inflections = {
adj_field_inf,
adj_field_obl,
adj_field_def,
{field = "f", label = "feminin", generate_default = default_feminine,
default_when_not_explicit = is_defaulting_adjective},
{field = "d", label = "maskulin duaan", generate_default = default_masculine_dual},
{field = "fd", label = "feminin duaan", generate_default = default_feminine_dual},
{field = "cpl", label = "am jamak"},
{field = "pl", label = "maskulin jamak", generate_default = default_masculine_plural,
default_when_not_explicit = is_defaulting_adjective},
{field = "fpl", label = "feminin jamak", generate_default = default_feminine_plural,
default_when_not_explicit = is_defaulting_adjective},
}
local function get_adj_params()
local params = {}
add_infl_list_params(params, adj_inflections)
add_infl_params(params, "el")
params.nisba = boolean_param
return params
end
local function handle_adj_args(data, args)
handle_infl_list_args(data, args, adj_inflections)
handle_infl(data, args, adj_field_elative)
for _, headobj in ipairs(data.heads) do
if term_is_nisba(headobj.term, args.nisba) then
insert(data.categories, "Kata sifat relatif (nisba) bahasa " .. langname)
break
end
end
end
pos_functions["Kata sifat"] = {
params = get_adj_params,
func = handle_adj_args,
}
pos_functions["defaulting adjectives"] = {
params = get_adj_params,
func = function(data, args)
data.pos_category = "Kata sifat"
handle_adj_args(data, args)
end,
}
-----------------------------------------------------------------------------------------
-- Nouns, etc. --
-----------------------------------------------------------------------------------------
local function get_masc_or_feminine_gender(data, default_type)
local saw_m, saw_f, saw_mf
for _, gender in ipairs(data.genders) do
if is_masc_sg(gender.spec) then
saw_m = true
elseif is_fem_sg(gender.spec) then
saw_f = true
elseif is_masc_fem_sg(gender.spec) then
saw_mf = true
end
end
if saw_mf or saw_m and saw_f then
error("Can't generate default for " .. default_type .. " when gender is both masculine and feminine")
elseif saw_m then
return "m"
elseif saw_f then
return "f"
else
error("Can't generate default for " .. default_type .. " when gender is not specified as " ..
"maskulin or feminin tunggal")
end
end
local function is_defaulting_noun(data, args)
return data.orig_pos_category == "defaulting nouns"
end
local noun_field_dual = {
field = "d", label = "dual",
generate_default = function(data, args)
local gender = get_masc_or_feminine_gender(data, "noun dual")
if gender == "m" then
return default_masculine_dual(data, args)
else
return default_feminine_dual(data, args)
end
end,
}
local noun_field_plural = {
field = "pl", label = "jamak",
generate_default = function(data, args, defspec)
local gender = get_masc_or_feminine_gender(data, "noun plural")
if gender == "m" then
if defspec == "+f" then
return default_feminine_plural(data, args)
else
return default_masculine_plural(data, args)
end
elseif defspec == "+f" then
error("Can't specify '+f' with feminine gender; just use '+'")
else
return default_feminine_plural(data, args)
end
end,
-- Handle the case where pl=-, indicating an uncountable noun.
handle = function(data, args, terms)
if terms[1] and terms[1] == "-" then
insert(data.categories, langname .. " uncountable nouns")
if args.pauc and args.pauc[1] then
error("Can't specify paucals when pl=-")
end
end
end,
allowed_defspecs = {["+"] = true, ["+f"] = true},
default_when_not_explicit = is_defaulting_noun,
no_label = "<<uncountable>>",
usually_no_label = "usually <<uncountable>>",
}
local noun_field_paucal = {
field = "pauc", label = "<<paucal>>", generate_default = default_feminine_plural,
}
local noun_field_feminine = {
field = "f", label = "feminin", generate_default = default_feminine,
default_when_not_explicit = function(data, args)
if data.orig_pos_category ~= "defaulting nouns" then
return nil
end
local gender = get_masc_or_feminine_gender(data, "defaulting-if-masculine noun feminine")
return gender == "m"
end,
}
local noun_field_masculine = {
field = "m", label = "maskulin", generate_default = default_masculine,
default_when_not_explicit = function(data, args)
if data.orig_pos_category ~= "defaulting nouns" then
return nil
end
local gender = get_masc_or_feminine_gender(data, "defaulting-if-feminine noun masculine")
return gender == "f"
end,
}
local noun_basic_inflections = {
noun_field_cons,
noun_field_inf,
noun_field_obl,
noun_field_def,
}
local noun_shared_inflections = {
noun_field_dual,
noun_field_plural,
}
local noun_extra_inflections = {
noun_field_paucal,
noun_field_feminine,
noun_field_masculine,
}
local function get_noun_params()
local params = {}
add_gender_params(params)
add_infl_list_params(params, noun_basic_inflections)
add_infl_list_params(params, noun_shared_inflections)
add_infl_list_params(params, noun_extra_inflections)
params.nisba = boolean_param
return params
end
local function handle_noun_args(data, args)
handle_gender(data, args)
handle_infl_list_args(data, args, noun_basic_inflections)
handle_infl_list_args(data, args, noun_shared_inflections)
handle_infl_list_args(data, args, noun_extra_inflections)
for _, headobj in ipairs(data.heads) do
if term_is_nisba(headobj.term, args.nisba) then
insert(data.categories, langname .. " relative nouns (nisba)")
break
end
end
end
pos_functions["Kata nama"] = {
params = get_noun_params,
func = handle_noun_args,
}
pos_functions["defaulting nouns"] = {
params = get_noun_params,
func = function(data, args)
data.pos_category = "Kata nama"
handle_noun_args(data, args)
end,
}
local noun_field_singulative = {field = "sing", label = "<<singulative>>", defgender = "f", generate_default = default_feminine}
local noun_field_collective = {field = "coll", label = "<<collective>>", defgender = "m", generate_default = default_masculine}
local function handle_sing_coll_noun_infls(data, args, otherinfl, otherlabel, othergender)
-- Handle sing= (corresponding singulative noun) or coll= (corresponding collective noun) and their gender
handle_infl(data, args, otherinfl, otherlabel, nil, othergender)
handle_infl_list_args(data, args, sing_coll_noun_inflections)
end
local function get_singulative_collective_noun_params(defgender, otherinfl)
local params = {}
add_gender_params(params, defgender)
add_infl_list_params(params, noun_basic_inflections)
add_infl_params(params, otherinfl)
add_infl_list_params(params, noun_shared_inflections)
add_infl_params(params, "pauc")
return params
end
pos_functions["collective nouns"] = {
params = function() return get_singulative_collective_noun_params("m", "sing") end,
func = function(data, args)
data.pos_category = "Kata nama"
insert(data.categories, langname .. " collective nouns")
m_headword_utilities.insert_fixed_inflection {
headdata = data,
label = "<<collective>>",
}
handle_gender(data, args)
handle_infl_list_args(data, args, noun_basic_inflections)
handle_infl(data, args, noun_field_singulative)
handle_infl_list_args(data, args, noun_shared_inflections)
handle_infl(data, args, noun_field_paucal)
end
}
pos_functions["singulative nouns"] = {
params = function() return get_singulative_collective_noun_params("f", "coll") end,
func = function(data, args)
data.pos_category = "Kata nama"
insert(data.categories, langname .. " singulative nouns")
m_headword_utilities.insert_fixed_inflection {
headdata = data,
label = "<<singulative>>",
}
handle_gender(data, args)
handle_infl_list_args(data, args, noun_basic_inflections)
handle_infl(data, args, noun_field_collective)
handle_infl_list_args(data, args, noun_shared_inflections)
handle_infl(data, args, noun_field_paucal)
end
}
-- FIXME: Do numerals really behave almost as nouns? They vary by masc/fem.
pos_functions["numerals"] = {
params = get_noun_params,
func = function(data, args)
insert(data.categories, langname .. " cardinal numbers")
handle_noun_args(data, args)
end
}
pos_functions["Kata nama khas"] = {
params = get_noun_params,
func = handle_noun_args,
}
local function get_pronoun_params()
local params = {}
add_gender_params(params, defgender)
add_infl_list_params(params, noun_basic_inflections)
add_infl_list_params(params, noun_shared_inflections)
add_infl_params(params, "f")
return params
end
pos_functions["pronouns"] = {
params = get_pronoun_params,
func = function(data, args)
handle_gender(data, args)
handle_infl_list_args(data, args, noun_basic_inflections)
handle_infl_list_args(data, args, noun_shared_inflections)
handle_infl(data, args, noun_field_feminine)
end
}
-----------------------------------------------------------------------------------------
-- Non-lemma forms --
-----------------------------------------------------------------------------------------
local valid_forms = list_to_set(
{ "I", "II", "III", "IV", "V", "VI", "VII", "VIII", "IX", "X", "XI", "XII",
"XIII", "XIV", "XV", "Iq", "IIq", "IIIq", "IVq" })
-- FIXME: Partly duplicated in [[Module:ar-inflections]].
local function handle_conj_form(data, args)
local form = args[2]
if form then
if not valid_forms[form] then
error("Invalid verb conjugation form " .. form)
end
insert(data.inflections, { label = "[[Appendix:Arabic verbs#Form " .. form .. "|form " .. form .. "]]" })
end
end
pos_functions["verb forms"] = {
params = function()
return {
[2] = {},
}
end,
func = function(data, args)
handle_conj_form(data, args)
end
}
local function get_participle_params()
local params = get_adj_params()
params[2] = {}
return params
end
pos_functions["active participles"] = {
params = get_participle_params,
func = function(data, args)
data.pos_category = "participles"
insert(data.categories, langname .. " active participles")
handle_conj_form(data, args)
handle_infl_list_args(data, args, adj_inflections)
end
}
pos_functions["passive participles"] = {
params = get_participle_params,
func = function(data, args)
data.pos_category = "participles"
insert(data.categories, langname .. " passive participles")
handle_conj_form(data, args)
handle_infl_list_args(data, args, adj_inflections)
end
}
-----------------------------------------------------------------------------------------
-- Verbs --
-----------------------------------------------------------------------------------------
pos_functions["Kata kerja"] = {
head_is_not_1 = true,
params = function() return {
[1] = {},
-- Comma-separated lists with possible inline modifiers
["past"] = {},
["past1s"] = {},
["nonpast"] = {},
["vn"] = {},
["noautolinktext"] = {type = "boolean"},
["noautolinkverb"] = {type = "boolean"},
} end,
func = function(data, args)
local ar_verb = require(ar_verb_module)
local alternant_multiword_spec =
args[1] ~= "-" and ar_verb.do_generate_forms(args, "ar-verb", data.pagename) or nil
local function do_slot(slots_to_check, override, label, slot_is_headword)
-- Do this even with an override so we can return the correct filled slot.
local slot, slotval
if alternant_multiword_spec then
for _, potential_slot in ipairs(slots_to_check) do
slotval = alternant_multiword_spec.forms[potential_slot]
if slotval then
slot = potential_slot
break
end
end
end
local function get_slot_values()
local terms = {}
for _, form in ipairs(slotval) do
local term = {
term = form.form,
id = form.id,
genders = form.genders,
pos = form.pos,
lit = form.lit,
}
term.tr = form.translit
if form.footnotes then
local quals, refs = require(inflection_utilities_module).
convert_footnotes_to_qualifiers_and_references(form.footnotes)
term.q = quals
term.refs = refs
end
insert(terms, term)
end
return terms
end
if override then
local override_param_mods = {
alt = {},
t = {
-- [[Module:headword]] expects the gloss in "gloss".
item_dest = "gloss",
},
gloss = {},
g = {
-- [[Module:headword]] expects the genders in "genders".
item_dest = "genders",
type = "genders",
},
pos = {},
lit = {},
id = {},
-- Qualifiers and labels
q = {
type = "qualifier",
},
qq = {
type = "qualifier",
},
l = {
type = "labels",
},
ll = {
type = "labels",
},
ref = {
-- [[Module:headword]] expects the references in "refs".
item_dest = "refs",
type = "references",
},
}
local function generate_obj(formval, parse_err)
if formval == "+" then
return {term = "+", underlying_terms = get_slot_values()}
end
local val, uncertain = formval:match("^(.*)(%?)$")
val = val or formval
uncertain = not not uncertain
local ar, translit = val:match("^(.*)//(.*)$")
if not ar then
ar = formval
end
local retval = {term = ar, uncertain = uncertain}
retval.tr = translit
end
local terms
if override:find("<") then
terms = require(parse_utilities_module).parse_inline_modifiers(override, {
paramname = paramname,
param_mods = override_param_mods,
generate_obj = generate_obj,
splitchar = "[,،]",
escape_fun = escape_comma_whitespace,
unescape_fun = unescape_comma_whitespace,
})
else
terms = split_on_comma(override)
for i, split in ipairs(terms) do
terms[i] = generate_obj(split)
end
end
-- See if + was supplied and we have to potentially flatten multiple default terms and harmonize
-- default properties with override properties.
local saw_underlying_terms = false
for _, term in ipairs(terms) do
if term.underlying_terms then
saw_underlying_terms = true
break
end
end
if saw_underlying_terms then
-- Flatten any default terms, copying the corresponding override properties over the default
-- properties. Non-default terms get inserted directly.
local flattened = {}
for _, term in ipairs(terms) do
if term.underlying_terms then
for _, underlying in ipairs(term.underlying_terms) do
for k, v in pairs(term) do
if k ~= "term" and k ~= "underlying_terms" then
if k == "uncertain" then
underlying.uncertain = underlying.uncertain or v
elseif type(v) ~= "table" or v[1] then
-- Don't copy empty lists (which are the default) over possibly non-empty
-- lists.
underlying[k] = v
end
end
end
insert(flattened, underlying)
end
else
insert(flattened, term)
end
end
terms = flattened
end
if not slot_is_headword then
terms.label = label
end
return terms, slot
elseif not alternant_multiword_spec then
return nil, slot
else
if not slotval then
if slot_is_headword then
-- FIXME, put "uncertain" as qualifier? Does this ever happen?
return nil, slot
elseif alternant_multiword_spec.slot_uncertain[slot] then
return {label = label .. " uncertain"}, slot
elseif alternant_multiword_spec.slot_explicitly_missing[slot] then
return {label = "no " .. label}, slot
else
-- just say nothing about this slot
return nil, slot
end
end
local terms = get_slot_values()
if not slot_is_headword then
terms.label = label
end
return terms, slot
end
end
local gloss_parts = {}
for _, vform in ipairs(alternant_multiword_spec.verb_forms) do
insert(gloss_parts, "[[Appendix:Arabic verbs#Form " .. vform .. "|" .. vform .. "]]")
end
if gloss_parts[1] then
data.gloss = concat(gloss_parts, ", ")
end
if data.heads[1] and args.past then
error("Can't specify both head= and past= to {{ar-verb}}; prefer past=")
end
if not alternant_multiword_spec.has_active then
insert(data.inflections, {label = "passive-only"})
end
-- Do this always so `past_slot` is correctly filled.
local past, past_slot = do_slot(ar_verb.potential_lemma_slots, args.past, "-", "slot is headword")
if data.heads[1] then
-- user specified head=; don't override with past= or slot 'past_3sm' etc.
else
if past then
data.heads = past
end
end
local should_do_past1s = not not args.past1s
if not should_do_past1s then
local is_form_I = false
for _, vform in ipairs(alternant_multiword_spec.verb_forms) do
if vform == "I" then
is_form_I = true
break
end
end
if is_form_I then
require(inflection_utilities_module).map_word_specs(alternant_multiword_spec, function(base)
if base.verb_form == "I" then
for _, vowel_spec in ipairs(base.conj_vowels) do
-- For form-I geminate verbs, the final vowel of the past is elided in the citation form.
-- We want to display it for all cases other than active a~u and a~i (the most common
-- cases).
if vowel_spec.weakness == "geminate" then
if ar_verb.is_passive_only(base.passive) then
should_do_past1s = true
break
end
local past_vowel = ar_verb.rget(vowel_spec.past)
local nonpast_vowel = ar_verb.rget(vowel_spec.nonpast)
if not (past_vowel == ar.A and (nonpast_vowel == ar.U or nonpast_vowel == ar.I)) then
should_do_past1s = true
break
end
end
end
-- FIXME, provide way of breaking early from map_word_specs().
end
end)
end
end
local past1s
if should_do_past1s then
past1s, _ = do_slot({"past_1s", "past_pass_1s"}, args.past1s, "first-person singular past")
if past1s then
insert(data.inflections, past1s)
end
end
local nonpast_slots
if not past_slot or past_slot:find("^past_") then
nonpast_slots = {"ind_3ms", "ind_pass_3ms", "imp_2ms"}
else
nonpast_slots = {}
end
local nonpast, _ = do_slot(nonpast_slots, args.nonpast, "non-past")
if nonpast then
insert(data.inflections, nonpast)
end
local vn, _ = do_slot({"vn"}, args.vn, "verbal noun")
if vn then
insert(data.inflections, vn)
end
-- FIXME: Should we insert categories? Conjugation also does it and is more likely to be accurate.
--for _, cat in ipairs(alternant_multiword_spec.categories) do
-- insert(data.categories, cat)
--end
--[=[
-- FIXME: Review this to see if we need to port it.
-- If the user didn't explicitly specify head=, or specified exactly one head (not 2+) and we were able to
-- incorporate any links in that head into the 1= specification, use the infinitive generated by
-- [[Module:pt-verb]] in place of the user-specified or auto-generated head. This was copied from
-- [[Module:it-headword]], where doing this gets accents marked on the verb(s). We don't have accents marked on
-- the verb but by doing this we do get any footnotes on the infinitive propagated here. Don't do this if the
-- user gave multiple heads or gave a head with a multiword-linked verbal expression such as Italian
-- '[[dare esca]] [[al]] [[fuoco]]' (FIXME: give Portuguese equivalent).
if not data.user_specified_heads[1] or (
not data.user_specified_heads[2] and alternant_multiword_spec.incorporated_headword_head_into_lemma
) then
data.heads = {}
for _, lemma_obj in ipairs(alternant_multiword_spec.forms.infinitive_linked) do
local quals, refs = require(inflection_utilities_module).
convert_footnotes_to_qualifiers_and_references(lemma_obj.footnotes)
insert(data.heads, {term = lemma_obj.form, q = quals, refs = refs})
end
end
]=]
end
}
-----------------------------------------------------------------------------------------
-- Generic parts of speech --
-----------------------------------------------------------------------------------------
pos_functions.head_with_gender = {
params = function()
return {
[3] = {type = "genders"},
}
end,
func = function(data, args)
handle_gender(data, args, "nonlemma", 3)
end,
}
return export
1mc21n9hd7rlm8q5tnhtlhefbmz6j3l
Modul:rhymes
828
11249
375346
228342
2026-09-22T03:11:54Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708593|92708593]])
375346
Scribunto
text/plain
local export = {}
local force_cat = false -- for testing
local decorations_module = "Module:decorations"
local IPA_module = "Module:IPA"
local parameters_module = "Module:parameters"
local parameter_utilities_module = "Module:parameter utilities"
local rhymes_styles_css_module = "Module:rhymes/styles.css"
local TemplateStyles_module = "Module:TemplateStyles"
local utilities_module = "Module:utilities"
local rhymes_data = require("Module:rhymes/data")
local concat = table.concat
local insert = table.insert
local function track(page)
require("Module:debug/track")("rhymes/" .. page)
return true
end
local function tag_rhyme(rhyme, lang)
local formatted_rhyme, cats, err
formatted_rhyme, cats, err = require(IPA_module).format_IPA(lang, rhyme, "raw")
return formatted_rhyme, cats, err
end
local function make_rhyme_link(lang, link_rhyme, display_rhyme)
local retval, cats
local prefix = "[[Rima:Bahasa "
if rhymes_data.link_to_category_langs[lang:getCode()] then
prefix = "[[:Kategori:Rima:Bahasa "
end
if not link_rhyme then
retval = concat{prefix, lang:getCanonicalName(), "|", lang:getCanonicalName(), "]]"}
cats = {}
else
local formatted_rhyme, err
formatted_rhyme, cats, err = tag_rhyme(display_rhyme or link_rhyme, lang)
retval = concat{prefix, lang:getCanonicalName(), "/", link_rhyme, "|", formatted_rhyme, "]]", err}
end
return retval, cats
end
--[==[
Implementation of {{tl|rhymes row}}.
]==]
function export.show_row(frame)
local args = require(parameters_module).process(
frame.getParent and frame:getParent().args or frame,
{
[1] = {required = true, type = "full language"},
[2] = {required = true},
[3] = {},
}
)
if not args[1] then
return "[[Rhymes:English/aɪmz|<span class=\"IPA\">-aɪmz</span>]]"
end
-- Discard cleanup categories from make_rhyme_link().
return (make_rhyme_link(args[1], args[2], "-" .. args[2])) .. (args[3] and (" (''" .. args[3] .. "'')") or "")
end
do
local function add_syllable_categories(categories, lang, rhyme, num_syl)
local prefix = "Rima:Bahasa " .. lang .. "/" .. rhyme
insert(categories, prefix)
if num_syl then
for _, n in ipairs(num_syl) do
local c
if n > 1 then
c = prefix .. "/" .. n .. " suku kata"
else
c = prefix .. "/1 suku kata"
end
insert(categories, c)
end
end
end
--[==[
Meant to be called from a module. `data` is a table containing the following fields:
* `lang`: language object for the rhymes;
* `rhymes`: a list of rhymes, each described by an object which specifies the rhyme, optional number of syllables, and
optional decoration fields:
** `rhyme`: the rhyme itself;
** `num_syl`: {nil} or a list of numbers, specifying the number of syllables of the word with this rhyme; optional and
currently used only for categorization; if omitted, defaults to the top-level `num_syl`;
** `separator`: {nil} or the string used to separate this rhyme from the preceding one when displayed; defaults to the
top-level `separator`;
** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the rhyme in question;
** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the rhyme in question;
** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed directly before
the rhyme in question;
** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the rhyme in question;
** `refs`: {nil} or a list of references or reference specs to add directly after the rhyme; the value of a list item is
either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or a
template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the
reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}})
and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or
{{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference
appropriately and insert a footnote number that hyperlinks to the actual reference, located in the
{{cd|<nowiki><references /></nowiki>}} section;
** `nocat`: if {true}, suppress categorization for this rhyme only;
* `num_syl`: {nil} or a list of numbers, specifying the number of syllables for all rhymes; optional and currently used
only for categorization; overridable at the individual rhyme level;
* `separator`: {nil} or a string, specifying the separator displayed before all rhymes but the first; by default,
{", "}; overridable at the individual rhyme level;
* `q`: {nil} or a list of overall left regular qualifier strings, displayed before the initial caption;
* `qq`: {nil} or a list of overall right regular qualifier strings, displayed after all rhymes;
* `a`: {nil} or a list of overall left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the
initial caption;
* `aa`: {nil} or a list of right accent qualifier strings, displayed after all rhymes;
* `sort`: {nil} or sort key;
* `caption`: {nil} or string specifying the caption to use, in place of {"Rhymes"}; a colon and space is automatically
added after the caption;
* `nocaption`: if {true}, suppress the caption display;
* `nocat`: if {true}, suppress categorization;
* `force_cat`: if {true}, force categorization even on non-mainspace pages.
If both regular and accent qualifiers on the same side and at the same level are specified, the accent qualifiers
precede the regular qualifiers on both left and right.
'''WARNING''': Destructively modifies the objects inside the `rhymes` field.
Note that the number of syllables is currently used only for categorization; if present, an extra category will
be added such as [[:Category:Rhymes:Italian/ino/3 syllables]] in addition to [[:Category:Rhymes:Italian/ino]].
]==]
function export.format_rhymes(data)
local langname = data.lang:getFullName()
local parts = {}
local categories = {}
local overall_sep = data.separator or ", "
for i, r in ipairs(data.rhymes) do
local rhyme = r.rhyme
local link, link_cats = make_rhyme_link(data.lang, rhyme, "-" .. rhyme)
if not r.nocat and not data.nocat then
for _, cat in ipairs(link_cats) do
insert(categories, cat)
end
end
if r.qualifiers then
-- FIXME; Added 2026-09-18; remove in a month
error("qualifiers= not allowed here; use q=")
end
if r.q and r.q[1] or r.qq and r.qq[1] or r.a and r.a[1] or r.aa and r.aa[1] or r.refs and r.refs[1] then
link = require(decorations_module).format_decorations {
lang = data.lang,
text = link,
q = r.q,
qq = r.qq,
a = r.a,
aa = r.aa,
refs = r.refs,
}
end
insert(parts, r.separator or i > 1 and overall_sep or "")
insert(parts, link)
if not r.nocat and not data.nocat then
add_syllable_categories(categories, langname, rhyme, r.num_syl or data.num_syl)
end
end
local text = concat(parts)
if not data.nocaption then
text = (data.caption or "Rima") .. ": " .. text
end
if data.qualifiers then
-- FIXME; Added 2026-09-18; remove in a month
error("overall qualifiers= not allowed here; use q=")
end
if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then
text = require(decorations_module).format_decorations {
lang = data.lang,
text = text,
q = data.q,
qq = data.qq,
a = data.a,
aa = data.aa,
}
end
if categories[1] then
local cat_text = require(utilities_module).format_categories(categories, data.lang, data.sort, nil,
force_cat or data.force_cat)
text = text .. cat_text
end
return text
end
end
--[==[
Implementation of {{tl|rhymes}}.
]==]
function export.show(frame)
local parent_args = frame:getParent().args
local compat = parent_args.lang
local offset = compat and 0 or 1
local lang_param = compat and "lang" or 1
local plain = {}
local boolean = {type = "boolean"}
local params = {
[lang_param] = {required = true, type = "language", default = "ms"},
[1 + offset] = {list = true, required = true, disallow_holes = true, default = "aɪmz"},
["caption"] = plain,
["nocaption"] = boolean,
["nocat"] = boolean,
["sort"] = plain,
}
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{
param = "s",
item_dest = "num_syl",
separate_no_index = true,
type = "number",
sublist = true,
},
{group = {"q", "a", "ref"}},
}
local rhymes, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 1 + offset,
term_dest = "rhyme",
track_module = "rhymes",
}
local lang = args[lang_param]
local data = {
lang = lang,
rhymes = rhymes,
num_syl = args.s.default,
caption = args.caption,
nocaption = args.nocaption,
nocat = args.nocat,
sort = args.sort,
q = args.q.default,
qq = args.qq.default,
a = args.a.default,
aa = args.aa.default,
}
return export.format_rhymes(data)
end
--[==[
Implementation of {{tl|rhymes nav}}.
]==]
function export.show_nav(frame)
local args = require(parameters_module).process(
frame:getParent().args,
{
[1] = {required = true, type = "full language", default = "und"},
[2] = {list = true, allow_holes = true},
["nocat"] = {type = "boolean"},
}
)
local lang = args[1]
local langname = lang:getCanonicalName()
local parts = args[2]
-- Create steps
-- FIXME: We should probably use format_categories() in [[Module:utilities]] rather than constructing categories
-- manually.
local categories = {}
-- Here and below, we ignore any cleanup categories coming out of make_rhyme_link() by adding an extra set of parens
-- around the call to make_rhyme_link() to cause the second argument (the categories) to be ignored. {{rhymes nav}}
-- is run on a rhymes page so it's not clear we want the page to be added to any such categories, if they exist.
local steps = {"[[Wikikamus:Rima|Rima]]", (make_rhyme_link(lang))}
if #parts > 0 then
local last = parts[#parts]
parts[#parts] = nil
local prefix = ""
for i, part in ipairs(parts) do
prefix = prefix .. part
parts[i] = prefix
end
for _, part in ipairs(parts) do
insert(steps, (make_rhyme_link(lang, part .. "-", "-" .. part .. "-")))
end
if last == "-" then
insert(steps, (make_rhyme_link(lang, prefix, "-" .. prefix)))
insert(categories, "[[Kategori:Rima bahasa " .. langname .. (prefix == "" and "" or "/" .. prefix .. "-") .. "| ]]")
elseif mw.title.getCurrentTitle().text == langname .. "/" .. prefix .. last .. "-" then -- DO NOT replace with mw.loadData("Module:headword/data").pagename as we need the root portion
insert(steps, (make_rhyme_link(lang, prefix .. last .. "-", "-" .. prefix .. last .. "-")))
insert(categories, "[[Kategori:Rima bahasa " .. langname .. "/" .. prefix .. last .. "-|-]]")
else
insert(steps, (make_rhyme_link(lang, prefix .. last, "-" .. prefix .. last)))
insert(categories, "[[Kategori:Rima bahasa " .. langname .. (prefix == "" and "" or "/" .. prefix .. "-") .. "|" .. last .. "]]")
end
elseif lang:getCode() ~= "und" then
insert(categories, "[[Kategori:Rima bahasa " .. langname .. "| ]]")
end
if mw.title.getCurrentTitle().nsText == "Rima" then
frame:callParserFunction("DISPLAYTITLE",
mw.title.getCurrentTitle().fullText:gsub(
"/(.+)$",
function (rhyme)
return "/" .. (tag_rhyme(rhyme, lang)) -- ignore cleanup categories
end))
end
local templateStyles = require(TemplateStyles_module)(rhymes_styles_css_module)
local ol = mw.html.create("ol")
for _, step in ipairs(steps) do
ol:node(mw.html.create("li"):wikitext(step))
end
local div = mw.html.create("div")
:attr("role", "navigation")
:attr("aria-label", "Breadcrumb")
:addClass("ts-rhymesBreadcrumbs")
:node(ol)
local formatted_cats = args.nocat and "" or concat(categories)
return templateStyles .. tostring(div) .. formatted_cats
end
return export
ht9j9li2mw2ejqi4l9yfaule57m75kd
Modul:etymology
828
11468
375371
373823
2026-09-22T05:09:38Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708527|92708527]])
375371
Scribunto
text/plain
local export = {}
-- For testing
local force_cat = false
local debug_track_module = "Module:debug/track"
local languages_module = "Module:languages"
local links_module = "Module:links"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local insert = table.insert
local new_title = mw.title.new
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_language_data_module_name(...)
get_language_data_module_name = require(languages_module).getDataModuleName
return get_language_data_module_name(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function serial_comma_join(...)
serial_comma_join = require(table_module).serialCommaJoin
return serial_comma_join(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function track(page, code)
local tracking_page = "etymology/" .. page
debug_track(tracking_page)
if code then
debug_track(tracking_page .. "/" .. code)
end
end
local function join_segs(segs, conj)
if not segs[2] then
return segs[1]
elseif conj == "and" or conj == "or" then
return serial_comma_join(segs, {conj = conj})
end
local sep
if conj == "," or conj == ";" then
sep = conj .. " "
elseif conj == "/" then
sep = "/"
elseif conj == "~" then
sep = " ~ "
elseif conj then
error(("Internal error: Unrecognized conjunction \"%s\""):format(conj))
else
error(("Internal error: No value supplied for conjunction"):format(conj))
end
return concat(segs, sep)
end
-- Returns true if `lang` is the same as `source`, or a variety of it.
local function lang_is_source(lang, source)
return lang:getCode() == source:getCode() or lang:hasParent(source)
end
--[==[
Format one or more links as specified in `termobjs`, a list of term objects of the format accepted by `full_link()` in
[[Module:links]], including decorations (qualifiers, labels and references). `conj` is used to join multiple terms
and must be specified if there is more than one term. `template_name` is the template name used in debug tracking and
must be specified. Optional `sourcetext` is text to prepend to the concatenated terms, separated by a space if the
concatenated terms are non-empty (which is always the case unless there is a single term with the value "-"). If
`decorations_on_outside` is given, any decorations specified in the first term go on the outside of (i.e before)
`sourcetext`; otherwise they will end up on the inside.
]==]
function export.format_links(termobjs, conj, template_name, sourcetext, decorations_on_outside)
if not template_name then
error("Internal error: Must specify `template_name` to format_links()")
end
for i, termobj in ipairs(termobjs) do
if termobj.lang:hasType("family") or termobj.lang:getFamilyCode() == "qfa-sub" then
if termobj.term and termobj.term ~= "-" then
debug_track(template_name .. "/family-with-term")
end
termobj.term = "-"
end
if termobj.term == "-" then
--[=[
[[Special:WhatLinksHere/Wiktionary:Tracking/cognate/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/derived/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/borrowed/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/calque/no-term]]
]=]
debug_track(template_name .. "/no-term")
termobjs[i] = i == 1 and sourcetext or ""
else
if i == 1 and decorations_on_outside and sourcetext then
termobj.pretext = sourcetext .. " "
sourcetext = nil
end
termobjs[i] = (i == 1 and sourcetext and sourcetext .. " " or "") .. full_link(termobj, "term")
end
end
return join_segs(termobjs, conj)
end
function export.get_display_and_cat_name(source, raw)
local display, cat_name
if source:getCode() == "und" then
display = "tidak ditentukan"
cat_name = "bahasa lain"
elseif source:getCode() == "mul" then
display = raw and "rentas bahasa" or "[[w:Translingualisme|rentas bahasa]]"
cat_name = "Rentas bahasa"
elseif source:getCode() == "mul-tax" then
display = raw and "taxonomic name" or "[[w:Tatanama biologi|nama taksonomi]]"
cat_name = "Nama taksonomi"
else
display = raw and source:getCanonicalName() or source:makeWikipediaLink()
cat_name = source:getDisplayForm()
end
return display, cat_name
end
function export.insert_source_cat_get_display(data)
local categories, lang, source = data.categories, data.lang, data.source
local display, cat_name = export.get_display_and_cat_name(source, data.raw)
if lang and not data.nocat then
-- Add the category, but only if there is a current language
if not categories then
categories = {}
end
local langname = lang:getFullName()
-- If `lang` is an etym-only language, we need to check both it and its parent full language against `source`.
-- Otherwise if e.g. `lang` is Medieval Latin and `source` is Latin, we'll end up wrongly constructing a
-- category 'Latin terms derived from Latin'.
insert(categories, "Perkataan bahasa " .. langname .. (
lang_is_source(lang, source) and " dipinjam balik ke dalam bahasa " .. cat_name or
" " .. (data.borrowing_type or "diterbitkan") .. " daripada bahasa " .. cat_name
))
end
return display, categories
end
function export.format_source(data)
local lang, sort_key = data.lang, data.sort_key
-- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/sortkey]]
if sort_key then
track("sortkey")
end
local display, categories = export.insert_source_cat_get_display(data)
if lang and not data.nocat then
-- Format categories, but only if there is a current language; {{cog}} currently gets no categories
categories = format_categories(categories, lang, sort_key, nil, data.force_cat or force_cat)
else
categories = ""
end
return "<span class=\"etyl\">" .. display .. categories .. "</span>"
end
--[==[
Format sources for etymology templates such as {{tl|bor}}, {{tl|der}}, {{tl|inh}} and {{tl|cog}}. There may potentially
be more than one source language (except currently {{tl|inh}}, which doesn't support it because it doesn't really
make sense). In that case, all but the last source language is linked to the first term, but only if there is such a
term and this linking makes sense, i.e. either (1) the term page exists after stripping diacritics according to the
source language in question, or (2) the result of stripping diacritics according to the source language in question
results in a different page from the same process applied with the last source language. For example, {{m|ru|соля́нка}}
will link to [[солянка]] but {{m|en|соля́нка}} will link to [[соля́нка]] with an accent, and since they are different
pages, the use of English as a non-final source with term 'соля́нка' will link to [[соля́нка]] even though it doesn't
exist, on the assumption that it is merely a redlink that might exist. If none of the above criteria apply, a non-final
source language will be linked to the Wikipedia entry for the language, just as final source languages always are.
`data` contains the following fields:
* `lang`: The destination language object into which the terms were borrowed, inherited or otherwise derived. Used for
categorization and can be nil, as with {{tl|cog}}.
* `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are
handled specially; see above.
* `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as
multiple term objects, the non-final source objects link to the first term object.
* `sort_key`: Sort key for categories. Usually nil.
* `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source
languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be
added.
* `nocat`: Don't add any categories to the page.
* `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized
values are `and`, `or`, `,`, `;`, `/` and `~`.
* `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}.
* `force_cat`: Force category generation on non-mainspace pages.
]==]
function export.format_sources(data)
local lang, sources, terms, borrowing_type, sort_key, categories, nocat =
data.lang, data.sources, data.terms, data.borrowing_type, data.sort_key, data.categories, data.nocat
local term1, sources_n, source_segs = terms[1], #sources, {}
local final_link_page
local term1_term, term1_sc = term1.term, term1.sc
if sources_n > 1 and term1_term and term1_term ~= "-" then
final_link_page = get_link_page(term1_term, sources[sources_n], term1_sc)
end
for i, source in ipairs(sources) do
local seg, display_term
if i < sources_n and term1_term and term1_term ~= "-" then
local link_page = get_link_page(term1_term, source, term1_sc)
display_term = (link_page ~= final_link_page) or (link_page and not not new_title(link_page):getContent())
end
-- TODO: if the display forms or transliterations are different, display the terms separately.
if display_term then
local display, this_cats = export.insert_source_cat_get_display{
lang = lang,
source = source,
borrowing_type = borrowing_type,
raw = true,
categories = categories,
nocat = nocat,
}
seg = language_link {
lang = source,
term = term1_term,
alt = display,
tr = "-",
}
if lang and not nocat then
-- Format categories, but only if there is a current language; {{cog}} currently gets no categories
this_cats = format_categories(this_cats, lang, sort_key, nil, data.force_cat or force_cat)
else
this_cats = ""
end
seg = "<span class=\"etyl\">" .. seg .. this_cats .. "</span>"
else
seg = export.format_source{
lang = lang,
source = source,
borrowing_type = borrowing_type,
sort_key = sort_key,
categories = categories,
nocat = nocat,
}
end
insert(source_segs, seg)
end
return join_segs(source_segs, data.sourceconj or "and")
end
-- Internal implementation of {{cognate}}/{{cog}} template.
function export.format_cognate(data)
return export.format_derived {
sources = data.sources,
terms = data.terms,
sort_key = data.sort_key,
sourceconj = data.sourceconj,
conj = data.conj,
template_name = "cognate",
force_cat = data.force_cat,
}
end
--[==[
Internal implementation of {{derived}}/{{der}} template. This is called externally from [[Module:affix]],
[[Module:affixusex]] and [[Module:see]] and needs to support decorations (qualifiers, labels and references) on the
outside of the sources for use by those modules.
`data` contains the following fields:
* `lang`: The destination language object into which the terms were derived. Used for categorization and can be nil, as
with {{tl|cog}}; in this case, no categories are added.
* `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are
handled specially; see `format_sources()`.
* `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as
multiple term objects, the non-final source objects link to the first term object.
* `conj`: Conjunction used to separate multiple terms. '''Required'''. Currently recognized values are `and`, `or`, `,`,
`;`, `/` and `~`.
* `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized
values are as for `conj` above.
* `decorations_on_outside`: If specified, any decorations (qualifiers, labels or references) in the first term in
`terms` will be displayed on the outside of (before) the source language(s) in `sources`. Normally this should be
specified if there is only one term possible in `terms`.
* `template_name`: Name of the template invoking this function. Must be specified. Only used for tracking pages.
* `sort_key`: Sort key for categories. Usually nil.
* `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source
languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be
added.
* `nocat`: Don't add any categories to the page.
* `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}.
* `force_cat`: Force category generation on non-mainspace pages.
]==]
function export.format_derived(data)
local terms = data.terms
local sourcetext = export.format_sources(data)
return export.format_links(terms, data.conj, data.template_name, sourcetext, data.decorations_on_outside)
end
function export.insert_borrowed_cat(categories, lang, source)
if lang_is_source(lang, source) then
return
end
-- If both are the same, we want e.g. [[:Category:English terms borrowed back into English]] not
-- [[:Category:English terms borrowed from English]]; the former is inserted automatically by format_source().
-- The second parameter here doesn't matter as it only affects `display`, which we don't use.
insert(categories, "Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada " .. select(2, export.get_display_and_cat_name(source, "raw")))
end
-- Internal implementation of {{borrowed}}/{{bor}} template.
function export.format_borrowed(data)
local categories = {}
if not data.nocat then
local lang = data.lang
for _, source in ipairs(data.sources) do
export.insert_borrowed_cat(categories, lang, source)
end
end
data = shallow_copy(data)
data.categories = categories
return export.format_links(data.terms, data.conj, "borrowed", export.format_sources(data))
end
do
-- Generate the non-ancestor error message.
local function show_language(lang)
local retval = ("%s (%s)"):format(lang:makeCategoryLink(), lang:getCode())
if lang:hasType("etymology-only") then
retval = retval .. (" (an etymology-only language whose regular parent is %s)"):format(
show_language(lang:getParent()))
end
return retval
end
-- Check that `lang` has `otherlang` (which may be an etymology-only language) as an ancestor. Throw an error if
-- not. When `lang` is a family, verifies that `otherlang` is a language in that family.
function export.check_ancestor(lang, otherlang)
-- When `lang` is a family, verify `otherlang` is in that family or in its parent family.
if lang.hasType and lang:hasType("family") then
local family_code = lang:getCode()
local function in_family_code(fcode, other)
if not fcode or fcode == "" then return false end
if other.inFamily and other:inFamily(fcode) then return true end
if other.getFamilyCode and other:getFamilyCode() == fcode then return true end
return false
end
local in_family = in_family_code(family_code, otherlang)
if not in_family then
local parent_code
if lang.getParent then
local parent_family = lang:getParent()
if parent_family and parent_family.getCode then
parent_code = parent_family:getCode()
end
end
if not parent_code and family_code:find("-", 1, true) then
parent_code = family_code:match("^(.+)-[^-]+$")
end
if parent_code then
in_family = in_family_code(parent_code, otherlang)
end
end
if not in_family then
local other_display = (otherlang.getCanonicalName and otherlang:getCanonicalName()) or (otherlang.getCode and otherlang:getCode()) or tostring(otherlang)
local fam_display = (lang.getCanonicalName and lang:getCanonicalName()) or family_code
error(("%s bukan dalam keluarga %s; leluhur diwarisi di bawah keluarga mestilah bahasa dalam keluarga tersebut atau keluarga induknya.")
:format(other_display, fam_display))
end
return
end
-- FIXME: I don't know if this function works correctly with etym-only languages in `lang`. I have fixed up
-- the module link code appropriately (June 2024) but the remaining logic is untouched.
if lang:hasAncestor(otherlang) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/variety]]
-- Track inheritance from varieties of Latin that shouldn't have any descendants (everything except Old Latin, Classical Latin and Vulgar Latin).
if otherlang:getFullCode() == "la" then
otherlang = otherlang:getCode()
if not (otherlang == "itc-ola" or otherlang == "la-cla" or otherlang == "la-vul") then
track("bad ancestor", otherlang)
end
end
return
end
local ancestors = lang:getAncestors()
local postscript
local etym_module_link = lang:hasType("etymology-only") and "[[Module:etymology languages/data]] or " or ""
local module_link = "[[" .. get_language_data_module_name(lang:getFullCode()) .. "]]"
if not ancestors[1] then
postscript = show_language(lang) .. " tidak mempunyai leluhur."
else
local ancestor_list = {}
for _, ancestor in ipairs(ancestors) do
insert(ancestor_list, show_language(ancestor))
end
postscript = ("Leluhur bahasa%s kepada %s %s %s."):format(
ancestors[2] and "" or "", lang:getCanonicalName(),
ancestors[2] and "adalah" or "adalah", concat(ancestor_list, " dan "))
end
error(("%s tidak ditetapkan sebagai luluhur kepada %s dalam %s%s. %s")
:format(show_language(otherlang), show_language(lang), etym_module_link, module_link, postscript))
end
end
-- Internal implementation of {{inherited}}/{{inh}} template.
function export.format_inherited(data)
local lang, terms, nocat = data.lang, data.terms, data.nocat
local source = terms[1].lang
local categories = {}
if not nocat then
insert(categories,"Perkataan bahasa " .. lang:getFullName() .. " diwariskan daripada bahasa " .. source:getCanonicalName())
end
export.check_ancestor(lang, source)
data = shallow_copy(data)
data.categories = categories
data.source = source
return export.format_links(terms, data.conj, "inherited", export.format_source(data))
end
-- Internal implementation of "misc variant" templates such as {{abbrev}}, {{clipping}}, {{reduplication}} and the like.
function export.format_misc_variant(data)
local lang, notext, terms, cats, parts = data.lang, data.notext, data.terms, data.cats, {}
if not notext then
insert(parts, data.text)
end
if terms[1] then
if not notext then
-- FIXME: If term is given as '-', we should consider displaying just "Clipping" not "Clipping of".
insert(parts, " " .. (data.oftext or "bagi"))
end
local termparts = {}
-- Make links out of all the parts.
for _, termobj in ipairs(terms) do
local result
if termobj.lang then
result = export.format_derived {
lang = lang,
terms = {termobj},
sources = termobj.termlangs or {termobj.lang},
template_name = "misc_variant",
decorations_on_outside = true,
force_cat = data.force_cat,
}
else
termobj.lang = lang
result = export.format_links({termobj}, nil, "misc_variant")
end
table.insert(termparts, result)
end
local linktext = join_segs(termparts, data.conj)
if not notext and linktext ~= "" then
insert(parts, " ")
end
insert(parts, linktext)
end
local categories = {}
if not data.nocat and cats then
for _, cat in ipairs(cats) do
insert(categories, cat .. " bahasa " .. lang:getFullName())
end
end
if categories[1] then
insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat))
end
return concat(parts)
end
-- Implementation of miscellaneous templates such as {{unknown}} and {{onomatopoeia}} that have no associated terms.
function export.format_misc_variant_no_term(data)
local parts = {}
if not data.notext then
insert(parts, data.title)
end
if not data.nocat and data.cat then
local lang, categories = data.lang, {}
insert(categories, data.cat .. " bahasa " .. lang:getFullName())
insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat))
end
return concat(parts)
end
return export
sp30xyxjrn9ihshbbbmeckrd2czv8nv
375373
375371
2026-09-22T05:18:56Z
Hakimi97
2668
tambah "bahasa" untuk kategori perkataan bahasa A dipinjam daripada "bahasa" B
375373
Scribunto
text/plain
local export = {}
-- For testing
local force_cat = false
local debug_track_module = "Module:debug/track"
local languages_module = "Module:languages"
local links_module = "Module:links"
local table_module = "Module:table"
local utilities_module = "Module:utilities"
local concat = table.concat
local insert = table.insert
local new_title = mw.title.new
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function format_categories(...)
format_categories = require(utilities_module).format_categories
return format_categories(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_language_data_module_name(...)
get_language_data_module_name = require(languages_module).getDataModuleName
return get_language_data_module_name(...)
end
local function get_link_page(...)
get_link_page = require(links_module).get_link_page
return get_link_page(...)
end
local function language_link(...)
language_link = require(links_module).language_link
return language_link(...)
end
local function serial_comma_join(...)
serial_comma_join = require(table_module).serialCommaJoin
return serial_comma_join(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function track(page, code)
local tracking_page = "etymology/" .. page
debug_track(tracking_page)
if code then
debug_track(tracking_page .. "/" .. code)
end
end
local function join_segs(segs, conj)
if not segs[2] then
return segs[1]
elseif conj == "and" or conj == "or" then
return serial_comma_join(segs, {conj = conj})
end
local sep
if conj == "," or conj == ";" then
sep = conj .. " "
elseif conj == "/" then
sep = "/"
elseif conj == "~" then
sep = " ~ "
elseif conj then
error(("Internal error: Unrecognized conjunction \"%s\""):format(conj))
else
error(("Internal error: No value supplied for conjunction"):format(conj))
end
return concat(segs, sep)
end
-- Returns true if `lang` is the same as `source`, or a variety of it.
local function lang_is_source(lang, source)
return lang:getCode() == source:getCode() or lang:hasParent(source)
end
--[==[
Format one or more links as specified in `termobjs`, a list of term objects of the format accepted by `full_link()` in
[[Module:links]], including decorations (qualifiers, labels and references). `conj` is used to join multiple terms
and must be specified if there is more than one term. `template_name` is the template name used in debug tracking and
must be specified. Optional `sourcetext` is text to prepend to the concatenated terms, separated by a space if the
concatenated terms are non-empty (which is always the case unless there is a single term with the value "-"). If
`decorations_on_outside` is given, any decorations specified in the first term go on the outside of (i.e before)
`sourcetext`; otherwise they will end up on the inside.
]==]
function export.format_links(termobjs, conj, template_name, sourcetext, decorations_on_outside)
if not template_name then
error("Internal error: Must specify `template_name` to format_links()")
end
for i, termobj in ipairs(termobjs) do
if termobj.lang:hasType("family") or termobj.lang:getFamilyCode() == "qfa-sub" then
if termobj.term and termobj.term ~= "-" then
debug_track(template_name .. "/family-with-term")
end
termobj.term = "-"
end
if termobj.term == "-" then
--[=[
[[Special:WhatLinksHere/Wiktionary:Tracking/cognate/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/derived/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/borrowed/no-term]]
[[Special:WhatLinksHere/Wiktionary:Tracking/calque/no-term]]
]=]
debug_track(template_name .. "/no-term")
termobjs[i] = i == 1 and sourcetext or ""
else
if i == 1 and decorations_on_outside and sourcetext then
termobj.pretext = sourcetext .. " "
sourcetext = nil
end
termobjs[i] = (i == 1 and sourcetext and sourcetext .. " " or "") .. full_link(termobj, "term")
end
end
return join_segs(termobjs, conj)
end
function export.get_display_and_cat_name(source, raw)
local display, cat_name
if source:getCode() == "und" then
display = "tidak ditentukan"
cat_name = "bahasa lain"
elseif source:getCode() == "mul" then
display = raw and "rentas bahasa" or "[[w:Translingualisme|rentas bahasa]]"
cat_name = "Rentas bahasa"
elseif source:getCode() == "mul-tax" then
display = raw and "taxonomic name" or "[[w:Tatanama biologi|nama taksonomi]]"
cat_name = "Nama taksonomi"
else
display = raw and source:getCanonicalName() or source:makeWikipediaLink()
cat_name = source:getDisplayForm()
end
return display, cat_name
end
function export.insert_source_cat_get_display(data)
local categories, lang, source = data.categories, data.lang, data.source
local display, cat_name = export.get_display_and_cat_name(source, data.raw)
if lang and not data.nocat then
-- Add the category, but only if there is a current language
if not categories then
categories = {}
end
local langname = lang:getFullName()
-- If `lang` is an etym-only language, we need to check both it and its parent full language against `source`.
-- Otherwise if e.g. `lang` is Medieval Latin and `source` is Latin, we'll end up wrongly constructing a
-- category 'Latin terms derived from Latin'.
insert(categories, "Perkataan bahasa " .. langname .. (
lang_is_source(lang, source) and " dipinjam balik ke dalam bahasa " .. cat_name or
" " .. (data.borrowing_type or "diterbitkan") .. " daripada bahasa " .. cat_name
))
end
return display, categories
end
function export.format_source(data)
local lang, sort_key = data.lang, data.sort_key
-- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/sortkey]]
if sort_key then
track("sortkey")
end
local display, categories = export.insert_source_cat_get_display(data)
if lang and not data.nocat then
-- Format categories, but only if there is a current language; {{cog}} currently gets no categories
categories = format_categories(categories, lang, sort_key, nil, data.force_cat or force_cat)
else
categories = ""
end
return "<span class=\"etyl\">" .. display .. categories .. "</span>"
end
--[==[
Format sources for etymology templates such as {{tl|bor}}, {{tl|der}}, {{tl|inh}} and {{tl|cog}}. There may potentially
be more than one source language (except currently {{tl|inh}}, which doesn't support it because it doesn't really
make sense). In that case, all but the last source language is linked to the first term, but only if there is such a
term and this linking makes sense, i.e. either (1) the term page exists after stripping diacritics according to the
source language in question, or (2) the result of stripping diacritics according to the source language in question
results in a different page from the same process applied with the last source language. For example, {{m|ru|соля́нка}}
will link to [[солянка]] but {{m|en|соля́нка}} will link to [[соля́нка]] with an accent, and since they are different
pages, the use of English as a non-final source with term 'соля́нка' will link to [[соля́нка]] even though it doesn't
exist, on the assumption that it is merely a redlink that might exist. If none of the above criteria apply, a non-final
source language will be linked to the Wikipedia entry for the language, just as final source languages always are.
`data` contains the following fields:
* `lang`: The destination language object into which the terms were borrowed, inherited or otherwise derived. Used for
categorization and can be nil, as with {{tl|cog}}.
* `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are
handled specially; see above.
* `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as
multiple term objects, the non-final source objects link to the first term object.
* `sort_key`: Sort key for categories. Usually nil.
* `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source
languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be
added.
* `nocat`: Don't add any categories to the page.
* `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized
values are `and`, `or`, `,`, `;`, `/` and `~`.
* `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}.
* `force_cat`: Force category generation on non-mainspace pages.
]==]
function export.format_sources(data)
local lang, sources, terms, borrowing_type, sort_key, categories, nocat =
data.lang, data.sources, data.terms, data.borrowing_type, data.sort_key, data.categories, data.nocat
local term1, sources_n, source_segs = terms[1], #sources, {}
local final_link_page
local term1_term, term1_sc = term1.term, term1.sc
if sources_n > 1 and term1_term and term1_term ~= "-" then
final_link_page = get_link_page(term1_term, sources[sources_n], term1_sc)
end
for i, source in ipairs(sources) do
local seg, display_term
if i < sources_n and term1_term and term1_term ~= "-" then
local link_page = get_link_page(term1_term, source, term1_sc)
display_term = (link_page ~= final_link_page) or (link_page and not not new_title(link_page):getContent())
end
-- TODO: if the display forms or transliterations are different, display the terms separately.
if display_term then
local display, this_cats = export.insert_source_cat_get_display{
lang = lang,
source = source,
borrowing_type = borrowing_type,
raw = true,
categories = categories,
nocat = nocat,
}
seg = language_link {
lang = source,
term = term1_term,
alt = display,
tr = "-",
}
if lang and not nocat then
-- Format categories, but only if there is a current language; {{cog}} currently gets no categories
this_cats = format_categories(this_cats, lang, sort_key, nil, data.force_cat or force_cat)
else
this_cats = ""
end
seg = "<span class=\"etyl\">" .. seg .. this_cats .. "</span>"
else
seg = export.format_source{
lang = lang,
source = source,
borrowing_type = borrowing_type,
sort_key = sort_key,
categories = categories,
nocat = nocat,
}
end
insert(source_segs, seg)
end
return join_segs(source_segs, data.sourceconj or "and")
end
-- Internal implementation of {{cognate}}/{{cog}} template.
function export.format_cognate(data)
return export.format_derived {
sources = data.sources,
terms = data.terms,
sort_key = data.sort_key,
sourceconj = data.sourceconj,
conj = data.conj,
template_name = "cognate",
force_cat = data.force_cat,
}
end
--[==[
Internal implementation of {{derived}}/{{der}} template. This is called externally from [[Module:affix]],
[[Module:affixusex]] and [[Module:see]] and needs to support decorations (qualifiers, labels and references) on the
outside of the sources for use by those modules.
`data` contains the following fields:
* `lang`: The destination language object into which the terms were derived. Used for categorization and can be nil, as
with {{tl|cog}}; in this case, no categories are added.
* `sources`: List of source objects. Most commonly there is only one. If there are multiple, the non-final ones are
handled specially; see `format_sources()`.
* `terms`: List of term objects. Most commonly there is only one. If there are multiple source objects as well as
multiple term objects, the non-final source objects link to the first term object.
* `conj`: Conjunction used to separate multiple terms. '''Required'''. Currently recognized values are `and`, `or`, `,`,
`;`, `/` and `~`.
* `sourceconj`: Conjunction used to separate multiple source languages. Defaults to {"and"}. Currently recognized
values are as for `conj` above.
* `decorations_on_outside`: If specified, any decorations (qualifiers, labels or references) in the first term in
`terms` will be displayed on the outside of (before) the source language(s) in `sources`. Normally this should be
specified if there is only one term possible in `terms`.
* `template_name`: Name of the template invoking this function. Must be specified. Only used for tracking pages.
* `sort_key`: Sort key for categories. Usually nil.
* `categories`: Categories to add to the page. Additional categories may be added to `categories` based on the source
languages ('''in which case `categories` is destructively modified'''). If `lang` is nil, no categories will be
added.
* `nocat`: Don't add any categories to the page.
* `borrowing_type`: Borrowing type used in categories, such as {"learned borrowings"}. Defaults to {"terms derived"}.
* `force_cat`: Force category generation on non-mainspace pages.
]==]
function export.format_derived(data)
local terms = data.terms
local sourcetext = export.format_sources(data)
return export.format_links(terms, data.conj, data.template_name, sourcetext, data.decorations_on_outside)
end
function export.insert_borrowed_cat(categories, lang, source)
if lang_is_source(lang, source) then
return
end
-- If both are the same, we want e.g. [[:Category:English terms borrowed back into English]] not
-- [[:Category:English terms borrowed from English]]; the former is inserted automatically by format_source().
-- The second parameter here doesn't matter as it only affects `display`, which we don't use.
insert(categories, "Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada bahasa " .. select(2, export.get_display_and_cat_name(source, "raw")))
end
-- Internal implementation of {{borrowed}}/{{bor}} template.
function export.format_borrowed(data)
local categories = {}
if not data.nocat then
local lang = data.lang
for _, source in ipairs(data.sources) do
export.insert_borrowed_cat(categories, lang, source)
end
end
data = shallow_copy(data)
data.categories = categories
return export.format_links(data.terms, data.conj, "borrowed", export.format_sources(data))
end
do
-- Generate the non-ancestor error message.
local function show_language(lang)
local retval = ("%s (%s)"):format(lang:makeCategoryLink(), lang:getCode())
if lang:hasType("etymology-only") then
retval = retval .. (" (an etymology-only language whose regular parent is %s)"):format(
show_language(lang:getParent()))
end
return retval
end
-- Check that `lang` has `otherlang` (which may be an etymology-only language) as an ancestor. Throw an error if
-- not. When `lang` is a family, verifies that `otherlang` is a language in that family.
function export.check_ancestor(lang, otherlang)
-- When `lang` is a family, verify `otherlang` is in that family or in its parent family.
if lang.hasType and lang:hasType("family") then
local family_code = lang:getCode()
local function in_family_code(fcode, other)
if not fcode or fcode == "" then return false end
if other.inFamily and other:inFamily(fcode) then return true end
if other.getFamilyCode and other:getFamilyCode() == fcode then return true end
return false
end
local in_family = in_family_code(family_code, otherlang)
if not in_family then
local parent_code
if lang.getParent then
local parent_family = lang:getParent()
if parent_family and parent_family.getCode then
parent_code = parent_family:getCode()
end
end
if not parent_code and family_code:find("-", 1, true) then
parent_code = family_code:match("^(.+)-[^-]+$")
end
if parent_code then
in_family = in_family_code(parent_code, otherlang)
end
end
if not in_family then
local other_display = (otherlang.getCanonicalName and otherlang:getCanonicalName()) or (otherlang.getCode and otherlang:getCode()) or tostring(otherlang)
local fam_display = (lang.getCanonicalName and lang:getCanonicalName()) or family_code
error(("%s bukan dalam keluarga %s; leluhur diwarisi di bawah keluarga mestilah bahasa dalam keluarga tersebut atau keluarga induknya.")
:format(other_display, fam_display))
end
return
end
-- FIXME: I don't know if this function works correctly with etym-only languages in `lang`. I have fixed up
-- the module link code appropriately (June 2024) but the remaining logic is untouched.
if lang:hasAncestor(otherlang) then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/etymology/variety]]
-- Track inheritance from varieties of Latin that shouldn't have any descendants (everything except Old Latin, Classical Latin and Vulgar Latin).
if otherlang:getFullCode() == "la" then
otherlang = otherlang:getCode()
if not (otherlang == "itc-ola" or otherlang == "la-cla" or otherlang == "la-vul") then
track("bad ancestor", otherlang)
end
end
return
end
local ancestors = lang:getAncestors()
local postscript
local etym_module_link = lang:hasType("etymology-only") and "[[Module:etymology languages/data]] or " or ""
local module_link = "[[" .. get_language_data_module_name(lang:getFullCode()) .. "]]"
if not ancestors[1] then
postscript = show_language(lang) .. " tidak mempunyai leluhur."
else
local ancestor_list = {}
for _, ancestor in ipairs(ancestors) do
insert(ancestor_list, show_language(ancestor))
end
postscript = ("Leluhur bahasa%s kepada %s %s %s."):format(
ancestors[2] and "" or "", lang:getCanonicalName(),
ancestors[2] and "adalah" or "adalah", concat(ancestor_list, " dan "))
end
error(("%s tidak ditetapkan sebagai luluhur kepada %s dalam %s%s. %s")
:format(show_language(otherlang), show_language(lang), etym_module_link, module_link, postscript))
end
end
-- Internal implementation of {{inherited}}/{{inh}} template.
function export.format_inherited(data)
local lang, terms, nocat = data.lang, data.terms, data.nocat
local source = terms[1].lang
local categories = {}
if not nocat then
insert(categories,"Perkataan bahasa " .. lang:getFullName() .. " diwariskan daripada bahasa " .. source:getCanonicalName())
end
export.check_ancestor(lang, source)
data = shallow_copy(data)
data.categories = categories
data.source = source
return export.format_links(terms, data.conj, "inherited", export.format_source(data))
end
-- Internal implementation of "misc variant" templates such as {{abbrev}}, {{clipping}}, {{reduplication}} and the like.
function export.format_misc_variant(data)
local lang, notext, terms, cats, parts = data.lang, data.notext, data.terms, data.cats, {}
if not notext then
insert(parts, data.text)
end
if terms[1] then
if not notext then
-- FIXME: If term is given as '-', we should consider displaying just "Clipping" not "Clipping of".
insert(parts, " " .. (data.oftext or "bagi"))
end
local termparts = {}
-- Make links out of all the parts.
for _, termobj in ipairs(terms) do
local result
if termobj.lang then
result = export.format_derived {
lang = lang,
terms = {termobj},
sources = termobj.termlangs or {termobj.lang},
template_name = "misc_variant",
decorations_on_outside = true,
force_cat = data.force_cat,
}
else
termobj.lang = lang
result = export.format_links({termobj}, nil, "misc_variant")
end
table.insert(termparts, result)
end
local linktext = join_segs(termparts, data.conj)
if not notext and linktext ~= "" then
insert(parts, " ")
end
insert(parts, linktext)
end
local categories = {}
if not data.nocat and cats then
for _, cat in ipairs(cats) do
insert(categories, cat .. " bahasa " .. lang:getFullName())
end
end
if categories[1] then
insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat))
end
return concat(parts)
end
-- Implementation of miscellaneous templates such as {{unknown}} and {{onomatopoeia}} that have no associated terms.
function export.format_misc_variant_no_term(data)
local parts = {}
if not data.notext then
insert(parts, data.title)
end
if not data.nocat and data.cat then
local lang, categories = data.lang, {}
insert(categories, data.cat .. " bahasa " .. lang:getFullName())
insert(parts, format_categories(categories, lang, data.sort_key, nil, data.force_cat or force_cat))
end
return concat(parts)
end
return export
bbu26wo843f49ze9uihvwyyik4ah21g
Modul:syllables
828
11549
375369
134546
2026-09-22T04:51:22Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91137329|91137329]])
375369
Scribunto
text/plain
local export = {}
local m_str_utils = require("Module:string utilities")
local gsub = m_str_utils.gsub
local match = m_str_utils.match
local toNFD = mw.ustring.toNFD
local U = m_str_utils.char
local diphthongs = mw.loadData("Module:IPA/data").diphthongs
local vowels = mw.loadData("Module:IPA/data/symbols").vowels .. "ᵻ" .. "ᵿ"
--[[ No use for this at the moment, though it is an interesting catalogue.
It might be usable for phonetic transcriptions.
Diacritics added to vowels:
inverted breve above, inverted breve below,
up tack, down tack,
left tack, right tack,
diaeresis (above), diaeresis below,
right half ring, left half ring,
plus sign below, minus sign below,
combining x above, rhotic hook,
tilde (above), tilde below
ligature tie (combining double breve), ligature tie below
]]
local diacritics = U(
0x311, 0x32F,
0x31D, 0x31E,
0x318, 0x319,
0x308, 0x324,
0x339, 0x31C,
0x31F, 0x320,
0x33D, 0x2DE,
0x303, 0x330,
0x361, 0x35C
)
--[[
combining acute and grave tone marks, circumflex
]]--
local tone = "[" .. U(0x341, 0x340, 0x302) .. "]"
local nonsyllabicDiacritics = U(0x311, 0x32F)
local syllabicDiacritics = U(0x0329, 0x030D)
local ties = U(0x361, 0x35C)
-- long, half-long, extra short
local lengthDiacritics = U(0x2D0, 0x2D1, 0x306)
local vowel = "[" .. vowels .. "]" .. tone .. "?"
local tie = "[" .. ties .. "]"
local nonsyllabicDiacritic = "[" .. nonsyllabicDiacritics .. "]"
local syllabicDiacritic = "[" .. syllabicDiacritics .. "]"
local UTF8Char = "[%z\1-\127\194-\244][\128-\191]*"
function export.getVowels(remainder, lang)
if string.find(remainder, "^[%[/]?%-") or string.find(remainder, "%-[%[/]?$") then
return nil
end -- If a hyphen is at the beginning or end of the transcription, do not count syllables.
local count = 0
local diphs = diphthongs[lang:getCode()] or {}
remainder = toNFD(remainder)
remainder = string.gsub(remainder, "%((.*)%)", "%1") -- Remove parentheses.
while remainder ~= "" do
-- Ignore nonsyllabic vowels
remainder = gsub(remainder, "^" .. vowel .. nonsyllabicDiacritic, "")
local m =
match(remainder, "^." .. syllabicDiacritic) or -- Syllabic consonant
match(remainder, "^" .. vowel .. tie .. vowel) -- Tie bar
-- Starts with a recognised diphthong?
for _, diph in ipairs(diphs) do
if m then
break
end
m = m or match(remainder, "^" .. diph)
end
-- If we haven't found anything yet, just match on a single vowel
m = m or match(remainder, "^" .. vowel)
if m then
-- Found a vowel, add it
count = count + 1
remainder = string.sub(remainder, #m + 1)
else
-- Found a non-vowel, skip it
remainder = string.gsub(remainder, "^" .. UTF8Char, "")
end
end
if count ~= 0 then return count end
return nil
end
function export.countVowels2Test(frame)
local params = {
[1] = {required = true},
[2] = {default = ""},
}
local args = require("Module:parameters").process(frame.args, params)
local lang = require("Module:languages").getByCode(args[1]) or require("Module:languages").err(args[1], 1)
local count = export.getVowels(args[2], lang)
return 'The text "' .. args[2] .. '" contains ' .. count .. ' vowels.'
end
local function countVowels(text)
text = toNFD(text) or error("Invalid UTF-8")
local _, count = gsub(text, vowel, "")
local _, sequenceCount = gsub(text, vowel.."+", "")
local _, nonsyllabicCount = gsub(text, vowel .. nonsyllabicDiacritic, "")
local _, tieCount = gsub(text, vowel .. tie .. vowel, "")
local diphthongCount = count - (nonsyllabicCount + tieCount)
return count, sequenceCount, diphthongCount
end
local function countDiphthongs(text, lang)
text = toNFD(text) or error("Invalid UTF-8")
local diphthongs = diphthongs[lang:getCode()] or {}
local _, count
local total = 0
if diphthongs then
for i, diphthong in pairs(diphthongs) do
_, count = gsub(text, diphthong, "")
total = total + count
end
end
return total
end
function export.countVowels(frame)
local params = {
[1] = {default = ""},
}
local args = require("Module:parameters").process(frame.args, params)
local count, sequenceCount, diphthongCount = countVowels(args[1])
local outputs = {}
table.insert(outputs, (count or 'an unknown number of') .. ' vowels')
table.insert(outputs, (sequenceCount or 'an unknown number of') .. ' vowel sequences')
table.insert(outputs, (diphthongCount or 'an unknown number of') .. ' vowels or vowels and diphthongs')
return 'The text "' .. args[1] .. '" contains ' .. mw.text.listToText(outputs) .. "."
end
function export.countVowelsDiphthongs(frame)
local params = {
[1] = {required = true},
[2] = {default = ""},
}
local args = require("Module:parameters").process(frame.args, params)
local lang = require("Module:languages").getByCode(args[1]) or require("Module:languages").err(args[1], 1)
local vowels = countVowels(args[2])
local count = vowels - countDiphthongs(args[2], lang) or 0
local out = 'The text "' .. args[2] .. '" contains ' .. (count or 'an unknown number of')
if count == 1 then
out = out .. ' vowel or diphthong.'
else
out = out .. ' vowels or diphthongs.'
end
return out
end
return export
2vcs9nzyxjbyq8j9l53t5ul8l7zkkz7
Modul:IPA/data/symbols
828
11596
375368
218693
2026-09-22T04:49:53Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/91727113|91727113]])
375368
Scribunto
text/plain
local data = {}
--[=[ Valid IPA symbols.
Currently almost all values of "title" and "link" keys
are just the comments that were used in [[Module:IPA]].
The "link" fields should be checked (those that start with an uppercase letter are checked). ]=]
--[=[
local phones = {}
-- Vowels.
phones["i"] = {
close = true,
front = true,
unrounded = true,
vowel = true,
}
phones["e"] = {
["close-mid"] = true,
front = true,
unrounded = true,
vowel = true,
}
phones["ɛ"] = {
["open-mid"] = true,
front = true,
unrounded = true,
vowel = true,
}
phones["æ"] = {
["near-open"] = true,
front = true,
unrounded = true,
vowel = true,
}
phones["a"] = {
open = true,
front = true,
unrounded = true,
vowel = true,
}
phones["y"] = {
close = true,
front = true,
rounded = true,
vowel = true,
}
phones["ø"] = {
["close-mid"] = true,
front = true,
rounded = true,
vowel = true,
}
phones["œ"] = {
["open-mid"] = true,
front = true,
rounded = true,
vowel = true,
}
phones["ɶ"] = {
open = true,
front = true,
rounded = true,
vowel = true,
}
phones["ɪ"] = {
["near-close"] = true,
["near-front"] = true,
unrounded = true,
vowel = true,
}
phones["ʏ"] = {
["near-close"] = true,
["near-front"] = true,
rounded = true,
vowel = true,
}
phones["ɨ"] = {
close = true,
central = true,
unrounded = true,
vowel = true,
}
phones["ᵻ"] = {
["near-close"] = true,
central = true,
unrounded = true,
vowel = true,
}
phones["ɘ"] = {
["close-mid"] = true,
central = true,
unrounded = true,
vowel = true,
}
phones["ɜ"] = {
["open-mid"] = true,
central = true,
unrounded = true,
vowel = true,
}
phones["ɝ"] = {
rhotic = true,
["open-mid"] = true,
central = true,
unrounded = true,
vowel = true,
}
phones["ə"] = {
mid = true,
central = true,
vowel = true,
}
phones["ɚ"] = {
rhotic = true,
mid = true,
central = true,
vowel = true,
}
phones["ɐ"] = {
["near-open"] = true,
central = true,
vowel = true,
}
phones["ʉ"] = {
close = true,
central = true,
rounded = true,
vowel = true,
}
phones["ᵿ"] = {
["near-close"] = true,
central = true,
rounded = true,
vowel = true,
}
phones["ɵ"] = {
["close-mid"] = true,
central = true,
rounded = true,
vowel = true,
}
phones["ɞ"] = {
["open-mid"] = true,
central = true,
rounded = true,
vowel = true,
}
phones["ʊ"] = {
["near-close"] = true,
["near-back"] = true,
rounded = true,
vowel = true,
}
phones["ɯ"] = {
close = true,
back = true,
unrounded = true,
vowel = true,
}
phones["ɤ"] = {
["close-mid"] = true,
back = true,
unrounded = true,
vowel = true,
}
phones["ʌ"] = {
["open-mid"] = true,
back = true,
unrounded = true,
vowel = true,
}
phones["ɑ"] = {
open = true,
back = true,
unrounded = true,
vowel = true,
}
phones["u"] = {
close = true,
back = true,
rounded = true,
vowel = true,
}
phones["o"] = {
["close-mid"] = true,
back = true,
rounded = true,
vowel = true,
}
phones["ɔ"] = {
["open-mid"] = true,
back = true,
rounded = true,
vowel = true,
}
phones["ɒ"] = {
open = true,
back = true,
rounded = true,
vowel = true,
}
-- Nasals.
phones["m"] = {
voiced = true,
bilabial = true,
nasal = true,
}
phones["ɱ"] = {
voiced = true,
labiodental = true,
nasal = true,
}
phones["n"] = {
voiced = true,
alveolar = true,
nasal = true,
}
phones["ɳ"] = {
voiced = true,
retroflex = true,
nasal = true,
}
phones["ɲ"] = {
voiced = true,
palatal = true,
nasal = true,
}
phones["ŋ"] = {
voiced = true,
velar = true,
nasal = true,
}
phones["𝼇"] = {
voiced = true,
velodorsal = true,
nasal = true,
}
phones["ɴ"] = {
voiced = true,
uvular = true,
nasal = true,
}
-- Plosives.
phones["p"] = {
voiceless = true,
bilabial = true,
plosive = true,
}
phones["b"] = {
voiced = true,
bilabial = true,
plosive = true,
}
phones["t"] = {
voiceless = true,
alveolar = true,
plosive = true,
}
phones["d"] = {
voiced = true,
alveolar = true,
plosive = true,
}
phones["ʈ"] = {
voiceless = true,
retroflex = true,
plosive = true,
}
phones["ɖ"] = {
voiced = true,
retroflex = true,
plosive = true,
}
phones["c"] = {
voiceless = true,
palatal = true,
plosive = true,
}
phones["ɟ"] = {
voiced = true,
palatal = true,
plosive = true,
}
phones["k"] = {
voiceless = true,
velar = true,
plosive = true,
}
phones["ɡ"] = {
voiced = true,
velar = true,
plosive = true,
}
phones["𝼃"] = {
voiceless = true,
velodorsal = true,
plosive = true,
}
phones["𝼁"] = {
voiced = true,
velodorsal = true,
plosive = true,
}
phones["q"] = {
voiceless = true,
uvular = true,
plosive = true,
}
phones["ɢ"] = {
voiced = true,
uvular = true,
plosive = true,
}
phones["ꞯ"] = {
voiceless = true,
["upper-pharyngeal"] = true,
plosive = true,
}
phones["𝼂"] = {
voiced = true,
["upper-pharyngeal"] = true,
plosive = true,
}
phones["ʡ"] = {
epiglottal = true,
plosive = true,
}
phones["ʔ"] = {
glottal = true,
plosive = true,
}
-- Fricatives.
phones["ɸ"] = {
voiceless = true,
bilabial = true,
fricative = true,
}
phones["β"] = {
voiced = true,
bilabial = true,
fricative = true,
}
phones["ʍ"] = {
voiceless = true,
["labial-velar"] = true,
fricative = true,
}
phones["f"] = {
voiceless = true,
labiodental = true,
fricative = true,
}
phones["v"] = {
voiced = true,
labiodental = true,
fricative = true,
}
phones["θ"] = {
voiceless = true,
dental = true,
["non-sibilant"] = true,
fricative = true,
}
phones["ð"] = {
voiced = true,
dental = true,
["non-sibilant"] = true,
fricative = true,
}
phones["s"] = {
voiceless = true,
alveolar = true,
sibilant = true,
fricative = true,
}
phones["z"] = {
voiced = true,
alveolar = true,
sibilant = true,
fricative = true,
}
phones["ɬ"] = {
voiceless = true,
alveolar = true,
lateral = true,
fricative = true,
}
phones["ɮ"] = {
voiced = true,
alveolar = true,
lateral = true,
fricative = true,
}
phones["ʃ"] = {
voiceless = true,
postalveolar = true,
sibilant = true,
fricative = true,
}
phones["ʒ"] = {
voiced = true,
postalveolar = true,
sibilant = true,
fricative = true,
}
phones["ʂ"] = {
voiceless = true,
retroflex = true,
sibilant = true,
fricative = true,
}
phones["ʐ"] = {
voiced = true,
retroflex = true,
sibilant = true,
fricative = true,
}
phones["ꞎ"] = {
voiceless = true,
retroflex = true,
lateral = true,
fricative = true,
}
phones["𝼅"] = {
voiced = true,
retroflex = true,
lateral = true,
fricative = true,
}
phones["ɕ"] = {
voiceless = true,
["alveolo-palatal"] = true,
sibilant = true,
fricative = true,
}
phones["ʑ"] = {
voiced = true,
["alveolo-palatal"] = true,
sibilant = true,
fricative = true,
}
phones["ç"] = {
voiceless = true,
palatal = true,
fricative = true,
}
phones["ʝ"] = {
voiced = true,
palatal = true,
fricative = true,
}
phones["𝼆"] = {
voiceless = true,
palatal = true,
lateral = true,
fricative = true,
}
phones["ɧ"] = {
voiceless = true,
["palatal-velar"] = true,
fricative = true,
}
phones["x"] = {
voiceless = true,
velar = true,
fricative = true,
}
phones["ɣ"] = {
voiced = true,
velar = true,
fricative = true,
}
phones["𝼄"] = {
voiceless = true,
velar = true,
lateral = true,
fricative = true,
}
phones["ʩ"] = {
voiceless = true,
velopharyngeal = true,
fricative = true,
}
phones["χ"] = {
voiceless = true,
uvular = true,
fricative = true,
}
phones["ʁ"] = {
voiced = true,
uvular = true,
fricative = true,
}
phones["ħ"] = {
voiceless = true,
pharyngeal = true,
fricative = true,
}
phones["ʕ"] = {
voiced = true,
pharyngeal = true,
fricative = true,
}
phones["ʜ"] = {
voiceless = true,
epiglottal = true,
fricative = true,
}
phones["ʢ"] = {
voiced = true,
epiglottal = true,
fricative = true,
}
phones["h"] = {
voiceless = true,
glottal = true,
fricative = true,
}
phones["ɦ"] = {
voiced = true,
glottal = true,
fricative = true,
}
-- Approximants.
phones["ʋ"] = {
voiced = true,
labiodental = true,
approximant = true,
}
phones["ɥ"] = {
voiced = true,
["labial–palatal"] = true,
approximant = true,
}
phones["w"] = {
voiced = true,
["labial–velar"] = true,
approximant = true,
}
phones["ɹ"] = {
voiced = true,
alveolar = true,
approximant = true,
}
phones["ꭨ"] = {
["velarized or pharyngealized"] = true,
voiced = true,
alveolar = true,
approximant = true,
}
phones["l"] = {
voiced = true,
alveolar = true,
lateral = true,
approximant = true,
}
phones["ɫ"] = {
["velarized or pharyngealized"] = true,
voiced = true,
alveolar = true,
lateral = true,
approximant = true,
}
phones["ɻ"] = {
voiced = true,
retroflex = true,
approximant = true,
}
phones["ɭ"] = {
voiced = true,
retroflex = true,
lateral = true,
approximant = true,
}
phones["j"] = {
voiced = true,
palatal = true,
approximant = true,
}
phones["ʎ"] = {
voiced = true,
palatal = true,
lateral = true,
approximant = true,
}
phones["ɰ"] = {
voiced = true,
velar = true,
approximant = true,
}
phones["ʟ"] = {
voiced = true,
velar = true,
lateral = true,
approximant = true,
}
-- Flaps.
phones["ⱱ"] = {
voiced = true,
labiodental = true,
flap = true,
}
phones["ɾ"] = {
voiced = true,
alveolar = true,
flap = true,
}
phones["ɺ"] = {
voiced = true,
alveolar = true,
lateral = true,
flap = true,
}
phones["ɽ"] = {
voiced = true,
retroflex = true,
flap = true,
}
phones["𝼈"] = {
voiced = true,
retroflex = true,
lateral = true,
flap = true,
}
-- Trills.
phones["ʙ"] = {
voiced = true,
bilabial = true,
trill = true,
}
phones["r"] = {
voiced = true,
alveolar = true,
trill = true,
}
phones["𝼀"] = {
voiceless = true,
velopharyngeal = true,
trill = true,
}
phones["ʀ"] = {
voiced = true,
uvular = true,
trill = true,
}
phones["ᴙ"] = {
voiced = true,
pharyngeal = true,
trill = true,
}
-- Clicks.
phones["ʘ"] = {
bilabial = true,
click = true,
}
phones["ǀ"] = {
dental = true,
click = true,
}
phones["ǃ"] = {
alveolar = true,
click = true,
}
phones["𝼊"] = {
retroflex = true,
click = true,
}
phones["ǂ"] = {
palatal = true,
click = true,
}
phones["ʞ"] = {
velar = true,
click = true,
}
phones["ǁ"] = {
lateral = true,
click = true,
}
-- Implosives.
phones["ɓ"] = {
voiced = true,
bilabial = true,
implosive = true,
}
phones["ɗ"] = {
voiced = true,
alveolar = true,
implosive = true,
}
phones["ᶑ"] = {
voiced = true,
retroflex = true,
implosive = true,
}
phones["ʄ"] = {
voiced = true,
palatal = true,
implosive = true,
}
phones["ɠ"] = {
voiced = true,
velar = true,
implosive = true,
}
phones["ʛ"] = {
voiced = true,
uvular = true,
implosive = true,
}
-- Percussives.
phones["ʬ"] = {
bilabial = true,
percussive = true,
}
phones["ʭ"] = {
bidental = true,
percussive = true,
}
phones["¡"] = {
sublaminal = true,
["lower-alveolar"] = true,
percussive = true,
}
]=]
local u = require("Module:string/char")
data[1] = {
-- PULMONIC CONSONANTS
-- nasal
["m"] = {
title = "bilabial nasal",
link = "w:Bilabial nasal",
},
["ɱ"] = {
title = "labiodental nasal",
link = "w:Labiodental nasal",
},
["n"] = {
title = "alveolar nasal",
link = "w:Alveolar nasal",
},
["ɳ"] = {
title = "retroflex nasal",
link = "w:Retroflex nasal",
},
["ɲ"] = {
title = "palatal nasal",
link = "w:Palatal nasal",
},
["ŋ"] = {
title = "velar nasal",
link = "w:Velar nasal",
},
["ɴ"] = {
title = "uvular nasal",
link = "w:Uvular nasal",
},
-- plosive
["p"] = {
title = "voiceless bilabial plosive",
link = "w:Voiceless bilabial stop",
},
["b"] = {
title = "voiced bilabial plosive",
link = "w:Voiced bilabial stop",
},
["t"] = {
title = "voiceless alveolar plosive",
link = "w:Voiceless alveolar stop",
},
["d"] = {
title = "voiced alveolar plosive",
link = "w:Voiced alveolar stop",
},
["ʈ"] = {
title = "voiceless retroflex plosive",
link = "w:Voiceless retroflex stop",
},
["ɖ"] = {
title = "voiced retroflex plosive",
link = "w:Voiced retroflex stop",
},
["c"] = {
title = "voiceless palatal plosive",
link = "w:Voiceless palatal stop",
},
["ɟ"] = {
title = "voiced palatal plosive",
link = "w:Voiced palatal stop",
},
["k"] = {
title = "voiceless velar plosive",
link = "w:Voiceless velar stop",
},
["ɡ"] = {
title = "voiced velar plosive",
link = "w:Voiced velar stop",
},
["q"] = {
title = "voiceless uvular plosive",
link = "w:Voiceless uvular stop",
},
["ɢ"] = {
title = "voiced uvular plosive",
link = "w:Voiced uvular stop",
},
["ʡ"] = {
title = "epiglottal plosive",
link = "w:Epiglottal stop",
},
["ʔ"] = {
title = "glottal stop",
link = "w:Glottal stop",
},
-- fricative
["ɸ"] = {
title = "voiceless bilabial fricative",
link = "w:Voiceless bilabial fricative",
},
["β"] = {
title = "voiced bilabial fricative",
link = "w:Voiced bilabial fricative",
},
["f"] = {
title = "voiceless labiodental fricative",
link = "w:Voiceless labiodental fricative",
},
["v"] = {
title = "voiced labiodental fricative",
link = "w:Voiced labiodental fricative",
},
["θ"] = {
title = "voiceless dental fricative",
link = "w:Voiceless dental fricative",
},
["ð"] = {
title = "voiced dental fricative",
link = "w:Voiced dental fricative",
},
["s"] = {
title = "voiceless alveolar fricative",
link = "w:Voiceless alveolar fricative",
},
["z"] = {
title = "voiced alveolar fricative",
link = "w:Voiced alveolar fricative",
},
["ʃ"] = {
title = "voiceless postalveolar fricative",
link = "w:Voiceless palato-alveolar sibilant",
},
["ʒ"] = {
title = "voiced postalveolar fricative",
link = "w:Voiced palato-alveolar sibilant",
},
["ʂ"] = {
title = "voiceless retroflex fricative",
link = "w:Voiceless retroflex sibilant",
},
["ʐ"] = {
title = "voiced retroflex fricative",
link = "w:Voiced retroflex sibilant",
},
["ɕ"] = {
title = "voiceless alveolo-palatal fricative",
link = "w:Voiceless alveolo-palatal sibilant",
},
["ʑ"] = {
title = "voiced alveolo-palatal fricative",
link = "w:Voiced alveolo-palatal sibilant",
},
["ç"] = {
title = "voiceless palatal fricative",
link = "w:Voiceless palatal fricative",
},
["ʝ"] = {
title = "voiced palatal fricative",
link = "w:Voiced palatal fricative",
},
["x"] = {
title = "voiceless velar fricative",
link = "w:Voiceless velar fricative",
},
["ɣ"] = {
title = "voiced velar fricative",
link = "w:Voiced velar fricative",
},
["χ"] = {
title = "voiceless uvular fricative",
link = "w:Voiceless uvular fricative",
},
["ʁ"] = {
title = "voiced uvular fricative",
link = "w:Voiced uvular fricative",
},
["ħ"] = {
title = "voiceless pharyngeal fricative",
link = "w:Voiceless pharyngeal fricative",
},
["ʕ"] = {
title = "voiced pharyngeal fricative",
link = "w:Voiced pharyngeal fricative",
},
["ʜ"] = {
title = "voiceless epiglottal fricative",
link = "w:Voiceless epiglottal fricative",
},
["ʢ"] = {
title = "voiced epiglottal fricative",
link = "w:Voiced epiglottal fricative",
},
["h"] = {
title = "voiceless glottal fricative",
link = "w:Voiceless glottal fricative",
},
["ɦ"] = {
title = "voiced glottal fricative",
link = "w:Voiced glottal fricative",
},
-- approximant
["ʋ"] = {
title = "labiodental approximant",
link = "w:Labiodental approximant",
},
["ɹ"] = {
title = "alveolar approximant",
link = "w:Alveolar approximant",
},
["ɻ"] = {
title = "retroflex approximant",
link = "w:Retroflex approximant",
},
["j"] = {
title = "palatal approximant",
link = "w:Palatal approximant",
},
["ɰ"] = {
title = "velar approximant",
link = "w:Velar approximant",
},
-- tap, flap
["ⱱ"] = {
title = "labiodental tap",
link = "w:Labiodental flap",
},
["ɾ"] = {
title = "alveolar flap",
link = "w:Alveolar flap",
},
["ɽ"] = {
title = "retroflex flap",
link = "w:Retroflex flap",
},
-- trill
["ʙ"] = {
title = "bilabial trill",
link = "w:Bilabial trill",
},
["r"] = {
title = "alveolar trill",
link = "w:Alveolar trill",
},
["ʀ"] = {
title = "uvular trill",
link = "w:Uvular trill",
},
["ᴙ"] = {
title = "epiglottal trill",
link = "w:Epiglottal trill",
},
-- lateral fricative
["ɬ"] = {
title = "voiceless alveolar lateral fricative",
link = "w:Voiceless alveolar lateral fricative",
},
["ɮ"] = {
title = "voiced alveolar lateral fricative",
link = "w:Voiced alveolar lateral fricative",
},
-- no precomposed Unicode character --TOMOVE
--["ɬ̢"] = {title = "voiceless retroflex lateral fricative", link = "w:voiceless retroflex lateral fricative"},
-- no precomposed Unicode character --TOMOVE:3
--["ʎ̝̊"] = {title = "voiceless palatal lateral fricative", link = "w:voiceless palatal lateral fricative"},
-- no precomposed Unicode character --TOMOVE:3
--["ʟ̝̊"] = {title = "voiceless velar lateral fricative", link = "w:voiceless velar lateral fricative"},
-- no precomposed Unicode character --TOMOVE
--["ʟ̝"] = {title = "voiced velar lateral fricative", link = "w:voiced velar lateral fricative"},
-- lateral approximant
["l"] = {
title = "alveolar lateral approximant",
link = "w:Alveolar lateral approximant",
},
["ɭ"] = {
title = "retroflex lateral approximant",
link = "w:Retroflex lateral approximant",
},
["ʎ"] = {
title = "palatal lateral approximant",
link = "w:Palatal lateral approximant",
},
["ʟ"] = {
title = "velar lateral approximant",
link = "w:Velar lateral approximant",
},
-- lateral flap
["ɺ"] = {
title = "alveolar lateral flap",
link = "w:Alveolar lateral flap",
},
--["ɭ̆"] = {title = "retroflex lateral flap", link = "w:retroflex lateral flap"}, -- no precomposed Unicode character --TOMOVE
--["ɺ˞"] = {title = "retroflex lateral flap", link = "w:retroflex lateral flap"}, -- no precomposed Unicode character --TOMOVE
-- NON-PULMONIC CONSONANTS
-- clicks
["ʘ"] = {
title = "bilabial click",
link = "w:Bilabial clicks",
},
["ǀ"] = {
title = "dental click",
link = "w:Dental clicks",
},
["ǃ"] = {
title = "postalveolar click",
link = "w:Alveolar clicks",
},
["𝼊"] = {
title = "subapical retroflex",
link = "w:Retroflex clicks",
}, -- NOT IN X-SAMPA
["ǂ"] = {
title = "palatal click",
link = "w:Palatal clicks",
},
["ǁ"] = {
title = "alveolar lateral click",
link = "w:Lateral clicks",
},
-- implosives
["ɓ"] = {
title = "voiced bilabial implosive",
link = "w:Voiced bilabial implosive",
},
["ɗ"] = {
title = "voiced alveolar implosive",
link = "w:Voiced alveolar implosive",
},
-- NOT IN X-SAMPA
["ᶑ"] = {
title = "retroflex implosive",
link = "w:Voiced retroflex implosive",
},
["ʄ"] = {
title = "voiced palatal implosive",
link = "w:Voiced palatal implosive",
},
["ɠ"] = {
title = "voiced velar implosive",
link = "w:Voiced velar implosive",
},
["ʛ"] = {
title = "voiced uvular implosive",
link = "w:Voiced uvular implosive",
},
-- ejectives
["ʼ"] = {
title = "ejective",
link = "w:Ejective consonant",
},
-- CO-ARTICULATED CONSONANTS
["ʍ"] = {
title = "voiceless labial-velar fricative",
link = "w:Voiceless labio-velar approximant",
},
["w"] = {
title = "labial-velar approximant",
link = "w:Labio-velar approximant",
},
["ɥ"] = {
title = "labial-palatal approximant",
link = "w:Labialized palatal approximant",
},
["ɧ"] = {
title = "voiceless palatal-velar fricative",
link = "w:Sj-sound",
},
-- should be handled in [[Module:IPA]] and not through this table
-- BRACKETS
--[[
-- ["//"] = {
title = "morphophonemic",
link = "w:morphophonemic",
},
["/"] = {
title = "phonemic",
link = "w:phonemic",
},
["["] = {
title = "phonetic",
link = "w:phonetic",
},
["["] = {
title = "phonetic",
link = "w:phonetic",
},
["〈"] = {
title = "orthographic",
link = "w:orthographic",
},
["〉"] = {
title = "orthographic",
link = "w:orthographic",
},
["⟨"] = {
title = "orthographic",
link = "w:orthographic",
},
["⟩"] = {
title = "orthographic",
link = "w:orthographic",
},
]]
-- VOWELS
-- close
["i"] = {
title = "close front unrounded vowel",
link = "w:Close front unrounded vowel",
},
["y"] = {
title = "close front rounded vowel",
link = "w:Close front rounded vowel",
},
["ɨ"] = {
title = "close central unrounded vowel",
link = "w:Close central unrounded vowel",
},
["ʉ"] = {
title = "close central rounded vowel",
link = "w:Close central rounded vowel",
},
["ɯ"] = {
title = "close back unrounded vowel",
link = "w:Close back unrounded vowel",
},
["u"] = {
title = "close back rounded vowel",
link = "w:Close back rounded vowel",
},
-- near close
["ɪ"] = {
title = "near-close near-front unrounded vowel",
link = "w:Near-close near-front unrounded vowel",
},
["ʏ"] = {
title = "near-close near-front rounded vowel",
link = "w:Near-close near-front rounded vowel",
},
["ᵻ"] = {
title = "near-close central unrounded vowel",
link = "w:Near-close central unrounded vowel",
},
-- (alternative) --TOMOVE
--[[
["ɪ̈"] = {
title = "near-close central unrounded vowel",
link = "w:near-close central unrounded vowel",
}, ]]
["ᵿ"] = {
title = "near-close central rounded vowel",
link = "w:Near-close central rounded vowel",
},
--[[
(alternative) TOMOVE
["ʊ̈"] = {
title = "near-close central rounded vowel",
link = "w:near-close central rounded vowel",
},
]]
["ʊ"] = {
title = "near-close near-back rounded vowel",
link = "w:Near-close near-back rounded vowel",
},
--close mid
["e"] = {
title = "close-mid front unrounded vowel",
link = "w:Close-mid front unrounded vowel",
},
["ø"] = {
title = "close-mid front rounded vowel",
link = "w:Close-mid front rounded vowel",
},
["ɘ"] = {
title = "close-mid central unrounded vowel",
link = "w:Close-mid central unrounded vowel",
},
["ɵ"] = {
title = "close-mid central rounded vowel",
link = "w:Close-mid central rounded vowel",
},
["ɤ"] = {
title = "close-mid back unrounded vowel",
link = "w:Close-mid back unrounded vowel",
},
["o"] = {
title = "close-mid back rounded vowel",
link = "w:Close-mid back rounded vowel",
},
-- mid
["ə"] = {
title = "schwa",
link = "w:Schwa",
},
["ɚ"] = {
title = "schwa+r",
link = "w:R-colored vowel",
},
-- open mid
["ɛ"] = {
title = "open-mid front unrounded vowel",
link = "w:Open-mid front unrounded vowel",
},
["œ"] = {
title = "open-mid front rounded vowel",
link = "w:Open-mid front rounded vowel",
},
["ɜ"] = {
title = "open-mid central unrounded vowel",
link = "w:Open-mid central unrounded vowel",
},
["ɝ"] = {
title = "open-mid central unrounded vowel+r",
link = "w:R-colored vowel",
},
["ɞ"] = {
title = "open-mid central rounded vowel",
link = "w:Open-mid central rounded vowel",
},
["ʌ"] = {
title = "open-mid back unrounded vowel",
link = "w:Open-mid back unrounded vowel",
},
["ɔ"] = {
title = "open-mid back rounded vowel",
link = "w:Open-mid back rounded vowel",
},
-- near open
["æ"] = {
title = "near-open front unrounded vowel",
link = "w:Near-open front unrounded vowel",
},
["ɐ"] = {
title = "near-open central vowel",
link = "w:Near-open central vowel",
},
-- open
["a"] = {
title = "open front unrounded vowel",
link = "w:Open front unrounded vowel",
},
["ɶ"] = {
title = "open front rounded vowel",
link = "w:Open front rounded vowel",
},
["ɑ"] = {
title = "open back unrounded vowel",
link = "w:Open back unrounded vowel",
},
["ɒ"] = {
title = "open back rounded vowel",
link = "w:Open back rounded vowel",
},
-- SUPRASEGMENTALS
["ˈ"] = {title = "primary stress", link = "w:Stress (linguistics)", XSAMPA = "\""},
--[[
["???"] = {
title = "extra stress: no Unicode char; double primary stress instead",
link = "w:extra stress: no Unicode char; double primary stress instead",
XSAMPA = ""
}, --TOMOVE:3 ]]
["ˌ"] = {
title = "secondary stress",
link = "w:Secondary stress",
},
["ː"] = {
title = "long",
link = "w:Length (phonetics)",
},
["ˑ"] = {
title = "half long",
link = "w:Length (phonetics)",
},
["̆"] = {
title = "extra-short",
link = "w:Length (phonetics)",
},
--[[
["%."] = {
title = "syllable break",
link = "w:syllable break",
},
]]
--TOMOVE
["‿"] = {
title = "linking mark (absence of a break)",
link = "w:Tie (typography)#International_Phonetic_Alphabet",
},
[" "] = {
title = "separator",
link = "w:separator",
},
-- TONE
-- level tones
["˥"] = {
title = "top",
link = "w:Tone letter",
},
["˦"] = {
title = "high",
link = "w:Tone letter",
},
["˧"] = {
title = "mid",
link = "w:Tone letter",
},
["˨"] = {
title = "low",
link = "w:Tone letter",
},
["˩"] = {
title = "bottom",
link = "w:Tone letter",
},
["̋"] = {
title = "extra high tone",
link = "w:Tone letter",
},
["́"] = {
title = "high tone",
link = "w:Tone letter",
},
["̄"] = {
title = "mid tone",
link = "w:Tone letter",
},
["̀"] = {
title = "low tone",
link = "w:Tone letter",
},
["̏"] = {
title = "extra low tone",
link = "w:Tone letter",
},
-- tone terracing
["ꜛ"] = {
title = "upstep",
link = "w:Upstep",
},
["ꜜ"] = {
title = "downstep",
link = "w:Downstep",
},
-- contour tones
["̌"] = {
title = "rising tone",
link = "w:Tone (linguistics)",
},
["̂"] = {
title = "falling tone",
link = "w:Tone (linguistics)",
},
["᷄"] = {
title = "high rising tone",
link = "w:Tone (linguistics)",
},
["᷅"] = {
title = "low rising tone",
link = "w:Tone (linguistics)",
},
["᷇"] = {
title = "high falling tone",
link = "w:Tone (linguistics)",
},
["᷆"] = {
title = "low falling tone",
link = "w:Tone (linguistics)",
},
["᷈"] = {
title = "rising falling tone (peaking)",
link = "w:Tone (linguistics)",
},
["᷉"] = {
title = "dipping",
link = "w:Tone (linguistics)",
}, -- [extrapolated from the chart -- please confirm]
-- intonation
["|"] = {
title = "minor (foot) group",
link = "w:Prosodic unit",
},
["‖"] = {
title = "major (intonation) group",
link = "w:Prosodic unit",
},
["↗"] = {
title = "global rise",
link = "w:Intonation (linguistics)",
},
["↘"] = {
title = "global fall",
link = "w:Intonation (linguistics)",
},
-- DIACRITICS
-- syllabicity & releases
["̩"] = {
title = "syllabi ",
link = "w:Syllabic consonant",
withdescender = "̍"
}, -- (or "_="
["̯"] = {
title = "non-syllabic",
link = "w:Semivowel",
withdescender = "̑"
},
["ʰ"] = {
title = "aspirated",
link = "w:Aspirated consonant",
},
["ⁿ"] = {
title = "nasal release",
link = "w:Nasal release",
},
["ˡ"] = {
title = "lateral release",
link = "w:Lateral release (phonetics)",
},
["̚"] = {
title = "no audible release",
link = "w:No audible release",
},
-- phonation
["̥"] = {
title = "voiceless",
link = "w:Voicelessness",
withdescender = "̊"
},
["̬"] = {
title = "voiced",
link = "w:Voice (phonetics)",
},
["̤"] = {
title = "breathy voice",
link = "w:Breathy voice",
},
["̰"] = {
title = "creaky voice",
link = "w:Creaky voice",
},
["᷽"] = {
title = "strident",
link = "w:Strident vowel",
},
-- primary articulation
["̪"] = {
title = "dental",
link = "w:Dental consonant",
},
["̺"] = {
title = "apical",
link = "w:Apical consonant",
},
["̻"] = {
title = "laminal",
link = "w:Laminal consonant",
},
["̟"] = {
title = "advanced",
link = "w:Relative articulation#Advanced_and_retracted",
withdescender = "˖"
},
["̠"] = {
title = "retracted",
link = "w:Relative articulation#Retracted",
withdescender = "˗"
},
["̼"] = {
title = "linguolabial",
link = "w:Linguolabial consonant",
},
["̈"] = {
title = "centralized",
link = "w:Relative articulation#Centralized_vowels",
XSAMPA = "_\""
},
["̽"] = {
title = "mid-centralized",
link = "Relative articulation#Mid-centralized_vowel",
},
["̞"] = {
title = "lowered",
link = "w:Relative articulation#Raised_and_lowered",
withdescender = "˕"
},
["̝"] = {
title = "raised",
link = "w:Relative articulation#Raised_and_lowered",
withdescender = "˔"
},
["͡"] = {
title = "coarticulated",
link = "w:Co-articulated consonant",
},
["͈"] = {
title = "strong articulation",
link = "w:Fortis and lenis",
},
-- secondary articulation
["ʷ"] = {
title = "labialized",
link = "w:Labialization",
},
["ʲ"] = {
title = "palatalized",
link = "w:Palatalization (phonetics)",
},
["ˠ"] = {
title = "velarized",
link = "w:Velarization",
},
["ˤ"] = {
title = "pharyngealized",
link = "w:Pharyngealization",
},
-- also see _e
["ɫ"] = {
title = "velarized alveolar lateral approximant",
link = "w:Alveolar lateral approximant",
},
["̴"] = {
title = "velarized or pharyngealized; also see 5",
link = "w:Velarization",
},
["̹"] = {
title = "more rounded",
link = "w:Roundedness",
},
["̜"] = {
title = "less rounded",
link = "w:Roundedness",
},
["̃"] = {
title = "nasalization",
link = "w:Nasalization",
},
["˞"] = {
title = "rhotacization in vowels, retroflexion in consonants",
link = "w:R-colored vowel",
},
["̘"] = {
title = "advanced tongue root",
link = "w:Advanced and retracted tongue root",
},
["̙"] = {
title = "retracted tongue root",
link = "w:Advanced and retracted tongue root",
},
}
data[2] = {
-- TODO
--["%("] = {},
--["%)"] = {},
["ːː"] = {
title = "extra long",
link = "w:Length (phonetics)",
},
["r̥"] = {title = "voiceless alveolar trill", link = "w:Voiceless alveolar trill"},
["ɬ’"] = {title = "alveolar lateral ejective fricative", link = "w:Alveolar lateral ejective fricative"},
}
data[3] = {
["t͡s"] = {title = "voiceless alveolar sibilant affricate", link = "w:Voiceless alveolar affricate"},
["d͡z"] = {title = "voiced alveolar sibilant affricate", link = "w:Voiced alveolar affricate"},
["t͡ʃ"] = {title = "voiceless palato-alveolar affricate", link = "w:Voiceless palato-alveolar affricate", descender = true},
["d͡ʒ"] = {title = "voiced palato-alveolar affricate", link = "w:Voiced palato-alveolar affricate"},
["ʈ͡ʂ"] = {title = "voiceless retroflex affricate", link = "w:Voiceless retroflex affricate", descender = true},
["ɖ͡ʐ"] = {title = "voiced retroflex affricate", link = "w:Voiced retroflex affricate, descender = true"},
["t͡ɕ"] = {title = "voiceless alveolo-palatal affricate", link = "w:Voiceless alveolo-palatal affricate"},
["d͡ʑ"] = {title = "voiced alveolo-palatal affricate", link = "w:Voiced alveolo-palatal affricate"},
["c͡ç"] = {title = "voiceless palatal affricate", link = "w:Voiceless palatal affricate, descender = true"},
["ɟ͡ʝ"] = {title = "voiced palatal affricate", link = "w:Voiced palatal affricate, descender = true"},
["k͡x"] = {title = "voiceless velar affricate", link = "w:Voiceless velar affricate"},
["ɡ͡ɣ"] = {title = "voiced velar affricate", link = "w:Voiced velar affricate, descender = true"},
}
data[4] = {
["ǃ͡qʼ"] = {title = "alveolar linguo-glottalic stop", link = "w:Ejective-contour clicks, descender = true"},
["ǁ͡χʼ"] = {title = "lateral linguo-glottalic affricate (homorganic)", link = "w:Ejective-contour clicks", descender = true},
}
data[5] = {
["k͡ʟ̝̊"] = {title = "voiceless velar lateral affricate", link = "w:Voiceless velar lateral affricate"},
["ᶢǀ͡qʼ"] = {title = "voiced dental linguo-glottalic stop", link = "w:Ejective-contour clicks"},
["ǂ͡kxʼ"] = {title = "palatal linguo-glottalic affricate (heterorganic)", link = "w:Ejective-contour clicks"},
}
data[6] = {
["k͡ʟ̝̊ʼ"] = {title = "velar lateral ejective affricate", link = "w:Velar lateral ejective affricate"},
["ᶢʘ͡kxʼ"] = {title = "voiced labial linguo-glottalic affricate", link = "w:Ejective-contour clicks"},
}
data.separator_escapes = {
["⁽"] = "(", ["⁾"] = ")",
["₍"] = "(", ["₎"] = ")",
["ˈ"] = "\1", ["ˌ"] = "\2",
["ː"] = ":", ["ˑ"] = ";",
}
-- acute and grave tone marks
local diacritics = u(
-- grave, acute, circumflex, tilde, macron, breve
0x300, 0x301, 0x302, 0x303, 0x304, 0x306,
-- diaeresis, ring above, double acute, caron, vertical line above, double grave, left tack
0x308, 0x30A, 0x30B, 0x30C, 0x30D, 0x30F, 0x318,
-- right tack, left angle, left half ring below, up tack below, down tack below, plus sign below
0x319, 0x31A, 0x31C, 0x31D, 0x31E, 0x31F,
-- minus sign below, rhotic hook below, dot below, diaeresis below, ring below, vertical line below, bridge below
0x320, 0x322, 0x323, 0x324, 0x325, 0x329, 0x32A,
-- caron below, inverted breve below
0x32C, 0x32F,
-- tilde below, combining tilde overlay, right half ring below, inverted bridge below, square below, seagull below, x above
0x330, 0x334, 0x339, 0x33A, 0x33B, 0x33C, 0x33D,
-- grave tone mark, acute tone mark, bridge above, equals sign below, double vertical line below
0x340, 0x341, 0x346, 0x347, 0x348,
-- left angle below, not tilde above, homothetic above, almost equal above, left right arrow below
0x349, 0x34A, 0x34B, 0x34C, 0x34D,
-- upwards arrow below, left arrowhead below, right arrowhead below
0x34E, 0x354, 0x355,
-- double rightwards arrow below, combining Latin small letter a
0x362, 0x361,
-- macron–acute, grave–macron, macron–grave, acute–macron, grave–acute–grave, acute–grave–acute
0x1DC4, 0x1DC5, 0x1DC6, 0x1DC7, 0x1DC8, 0x1DC9)
data.diacritics = diacritics
data.vowels = "iyɨʉɯuɪᵻʏʊᵿeøɘɵɤoəɚɛœɜɝɞʌɔæɐaɶɑɒäëïöüÿ"
local tones = "˥˦˧˨˩꜒꜓꜔꜕꜖꜈꜉꜊꜋꜌꜍꜎꜏꜐꜑¹²³⁴⁵⁶⁷⁸⁹⁰"
data.tones = tones
local superscripts = u(0xA0) .. " ⁰¹²³⁴⁵⁶⁷⁸⁹ᵃ𐞃ᵄᵅᶛᵇ𐞄𐞅ᶜᶝᵈᶞ𐞋𐞌𐞍ᵉᵊᵋ𐞎ᶟᵌ𐞏𐞑ᶠᶢ𐞒𐞓𐞔ˠʰ𐞕𐞖ʱ𐞗ⁱᶦᶤʲᶨᶡ𐞘ᵏˡᶫꭞ𐞛ᶩ𐞞𐞠𐞡ᵐᶬⁿᶰᶮᶯᵑᵒ𐞢ꟹ𐞣ᵓᶱᵖᶲ𐞥ʳ𐞪ʴ𐞦𐞧ʵ𐞨𐞩ʶˢᶳᶴᵗ𐞯ᵘᶶᶣᵚᶭᶷᵛᶹ𐞰ᶺʷꭩˣʸ𐞲ᶻᶼᶽᶾˀˤ𐞳𐞴𐞶𐞷𐞸𐞹𐞵ᵝᶿᵡ˞⁻𐞁𐞂"
data.superscripts = superscripts
-- An array of patterns of valid character sequences.
data.valid = {
"⁽[" .. superscripts .. "]+⁾",
"[ %(%)%%<>{|}%-→~⁓%.◌abcdefhijklmnopqrstuvwxyz¡àáâãāăēäæçèéêëĕěħìíîïĩīĭĺḿǹńňðòóôõöōŏőœøŕùúûüũūŭűýÿŷŋ"
.. "ǀǁǂǃǎǐǒǔřǖǘǚǜǟǣǽǿȁȅȉȍȕȫȭȳɐɑɒɓɔɕɖɗɘəɚɛɜɝɞɟɠɡɢɣɤɥɦɧɨɪᵻɫɬɭɮɯɰɱɲɳɴɵɶɸɹɺ𝼈ɻɽɾʀʁʂʃʄʈʉʊᵿʋṽʌʍʎ𝼆ʏʐʑʒʔʕʘʞʙʛʜʝʟʡʢ𝼊ʬʭ"
.. "ʼˈˌːˑˣ˔˕ˬ͗˭ˇ˖β͜θχᴙᶑ᷽ḁḛḭḯṍṏṳṵṹṻạẹẽịọụỳỵỹ‖․‥…‿↑↓↗↘ⱱꜛꜜꟸ𝆏𝆑˗ˋˊ–⸨⸩⁽⁾" .. diacritics .. tones .. superscripts .. "]+"
}
-- Character sequences which are valid only in a particular language.
-- These can be either a single pattern (as a string), or an array of patterns (as a table).
data.per_lang_valid = {
["egy"] = "V+", -- V for uncertain vowel
["okm"] = "[LHR!WT]+", -- irregular verb morphophonemes
}
-- Characters to add VARIATION SELECTOR-15 (U+FE0E) after.
-- These are characters with emoji variants that are used by default by some clients.
-- Adding VS15 after them instructs them to draw the characters as text instead.
data.add_vs15 = "↗↘"
data.invalid = {
["!"] = "ǃ",
["ꜝ"] = "ꜜ",
["ꜞ"] = "ꜛ",
["ꜟ"] = "ꜛ",
["'"] = "ˈ",
["’"] = "ʼ",
[":"] = "ː",
-- Confusable Latin letters
["B"] = "ʙ",
["g"] = "ɡ",
["G"] = "ɢ",
["Ɠ"] = "ʛ",
["H"] = "ʜ",
["ı"] = "ɪ",
["I"] = "ɪ",
["L"] = "ʟ",
["N"] = "ɴ",
["Œ"] = "ɶ",
["Q"] = "ꞯ",
["R"] = "ʀ",
["∫"] = "ʃ",
["⨎"] = "ǂ", -- due to confusion with obsolete 𝼋 below
["ß"] = "β",
["ẞ"] = "β",
["Y"] = "ʏ",
["Ə"] = "ə",
["ǝ"] = "ə",
["Ɂ"] = "ʔ",
["ɂ"] = "ʔ",
["ˁ"] = "ˤ",
-- Confusable Greek letters
["α"] = "ɑ",
["γ"] = "ɣ",
["δ"] = "ð",
["ε"] = "ɛ",
["Η"] = "ʜ",
["η"] = "ŋ",
["ι"] = "ɪ",
["λ"] = "ʎ",
["υ"] = "ʋ",
["Ψ"] = "𝼊",
["ψ"] = "𝼊",
["Φ"] = "ɸ",
["ϕ"] = "ɸ",
["ꭓ"] = "χ", -- Actually Latin, since IPA uses the Greek letter(!)
-- Confusable Cyrillic letters
["ӕ"] = "æ",
["Ә"] = "ə",
["ә"] = "ə",
["В"] = "ʙ",
["в"] = "ʙ",
["е"] = "e",
["З"] = "ɜ",
["з"] = "ɜ",
["Ѕ"] = "s",
["ѕ"] = "s",
["і"] = "i",
["ј"] = "j",
["Н"] = "ʜ",
["н"] = "ʜ",
["О"] = "o",
["о"] = "o",
["р"] = "p",
["с"] = "c",
["у"] = "y",
["Ү"] = "ʏ",
["ү"] = "ʏ",
["Ф"] = "ɸ",
["ф"] = "ɸ",
["х"] = "x",
["Һ"] = "h",
["һ"] = "h",
["Я"] = "ᴙ",
["я"] = "ᴙ",
["Ѱ"] = "𝼊",
["ѱ"] = "𝼊",
["Ѵ"] = "ⱱ",
["ѵ"] = "ⱱ",
["Ҁ"] = "ʕ",
["ҁ"] = "ʕ",
-- Palatalization
["ᶀ"] = "bʲ",
["ꞔ"] = "cʲ",
["ᶁ"] = "dʲ",
["ȡ"] = "d̠ʲ",
["d̂"] = "d̠ʲ",
["ᶂ"] = "fʲ",
["ᶃ"] = "ɡʲ",
["ꞕ"] = "hʲ",
["ᶄ"] = "kʲ",
["ᶅ"] = "lʲ",
["ȴ"] = "l̠ʲ",
["l̂"] = "l̠ʲ",
["𝼓"] = "ɬʲ",
["ᶆ"] = "mʲ",
["ᶇ"] = "nʲ",
["ȵ"] = "n̠ʲ",
["n̂"] = "n̠ʲ",
["𝼔"] = "ŋʲ",
["ᶈ"] = "pʲ",
["ᶉ"] = "rʲ",
["𝼕"] = "ɹʲ",
["𝼖"] = "ɾʲ",
["ᶊ"] = "sʲ",
["𝼞"] = "ɕ",
["𐞺"] = "ᶝ",
["ᶋ"] = "ʃʲ",
["ʆ"] = "ʃʲ",
["ƫ"] = "tʲ",
["ȶ"] = "t̠ʲ",
["t̂"] = "t̠ʲ",
["ᶌ"] = "vʲ",
["ᶍ"] = "xʲ",
["ᶎ"] = "zʲ",
["𝼘"] = "ʒʲ",
["ʓ"] = "ʒʲ",
-- Retroflex
["𝼝"] = "ʈ͡ʂ",
["𝼥"] = "ɖ",
["𝼦"] = "ɭ",
["𝼧"] = "ɳ",
["𝼨"] = "ɽ",
["𝼩"] = "ʂ",
["𝼪"] = "ʈ",
-- Rhotic vowels
["ᶏ"] = "a˞",
["ᶐ"] = "ɑ˞",
["ᶒ"] = "e˞",
["ə˞"] = "ɚ",
["ᶕ"] = "ɚ",
["ᶓ"] = "ɛ˞",
["ɜ˞"] = "ɝ",
["ᶔ"] = "ɝ",
["ᶖ"] = "i˞",
["𝼚"] = "ɨ˞",
["𝼛"] = "o˞",
["ᶗ"] = "ɔ˞",
["ᶙ"] = "u˞",
-- Syllabic approximants
["ɿ"] = "ɹ̩",
["ʅ"] = "ɻ̩",
["ʮ"] = "ɹ̩ʷ",
["ʯ"] = "ɻ̩ʷ",
-- Clicks
["ʗ"] = "ǃ",
["𝼋"] = "ǂ",
["ʇ"] = "ǀ",
["ʖ"] = "ǁ",
["‼"] = "𝼊",
-- Voiceless implosives
["ƈ"] = "ʄ̊",
["ƙ"] = "ɠ̊",
["ƥ"] = "ɓ̥",
["ʠ"] = "ʛ̥",
["ƭ"] = "ɗ̥",
["𝼉"] = "ᶑ̥",
-- Monographs
["ꜰ"] = "ɸ",
["ɩ"] = "ɪ",
["ɼ"] = "r̝",
["ᴜ"] = "ʊ",
["ɷ"] = "ʊ",
["𐞤"] = "ᶷ",
["ƛ"] = "t͡ɬ",
["ƻ"] = "d͡z",
["ƾ"] = "t͡s",
-- Digraphs
["ȸ"] = "b̪",
["ʣ"] = "d͡z",
["ʥ"] = "d͡ʑ",
["ꭦ"] = "ɖ͡ʐ",
["ʤ"] = "d͡ʒ",
["𝼒"] = "d͡ʒʲ",
["𝼙"] = "d͡ᶚ",
["ʪ"] = "ɬ͡s",
["ʫ"] = "ɮ͡z",
["ȹ"] = "p̪",
["ʦ"] = "t͡s",
["ʨ"] = "t͡ɕ",
["ꭧ"] = "ʈ͡ʂ",
["ʧ"] = "t͡ʃ",
["𝼗"] = "t͡ʃʲ",
["𝼜"] = "t͡ᶘ",
-- Deprecated or confusable diacritics
["̫"] = "ʷ",
["͂"] = "̃",
["᫇"] = "ʷ",
["⸋"] = "̚",
["̱"] = "̠", -- COMBINING MACRON BELOW (U+0331) -> COMBINING MINUS SIGN BELOW (U+0320)
-- Precomposed characters with deprecated or confusable diacritics; the left is a precomposed
-- version of a lowercase letter with COMBINING MACRON BELOW and the right is the equivalent
-- using COMBINING MINUS SIGN BELOW
["ḇ"] = "b̠",
["ḏ"] = "d̠",
["ẖ"] = "h̠",
["ḵ"] = "k̠",
["ḻ"] = "l̠",
["ṉ"] = "n̠",
["ṟ"] = "r̠",
["ṯ"] = "t̠",
["ẕ"] = "z̠",
}
return data
mjopki8sklky9m3kv63kzgka1usuiki
Modul:nyms
828
12091
375347
363508
2026-09-22T03:12:08Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708739|92708739]])
375347
Scribunto
text/plain
local export = {}
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local labels_module = "Module:labels"
local links_module = "Module:links"
local parameter_utilities_module = "Module:parameter utilities"
local parse_utilities_module = "Module:parse utilities"
local function track(page)
return require(debug_track_module)("nyms/" .. page)
end
local function wrap_span(text, lang, sc)
return '<span class="' .. sc .. '" lang="' .. lang .. '">' .. text .. '</span>'
end
local function term_already_linked(term)
-- optimization to avoid unnecessarily loading [[Module:parse utilities]]
return term:find("[<{]") and require(parse_utilities_module).term_already_linked(term)
end
function export.nyms(frame)
local parent_args = frame:getParent().args
-- FIXME: Temporary error message and tracking.
for arg, _ in pairs(parent_args) do
if type(arg) == "string" and arg:find("^lb[0-9]*$") then
local llarg = arg:gsub("^lb", "ll")
error(("%s= is deprecated; use %s= instead, per the documentation"):format(arg, llarg))
end
if arg == "q" or arg == "qq" then
track(arg)
end
if type(arg) == "string" and arg:find("^tag[0-9]*$") then
local larg = arg:gsub("^tag", "l")
local llarg = arg:gsub("^tag", "ll")
error(("Use %s= (on the left) or %s= (on the right) instead of %s="):format(larg, llarg, arg))
end
end
local params = {
[1] = {required = true, type = "language", default = "und"},
[2] = {list = true, allow_holes = true, required = true},
}
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{group = {"link", "ref", "l"}},
-- For compatibility, we don't distinguish q= from q1= and qq= from q1=. FIXME: Maybe we should change this.
{group = "q", separate_no_index = false},
{param = "lb", deprecated = true},
}
local special_separators = mw.clone(m_param_utils.default_special_separators)
special_separators["<"] = " < "
local items, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 2,
parse_lang_prefix = true,
track_module = "nyms",
lang = 1,
special_separators = special_separators,
sc = "sc.default",
-- because we show them ourselves below so e.g. we can handle already-linked terms
no_show_decorations = true,
}
local nym_type = frame.args[1]
local nym_type_class = string.gsub(nym_type, "%s", "-")
local lang = args[1]
local langcode = lang:getCode()
local data = {
lang = lang,
items = items,
sc = args.sc.default,
l = args.l.default,
ll = args.ll.default,
q = args.q.default,
qq = args.qq.default,
}
local parts = {}
local thesaurus_parts = {}
for i, item in ipairs(data.items) do
if item.lb then
error("Inline modifier <lb:...> is deprecated; use <ll:...> per the documentation")
end
local explicit_item_lang = item.lang
item.lang = item.lang or data.lang
item.sc = item.sc or data.sc
local text
local is_thesaurus
if item.term and item.term:find("^Tesaurus:") then
is_thesaurus = true
for k, _ in pairs(item) do
if m_param_utils.item_key_is_property(k) and k ~= "lang" and k ~= "sc" and k ~= "q" and k ~= "qq" and
k ~= "l" and k ~= "ll" and k ~= "refs" then
error(("You cannot use most named parameters and inline modifiers with Thesaurus links, but saw %s%s= or its equivalent inline modifier <%s:...>"):format(
k, item.itemno, k))
end
end
local term = item.term:match("^Tesaurus:(.*)$")
-- Chop off fragment
term = term:match("^(.-)#.*$") or term
local lang = item.lang
local sccode = (item.sc or lang:findBestScript(term)):getCode()
-- FIXME: I assume it's better to include full-language codes in the CSS rather than etym-language codes,
-- which are generally specific to Wiktionary. However, we should probably instead be using the functions
-- from [[Module:script utilities]] in preference to rolling our own.
text = "[[" .. item.term .. "#Bahasa " .. lang:getFullName() .. "|Tesaurus:" .. wrap_span(term, lang:getFullCode(), sccode) .. "]]"
else
if thesaurus_parts[1] then
error("Links to the Thesaurus must follow all non-Thesaurus links")
end
local raw_term = item.alt or item.term
if raw_term and term_already_linked(raw_term) then
text = raw_term
else
text = require(links_module).full_link(item)
end
end
local qq = item.qq
-- If a separate language code was given for the term, display the language name as a right qualifier.
-- Otherwise it may not be obvious that the term is in a separate language (e.g. if the main language is 'zh'
-- and the term language is a Chinese lect such as Min Nan). But don't do this for Translingual terms, which
-- are often added to the list of English and other-language terms.
if explicit_item_lang then
local explicit_code = explicit_item_lang:getCode()
if explicit_code ~= langcode and explicit_code ~= "mul" then
qq = mw.clone(qq) or {}
table.insert(qq, 1, explicit_item_lang:getCanonicalName())
end
end
if item.q and item.q[1] or qq and qq[1] or item.l and item.l[1] or item.ll and item.ll[1] or
item.refs and item.refs[1] then
text = require(decorations_module).format_decorations {
lang = item.lang,
text = text,
q = item.q,
qq = qq,
l = item.l,
ll = item.ll,
refs = item.refs,
}
end
local insert_place = is_thesaurus and thesaurus_parts or parts
-- Don't include the separator if this is the first item of this class that we're inserting.
table.insert(insert_place, insert_place[1] and item.separator or "")
table.insert(insert_place, text)
end
local text = table.concat(parts)
local thesaurus_text = table.concat(thesaurus_parts)
local caption = "<span style=\"font-size: smaller\">" .. mw.getContentLanguage():ucfirst(nym_type) ..
((#items > 1 or thesaurus_text ~= "") and "s" or "") .. ":</span> "
text = caption .. text
local function decoration_error_if_no_terms()
if not parts[1] then
error("Cannot specify overall decorations if no non-Thesaurus terms given")
end
end
if data.q and data.q[1] or data.qq and data.qq[1] or data.l and data.l[1] then
decoration_error_if_no_terms()
text = require(decorations_module).format_decorations {
lang = data.lang,
text = text,
q = data.q,
qq = data.qq,
l = data.l,
-- ll handled specially for compatibility's sake
}
end
if data.ll and data.ll[1] then
decoration_error_if_no_terms()
text = text .. " — " .. require(labels_module).show_labels {
lang = data.lang,
labels = data.ll,
nocat = true,
open = false,
close = false,
no_track_already_seen = true,
}
end
if thesaurus_text ~= "" then
local thesaurus_intro = parts[1] and "; ''lihat juga'' " or "''lihat'' "
text = text .. thesaurus_intro .. thesaurus_text
end
return "<span class=\"nyms " .. nym_type_class .. "\">" .. text .. "</span>"
end
return export
izsdelhiyeoailt2eqmxaihby4w3q4o
Modul:ar-pronunciation
828
12378
375390
365593
2026-09-22T07:32:09Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92717545|92717545]])
375390
Scribunto
text/plain
local export = {}
local m_str_utils = require("Module:string utilities")
local m_table = require("Module:table")
local audio_module = "Module:audio"
local parse_utilities_module = "Module:parse utilities"
local rfind = m_str_utils.find
local rsplit = m_str_utils.split
local ugsub = m_str_utils.gsub
local ulen = m_str_utils.len
local ulower = m_str_utils.lower
local usub = m_str_utils.sub
local concat = table.concat
local insert = table.insert
local lang = require("Module:languages").getByCode("ar")
local sc = require("Module:scripts").getByCode("Arab")
local correspondences = {
["ʾ"] = "ʔ",
["ṯ"] = "θ",
["j"] = "d͡ʒ",
["ḥ"] = "ħ",
["ḵ"] = "x",
["ḏ"] = "ð",
["š"] = "ʃ",
["ṣ"] = "sˤ",
["ḍ"] = "dˤ",
["ṭ"] = "tˤ",
["ẓ"] = "ðˤ",
["ž"] = "ʒ",
["ʿ"] = "ʕ",
["ḡ"] = "ɣ",
["ḷ"] = "lˤ",
["ū"] = "uː",
["ī"] = "iː",
["ā"] = "aː",
["y"] = "j",
["g"] = "ɡ",
["ē"] = "eː",
["ō"] = "oː",
[""] = "",
}
local vowels = "aāeēiīoōuū"
local vowel = "[" .. vowels .. "]"
local long_vowels = "āēīōū"
local long_vowel = "[" .. long_vowels .. "]"
local consonant = "[^" .. vowels .. ". -]"
local syllabify_pattern = "(" .. vowel .. ")(" .. consonant .. "?)(" .. consonant .. "?)(" .. vowel .. ")"
local tie = "‿"
local closed_syllable_shortening_pattern = "(" .. long_vowel .. ")(" .. tie .. ")" .. "(" .. consonant .. ")"
local function rsub(term, foo, bar)
local retval = ugsub(term, foo, bar)
return retval
end
local function generate_obj(respelling)
return { respelling = respelling }
end
local function split_on_comma(term)
if not term then
return nil
end
if term:find(",%s") or term:find("\\") then
return require(parse_utilities_module).split_on_comma(term)
else
return rsplit(term, ",")
end
end
local function parse_respellings_with_modifiers(respelling, paramname)
if respelling:find("[<%[]") then
local put = require(parse_utilities_module)
local segments = put.parse_multi_delimiter_balanced_segment_run(respelling, { { "<", ">" }, { "[", "]" } })
local comma_separated_groups = put.split_alternating_runs_on_comma(segments)
local retval = {}
for _, group in ipairs(comma_separated_groups) do
local j = 2
while j <= #group do
if not group[j]:find("^<.*>$") then
group[j - 1] = group[j - 1] .. group[j] .. group[j + 1]
table.remove(group, j)
table.remove(group, j)
else
j = j + 2
end
end
local param_mods = {
q = { type = "qualifier" },
qq = { type = "qualifier" },
a = { type = "labels" },
aa = { type = "labels" },
ref = { item_dest = "refs", type = "references" },
}
table.insert(retval, put.parse_inline_modifiers_from_segments {
group = group,
arg = respelling,
props = {
paramname = paramname,
param_mods = param_mods,
generate_obj = generate_obj,
},
})
end
return retval
else
local retval = {}
for _, item in ipairs(split_on_comma(respelling)) do
table.insert(retval, generate_obj(item))
end
return retval
end
end
local function parse_pron_modifier(arg, paramname, generate_obj, param_mods, splitchar)
splitchar = splitchar or ","
if arg:find("<") then
param_mods.q = { type = "qualifier" }
param_mods.qq = { type = "qualifier" }
param_mods.a = { type = "labels" }
param_mods.aa = { type = "labels" }
param_mods.ref = { item_dest = "refs", type = "references" }
return require(parse_utilities_module).parse_inline_modifiers(arg, {
param_mods = param_mods,
generate_obj = generate_obj,
paramname = paramname,
splitchar = splitchar,
})
else
local retval = {}
local split_arg = splitchar == "," and split_on_comma(arg) or rsplit(arg, splitchar)
for _, term in ipairs(split_arg) do
table.insert(retval, generate_obj(term))
end
return retval
end
end
local function parse_audio(lang, arg, pagename, paramname)
local param_mods = {
IPA = { sublist = true },
text = {},
t = { item_dest = "gloss" },
gloss = {},
pos = {},
lit = {},
g = { item_dest = "genders", sublist = true },
bad = {},
cap = { item_dest = "caption" },
}
local function process_special_chars(val)
if not val then
return val
end
return (val:gsub("#", pagename))
end
local function generate_audio_obj(arg)
return { file = process_special_chars(arg) }
end
local retvals = parse_pron_modifier(arg, paramname, generate_audio_obj, param_mods, "%s*;%s*")
for _, retval in ipairs(retvals) do
retval.lang = lang
retval.text = process_special_chars(retval.text)
retval.caption = process_special_chars(retval.caption)
local textobj = require(audio_module).construct_audio_textobj(retval)
retval.text = textobj
retval.gloss = nil
retval.pos = nil
retval.lit = nil
retval.genders = nil
end
return retvals
end
local function parse_regional_phonetics(ph_arg, pagename)
if not ph_arg or ph_arg == "" then
return {}
end
local regionals = {}
for _, item in ipairs(rsplit(ph_arg, "%s*;%s*")) do
local audio = nil
local item_no_mod = item:gsub("<a:([^>]+)>", function(a)
audio = a:gsub("#", pagename)
return ""
end)
local region, ipa = item_no_mod:match("^([^:]+):(.+)$")
if region and ipa then
local regions = rsplit(region, "%s*,%s*")
table.insert(regionals, { regions = regions, ipa = ipa, audio = audio })
end
end
return regionals
end
local function syllabify(text)
text = ugsub(text, "%-(" .. consonant .. ")%-(" .. consonant .. ")", "%1.%2")
text = ugsub(text, "%-", ".")
for _ = 1, 2 do
text = ugsub(
text,
syllabify_pattern,
function(a, b, c, d)
if c == "" and b ~= "" then
c, b = b, ""
end
return a .. b .. "." .. c .. d
end
)
end
text = ugsub(text, "(" .. vowel .. ") (" .. consonant .. ")%.?(" ..
consonant .. ")", "%1" .. tie .. "%2.%3")
return text
end
local function closed_syllable_shortening(text)
local shorten = {
["ā"] = "a",
["ē"] = "e",
["ī"] = "i",
["ō"] = "o",
["ū"] = "u",
}
text = ugsub(text,
closed_syllable_shortening_pattern,
function(vowel, tie, consonant)
return shorten[vowel] .. tie .. consonant
end)
return text
end
function export.link(term)
return require("Module:links").full_link { term = term, lang = lang, sc = sc }
end
function export.toIPA(list, silent_error)
local translit
if list.tr then
translit = list.tr
elseif list.term then
require("Module:script utilities").checkScript(list.term, "Arab")
translit = lang:transliterate(list.term)
if not translit then
if silent_error then
return ''
else
error('Module:ar-translit failed to generate a transliteration from "' .. list.term .. '".')
end
end
else
if silent_error then
return ''
else
error('No Arabic text or transliteration was provided to the function "toIPA".')
end
end
translit = ugsub(translit, "llāh", "ḷḷāh")
translit = ugsub(translit, "([iī] ?)ḷḷ", "%1ll")
translit = ugsub(translit, "%(t%)", "")
translit = ugsub(translit, "(" .. vowel .. ") " .. vowel, "%1 ")
translit = ugsub(translit, "%-?l%-?", "l")
translit = syllabify(translit)
translit = closed_syllable_shortening(translit)
local output = ugsub(translit, ".", correspondences)
output = ugsub(output, "%-", "")
return output
end
function export.get_pron_info(terms, pagename, paramname)
if #terms == 1 and terms[1].respelling == "-" then
return { pron_list = nil }
end
local pron_list = {}
local brackets = "/%s/"
for _, term in ipairs(terms) do
local respelling = term.respelling
local ar_term, tr
if not respelling or respelling == "" or respelling == "#" then
ar_term = pagename
elseif rfind(respelling, "[a-zA-Z]") then
tr = respelling
elseif respelling:find("[ء-ي]") then
ar_term = respelling
else
tr = respelling
end
local pron = export.toIPA({ term = ar_term, tr = tr }, false)
if pron and pron ~= "" then
local bracketed_pron = brackets:format(pron)
table.insert(pron_list, {
pron = bracketed_pron,
q = term.q,
qq = term.qq,
a = term.a,
aa = term.aa,
refs = term.refs,
})
end
end
return { pron_list = pron_list }
end
function export.show_old(frame)
local params = {
[1] = { list = true, allow_holes = true },
["tr"] = { list = true, allow_holes = true },
["qual"] = { list = true, allow_holes = true },
["nl"] = { type = "boolean" },
["ann"] = {},
}
local args = require("Module:parameters").process(frame:getParent().args, params)
local ar_terms = args[1]
local transliterations = args.tr
local qualifiers = args.qual
local nl = args.nl
if not (ar_terms.maxindex > 0 or transliterations.maxindex > 0) then
if mw.title.getCurrentTitle().nsText == "Template" then
ar_terms[1] = "كَلِمَة"
ar_terms.maxindex = 1
else
error(
'Please provide vocalized Arabic in the first parameter of {{[[Template:ar-IPA|ar-IPA]]}}, or transliteration in the "tr" parameter.')
end
end
local pronunciations = {}
for i = 1, math.max(ar_terms.maxindex, transliterations.maxindex) do
local ar_term = ar_terms[i]
local tr = transliterations[i]
local qual = qualifiers[i]
if not (ar_term or tr) then
error("There is a gap in the parameters. Provide either |" .. i .. "= or |tr" .. i .. "=.")
elseif ar_term and tr then
mw.logObject("Duplicate parameters |" .. i .. "= and |tr" .. i .. "= in {{ar-IPA}},")
end
local pron = export.toIPA { term = ar_term, tr = tr }
table.insert(pronunciations, { pron = "/" .. pron .. "/", q = qual and { qual } or nil })
end
local anntext = ""
if args.ann then
anntext = args.ann
if args.ann:find("%+") then
local anndefs = {}
for i = 1, ar_terms.maxindex do
local ar_term = ar_terms[i]
if ar_term then
table.insert(anndefs, "'''" .. ar_term .. "'''")
end
end
if anndefs[1] then
anndefs = table.concat(anndefs, ", ")
anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs))
end
end
anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ": "
end
if nl then
return anntext .. require("Module:IPA").format_IPA_multiple(lang, pronunciations)
else
return anntext .. require("Module:IPA").format_IPA_full { lang = lang, items = pronunciations }
end
end
function export.show(frame)
local parent_args = frame:getParent().args
local process = require("Module:parameters").process
local params = {
[1] = {},
["audios"] = {},
["a"] = { alias_of = "audios" },
["ph"] = {},
["pagename"] = {},
["indent"] = {},
["ann"] = {},
}
local args = process(parent_args, params)
local pagename = args.pagename or mw.loadData("Module:headword/data").pagename
local indent = args.indent or "*"
local termspec = args[1] or "#"
local terms = parse_respellings_with_modifiers(termspec, 1)
local pronobj = export.get_pron_info(terms, pagename, 1)
local regional_phonetics = parse_regional_phonetics(args.ph, pagename)
local parts = {}
local function ins(text)
table.insert(parts, text)
end
local anntext = ""
if args.ann then
anntext = args.ann
if args.ann:find("%+") then
local anndefs = {}
for _, term in ipairs(terms) do
local respelling = term.respelling
if respelling and respelling:find("[ء-ي]") then
table.insert(anndefs, "'''" .. respelling .. "'''")
end
end
if anndefs[1] then
anndefs = table.concat(anndefs, ", ")
anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs))
end
end
anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ": "
end
if pronobj.pron_list and #pronobj.pron_list > 0 then
local formatted = require("Module:IPA").format_IPA_full { lang = lang, items = pronobj.pron_list }
ins(indent .. anntext .. mw.ustring.toNFC(formatted))
end
if args.audios then
local format_audio = require("Module:audio").format_audio
local audio_objs = parse_audio(lang, args.audios, pagename, "audios")
for i, audio_obj in ipairs(audio_objs) do
if #audio_objs > 1 and not audio_obj.caption then
audio_obj.caption = "Audio " .. i
end
ins("\n" .. indent .. " " .. format_audio(audio_obj))
end
end
if #regional_phonetics > 0 then
local m_IPA = require("Module:IPA")
local m_accent = require("Module:accent qualifier")
for _, regional in ipairs(regional_phonetics) do
local regions = regional.regions
local ipa = regional.ipa
local pron_item = { pron = "[" .. ipa .. "]" }
local formatted_ipa = m_IPA.format_IPA_full { lang = lang, items = { pron_item } }
local formatted_region = m_accent.format_qualifiers(lang, regions)
local line = "\n" .. indent .. indent .. " " .. formatted_region .. " " .. mw.ustring.toNFC(formatted_ipa)
if regional.audio then
local audio_obj = {
lang = lang,
file = regional.audio,
}
local textobj = require(audio_module).construct_audio_textobj(audio_obj)
audio_obj.text = textobj
line = line .. " " .. require("Module:audio").format_audio(audio_obj)
end
ins(line)
end
end
return concat(parts)
end
return export
5fgfg64x6pqtsaczm7x9u3q3t4t78ua
375392
375390
2026-09-22T10:18:09Z
Hakimi97
2668
Minor correction
375392
Scribunto
text/plain
local export = {}
local m_str_utils = require("Module:string utilities")
local m_table = require("Module:table")
local audio_module = "Module:audio"
local parse_utilities_module = "Module:parse utilities"
local rfind = m_str_utils.find
local rsplit = m_str_utils.split
local ugsub = m_str_utils.gsub
local ulen = m_str_utils.len
local ulower = m_str_utils.lower
local usub = m_str_utils.sub
local concat = table.concat
local insert = table.insert
local lang = require("Module:languages").getByCode("ar")
local sc = require("Module:scripts").getByCode("Arab")
local correspondences = {
["ʾ"] = "ʔ",
["ṯ"] = "θ",
["j"] = "d͡ʒ",
["ḥ"] = "ħ",
["ḵ"] = "x",
["ḏ"] = "ð",
["š"] = "ʃ",
["ṣ"] = "sˤ",
["ḍ"] = "dˤ",
["ṭ"] = "tˤ",
["ẓ"] = "ðˤ",
["ž"] = "ʒ",
["ʿ"] = "ʕ",
["ḡ"] = "ɣ",
["ḷ"] = "lˤ",
["ū"] = "uː",
["ī"] = "iː",
["ā"] = "aː",
["y"] = "j",
["g"] = "ɡ",
["ē"] = "eː",
["ō"] = "oː",
[""] = "",
}
local vowels = "aāeēiīoōuū"
local vowel = "[" .. vowels .. "]"
local long_vowels = "āēīōū"
local long_vowel = "[" .. long_vowels .. "]"
local consonant = "[^" .. vowels .. ". -]"
local syllabify_pattern = "(" .. vowel .. ")(" .. consonant .. "?)(" .. consonant .. "?)(" .. vowel .. ")"
local tie = "‿"
local closed_syllable_shortening_pattern = "(" .. long_vowel .. ")(" .. tie .. ")" .. "(" .. consonant .. ")"
local function rsub(term, foo, bar)
local retval = ugsub(term, foo, bar)
return retval
end
local function generate_obj(respelling)
return { respelling = respelling }
end
local function split_on_comma(term)
if not term then
return nil
end
if term:find(",%s") or term:find("\\") then
return require(parse_utilities_module).split_on_comma(term)
else
return rsplit(term, ",")
end
end
local function parse_respellings_with_modifiers(respelling, paramname)
if respelling:find("[<%[]") then
local put = require(parse_utilities_module)
local segments = put.parse_multi_delimiter_balanced_segment_run(respelling, { { "<", ">" }, { "[", "]" } })
local comma_separated_groups = put.split_alternating_runs_on_comma(segments)
local retval = {}
for _, group in ipairs(comma_separated_groups) do
local j = 2
while j <= #group do
if not group[j]:find("^<.*>$") then
group[j - 1] = group[j - 1] .. group[j] .. group[j + 1]
table.remove(group, j)
table.remove(group, j)
else
j = j + 2
end
end
local param_mods = {
q = { type = "qualifier" },
qq = { type = "qualifier" },
a = { type = "labels" },
aa = { type = "labels" },
ref = { item_dest = "refs", type = "references" },
}
table.insert(retval, put.parse_inline_modifiers_from_segments {
group = group,
arg = respelling,
props = {
paramname = paramname,
param_mods = param_mods,
generate_obj = generate_obj,
},
})
end
return retval
else
local retval = {}
for _, item in ipairs(split_on_comma(respelling)) do
table.insert(retval, generate_obj(item))
end
return retval
end
end
local function parse_pron_modifier(arg, paramname, generate_obj, param_mods, splitchar)
splitchar = splitchar or ","
if arg:find("<") then
param_mods.q = { type = "qualifier" }
param_mods.qq = { type = "qualifier" }
param_mods.a = { type = "labels" }
param_mods.aa = { type = "labels" }
param_mods.ref = { item_dest = "refs", type = "references" }
return require(parse_utilities_module).parse_inline_modifiers(arg, {
param_mods = param_mods,
generate_obj = generate_obj,
paramname = paramname,
splitchar = splitchar,
})
else
local retval = {}
local split_arg = splitchar == "," and split_on_comma(arg) or rsplit(arg, splitchar)
for _, term in ipairs(split_arg) do
table.insert(retval, generate_obj(term))
end
return retval
end
end
local function parse_audio(lang, arg, pagename, paramname)
local param_mods = {
IPA = { sublist = true },
text = {},
t = { item_dest = "gloss" },
gloss = {},
pos = {},
lit = {},
g = { item_dest = "genders", sublist = true },
bad = {},
cap = { item_dest = "caption" },
}
local function process_special_chars(val)
if not val then
return val
end
return (val:gsub("#", pagename))
end
local function generate_audio_obj(arg)
return { file = process_special_chars(arg) }
end
local retvals = parse_pron_modifier(arg, paramname, generate_audio_obj, param_mods, "%s*;%s*")
for _, retval in ipairs(retvals) do
retval.lang = lang
retval.text = process_special_chars(retval.text)
retval.caption = process_special_chars(retval.caption)
local textobj = require(audio_module).construct_audio_textobj(retval)
retval.text = textobj
retval.gloss = nil
retval.pos = nil
retval.lit = nil
retval.genders = nil
end
return retvals
end
local function parse_regional_phonetics(ph_arg, pagename)
if not ph_arg or ph_arg == "" then
return {}
end
local regionals = {}
for _, item in ipairs(rsplit(ph_arg, "%s*;%s*")) do
local audio = nil
local item_no_mod = item:gsub("<a:([^>]+)>", function(a)
audio = a:gsub("#", pagename)
return ""
end)
local region, ipa = item_no_mod:match("^([^:]+):(.+)$")
if region and ipa then
local regions = rsplit(region, "%s*,%s*")
table.insert(regionals, { regions = regions, ipa = ipa, audio = audio })
end
end
return regionals
end
local function syllabify(text)
text = ugsub(text, "%-(" .. consonant .. ")%-(" .. consonant .. ")", "%1.%2")
text = ugsub(text, "%-", ".")
for _ = 1, 2 do
text = ugsub(
text,
syllabify_pattern,
function(a, b, c, d)
if c == "" and b ~= "" then
c, b = b, ""
end
return a .. b .. "." .. c .. d
end
)
end
text = ugsub(text, "(" .. vowel .. ") (" .. consonant .. ")%.?(" ..
consonant .. ")", "%1" .. tie .. "%2.%3")
return text
end
local function closed_syllable_shortening(text)
local shorten = {
["ā"] = "a",
["ē"] = "e",
["ī"] = "i",
["ō"] = "o",
["ū"] = "u",
}
text = ugsub(text,
closed_syllable_shortening_pattern,
function(vowel, tie, consonant)
return shorten[vowel] .. tie .. consonant
end)
return text
end
function export.link(term)
return require("Module:links").full_link { term = term, lang = lang, sc = sc }
end
function export.toIPA(list, silent_error)
local translit
if list.tr then
translit = list.tr
elseif list.term then
require("Module:script utilities").checkScript(list.term, "Arab")
translit = lang:transliterate(list.term)
if not translit then
if silent_error then
return ''
else
error('Module:ar-translit failed to generate a transliteration from "' .. list.term .. '".')
end
end
else
if silent_error then
return ''
else
error('No Arabic text or transliteration was provided to the function "toIPA".')
end
end
translit = ugsub(translit, "llāh", "ḷḷāh")
translit = ugsub(translit, "([iī] ?)ḷḷ", "%1ll")
translit = ugsub(translit, "%(t%)", "")
translit = ugsub(translit, "(" .. vowel .. ") " .. vowel, "%1 ")
translit = ugsub(translit, "%-?l%-?", "l")
translit = syllabify(translit)
translit = closed_syllable_shortening(translit)
local output = ugsub(translit, ".", correspondences)
output = ugsub(output, "%-", "")
return output
end
function export.get_pron_info(terms, pagename, paramname)
if #terms == 1 and terms[1].respelling == "-" then
return { pron_list = nil }
end
local pron_list = {}
local brackets = "/%s/"
for _, term in ipairs(terms) do
local respelling = term.respelling
local ar_term, tr
if not respelling or respelling == "" or respelling == "#" then
ar_term = pagename
elseif rfind(respelling, "[a-zA-Z]") then
tr = respelling
elseif respelling:find("[ء-ي]") then
ar_term = respelling
else
tr = respelling
end
local pron = export.toIPA({ term = ar_term, tr = tr }, false)
if pron and pron ~= "" then
local bracketed_pron = brackets:format(pron)
table.insert(pron_list, {
pron = bracketed_pron,
q = term.q,
qq = term.qq,
a = term.a,
aa = term.aa,
refs = term.refs,
})
end
end
return { pron_list = pron_list }
end
function export.show_old(frame)
local params = {
[1] = { list = true, allow_holes = true },
["tr"] = { list = true, allow_holes = true },
["qual"] = { list = true, allow_holes = true },
["nl"] = { type = "boolean" },
["ann"] = {},
}
local args = require("Module:parameters").process(frame:getParent().args, params)
local ar_terms = args[1]
local transliterations = args.tr
local qualifiers = args.qual
local nl = args.nl
if not (ar_terms.maxindex > 0 or transliterations.maxindex > 0) then
if mw.title.getCurrentTitle().nsText == "Templat" then
ar_terms[1] = "كَلِمَة"
ar_terms.maxindex = 1
else
error(
'Please provide vocalized Arabic in the first parameter of {{[[Templat:ar-IPA|ar-IPA]]}}, or transliteration in the "tr" parameter.')
end
end
local pronunciations = {}
for i = 1, math.max(ar_terms.maxindex, transliterations.maxindex) do
local ar_term = ar_terms[i]
local tr = transliterations[i]
local qual = qualifiers[i]
if not (ar_term or tr) then
error("There is a gap in the parameters. Provide either |" .. i .. "= or |tr" .. i .. "=.")
elseif ar_term and tr then
mw.logObject("Duplicate parameters |" .. i .. "= and |tr" .. i .. "= in {{ar-IPA}},")
end
local pron = export.toIPA { term = ar_term, tr = tr }
table.insert(pronunciations, { pron = "/" .. pron .. "/", q = qual and { qual } or nil })
end
local anntext = ""
if args.ann then
anntext = args.ann
if args.ann:find("%+") then
local anndefs = {}
for i = 1, ar_terms.maxindex do
local ar_term = ar_terms[i]
if ar_term then
table.insert(anndefs, "'''" .. ar_term .. "'''")
end
end
if anndefs[1] then
anndefs = table.concat(anndefs, ", ")
anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs))
end
end
anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ": "
end
if nl then
return anntext .. require("Module:IPA").format_IPA_multiple(lang, pronunciations)
else
return anntext .. require("Module:IPA").format_IPA_full { lang = lang, items = pronunciations }
end
end
function export.show(frame)
local parent_args = frame:getParent().args
local process = require("Module:parameters").process
local params = {
[1] = {},
["audios"] = {},
["a"] = { alias_of = "audios" },
["ph"] = {},
["pagename"] = {},
["indent"] = {},
["ann"] = {},
}
local args = process(parent_args, params)
local pagename = args.pagename or mw.loadData("Module:headword/data").pagename
local indent = args.indent or "*"
local termspec = args[1] or "#"
local terms = parse_respellings_with_modifiers(termspec, 1)
local pronobj = export.get_pron_info(terms, pagename, 1)
local regional_phonetics = parse_regional_phonetics(args.ph, pagename)
local parts = {}
local function ins(text)
table.insert(parts, text)
end
local anntext = ""
if args.ann then
anntext = args.ann
if args.ann:find("%+") then
local anndefs = {}
for _, term in ipairs(terms) do
local respelling = term.respelling
if respelling and respelling:find("[ء-ي]") then
table.insert(anndefs, "'''" .. respelling .. "'''")
end
end
if anndefs[1] then
anndefs = table.concat(anndefs, ", ")
anntext = anntext:gsub("%+", require("Module:string utilities").replacement_escape(anndefs))
end
end
anntext = require("Module:qualifier").format_qualifier(anntext, "", "") .. ": "
end
if pronobj.pron_list and #pronobj.pron_list > 0 then
local formatted = require("Module:IPA").format_IPA_full { lang = lang, items = pronobj.pron_list }
ins(indent .. anntext .. mw.ustring.toNFC(formatted))
end
if args.audios then
local format_audio = require("Module:audio").format_audio
local audio_objs = parse_audio(lang, args.audios, pagename, "audios")
for i, audio_obj in ipairs(audio_objs) do
if #audio_objs > 1 and not audio_obj.caption then
audio_obj.caption = "Audio " .. i
end
ins("\n" .. indent .. " " .. format_audio(audio_obj))
end
end
if #regional_phonetics > 0 then
local m_IPA = require("Module:IPA")
local m_accent = require("Module:accent qualifier")
for _, regional in ipairs(regional_phonetics) do
local regions = regional.regions
local ipa = regional.ipa
local pron_item = { pron = "[" .. ipa .. "]" }
local formatted_ipa = m_IPA.format_IPA_full { lang = lang, items = { pron_item } }
local formatted_region = m_accent.format_qualifiers(lang, regions)
local line = "\n" .. indent .. indent .. " " .. formatted_region .. " " .. mw.ustring.toNFC(formatted_ipa)
if regional.audio then
local audio_obj = {
lang = lang,
file = regional.audio,
}
local textobj = require(audio_module).construct_audio_textobj(audio_obj)
audio_obj.text = textobj
line = line .. " " .. require("Module:audio").format_audio(audio_obj)
end
ins(line)
end
end
return concat(parts)
end
return export
psx2d0cilm9hbxy6gfbf7esbqrthwdf
Templat:ar-AFA
10
12379
375389
109891
2026-09-22T07:31:04Z
Hakimi97
2668
Kemas kini
375389
wikitext
text/x-wiki
{{#invoke:ar-pronunciation|show_old}}<noinclude>{{documentation}}</noinclude>
qeclh9o7e49x3yzowzjt9uk3z1pwafp
Australia
0
12875
375381
334113
2026-09-22T06:59:37Z
Hakimi97
2668
Pembetulan
375381
wikitext
text/x-wiki
==Bahasa Melayu==
{{wikipedia}}
[[Fail:Australia satellite plane.jpg|thumb|Gambar satelit Australia]]
===Kata nama khas===
{{ms-knk}}
# Sebuah negara di [[Oceania]] dengan ibu negaranya ialah [[Canberra]].
===Etimologi===
Daripada {{der|ms|en|Australia|}}.
===Sebutan===
* {{a|baku}} {{AFA|ms|/ɔ.stra.li.ja/}}
* {{a|Johor-Selangor}} {{AFA|ms|/au̯straliə/}}
* {{a|Riau-Lingga}} {{AFA|ms|/au̯stralia/}}
* {{penyempangan|ms|Au|stra|lia}}
===Tulisan Jawi===
{{ARchar|اوستراليا}}
===Sinonim===
* [[Komanwel Australia]], nama rasmi.
[[Kategori:ms:Negara]]
ey6m4e4opemvk21nfi74rb8334yt07m
Templat:ms-jawi
10
13235
375384
228927
2026-09-22T07:08:51Z
Hakimi97
2668
Kemas kini templat
375384
wikitext
text/x-wiki
{{#invoke:form of/templates|form_of_t|lang=ms|Ejaan [[w:Tulisan Jawi|Jawi]] bagi|withdot=true}}<noinclude>{{documentation}}</noinclude>
a33znw28lum5b9r9mse6rv40vwiycbm
saro
0
14248
375344
190601
2026-09-21T16:04:52Z
Muhamad Izzul Fiqih
7922
/* Bahasa Batak Mandailing */
375344
wikitext
text/x-wiki
== Bahasa Batak Mandailing ==
=== Takrifan ===
==== Kata nama ====
{{head|btm|kata nama}}
# [[bahasa]]
=== Sebutan ===
* {{penyempangan|btm|sa|ro}}
== Bahasa Melayu Jambi ==
===Takrifan===
====Kata nama====
{{inti|jax|kata sifat}}
# [[payah]]
guy3762r3vyymg4ihwctfftjkk99r38
375345
375344
2026-09-21T16:05:16Z
Muhamad Izzul Fiqih
7922
/* Kata nama */
375345
wikitext
text/x-wiki
== Bahasa Batak Mandailing ==
=== Takrifan ===
==== Kata nama ====
{{head|btm|kata nama}}
# [[bahasa]]
=== Sebutan ===
* {{penyempangan|btm|sa|ro}}
== Bahasa Melayu Jambi ==
===Takrifan===
====Kata nama====
{{inti|jax|kata sifat}}
# [[susah]]
es07wyeaoc32vv6bhkljvvssw7ycbhn
Modul:hyphenation
828
14282
375350
236146
2026-09-22T03:12:42Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708593|92708593]])
375350
Scribunto
text/plain
local export = {}
local decorations_module = "Module:decorations"
local links_module = "Module:links"
local parameters_module = "Module:parameters"
local function track(page)
require("Module:debug/track")("hyphenation/" .. page)
return true
end
--[==[
Meant to be called from a module. `data` is a table containing the following fields:
* `lang`: language object for the hyphenations or syllabifications;
* `hyphs`: a list of hyphenations/syllabifications, each described by an object which can contain the following fields:
** `hyph`: list of syllables comprising the hyphenation or syllabification, each a string;
** `q`: {nil} or a list of left regular qualifier strings, displayed directly before the hyphenation in question;
** `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the hyphenation in question;
** `a`: {nil} or a list of left accent qualifier strings (see [[Module:accent qualifier]]), displayed directly before
the hyphenation in question;
** `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the hyphenation in question;
** `refs`: {nil} or a list of references or reference specs to add directly after the hyphenation; the value of a list
item is either a string containing the reference text (typically a call to a citation template such as
{{tl|cite-book}}, or a template wrapping such a call), or an object with fields `text` (the reference text), `name`
(the name of the reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}})
and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or
{{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference
appropriately and insert a footnote number that hyperlinks to the actual reference, located in the
{{cd|<nowiki><references /></nowiki>}} section;
** `sc`: {nil} or script object for this particular hyphenation/syllabification;
* `q`: {nil} or a list of overall left regular qualifier strings, displayed before the initial caption;
* `qq`: {nil} or a list of overall right regular qualifier strings, displayed after all hyphenations;
* `a`: {nil} or a list of overall left accent qualifier strings (see [[Module:accent qualifier]]), displayed before the
initial caption;
* `aa`: {nil} or a list of right accent qualifier strings, displayed after all hyphenations;
* `sc`: {nil} or script object for the hyphenations/syllabifications;
* `caption`, {nil} or a string specifying the caption to use, in place of {"Hyphenation"}; e.g. use {"Syllabification"}
if what is passed in is actually a syllabification, as is common; a colon and space is automatically added after the
caption;
* `nocaption`: if true, suppress the caption display.
]==]
function export.format_hyphenations(data)
local hyphtexts = {}
for _, hyph in ipairs(data.hyphs) do
if #hyph.hyph == 0 then
error("Saw empty hyphenation; use || to separate hyphenations")
end
if hyph.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("`.qualifiers` is no longer supported; change the code to use `.qq` or `.q`")
end
local text = require(links_module).full_link {
lang = data.lang,
sc = hyph.sc or data.sc,
alt = table.concat(hyph.hyph, "‧"),
tr = "-",
q = hyph.q,
qq = hyph.qq,
a = hyph.a,
aa = hyph.aa,
refs = hyph.refs,
show_decorations = true,
}
table.insert(hyphtexts, text)
end
local text = ((data.nocaption and "") or ((data.caption or "Penyempangan") .. ": ")) .. table.concat(hyphtexts, ", ")
if data.qualifiers then
-- FIXME: added 2026-09-18; consider removing eventually.
error("overall `.qualifiers` is no longer supported; change the code to use `.qq` or `.q`")
end
if data.q and data.q[1] or data.qq and data.qq[1] or data.a and data.a[1] or data.aa and data.aa[1] then
text = require(decorations_module).format_decorations {
lang = data.lang,
text = text,
q = data.q,
qq = data.qq,
a = data.a,
aa = data.aa,
}
end
return text
end
--[==[
Entry point for {{tl|hyphenation}} template (also written {{tl|hyph}}).
]==]
function export.hyphenation(frame)
local parent_args = frame:getParent().args
local compat = parent_args.lang
local lang_param = compat and "lang" or 1
local offset = compat and 0 or 1
local params = {
[lang_param] = {required = true, type = "language", default = "und"},
[1 + offset] = {list = true, required = true, allow_holes = true, default = "{{{2}}}"},
-- FIXME: For compatibility, q= and qq= refer to the first hyphenation rather than overall. Consider tracking and
-- changing this.
["q"] = {list = true, allow_holes = true, type = "qualifier"},
["qq"] = {list = true, allow_holes = true, type = "qualifier"},
["a"] = {list = true, allow_holes = true, separate_no_index = true, type = "labels"},
["aa"] = {list = true, allow_holes = true, separate_no_index = true, type = "labels"},
["ref"] = {list = true, allow_holes = true, type = "references"},
["caption"] = {},
["nocaption"] = {type = "boolean"},
["sc"] = {list = true, allow_holes = true, separate_no_index = true, type = "script"},
}
local args = require(parameters_module).process(parent_args, params)
local lang = args[lang_param]
local sc = args.sc.default
local data = {
lang = lang,
sc = sc,
hyphs = {},
caption = args.caption,
nocaption = args.nocaption,
a = args.a.default,
aa = args.aa.default,
}
local this_hyph = {hyph = {}}
local maxindex = args[1 + offset].maxindex
local function insert_hyph()
local hyphnum = #data.hyphs + 1
this_hyph.q = args.q[hyphnum]
this_hyph.qq = args.qq[hyphnum]
this_hyph.a = args.a[hyphnum]
this_hyph.aa = args.aa[hyphnum]
this_hyph.refs = args.ref[hyphnum]
table.insert(data.hyphs, this_hyph)
end
for i=1, maxindex do
local syl = args[1 + offset][i]
if not syl then
insert_hyph()
this_hyph = {hyph = {}}
else
table.insert(this_hyph.hyph, syl)
end
end
insert_hyph()
return export.format_hyphenations(data)
end
return export
noi6hiwft43y5yinfixbo8jrkhbmo8m
Modul:affixusex
828
22134
375357
223660
2026-09-22T03:14:12Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708532|92708532]])
375357
Scribunto
text/plain
local export = {}
local decorations_module = "Module:decorations"
local etymology_module = "Module:etymology"
local links_module = "Module:links"
-- main function
--[==[
Format the affix usexes in `data`. We more or less simply call full_link() on each item, along with the
associated params, to format the link, but need some special-casing for affixes. On input, the `data`
object contains the following fields:
* `lang` ('''required'''): Overall language object; default for items not specifying their own language.
* `sc`: Overall script object; default for items not specifying their own script.
* `items`: List of items. Each is an object with the following fields:
** `term`: The term (affix or resulting term).
** `gloss`, `tr`, `ts`, `genders`, `alt`, `id`, `lit`, `pos`, `ng`: The same as for `full_links()` in [[Module:links]].
** `lang`: Language of the term. Should only be set when the term has its own language, and will cause the
language to be displayed before the term. Defaults to the overall `lang`.
** `sc`: Script of the term. Defaults to the overall `sc`.
** `fulljoiner`: Text of the separator appearing before the item, including spaces. Takes precedence over `joiner`
and `arrow`.
** `joiner`: Text of the separator appearing before the item, not including spaces. Takes precedence over `arrow`.
** `arrow`: If specified, the separator is a right arrow. If none of `fulljoiner`, `joiner` and `arrow` are given,
the separator is a right arrow if it's the last item, otherwise a plus sign if it's not the first item, otherwise
there's no displayed separator.
** `l`: Left labels for the term.
** `ll`: Right labels for the term.
** `q`: Left regular qualifier(s) for the term.
** `qq`: Right regular qualifier(s) for the term.
** `refs`: References for the term, in the structure expected by [[Module:references]].
* `lit`: Overall literal meaning.
* `l`: Overall left labels.
* `ll`: Overall right labels.
* `q`: Overall left regular qualifier(s).
* `qq`: Overall right regular qualifier(s).
'''WARNING:''' This destructively modifies the `items` objects (specifically by adding default values for `lang` and
`sc`).
]==]
function export.format_affixusex(data)
local result = {}
for index, item in ipairs(data.items) do
if item.fulljoiner then
table.insert(result, item.fulljoiner)
elseif item.joiner then
table.insert(result, " " .. item.joiner .. " ")
elseif index == #data.items or item.arrow then
table.insert(result, " → ")
elseif index > 1 then
table.insert(result, " + ")
end
table.insert(result, "‎")
local text
local item_lang_specific = item.lang
item.lang = item.lang or data.lang
item.sc = item.sc or data.sc
if item_lang_specific then
text = require(etymology_module).format_derived {
sources = {item.lang},
terms = {item},
-- Don't need to specify `nocat = true` because we don't pass in `lang`.
template_name = "affixusex",
decorations_on_outside = true,
}
else
text = require(links_module).full_link(item, "term")
end
table.insert(result, text)
end
result = table.concat(result) .. (data.lit and ", secara harfiah " ..
require(links_module).mark(data.lit, "gloss") or "")
if data.q and data.q[1] or data.qq and data.qq[1] or data.l and data.l[1] or data.ll and data.ll[1] then
result = require(decorations_module).format_decorations {
lang = data.lang,
text = result,
q = data.q,
qq = data.qq,
l = data.l,
ll = data.ll,
}
end
return result
end
return export
lyzn6lkv4gd5qxbwwauzu6e1iowi4c8
Modul:pendokumenan
828
24368
375361
335738
2026-09-22T03:50:53Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92487256|92487256]])
375361
Scribunto
text/plain
local export = {}
local debug_track_module = "Modul:debug/track"
local frame_module = "Modul:frame"
local fun_is_callable_module = "Modul:fun/isCallable"
local languages_module = "Modul:languages"
local links_module = "Modul:links"
local load_module = "Modul:load"
local module_categorization_module = "Modul:module categorization"
local number_list_show_module = "Modul:number list/show"
local chemical_element_list_show_module = "Modul:chemical element list/show"
local pages_module = "Modul:pages"
local parameters_module = "Modul:parameters"
local scripts_module = "Modul:scripts"
local string_endswith_module = "Modul:string/endswith"
local string_gline_module = "Modul:string/gline"
local string_startswith_module = "Modul:string/startswith"
local string_utilities_module = "Modul:string utilities"
local template_parser_module = "Modul:template parser"
local title_exists_module = "Modul:title/exists"
local title_new_title_module = "Modul:title/newTitle"
local concat = table.concat
local error = error
local full_url = mw.uri.fullUrl
local get_current_title = mw.title.getCurrentTitle
local insert = table.insert
local ipairs = ipairs
local list_to_text = mw.text.listToText
local new_message = mw.message.new
local pcall = pcall
local require = require
local tonumber = tonumber
local tostring = tostring
local type = type
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local function categorize_module(...)
categorize_module = require(module_categorization_module).categorize_module
return categorize_module(...)
end
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function endswith(...)
endswith = require(string_endswith_module)
return endswith(...)
end
local function expand_template(...)
expand_template = require(frame_module).expandTemplate
return expand_template(...)
end
local function find_templates(...)
find_templates = require(template_parser_module).find_templates
return find_templates(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_lang(...)
get_lang = require(languages_module).getByCode
return get_lang(...)
end
local function get_pagetype(...)
get_pagetype = require(pages_module).get_pagetype
return get_pagetype(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function gline(...)
gline = require(string_gline_module)
return gline(...)
end
local function is_callable(...)
is_callable = require(fun_is_callable_module)
return is_callable(...)
end
local function is_documentation(...)
is_documentation = require(pages_module).is_documentation
return is_documentation(...)
end
local function is_sandbox(...)
is_sandbox = require(pages_module).is_sandbox
return is_sandbox(...)
end
local function new_title(...)
new_title = require(title_new_title_module)
return new_title(...)
end
local function number_list_show_table(...)
number_list_show_table = require(number_list_show_module).table
return number_list_show_table(...)
end
local function chemical_element_list_show_table(...)
chemical_element_list_show_table = require(chemical_element_list_show_module).table
return chemical_element_list_show_table(...)
end
local function preprocess(...)
preprocess = require(frame_module).preprocess
return preprocess(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function safe_load_data(...)
safe_load_data = require(load_module).safe_load_data
return safe_load_data(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local function startswith(...)
startswith = require(string_startswith_module)
return startswith(...)
end
local function title_exists(...)
title_exists = require(title_exists_module)
return title_exists(...)
end
local function ugsub(...)
ugsub = require(string_utilities_module).gsub
return ugsub(...)
end
local function umatch(...)
umatch = require(string_utilities_module).match
return umatch(...)
end
local skins = {
["common"] = "",
["vector"] = "Vector",
["monobook"] = "Monobook",
["cologneblue"] = "Cologne Blue",
["modern"] = "Modern",
}
local function track(page)
debug_track("documentation/" .. page)
return true
end
local function compare_pages(page1, page2, text)
return "[" .. tostring(
full_url("Khas:ComparePages", { page1 = page1, page2 = page2 }))
.. " " .. text .. "]"
end
-- Avoid transcluding [[Modul:languages/cache]] everywhere.
local lang_cache = setmetatable({}, {
__index = function(self, k)
return require("Modul:languages/cache")[k]
end
})
local function zh_link(word)
return full_link {
lang = lang_cache.zh,
term = word
}
end
local function make_languages_data_documentation(_title, cats, division)
local doc_template, module_cat
if endswith(division, "/extra") then
division = division:sub(1, -7)
doc_template = "language extradata documentation"
module_cat = "Modul data ekstra bahasa"
else
doc_template = "language data documentation"
module_cat = "Modul data bahasa"
end
local sort_key
if division == "exceptional" then
sort_key = "x"
else
sort_key = division:gsub("/", "")
end
insert(cats, module_cat .. "|" .. sort_key)
return {
title = doc_template
}
end
local function make_Unicode_data_documentation(title, _cats)
local subpage, first_three_of_code_point
= title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$")
if subpage == "names" or subpage == "images" or subpage == "emoji images" then
local low, high =
tonumber(first_three_of_code_point .. "000", 16),
tonumber(first_three_of_code_point .. "FFF", 16)
local text, text_type
if subpage == "names" then
text_type = "titles of images"
elseif subpage == "images" then
text_type = "titles of images"
elseif subpage == "emoji images" then
text_type = "emoji-style images"
end
text = string.format(
"Modul data ini mengandungi " .. text_type .. " kepada " ..
"titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.",
low, high)
if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then
text = text ..
" Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" ..
first_three_of_code_point .. "]]."
elseif subpage == "emoji images" then
text = text ..
" Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]."
end
return text
end
end
local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat)
local lang = lang_cache[langcode]
if lang then
local langname
if lang._fullCode then
langname = lang_cache[lang._fullCode]:getCanonicalName()
else
langname = lang:getCanonicalName()
end
insert(cats, overall_data_module_cat .. "|" .. langname)
insert(cats, "Modul bahasa " .. langname)
insert(cats, "Modul data bahasa " .. langname)
return lang, langname
end
end
--[=[
This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't
unnecessarily cluttered. It is a list of tables, each of which have the following possible fields:
`regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used.
Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the
second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the
`process` function as the third and subsequent parameters.
`process` (optional): This may be a function or a string. If it is a function, it is called as follows:
`process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)`
where:
* TITLE is a title object describing the module's title; see
[https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects].
* CATS is an list of categories that the module will be added to.
* CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field.
The return value of `process` should either be a string (which will be used as the module's documentation), or a
table specifying the name of a template to expand to get the documentation, along with the arguments to that
template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title`
field, and any arguments should be in `args; in this case, the template name will be listed above the generated
documentation as the source of the documentation, along with an edit button to edit the template's contents.
If, however, the return value of the `process` function is a string, any template invocations will be expanded
using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation.
If `process` itself is a string rather than a function, it should name a submodule under
[[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule
will be specified as the source of the documentation (unless it returns a table naming a template to expand to get
the documentation, as described above).
If `process` is omitted entirely, the module will have no documentation.
`cat` (optional): A string naming the category into which the module should be placed, or a list of such strings.
Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second,
etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS
list (the second parameter).
]=]
local module_regex = {
{
regex = "^Modul:languages/data/(3/%l/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(3/%l)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:scripts/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:data tables/data..?.?.?$",
cat = "Jadual data berpecah modul rujukan",
},
{
regex = "^Modul:zh/data/dial%-pron/.+$",
cat = "Modul data sebutan dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/dial%-syn/.+$",
cat = "Modul data sinonim dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/glyph%-data/.+$",
cat = "Modul data bentuk aksara Cina bersejarah",
process = function(title, _cats)
local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)")
if character then
return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.")
:format(zh_link(character))
end
end,
},
{
regex = "^Modul:zh/data/ltc%-pron/(.+)$",
cat = "Modul data sebutan bahasa Cina Pertengahan|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-BS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1",
process = "zh data",
},
{
-- capture rest of zh/data submodules
regex = "^Modul:zh/data/(.+)$",
cat = "Modul data bahasa Cina|%1",
},
{
regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$",
process = "guoxue-data",
},
{
regex = "^Modul:Unicode data/(.+)$",
cat = "Modul data Unicode|%1",
process = make_Unicode_data_documentation,
},
{
regex = "^Modul:number list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor")
if lang then
return ("This module contains data on various types of numbers in %s.\n%s")
:format(lang:makeCategoryLink(), number_list_show_table() or "")
end
end,
},
{
regex = "^Modul:chemical element list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia")
if lang then
return ("This module contains data on chemical elements in %s.\n%s")
:format(lang:makeCategoryLink(), chemical_element_list_show_table() or "")
end
end,
},
{
regex = "^Modul:accel/(.+)$",
process = function(title, cats)
local lang_code = title.subpageText
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel")
insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName()))
return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.")
:format(lang:makeCategoryLink())
end
end,
},
{
regex = "^Modul:inc%-ash/dial/data/(.+)$",
cat = "Modul Prakrit Ashoka|%1",
process = function(title, _cats)
local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$")
if word then
local lang = lang_cache["inc-ash"]
return ("This module contains data on the pronunciation of %s in dialects of %s.")
:format(full_link({ term = word, lang = lang }, "term"),
lang:makeCategoryLink())
end
end,
},
{
regex = "^.+%-translit$",
process = function(title, _cats)
return require("Modul:documentation/translit-like").documentation {
operation = "translit",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-sortkey$",
process = function(title, _cats)
return require("Modul:documentation/translit-like").documentation {
operation = "sortkey",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-stripdiacritics$",
process = function(title, _cats)
return require("Modul:documentation/translit-like").documentation {
operation = "strip diacritics",
title_without_namespace = title.text,
}
end,
},
{
regex = "^Modul:form of/lang%-data/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa")
if lang then
-- FIXME, display more info.
return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " ..
langname .. "."
end
end
},
{
regex = "^Modul:labels/data/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa")
if lang then
return {
title = "label language-specific data documentation",
args = { [1] = lang_code },
}
end
end
},
{
regex = "^Modul:category tree/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang")
if lang then
return "This module handles generating the descriptions and categorization for " ..
langname .. " category pages "
.. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are "
.. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. "
.. "This module is part of the category tree system, which is a general framework for generating the "
.. "descriptions and categorization of category pages.\n\n"
.. "For more information, see [[Modul:category tree/lang/documentation]].\n\n"
.. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the "
.. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized."
end
end
},
{
regex = "^Modul:category tree/topic/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/topic| ")
return {
title = "topic cat data submodule documentation"
}
end
},
{
regex = "^Modul:category tree/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/grammar| ")
return {
title = "category tree data submodule documentation"
}
end
},
{
regex = "^Modul:ja/data/(.+)$",
cat = "Modul data bahasa Jepun|%1",
},
{
regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$",
cat = "Modul atlas data dialek Finland|%1",
process = function(_title, _cats, shard)
return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " ..
"''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " ..
"It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " ..
"''Kotimaisten kielten keskus'' under the CC BY 4.0 license."
end
},
{
regex = "^Modul:fi%-dialects/data/feature/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:fi%-dialects/data/word/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:Swadesh/data/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "."
end
end
},
{
regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$",
process = function(_title, cats, lang_code, variety)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the "
local etym_lang = get_lang(variety, nil, "allow etym")
if etym_lang then
return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname)
end
local script = get_script(variety)
if script then
return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang))
end
return ("%s %s variety of %s."):format(prefix, variety, langname)
end
end
},
{
regex = "^Modul:typing%-aids",
process = function(title, cats)
local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$")
local sortkey
if data_suffix then
if data_suffix:find "^[%l-]+$" then
local lang = get_lang(data_suffix)
if lang then
sortkey = lang:getCanonicalName()
insert(cats, "Modul data bahasa " .. sortkey)
end
elseif data_suffix:find "^%u%l%l%l$" then
local script = get_script(data_suffix)
if script then
-- FIXME: no lang to pass here
sortkey = script:getCanonicalName()
insert(cats, script:getCategoryName())
end
end
insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix))
end
end,
},
{
regex = "^Modul:R:([%l-]+):(.+)$",
process = function(_title, cats, lang_code, refname)
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname)
insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName()))
return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}."
end
end,
},
{
regex = "^Modul:Quotations/([%l-]+)/?(.*)",
process = "Quotation",
},
{
regex = "^Modul:affix/lang%-data/([%l-]+)",
process = "affix lang-data",
},
{
regex = "^Modul:dialect synonyms/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "| ")
return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " ..
"{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" ..
"==== Struktur modul data bahasa ====\n" ..
"* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" ..
"* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" ..
"* <code>export.notes</code> — optional; table of note keys to text.\n" ..
"* <code>export.sources</code> — optional; table of source keys to text.\n" ..
"* <code>export.note_aliases</code> — optional; alias map for notes.\n" ..
"* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" ..
expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } })
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$",
process = function(_title, cats, lang_code, term)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term } }))
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$",
process = function(_title, cats, lang_code, term, id)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term, id = id } }))
end
end,
},
{
regex = "^Modul:bibliography/data/([%l-]+)$",
process = function(title, cats, lang_code)
if lang_code == "preload" then
return 'Used as a base model for other languages when the button "create new language submodule" is clicked.'
end
local page = require(title.fullText).bib_page
if not page then
page = lang_cache[lang_code]:getCanonicalName()
if page then
insert(cats, "Modul bahasa " .. page)
end
end
insert(cats, "Modul rujukan")
return "This module holds bibliographical data for " ..
page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''."
end,
},
}
function export.show(frame)
local boolean_default_false = { type = "boolean", default = false }
local args = process_params(frame.args, {
["hr"] = true,
["for"] = true,
["from"] = true,
["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage.
["notsubpage"] = boolean_default_false,
["nodoc"] = boolean_default_false,
["nolinks"] = boolean_default_false, -- suppress all "Useful links"
["nosandbox"] = boolean_default_false, -- supress sandbox
})
local output = {'\n<div class="documentation" style="display:block; clear:both">\n'}
local function ins(txt)
insert(output, txt)
end
local cats = {}
local function inscat(cat)
insert(cats, cat)
end
local nodoc = args.nodoc
if (not args.hr) or (args.hr == "above") then
ins("----\n")
end
local title = args["for"] and new_title(args["for"]) or get_current_title()
local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil
local contentModel = title.contentModel
local pagetype, is_script_or_stylesheet = get_pagetype(title)
local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc
local doc_content_source = "Modul:pendokumenan"
local auto_generated_cat_source
local cats_auto_generated = false
if not args.allowondoc and is_documentation(title) then
-- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type.
error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].")
elseif is_sandbox(title) then
local sandbox_ns = title.nsText
preload = ("Templat:pendokumenan/preload%s%sSandbox"):format(
sandbox_ns == "Modul" and sandbox_ns or "Templat",
title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or ""
)
elseif pagetype:match("%f[%w]gadget%f[%W]") then
preload = "Templat:pendokumenan/preloadGadget"
elseif pagetype:match("%f[%w]script%f[%W]") then -- .js
if title.nsText == "MediaWiki" then
preload = "Templat:pendokumenan/preloadMediaWikiJavaScript"
else
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
end
is_script_or_stylesheet = true
elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
is_script_or_stylesheet = true
elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto.
preload = "Templat:pendokumenan/preloadModule"
elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then
preload = "Templat:pendokumenan/preloadTemplate"
end
if doc_title and doc_title.isRedirect then
old_doc_title = doc_title
doc_title = doc_title.redirectTarget
end
ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">")
local function get_module_doc_and_cats(categories_only)
cats_auto_generated = true
local automatic_cats = nil
if user_name then
fallback_docs = "pendokumenan/fallback/user module"
automatic_cats = { "Modul kotak pasir pengguna" }
else
for _, data in ipairs(module_regex) do
local captures = { umatch(title.fullText, data.regex) }
if #captures > 0 then
local cat, process_function
if is_callable(data.process) then
process_function = data.process
elseif type(data.process) == "string" then
doc_content_source = "Modul:pendokumenan/functions/" .. data.process
process_function = require(doc_content_source)
end
if process_function then
doc_content = process_function(title, cats, unpack(captures))
end
if type(doc_content) == "table" then
doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source
doc_content = expand_template(doc_content)
elseif doc_content ~= nil then
doc_content = preprocess(doc_content)
end
cat = data.cat
if cat then
if type(cat) == "string" then
cat = { cat }
end
for _, c in ipairs(cat) do
insert(cats, (ugsub(title.fullText, data.regex, c)))
end
end
break
end
end
end
if title.subpageText == "templates" then
inscat("Modul antara muka templat")
end
if automatic_cats then
for _, c in ipairs(automatic_cats) do
inscat(c)
end
end
if #cats == 0 then
local auto_cats = categorize_module {
return_raw = true,
noerror = true,
}
if #auto_cats > 0 then
auto_generated_cat_source = "Modul:module categorization"
end
for _, category in ipairs(auto_cats) do
inscat(category)
end
end
-- meaning module is not in user’s sandbox or one of many datamodule boring series
needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs)
end
-- Override automatic documentation, if present.
if doc_title and doc_title.exists then
local cats_auto_generated_text = ""
if contentModel == "Scribunto" then
local doc_page_content = doc_title.content
-- Track then do nothing if there are uses of includeonly. The
-- pattern is slightly too permissive, but any false-positives are
-- obvious typos that should be corrected.
if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then
track("module-includeonly")
else
-- Check for uses of {{module cat}}. find_templates treats the
-- input as transcluded by default (i.e. it parses the wikitext
-- which will be transcluded through to the module page).
local module_cat
for template in find_templates(doc_page_content) do
if Templat:get_name() == "module cat" then
module_cat = true
break
end
end
if not module_cat then
get_module_doc_and_cats("categories only")
auto_generated_cat_source = auto_generated_cat_source or doc_content_source
cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" ..
auto_generated_cat_source .. "]]. <sup>[[" ..
new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>"
end
end
end
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" ..
doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" ..
cats_auto_generated_text .. "</i></dd>")
else
if contentModel == "Scribunto" then
get_module_doc_and_cats(false)
elseif title.nsText == "Templat" then
--inscat("Uncategorized templates")
needs_doc = not (fallback_docs or nodoc)
elseif user_name and is_script_or_stylesheet then
skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")]
if skin_name then
fallback_docs = "pendokumenan/fallback/user " .. contentModel
end
end
if doc_content then
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " ..
"dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" ..
new_title(doc_content_source):fullUrl { action = "edit" } ..
" sunting]]</sup> </i></dd>")
elseif not nodoc then
if doc_title then
ins(
"<dd><i style=\"font-size: larger;\">Laman " .. pagetype ..
" ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " ..
(fallback_docs and "Anda boleh " or "Minta tolong ") ..
"[" .. doc_title:fullUrl { action = "edit", preload = preload }
.. " ciptakan laman pendokumenan tersebut].</i></dd>\n")
else
ins(
"<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " ..
"pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n")
end
end
end
if startswith(title.fullText, "MediaWiki:Gadget-") then
local is_gadget = false
for line in gline(new_title("MediaWiki:Gadgets-definition").content) do
local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$")
if not gadget then
gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$")
end
if gadget then
items = split(items, "|")
for i, item in ipairs(items) do
if title.fullText == ("MediaWiki:Gadget-" .. item) then
is_gadget = true
ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>")
ins(gadget)
ins("</code> gajet ([")
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(" sunting takrifan])'' <dl>")
ins("<dd> ''Huraian ([")
ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" })))
ins(" sunting])'': ")
ins(preprocess(new_message('Gadget-' .. gadget):plain()))
ins(" </dd>")
table.remove(items, i)
if #items > 0 then
for j, item in ipairs(items) do
items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]'
end
ins("<dd> ''Bahagian lain'': ")
ins(list_to_text(items))
ins("</dd>")
end
ins("</dl></dd>")
break
end
end
end
end
if not is_gadget then
ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [")
ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" })))
ins(' gajet] ([')
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(' sunting takrifan]).</dd>')
-- else
-- inscat("Wiktionary gadgets")
end
end
if old_doc_title then
ins("<dd> ''Dilencong daripada'' [")
ins(old_doc_title:fullUrl { redirect = "no" })
ins(" ")
ins(old_doc_title.fullText)
ins("] ([")
ins(old_doc_title:fullUrl { action = "edit" })
ins(" sunting]).</dd>\n")
end
if not args.nolinks then
local links = {}
local function inslinks(txt)
insert(links, txt)
end
if title.isSubpage and not args.notsubpage then
inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]")
inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]")
else
inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]")
end
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) ..
" pautan]")
if contentModel ~= "Scribunto" then
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) ..
" lencongan]")
end
if is_script_or_stylesheet then
if user_name then
inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]")
end
else
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) ..
" transklusi]")
end
if contentModel == "Scribunto" then
local is_testcases = title.isSubpage and title.subpageText == "testcases"
local without_subpage = title.nsText .. ":" .. title.baseText
if is_testcases then
inslinks("[[:" .. without_subpage .. "|modul ujian]]")
else
inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]")
end
if user_name then
inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]")
inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]")
inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]")
-- If sandbox module, add a link to the module that this is a sandbox of.
-- Exclude user sandbox modules like [[User:Dine2016/sandbox]].
elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then
inscat("Modul kotak pasir")
-- Sandbox modules don’t really need documentation.
needs_doc = false
-- Don't track user sandbox modules.
local text_title = new_title(title.text)
if not (text_title and text_title.nsText == "Pengguna") then
local diff
local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]")
if sandbox_of then
track("sandbox to be moved")
else
sandbox_of = title.text:match("^sandbox%d*/(.*)$")
end
if not sandbox_of then
error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'")
:format(title.text))
end
sandbox_of = title.nsText .. ":" .. sandbox_of
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
end
-- If not a sandbox module, add link to sandbox module.
-- Sometimes there are multiple sandboxes for a single Modul:
-- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]].
else
local sandbox_title
local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$")
if not user_prefix then
user_prefix = ""
user_rest = title.text
end
sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest
local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks(sandbox_link .. (diff or ""))
end
end
if title.nsText == "Templat" then
-- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]]
-- apparently same with/without: &profile=advanced&fulltext=1
local errorq = 'searchengineselect=mediawiki&search=all: hasTemplat:\"' ..
title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:'
local eincategory =
"Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors"
inslinks(
'[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]'
.. ' (' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]'
.. '/' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]'
.. ')'
)
if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template.
-- At the moment there are no user sandbox templates with subpage
-- “/sandbox”.
inscat("Templat kotak pasir")
-- Sandbox templates don’t really need documentation.
needs_doc = false
-- Will behave badly if “/sandbox” occurs twice in title!
local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "")
local diff
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
-- This is a template that can have a sandbox.
elseif not args.nosandbox then -- unless we tell it not to
local sandbox_title = title.fullText .. "/sandbox"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or ""))
end
end
if #links > 0 then
ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>")
end
end
ins("</dl>\n")
-- Show error from [[Modul:category tree/topic cat/data]] on its submodules'
-- documentation to, for instance, warn about duplicate labels.
if startswith(title.fullText, "Modul:category tree/topic/") then
local ok, err = pcall(require, "Modul:category tree/topic/data")
if not ok then
ins('<span class="error">' .. err .. '</span>\n\n')
end
end
if doc_title and doc_title.exists then
-- Override automatic documentation, if present.
doc_content = expand_template { title = doc_title.fullText }
elseif not doc_content and fallback_docs then
doc_content = expand_template {
title = fallback_docs,
args = {
['user'] = user_name,
['page'] = title.fullText,
['skin name'] = skin_name,
},
}
end
if doc_content then
ins(doc_content)
end
ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br"))
if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then
if contentModel == "Scribunto" then
inscat("Modul belum dikategorikan")
-- elseif title.nsText == "Templat" then
-- inscat("Templat belum dikategorikan")
end
end
if needs_doc then
inscat("Templat dan modul yang memerlukan pendokumenan")
end
for _, cat in ipairs(cats) do
ins("[[Kategori:" .. cat .. "]]")
end
ins("</div>\n")
return concat(output)
end
function export.module_auto_doc_table()
local parts = {}
local function ins(text)
insert(parts, text)
end
ins('{|class="wikitable"')
ins("! Regex !! Kategori !! Modul yang dikendalikan")
for _, spec in ipairs(module_regex) do
local cat_text
local cats = spec.cat
if cats then
local cat_parts = {}
if type(cats) == "string" then
cats = { cats }
end
for _, cat in ipairs(cats) do
insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "|"))))
end
cat_text = concat(cat_parts, ", ")
else
cat_text = "''(tidak dinyatakan secara khusus)''"
end
ins("|-")
ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text,
is_callable(spec.process) and "''(dikendali secara dalaman)''" or
type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or
"''(tiada penjana pendokumenan)''"))
end
ins("|}")
return concat(parts, "\n")
end
return export
pllus5uhsp0qm036oww1d624tmv9bxr
375362
375361
2026-09-22T04:13:41Z
Hakimi97
2668
Pembetulan kecil
375362
Scribunto
text/plain
local export = {}
local debug_track_module = "Modul:debug/track"
local frame_module = "Modul:frame"
local fun_is_callable_module = "Modul:fun/isCallable"
local languages_module = "Modul:languages"
local links_module = "Modul:links"
local load_module = "Modul:load"
local module_categorization_module = "Modul:module categorization"
local number_list_show_module = "Modul:number list/show"
local chemical_element_list_show_module = "Modul:chemical element list/show"
local pages_module = "Modul:pages"
local parameters_module = "Modul:parameters"
local scripts_module = "Modul:scripts"
local string_endswith_module = "Modul:string/endswith"
local string_gline_module = "Modul:string/gline"
local string_startswith_module = "Modul:string/startswith"
local string_utilities_module = "Modul:string utilities"
local template_parser_module = "Modul:template parser"
local title_exists_module = "Modul:title/exists"
local title_new_title_module = "Modul:title/newTitle"
local concat = table.concat
local error = error
local full_url = mw.uri.fullUrl
local get_current_title = mw.title.getCurrentTitle
local insert = table.insert
local ipairs = ipairs
local list_to_text = mw.text.listToText
local new_message = mw.message.new
local pcall = pcall
local require = require
local tonumber = tonumber
local tostring = tostring
local type = type
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local function categorize_module(...)
categorize_module = require(module_categorization_module).categorize_module
return categorize_module(...)
end
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function endswith(...)
endswith = require(string_endswith_module)
return endswith(...)
end
local function expand_template(...)
expand_template = require(frame_module).expandTemplate
return expand_template(...)
end
local function find_templates(...)
find_templates = require(template_parser_module).find_templates
return find_templates(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_lang(...)
get_lang = require(languages_module).getByCode
return get_lang(...)
end
local function get_pagetype(...)
get_pagetype = require(pages_module).get_pagetype
return get_pagetype(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function gline(...)
gline = require(string_gline_module)
return gline(...)
end
local function is_callable(...)
is_callable = require(fun_is_callable_module)
return is_callable(...)
end
local function is_documentation(...)
is_documentation = require(pages_module).is_documentation
return is_documentation(...)
end
local function is_sandbox(...)
is_sandbox = require(pages_module).is_sandbox
return is_sandbox(...)
end
local function new_title(...)
new_title = require(title_new_title_module)
return new_title(...)
end
local function number_list_show_table(...)
number_list_show_table = require(number_list_show_module).table
return number_list_show_table(...)
end
local function chemical_element_list_show_table(...)
chemical_element_list_show_table = require(chemical_element_list_show_module).table
return chemical_element_list_show_table(...)
end
local function preprocess(...)
preprocess = require(frame_module).preprocess
return preprocess(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function safe_load_data(...)
safe_load_data = require(load_module).safe_load_data
return safe_load_data(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local function startswith(...)
startswith = require(string_startswith_module)
return startswith(...)
end
local function title_exists(...)
title_exists = require(title_exists_module)
return title_exists(...)
end
local function ugsub(...)
ugsub = require(string_utilities_module).gsub
return ugsub(...)
end
local function umatch(...)
umatch = require(string_utilities_module).match
return umatch(...)
end
local skins = {
["common"] = "",
["vector"] = "Vector",
["monobook"] = "Monobook",
["cologneblue"] = "Cologne Blue",
["modern"] = "Modern",
}
local function track(page)
debug_track("pendokumenan/" .. page)
return true
end
local function compare_pages(page1, page2, text)
return "[" .. tostring(
full_url("Khas:ComparePages", { page1 = page1, page2 = page2 }))
.. " " .. text .. "]"
end
-- Avoid transcluding [[Modul:languages/cache]] everywhere.
local lang_cache = setmetatable({}, {
__index = function(self, k)
return require("Modul:languages/cache")[k]
end
})
local function zh_link(word)
return full_link {
lang = lang_cache.zh,
term = word
}
end
local function make_languages_data_documentation(_title, cats, division)
local doc_template, module_cat
if endswith(division, "/extra") then
division = division:sub(1, -7)
doc_template = "language extradata documentation"
module_cat = "Modul data ekstra bahasa"
else
doc_template = "language data documentation"
module_cat = "Modul data bahasa"
end
local sort_key
if division == "exceptional" then
sort_key = "x"
else
sort_key = division:gsub("/", "")
end
insert(cats, module_cat .. "|" .. sort_key)
return {
title = doc_template
}
end
local function make_Unicode_data_documentation(title, _cats)
local subpage, first_three_of_code_point
= title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$")
if subpage == "names" or subpage == "images" or subpage == "emoji images" then
local low, high =
tonumber(first_three_of_code_point .. "000", 16),
tonumber(first_three_of_code_point .. "FFF", 16)
local text, text_type
if subpage == "names" then
text_type = "titles of images"
elseif subpage == "images" then
text_type = "titles of images"
elseif subpage == "emoji images" then
text_type = "emoji-style images"
end
text = string.format(
"Modul data ini mengandungi " .. text_type .. " kepada " ..
"titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.",
low, high)
if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then
text = text ..
" Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" ..
first_three_of_code_point .. "]]."
elseif subpage == "emoji images" then
text = text ..
" Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]."
end
return text
end
end
local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat)
local lang = lang_cache[langcode]
if lang then
local langname
if lang._fullCode then
langname = lang_cache[lang._fullCode]:getCanonicalName()
else
langname = lang:getCanonicalName()
end
insert(cats, overall_data_module_cat .. "|" .. langname)
insert(cats, "Modul bahasa " .. langname)
insert(cats, "Modul data bahasa " .. langname)
return lang, langname
end
end
--[=[
This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't
unnecessarily cluttered. It is a list of tables, each of which have the following possible fields:
`regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used.
Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the
second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the
`process` function as the third and subsequent parameters.
`process` (optional): This may be a function or a string. If it is a function, it is called as follows:
`process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)`
where:
* TITLE is a title object describing the module's title; see
[https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects].
* CATS is an list of categories that the module will be added to.
* CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field.
The return value of `process` should either be a string (which will be used as the module's documentation), or a
table specifying the name of a template to expand to get the documentation, along with the arguments to that
template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title`
field, and any arguments should be in `args; in this case, the template name will be listed above the generated
documentation as the source of the documentation, along with an edit button to edit the template's contents.
If, however, the return value of the `process` function is a string, any template invocations will be expanded
using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation.
If `process` itself is a string rather than a function, it should name a submodule under
[[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule
will be specified as the source of the documentation (unless it returns a table naming a template to expand to get
the documentation, as described above).
If `process` is omitted entirely, the module will have no documentation.
`cat` (optional): A string naming the category into which the module should be placed, or a list of such strings.
Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second,
etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS
list (the second parameter).
]=]
local module_regex = {
{
regex = "^Modul:languages/data/(3/%l/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(3/%l)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:scripts/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:data tables/data..?.?.?$",
cat = "Jadual data berpecah modul rujukan",
},
{
regex = "^Modul:zh/data/dial%-pron/.+$",
cat = "Modul data sebutan dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/dial%-syn/.+$",
cat = "Modul data sinonim dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/glyph%-data/.+$",
cat = "Modul data bentuk aksara Cina bersejarah",
process = function(title, _cats)
local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)")
if character then
return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.")
:format(zh_link(character))
end
end,
},
{
regex = "^Modul:zh/data/ltc%-pron/(.+)$",
cat = "Modul data sebutan bahasa Cina Pertengahan|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-BS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1",
process = "zh data",
},
{
-- capture rest of zh/data submodules
regex = "^Modul:zh/data/(.+)$",
cat = "Modul data bahasa Cina|%1",
},
{
regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$",
process = "guoxue-data",
},
{
regex = "^Modul:Unicode data/(.+)$",
cat = "Modul data Unicode|%1",
process = make_Unicode_data_documentation,
},
{
regex = "^Modul:number list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor")
if lang then
return ("This module contains data on various types of numbers in %s.\n%s")
:format(lang:makeCategoryLink(), number_list_show_table() or "")
end
end,
},
{
regex = "^Modul:chemical element list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia")
if lang then
return ("This module contains data on chemical elements in %s.\n%s")
:format(lang:makeCategoryLink(), chemical_element_list_show_table() or "")
end
end,
},
{
regex = "^Modul:accel/(.+)$",
process = function(title, cats)
local lang_code = title.subpageText
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel")
insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName()))
return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.")
:format(lang:makeCategoryLink())
end
end,
},
{
regex = "^Modul:inc%-ash/dial/data/(.+)$",
cat = "Modul Prakrit Ashoka|%1",
process = function(title, _cats)
local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$")
if word then
local lang = lang_cache["inc-ash"]
return ("This module contains data on the pronunciation of %s in dialects of %s.")
:format(full_link({ term = word, lang = lang }, "term"),
lang:makeCategoryLink())
end
end,
},
{
regex = "^.+%-translit$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "translit",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-sortkey$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "sortkey",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-stripdiacritics$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "strip diacritics",
title_without_namespace = title.text,
}
end,
},
{
regex = "^Modul:form of/lang%-data/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa")
if lang then
-- FIXME, display more info.
return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " ..
langname .. "."
end
end
},
{
regex = "^Modul:labels/data/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa")
if lang then
return {
title = "label language-specific data documentation",
args = { [1] = lang_code },
}
end
end
},
{
regex = "^Modul:category tree/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang")
if lang then
return "This module handles generating the descriptions and categorization for " ..
langname .. " category pages "
.. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are "
.. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. "
.. "This module is part of the category tree system, which is a general framework for generating the "
.. "descriptions and categorization of category pages.\n\n"
.. "For more information, see [[Modul:category tree/lang/documentation]].\n\n"
.. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the "
.. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized."
end
end
},
{
regex = "^Modul:category tree/topic/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/topic| ")
return {
title = "topic cat data submodule documentation"
}
end
},
{
regex = "^Modul:category tree/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/grammar| ")
return {
title = "category tree data submodule documentation"
}
end
},
{
regex = "^Modul:ja/data/(.+)$",
cat = "Modul data bahasa Jepun|%1",
},
{
regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$",
cat = "Modul atlas data dialek Finland|%1",
process = function(_title, _cats, shard)
return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " ..
"''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " ..
"It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " ..
"''Kotimaisten kielten keskus'' under the CC BY 4.0 license."
end
},
{
regex = "^Modul:fi%-dialects/data/feature/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:fi%-dialects/data/word/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:Swadesh/data/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "."
end
end
},
{
regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$",
process = function(_title, cats, lang_code, variety)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the "
local etym_lang = get_lang(variety, nil, "allow etym")
if etym_lang then
return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname)
end
local script = get_script(variety)
if script then
return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang))
end
return ("%s %s variety of %s."):format(prefix, variety, langname)
end
end
},
{
regex = "^Modul:typing%-aids",
process = function(title, cats)
local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$")
local sortkey
if data_suffix then
if data_suffix:find "^[%l-]+$" then
local lang = get_lang(data_suffix)
if lang then
sortkey = lang:getCanonicalName()
insert(cats, "Modul data bahasa " .. sortkey)
end
elseif data_suffix:find "^%u%l%l%l$" then
local script = get_script(data_suffix)
if script then
-- FIXME: no lang to pass here
sortkey = script:getCanonicalName()
insert(cats, script:getCategoryName())
end
end
insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix))
end
end,
},
{
regex = "^Modul:R:([%l-]+):(.+)$",
process = function(_title, cats, lang_code, refname)
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname)
insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName()))
return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}."
end
end,
},
{
regex = "^Modul:Quotations/([%l-]+)/?(.*)",
process = "Quotation",
},
{
regex = "^Modul:affix/lang%-data/([%l-]+)",
process = "affix lang-data",
},
{
regex = "^Modul:dialect synonyms/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "| ")
return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " ..
"{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" ..
"==== Struktur modul data bahasa ====\n" ..
"* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" ..
"* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" ..
"* <code>export.notes</code> — optional; table of note keys to text.\n" ..
"* <code>export.sources</code> — optional; table of source keys to text.\n" ..
"* <code>export.note_aliases</code> — optional; alias map for notes.\n" ..
"* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" ..
expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } })
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$",
process = function(_title, cats, lang_code, term)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term } }))
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$",
process = function(_title, cats, lang_code, term, id)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term, id = id } }))
end
end,
},
{
regex = "^Modul:bibliography/data/([%l-]+)$",
process = function(title, cats, lang_code)
if lang_code == "preload" then
return 'Used as a base model for other languages when the button "create new language submodule" is clicked.'
end
local page = require(title.fullText).bib_page
if not page then
page = lang_cache[lang_code]:getCanonicalName()
if page then
insert(cats, "Modul bahasa " .. page)
end
end
insert(cats, "Modul rujukan")
return "This module holds bibliographical data for " ..
page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''."
end,
},
}
function export.show(frame)
local boolean_default_false = { type = "boolean", default = false }
local args = process_params(frame.args, {
["hr"] = true,
["for"] = true,
["from"] = true,
["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage.
["notsubpage"] = boolean_default_false,
["nodoc"] = boolean_default_false,
["nolinks"] = boolean_default_false, -- suppress all "Useful links"
["nosandbox"] = boolean_default_false, -- supress sandbox
})
local output = {'\n<div class="documentation" style="display:block; clear:both">\n'}
local function ins(txt)
insert(output, txt)
end
local cats = {}
local function inscat(cat)
insert(cats, cat)
end
local nodoc = args.nodoc
if (not args.hr) or (args.hr == "above") then
ins("----\n")
end
local title = args["for"] and new_title(args["for"]) or get_current_title()
local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil
local contentModel = title.contentModel
local pagetype, is_script_or_stylesheet = get_pagetype(title)
local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc
local doc_content_source = "Modul:pendokumenan"
local auto_generated_cat_source
local cats_auto_generated = false
if not args.allowondoc and is_documentation(title) then
-- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type.
error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].")
elseif is_sandbox(title) then
local sandbox_ns = title.nsText
preload = ("Templat:pendokumenan/preload%s%sSandbox"):format(
sandbox_ns == "Modul" and sandbox_ns or "Templat",
title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or ""
)
elseif pagetype:match("%f[%w]gadget%f[%W]") then
preload = "Templat:pendokumenan/preloadGadget"
elseif pagetype:match("%f[%w]script%f[%W]") then -- .js
if title.nsText == "MediaWiki" then
preload = "Templat:pendokumenan/preloadMediaWikiJavaScript"
else
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
end
is_script_or_stylesheet = true
elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
is_script_or_stylesheet = true
elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto.
preload = "Templat:pendokumenan/preloadModule"
elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then
preload = "Templat:pendokumenan/preloadTemplate"
end
if doc_title and doc_title.isRedirect then
old_doc_title = doc_title
doc_title = doc_title.redirectTarget
end
ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">")
local function get_module_doc_and_cats(categories_only)
cats_auto_generated = true
local automatic_cats = nil
if user_name then
fallback_docs = "pendokumenan/fallback/user module"
automatic_cats = { "Modul kotak pasir pengguna" }
else
for _, data in ipairs(module_regex) do
local captures = { umatch(title.fullText, data.regex) }
if #captures > 0 then
local cat, process_function
if is_callable(data.process) then
process_function = data.process
elseif type(data.process) == "string" then
doc_content_source = "Modul:pendokumenan/functions/" .. data.process
process_function = require(doc_content_source)
end
if process_function then
doc_content = process_function(title, cats, unpack(captures))
end
if type(doc_content) == "table" then
doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source
doc_content = expand_template(doc_content)
elseif doc_content ~= nil then
doc_content = preprocess(doc_content)
end
cat = data.cat
if cat then
if type(cat) == "string" then
cat = { cat }
end
for _, c in ipairs(cat) do
insert(cats, (ugsub(title.fullText, data.regex, c)))
end
end
break
end
end
end
if title.subpageText == "templates" then
inscat("Modul antara muka templat")
end
if automatic_cats then
for _, c in ipairs(automatic_cats) do
inscat(c)
end
end
if #cats == 0 then
local auto_cats = categorize_module {
return_raw = true,
noerror = true,
}
if #auto_cats > 0 then
auto_generated_cat_source = "Modul:module categorization"
end
for _, category in ipairs(auto_cats) do
inscat(category)
end
end
-- meaning module is not in user’s sandbox or one of many datamodule boring series
needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs)
end
-- Override automatic documentation, if present.
if doc_title and doc_title.exists then
local cats_auto_generated_text = ""
if contentModel == "Scribunto" then
local doc_page_content = doc_title.content
-- Track then do nothing if there are uses of includeonly. The
-- pattern is slightly too permissive, but any false-positives are
-- obvious typos that should be corrected.
if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then
track("module-includeonly")
else
-- Check for uses of {{module cat}}. find_templates treats the
-- input as transcluded by default (i.e. it parses the wikitext
-- which will be transcluded through to the module page).
local module_cat
for template in find_templates(doc_page_content) do
if Templat:get_name() == "module cat" then
module_cat = true
break
end
end
if not module_cat then
get_module_doc_and_cats("categories only")
auto_generated_cat_source = auto_generated_cat_source or doc_content_source
cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" ..
auto_generated_cat_source .. "]]. <sup>[[" ..
new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>"
end
end
end
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" ..
doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" ..
cats_auto_generated_text .. "</i></dd>")
else
if contentModel == "Scribunto" then
get_module_doc_and_cats(false)
elseif title.nsText == "Templat" then
--inscat("Uncategorized templates")
needs_doc = not (fallback_docs or nodoc)
elseif user_name and is_script_or_stylesheet then
skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")]
if skin_name then
fallback_docs = "pendokumenan/fallback/user " .. contentModel
end
end
if doc_content then
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " ..
"dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" ..
new_title(doc_content_source):fullUrl { action = "edit" } ..
" sunting]]</sup> </i></dd>")
elseif not nodoc then
if doc_title then
ins(
"<dd><i style=\"font-size: larger;\">Laman " .. pagetype ..
" ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " ..
(fallback_docs and "Anda boleh " or "Minta tolong ") ..
"[" .. doc_title:fullUrl { action = "edit", preload = preload }
.. " ciptakan laman pendokumenan tersebut].</i></dd>\n")
else
ins(
"<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " ..
"pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n")
end
end
end
if startswith(title.fullText, "MediaWiki:Gadget-") then
local is_gadget = false
for line in gline(new_title("MediaWiki:Gadgets-definition").content) do
local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$")
if not gadget then
gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$")
end
if gadget then
items = split(items, "|")
for i, item in ipairs(items) do
if title.fullText == ("MediaWiki:Gadget-" .. item) then
is_gadget = true
ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>")
ins(gadget)
ins("</code> gajet ([")
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(" sunting takrifan])'' <dl>")
ins("<dd> ''Huraian ([")
ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" })))
ins(" sunting])'': ")
ins(preprocess(new_message('Gadget-' .. gadget):plain()))
ins(" </dd>")
table.remove(items, i)
if #items > 0 then
for j, item in ipairs(items) do
items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]'
end
ins("<dd> ''Bahagian lain'': ")
ins(list_to_text(items))
ins("</dd>")
end
ins("</dl></dd>")
break
end
end
end
end
if not is_gadget then
ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [")
ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" })))
ins(' gajet] ([')
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(' sunting takrifan]).</dd>')
-- else
-- inscat("Wiktionary gadgets")
end
end
if old_doc_title then
ins("<dd> ''Dilencong daripada'' [")
ins(old_doc_title:fullUrl { redirect = "no" })
ins(" ")
ins(old_doc_title.fullText)
ins("] ([")
ins(old_doc_title:fullUrl { action = "edit" })
ins(" sunting]).</dd>\n")
end
if not args.nolinks then
local links = {}
local function inslinks(txt)
insert(links, txt)
end
if title.isSubpage and not args.notsubpage then
inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]")
inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]")
else
inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]")
end
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) ..
" pautan]")
if contentModel ~= "Scribunto" then
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) ..
" lencongan]")
end
if is_script_or_stylesheet then
if user_name then
inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]")
end
else
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) ..
" transklusi]")
end
if contentModel == "Scribunto" then
local is_testcases = title.isSubpage and title.subpageText == "testcases"
local without_subpage = title.nsText .. ":" .. title.baseText
if is_testcases then
inslinks("[[:" .. without_subpage .. "|modul ujian]]")
else
inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]")
end
if user_name then
inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]")
inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]")
inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]")
-- If sandbox module, add a link to the module that this is a sandbox of.
-- Exclude user sandbox modules like [[User:Dine2016/sandbox]].
elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then
inscat("Modul kotak pasir")
-- Sandbox modules don’t really need documentation.
needs_doc = false
-- Don't track user sandbox modules.
local text_title = new_title(title.text)
if not (text_title and text_title.nsText == "Pengguna") then
local diff
local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]")
if sandbox_of then
track("sandbox to be moved")
else
sandbox_of = title.text:match("^sandbox%d*/(.*)$")
end
if not sandbox_of then
error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'")
:format(title.text))
end
sandbox_of = title.nsText .. ":" .. sandbox_of
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
end
-- If not a sandbox module, add link to sandbox module.
-- Sometimes there are multiple sandboxes for a single Modul:
-- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]].
else
local sandbox_title
local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$")
if not user_prefix then
user_prefix = ""
user_rest = title.text
end
sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest
local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks(sandbox_link .. (diff or ""))
end
end
if title.nsText == "Templat" then
-- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]]
-- apparently same with/without: &profile=advanced&fulltext=1
local errorq = 'searchengineselect=mediawiki&search=all: hasTemplat:\"' ..
title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:'
local eincategory =
"Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors"
inslinks(
'[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]'
.. ' (' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]'
.. '/' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]'
.. ')'
)
if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template.
-- At the moment there are no user sandbox templates with subpage
-- “/sandbox”.
inscat("Templat kotak pasir")
-- Sandbox templates don’t really need documentation.
needs_doc = false
-- Will behave badly if “/sandbox” occurs twice in title!
local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "")
local diff
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
-- This is a template that can have a sandbox.
elseif not args.nosandbox then -- unless we tell it not to
local sandbox_title = title.fullText .. "/sandbox"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or ""))
end
end
if #links > 0 then
ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>")
end
end
ins("</dl>\n")
-- Show error from [[Modul:category tree/topic cat/data]] on its submodules'
-- documentation to, for instance, warn about duplicate labels.
if startswith(title.fullText, "Modul:category tree/topic/") then
local ok, err = pcall(require, "Modul:category tree/topic/data")
if not ok then
ins('<span class="error">' .. err .. '</span>\n\n')
end
end
if doc_title and doc_title.exists then
-- Override automatic documentation, if present.
doc_content = expand_template { title = doc_title.fullText }
elseif not doc_content and fallback_docs then
doc_content = expand_template {
title = fallback_docs,
args = {
['user'] = user_name,
['page'] = title.fullText,
['skin name'] = skin_name,
},
}
end
if doc_content then
ins(doc_content)
end
ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br"))
if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then
if contentModel == "Scribunto" then
inscat("Modul belum dikategorikan")
-- elseif title.nsText == "Templat" then
-- inscat("Templat belum dikategorikan")
end
end
if needs_doc then
inscat("Templat dan modul yang memerlukan pendokumenan")
end
for _, cat in ipairs(cats) do
ins("[[Kategori:" .. cat .. "]]")
end
ins("</div>\n")
return concat(output)
end
function export.module_auto_doc_table()
local parts = {}
local function ins(text)
insert(parts, text)
end
ins('{|class="wikitable"')
ins("! Regex !! Kategori !! Modul yang dikendalikan")
for _, spec in ipairs(module_regex) do
local cat_text
local cats = spec.cat
if cats then
local cat_parts = {}
if type(cats) == "string" then
cats = { cats }
end
for _, cat in ipairs(cats) do
insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "|"))))
end
cat_text = concat(cat_parts, ", ")
else
cat_text = "''(tidak dinyatakan secara khusus)''"
end
ins("|-")
ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text,
is_callable(spec.process) and "''(dikendali secara dalaman)''" or
type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or
"''(tiada penjana pendokumenan)''"))
end
ins("|}")
return concat(parts, "\n")
end
return export
a6872eb4t58l8znugkpylbbqxanopxl
375366
375362
2026-09-22T04:42:04Z
Hakimi97
2668
Baiki ralat
375366
Scribunto
text/plain
local export = {}
local debug_track_module = "Modul:debug/track"
local frame_module = "Modul:frame"
local fun_is_callable_module = "Modul:fun/isCallable"
local languages_module = "Modul:languages"
local links_module = "Modul:links"
local load_module = "Modul:load"
local module_categorization_module = "Modul:module categorization"
local number_list_show_module = "Modul:number list/show"
local chemical_element_list_show_module = "Modul:chemical element list/show"
local pages_module = "Modul:pages"
local parameters_module = "Modul:parameters"
local scripts_module = "Modul:scripts"
local string_endswith_module = "Modul:string/endswith"
local string_gline_module = "Modul:string/gline"
local string_startswith_module = "Modul:string/startswith"
local string_utilities_module = "Modul:string utilities"
local template_parser_module = "Modul:template parser"
local title_exists_module = "Modul:title/exists"
local title_new_title_module = "Modul:title/newTitle"
local concat = table.concat
local error = error
local full_url = mw.uri.fullUrl
local get_current_title = mw.title.getCurrentTitle
local insert = table.insert
local ipairs = ipairs
local list_to_text = mw.text.listToText
local new_message = mw.message.new
local pcall = pcall
local require = require
local tonumber = tonumber
local tostring = tostring
local type = type
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local function categorize_module(...)
categorize_module = require(module_categorization_module).categorize_module
return categorize_module(...)
end
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function endswith(...)
endswith = require(string_endswith_module)
return endswith(...)
end
local function expand_template(...)
expand_template = require(frame_module).expandTemplate
return expand_template(...)
end
local function find_templates(...)
find_templates = require(template_parser_module).find_templates
return find_templates(...)
end
local function full_link(...)
full_link = require(links_module).full_link
return full_link(...)
end
local function get_lang(...)
get_lang = require(languages_module).getByCode
return get_lang(...)
end
local function get_pagetype(...)
get_pagetype = require(pages_module).get_pagetype
return get_pagetype(...)
end
local function get_script(...)
get_script = require(scripts_module).getByCode
return get_script(...)
end
local function gline(...)
gline = require(string_gline_module)
return gline(...)
end
local function is_callable(...)
is_callable = require(fun_is_callable_module)
return is_callable(...)
end
local function is_documentation(...)
is_documentation = require(pages_module).is_documentation
return is_documentation(...)
end
local function is_sandbox(...)
is_sandbox = require(pages_module).is_sandbox
return is_sandbox(...)
end
local function new_title(...)
new_title = require(title_new_title_module)
return new_title(...)
end
local function number_list_show_table(...)
number_list_show_table = require(number_list_show_module).table
return number_list_show_table(...)
end
local function chemical_element_list_show_table(...)
chemical_element_list_show_table = require(chemical_element_list_show_module).table
return chemical_element_list_show_table(...)
end
local function preprocess(...)
preprocess = require(frame_module).preprocess
return preprocess(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function safe_load_data(...)
safe_load_data = require(load_module).safe_load_data
return safe_load_data(...)
end
local function split(...)
split = require(string_utilities_module).split
return split(...)
end
local function startswith(...)
startswith = require(string_startswith_module)
return startswith(...)
end
local function title_exists(...)
title_exists = require(title_exists_module)
return title_exists(...)
end
local function ugsub(...)
ugsub = require(string_utilities_module).gsub
return ugsub(...)
end
local function umatch(...)
umatch = require(string_utilities_module).match
return umatch(...)
end
local skins = {
["common"] = "",
["vector"] = "Vector",
["monobook"] = "Monobook",
["cologneblue"] = "Cologne Blue",
["modern"] = "Modern",
}
local function track(page)
debug_track("pendokumenan/" .. page)
return true
end
local function compare_pages(page1, page2, text)
return "[" .. tostring(
full_url("Khas:ComparePages", { page1 = page1, page2 = page2 }))
.. " " .. text .. "]"
end
-- Avoid transcluding [[Modul:languages/cache]] everywhere.
local lang_cache = setmetatable({}, {
__index = function(self, k)
return require("Modul:languages/cache")[k]
end
})
local function zh_link(word)
return full_link {
lang = lang_cache.zh,
term = word
}
end
local function make_languages_data_documentation(_title, cats, division)
local doc_template, module_cat
if endswith(division, "/extra") then
division = division:sub(1, -7)
doc_template = "language extradata documentation"
module_cat = "Modul data ekstra bahasa"
else
doc_template = "language data documentation"
module_cat = "Modul data bahasa"
end
local sort_key
if division == "exceptional" then
sort_key = "x"
else
sort_key = division:gsub("/", "")
end
insert(cats, module_cat .. "|" .. sort_key)
return {
title = doc_template
}
end
local function make_Unicode_data_documentation(title, _cats)
local subpage, first_three_of_code_point
= title.fullText:match("^Modul:Unicode data/([^/]+)/(%x%x%x)$")
if subpage == "names" or subpage == "images" or subpage == "emoji images" then
local low, high =
tonumber(first_three_of_code_point .. "000", 16),
tonumber(first_three_of_code_point .. "FFF", 16)
local text, text_type
if subpage == "names" then
text_type = "titles of images"
elseif subpage == "images" then
text_type = "titles of images"
elseif subpage == "emoji images" then
text_type = "emoji-style images"
end
text = string.format(
"Modul data ini mengandungi " .. text_type .. " kepada " ..
"titik-titik kod [[Lampiran:Unicode|Unicode]] dalam julat U+%04X ke U+%04X.",
low, high)
if subpage == "images" and safe_load_data("Modul:Unicode data/emoji images/" .. first_three_of_code_point) then
text = text ..
" Senarai ini termasuk varian teks emoji. Untuk senarai varian emoji aksara tersebut, lihat [[Modul:Unicode data/emoji images/" ..
first_three_of_code_point .. "]]."
elseif subpage == "emoji images" then
text = text ..
" Untuk imej gaya teks, lihat [[Modul:Unicode data/images/" .. first_three_of_code_point .. "]]."
end
return text
end
end
local function insert_lang_data_module_cats(cats, langcode, overall_data_module_cat)
local lang = lang_cache[langcode]
if lang then
local langname
if lang._fullCode then
langname = lang_cache[lang._fullCode]:getCanonicalName()
else
langname = lang:getCanonicalName()
end
insert(cats, overall_data_module_cat .. "|" .. langname)
insert(cats, "Modul bahasa " .. langname)
insert(cats, "Modul data bahasa " .. langname)
return lang, langname
end
end
--[=[
This provides categories and documentation for various data modules, so that [[Category:Uncategorized modules]] isn't
unnecessarily cluttered. It is a list of tables, each of which have the following possible fields:
`regex` (required): A Lua pattern to match the module's title. If it matches, the data in this entry will be used.
Any captures in the pattern can by referenced in the `cat` field using %1 for the first capture, %2 for the
second, etc. (often used for creating the sortkey for the category). In addition, the captures are passed to the
`process` function as the third and subsequent parameters.
`process` (optional): This may be a function or a string. If it is a function, it is called as follows:
`process(TITLE, CATS, CAPTURE1, CAPTURE2, ...)`
where:
* TITLE is a title object describing the module's title; see
[https://www.mediawiki.org/wiki/Extension:Scribunto/Lua_reference_manual#Title_objects].
* CATS is an list of categories that the module will be added to.
* CAPTURE1, CAPTURE2, ... contain any captures in the `regex` field.
The return value of `process` should either be a string (which will be used as the module's documentation), or a
table specifying the name of a template to expand to get the documentation, along with the arguments to that
template. In the latter format, the template name (bare, without the "Templat:" prefix) should be in the `title`
field, and any arguments should be in `args; in this case, the template name will be listed above the generated
documentation as the source of the documentation, along with an edit button to edit the template's contents.
If, however, the return value of the `process` function is a string, any template invocations will be expanded
using frame:preprocess(), and [[Modul:documentation]] will be listed as the source of the documentation.
If `process` itself is a string rather than a function, it should name a submodule under
[[Modul:documentation/functions/]] which returns a function, of the same type as described above. This submodule
will be specified as the source of the documentation (unless it returns a table naming a template to expand to get
the documentation, as described above).
If `process` is omitted entirely, the module will have no documentation.
`cat` (optional): A string naming the category into which the module should be placed, or a list of such strings.
Captures specified in `regex` may be referenced in this string using %1 for the first capture, %2 for the second,
etc. It is also possible to add categories in the `process` function by inserting them into the passed-in CATS
list (the second parameter).
]=]
local module_regex = {
{
regex = "^Modul:languages/data/(3/%l/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(3/%l)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(2)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional/extra)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/data/(exceptional)$",
process = make_languages_data_documentation,
},
{
regex = "^Modul:languages/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:scripts/.+$",
cat = "Modul bahasa dan tulisan",
},
{
regex = "^Modul:data tables/data..?.?.?$",
cat = "Jadual data berpecah modul rujukan",
},
{
regex = "^Modul:zh/data/dial%-pron/.+$",
cat = "Modul data sebutan dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/dial%-syn/.+$",
cat = "Modul data sinonim dialek bahasa Cina",
process = "zh dial or syn",
},
{
regex = "^Modul:zh/data/glyph%-data/.+$",
cat = "Modul data bentuk aksara Cina bersejarah",
process = function(title, _cats)
local character = title.fullText:match("^Modul:zh/data/glyph%-data/(.+)")
if character then
return ("Modul ini mengandungi data tentang bentuk aksara Cina bersejarah %s.")
:format(zh_link(character))
end
end,
},
{
regex = "^Modul:zh/data/ltc%-pron/(.+)$",
cat = "Modul data sebutan bahasa Cina Pertengahan|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-BS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Baxter-Sagart)|%1",
process = "zh data",
},
{
regex = "^Modul:zh/data/och%-pron%-ZS/(.+)$",
cat = "Modul data sebutan bahasa Cina Kuno (Zhengzhang)|%1",
process = "zh data",
},
{
-- capture rest of zh/data submodules
regex = "^Modul:zh/data/(.+)$",
cat = "Modul data bahasa Cina|%1",
},
{
regex = "^Modul:mul/guoxue%-data/cjk%-?(.*)$",
process = "guoxue-data",
},
{
regex = "^Modul:Unicode data/(.+)$",
cat = "Modul data Unicode|%1",
process = make_Unicode_data_documentation,
},
{
regex = "^Modul:number list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data nombor")
if lang then
return ("This module contains data on various types of numbers in %s.\n%s")
:format(lang:makeCategoryLink(), number_list_show_table() or "")
end
end,
},
{
regex = "^Modul:chemical element list/data/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data unsur kimia")
if lang then
return ("This module contains data on chemical elements in %s.\n%s")
:format(lang:makeCategoryLink(), chemical_element_list_show_table() or "")
end
end,
},
{
regex = "^Modul:accel/(.+)$",
process = function(title, cats)
local lang_code = title.subpageText
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|accel")
insert(cats, ("Submodul accel|%s"):format(lang:getCanonicalName()))
return ("This module contains new entry creation rules for %s; see [[WT:ACCEL]] for an overview, and [[Modul:accel]] for information on creating new rules.")
:format(lang:makeCategoryLink())
end
end,
},
{
regex = "^Modul:inc%-ash/dial/data/(.+)$",
cat = "Modul Prakrit Ashoka|%1",
process = function(title, _cats)
local word = title.fullText:match("^Modul:inc%-ash/dial/data/(.+)$")
if word then
local lang = lang_cache["inc-ash"]
return ("This module contains data on the pronunciation of %s in dialects of %s.")
:format(full_link({ term = word, lang = lang }, "term"),
lang:makeCategoryLink())
end
end,
},
{
regex = "^.+%-translit$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "translit",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-sortkey$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "sortkey",
title_without_namespace = title.text,
}
end,
},
{
regex = "^.+%-stripdiacritics$",
process = function(title, _cats)
return require("Modul:pendokumenan/translit-like").documentation {
operation = "strip diacritics",
title_without_namespace = title.text,
}
end,
},
{
regex = "^Modul:form of/lang%-data/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul bentuk bagi khusus bahasa")
if lang then
-- FIXME, display more info.
return "This module contains language-specific form-of data (tags, shortcuts, base lemma params. etc.) for " ..
langname .. "."
end
end
},
{
regex = "^Modul:labels/data/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang = insert_lang_data_module_cats(cats, lang_code, "Modul data label khusus bahasa")
if lang then
return {
title = "label language-specific data documentation",
args = { [1] = lang_code },
}
end
end
},
{
regex = "^Modul:category tree/lang/(.+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul data category tree/lang")
if lang then
return "This module handles generating the descriptions and categorization for " ..
langname .. " category pages "
.. "of the format \"" .. langname .. " LABEL\" where LABEL can be any text. Examples are "
.. "[[:Category:Bulgarian conjugation 2.1 verbs]] and [[:Category:Russian velar-stem neuter-form nouns]]. "
.. "This module is part of the category tree system, which is a general framework for generating the "
.. "descriptions and categorization of category pages.\n\n"
.. "For more information, see [[Modul:category tree/lang/documentation]].\n\n"
.. "'''NOTE:''' If you add a new language-specific module, you must add the language code to the "
.. "list at the top of [[Modul:category tree/lang]] in order for the module to be recognized."
end
end
},
{
regex = "^Modul:category tree/topic/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/topic| ")
return {
title = "topic cat data submodule documentation"
}
end
},
{
regex = "^Modul:category tree/(.+)$",
process = function(_title, cats, _submodule)
insert(cats, "Modul data category tree/grammar| ")
return {
title = "category tree data submodule documentation"
}
end
},
{
regex = "^Modul:ja/data/(.+)$",
cat = "Modul data bahasa Jepun|%1",
},
{
regex = "^Modul:fi%-dialects/data/feature/Kettunen1940 ([0-9]+)$",
cat = "Modul atlas data dialek Finland|%1",
process = function(_title, _cats, shard)
return "This module contains shard " .. shard .. " of the online version of Lauri Kettunen's 1940 work " ..
"''Suomen murteet III A. Murrekartasto'' (\"Finnish dialects III A: Dialect atlas\"). " ..
"It was imported and converted from urn:nbn:fi:csc-kata20151130145346403821, published by the " ..
"''Kotimaisten kielten keskus'' under the CC BY 4.0 license."
end
},
{
regex = "^Modul:fi%-dialects/data/feature/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:fi%-dialects/data/word/(.+)",
cat = "Modul data dialek Finland|%1",
},
{
regex = "^Modul:Swadesh/data/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
return "This module contains the [[Swadesh list]] of basic vocabulary in " .. langname .. "."
end
end
},
{
regex = "^Modul:Swadesh/data/([%l-]+)/([^/]*)$",
process = function(_title, cats, lang_code, variety)
local lang, langname = insert_lang_data_module_cats(cats, lang_code, "Modul Swadesh")
if lang then
local prefix = "This module contains the [[Swadesh list]] of basic vocabulary in the "
local etym_lang = get_lang(variety, nil, "allow etym")
if etym_lang then
return ("%s %s variety of %s."):format(prefix, etym_lang:getCanonicalName(), langname)
end
local script = get_script(variety)
if script then
return ("%s %s %s script."):format(prefix, langname, script:getCanonicalName(lang))
end
return ("%s %s variety of %s."):format(prefix, variety, langname)
end
end
},
{
regex = "^Modul:typing%-aids",
process = function(title, cats)
local data_suffix = title.fullText:match("^Modul:typing%-aids/data/(.+)$")
local sortkey
if data_suffix then
if data_suffix:find "^[%l-]+$" then
local lang = get_lang(data_suffix)
if lang then
sortkey = lang:getCanonicalName()
insert(cats, "Modul data bahasa " .. sortkey)
end
elseif data_suffix:find "^%u%l%l%l$" then
local script = get_script(data_suffix)
if script then
-- FIXME: no lang to pass here
sortkey = script:getCanonicalName()
insert(cats, script:getCategoryName())
end
end
insert(cats, "Modul data kemasukan aksara|" .. (sortkey or data_suffix))
end
end,
},
{
regex = "^Modul:R:([%l-]+):(.+)$",
process = function(_title, cats, lang_code, refname)
local lang = lang_cache[lang_code]
if lang then
insert(cats, "Modul bahasa " .. lang:getCanonicalName() .. "|" .. refname)
insert(cats, ("Modul rujukan|%s"):format(lang:getCanonicalName()))
return "Modul ini menerapkan templat rujukan {{temp|R:" .. lang_code .. ":" .. refname .. "}}."
end
end,
},
{
regex = "^Modul:Quotations/([%l-]+)/?(.*)",
process = "Quotation",
},
{
regex = "^Modul:affix/lang%-data/([%l-]+)",
process = "affix lang-data",
},
{
regex = "^Modul:dialect synonyms/([%l-]+)$",
process = function(_title, cats, lang_code)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "| ")
return "Modul ini mengandungi data tentang kelainan bahasa " .. langname .. " tertentu, untuk kegunaan " ..
"{{tl|sinonim dialek}}. Sinonim sebenar itu sendiri terkandung dalam submodul.\n\n" ..
"==== Struktur modul data bahasa ====\n" ..
"* <code>export.title</code> — optional; table title template (e.g. \"Regional synonyms of %s\").\n" ..
"* <code>export.columns</code> — optional; list of column headers for location hierarchy (e.g. {\"Dialect group\", \"Dialect\", \"Location\"}).\n" ..
"* <code>export.notes</code> — optional; table of note keys to text.\n" ..
"* <code>export.sources</code> — optional; table of source keys to text.\n" ..
"* <code>export.note_aliases</code> — optional; alias map for notes.\n" ..
"* <code>export.varieties</code> — required; nested table of variety nodes. Each node must have <code>name</code>; list part holds children. Node keys can include <code>text_display</code>, <code>color</code>, <code>code</code>, <code>wikidata</code>, <code>lat</code>, <code>long</code>, and language-specific keys (e.g. <code>persian</code>, <code>armenian</code>, <code>chinese</code>).\n\n" ..
expand_template({ title = 'dial syn', args = { lang_code, ["demo mode"] = "y" } })
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)$",
process = function(_title, cats, lang_code, term)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term } }))
end
end,
},
{
regex = "^Modul:dialect synonyms/([%l-]+)/([^/]+)/([^/]+)$",
process = function(_title, cats, lang_code, term, id)
local lang = lang_cache[lang_code]
if lang then
local langname = "bahasa " .. lang:getCanonicalName()
insert(cats, "Modul data sinonim dialek|" .. langname)
insert(cats, "Modul data sinonim dialek " .. langname .. "|" .. term)
return ("%s\n\n%s"):format(
"==== Term/sense module structure ====\n" ..
"* <code>export.title</code> — optional; custom table title (e.g. \"Realization of 'strong R' between vowels\"). Overrides the language default.\n" ..
"* <code>export.meaning</code> — optional; meaning/gloss (alternative to <code>gloss</code>).\n" ..
"* <code>export.gloss</code> — optional; short meaning for the table.\n" ..
"* <code>export.note</code> — optional; single note key or string, or list of note keys.\n" ..
"* <code>export.notes</code> — optional; list of note keys.\n" ..
"* <code>export.source</code> / <code>export.sources</code> — optional; source keys.\n" ..
"* <code>export.last_column</code> — optional; label for the data column (default \"Words\"; e.g. \"Realization\").\n" ..
"* <code>export.syns</code> — required; table mapping variety/location names (keys from the language data module) to a list of term entries. Each entry can be a string or a table (e.g. <code>{ ipa = \"[ɽ]\" }</code> or <code>{ term = \"word\" }</code>).\n\n" ..
"Example (custom title and data column, IPA realizations):\n" ..
"<pre>\nlocal export = {}\n\nexport.title = \"Realization of 'strong R' between vowels\"\n" ..
"export.meaning = \"\"\nexport.note = \"realization of 'strong R' between vowels\"\n" ..
"export.last_column = \"Realization\"\n\nexport.syns = {\n\t[\"ALERS-158\"] = { { ipa = \"[ɽ]\" } },\n\t[\"ALERS-175\"] = { { ipa = \"[x]\" } },\n}\n\nreturn export\n</pre>\n\n",
expand_template({ title = 'dial syn', args = { lang_code, term, id = id } }))
end
end,
},
{
regex = "^Modul:bibliography/data/([%l-]+)$",
process = function(title, cats, lang_code)
if lang_code == "preload" then
return 'Used as a base model for other languages when the button "create new language submodule" is clicked.'
end
local page = require(title.fullText).bib_page
if not page then
page = lang_cache[lang_code]:getCanonicalName()
if page then
insert(cats, "Modul bahasa " .. page)
end
end
insert(cats, "Modul rujukan")
return "This module holds bibliographical data for " ..
page .. ". For the formatted bibliography see '''[[Appendix:Bibliography/" .. page .. "]]'''."
end,
},
}
function export.show(frame)
local boolean_default_false = { type = "boolean", default = false }
local args = process_params(frame.args, {
["hr"] = true,
["for"] = true,
["from"] = true,
["allowondoc"] = boolean_default_false, -- Don't throw an error if used on a documentation subpage.
["notsubpage"] = boolean_default_false,
["nodoc"] = boolean_default_false,
["nolinks"] = boolean_default_false, -- suppress all "Useful links"
["nosandbox"] = boolean_default_false, -- supress sandbox
})
local output = {'\n<div class="documentation" style="display:block; clear:both">\n'}
local function ins(txt)
insert(output, txt)
end
local cats = {}
local function inscat(cat)
insert(cats, cat)
end
local nodoc = args.nodoc
if (not args.hr) or (args.hr == "above") then
ins("----\n")
end
local title = args["for"] and new_title(args["for"]) or get_current_title()
local doc_title = args.from ~= "-" and new_title(args.from or title.fullText .. '/doc') or nil
local contentModel = title.contentModel
local pagetype, is_script_or_stylesheet = get_pagetype(title)
local preload, fallback_docs, doc_content, old_doc_title, user_name, skin_name, needs_doc
local doc_content_source = "Modul:pendokumenan"
local auto_generated_cat_source
local cats_auto_generated = false
if not args.allowondoc and is_documentation(title) then
-- TODO: merge with {{documentation subpage}}, and choose behaviour based on the page type.
error("This template should not be used on a documentation page. Please use [[Templat:documentation subpage]].")
elseif is_sandbox(title) then
local sandbox_ns = title.nsText
preload = ("Templat:pendokumenan/preload%s%sSandbox"):format(
sandbox_ns == "Modul" and sandbox_ns or "Templat",
title.rootText:match("^[Pp]engguna:(.+)") and "Pengguna" or ""
)
elseif pagetype:match("%f[%w]gadget%f[%W]") then
preload = "Templat:pendokumenan/preloadGadget"
elseif pagetype:match("%f[%w]script%f[%W]") then -- .js
if title.nsText == "MediaWiki" then
preload = "Templat:pendokumenan/preloadMediaWikiJavaScript"
else
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
end
is_script_or_stylesheet = true
elseif pagetype:match("%f[%w]stylesheet%f[%W]") then -- .css
preload = "Templat:pendokumenan/preloadTemplate" -- XXX
if title.nsText == "Pengguna" then
user_name = title.rootText
end
is_script_or_stylesheet = true
elseif contentModel == "Scribunto" then -- Exclude pages in Modul: which aren't Scribunto.
preload = "Templat:pendokumenan/preloadModule"
elseif pagetype:match("%f[%w]template%f[%W]") or pagetype:match("%f[%w]project%f[%W]") then
preload = "Templat:pendokumenan/preloadTemplate"
end
if doc_title and doc_title.isRedirect then
old_doc_title = doc_title
doc_title = doc_title.redirectTarget
end
ins("<dl class=\"plainlinks\" style=\"font-size: smaller;\">")
local function get_module_doc_and_cats(categories_only)
cats_auto_generated = true
local automatic_cats = nil
if user_name then
fallback_docs = "pendokumenan/fallback/user module"
automatic_cats = { "Modul kotak pasir pengguna" }
else
for _, data in ipairs(module_regex) do
local captures = { umatch(title.fullText, data.regex) }
if #captures > 0 then
local cat, process_function
if is_callable(data.process) then
process_function = data.process
elseif type(data.process) == "string" then
doc_content_source = "Modul:pendokumenan/functions/" .. data.process
process_function = require(doc_content_source)
end
if process_function then
doc_content = process_function(title, cats, unpack(captures))
end
if type(doc_content) == "table" then
doc_content_source = doc_content.title and "Templat:" .. doc_content.title or doc_content_source
doc_content = expand_template(doc_content)
elseif doc_content ~= nil then
doc_content = preprocess(doc_content)
end
cat = data.cat
if cat then
if type(cat) == "string" then
cat = { cat }
end
for _, c in ipairs(cat) do
insert(cats, (ugsub(title.fullText, data.regex, c)))
end
end
break
end
end
end
if title.subpageText == "templates" then
inscat("Modul antara muka templat")
end
if automatic_cats then
for _, c in ipairs(automatic_cats) do
inscat(c)
end
end
if #cats == 0 then
local auto_cats = categorize_module {
return_raw = true,
noerror = true,
}
if #auto_cats > 0 then
auto_generated_cat_source = "Modul:module categorization"
end
for _, category in ipairs(auto_cats) do
inscat(category)
end
end
-- meaning module is not in user’s sandbox or one of many datamodule boring series
needs_doc = not categories_only and not (automatic_cats or doc_content or fallback_docs)
end
-- Override automatic documentation, if present.
if doc_title and doc_title.exists then
local cats_auto_generated_text = ""
if contentModel == "Scribunto" then
local doc_page_content = doc_title.content
-- Track then do nothing if there are uses of includeonly. The
-- pattern is slightly too permissive, but any false-positives are
-- obvious typos that should be corrected.
if doc_page_content:lower():match("</?includeonly%f[%s/>][^>]*>") then
track("module-includeonly")
else
-- Check for uses of {{module cat}}. find_templates treats the
-- input as transcluded by default (i.e. it parses the wikitext
-- which will be transcluded through to the module page).
local module_cat
for template in find_templates(doc_page_content) do
if template:get_name() == "module cat" then
module_cat = true
break
end
end
if not module_cat then
get_module_doc_and_cats("categories only")
auto_generated_cat_source = auto_generated_cat_source or doc_content_source
cats_auto_generated_text = " Kategori dijana secara automatik oleh [[" ..
auto_generated_cat_source .. "]]. <sup>[[" ..
new_title(auto_generated_cat_source):fullUrl { action = "edit" } .. " sunting]]</sup>"
end
end
end
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang terletak di [[" ..
doc_title.fullText .. "]]. " .. "<sup>[[" .. doc_title:fullUrl { action = "edit" } .. " sunting]]</sup>" ..
cats_auto_generated_text .. "</i></dd>")
else
if contentModel == "Scribunto" then
get_module_doc_and_cats(false)
elseif title.nsText == "Templat" then
--inscat("Uncategorized templates")
needs_doc = not (fallback_docs or nodoc)
elseif user_name and is_script_or_stylesheet then
skin_name = skins[title.text:sub(#title.rootText + 1):match("^/(%l+)%.[jc]ss?$")]
if skin_name then
fallback_docs = "pendokumenan/fallback/user " .. contentModel
end
end
if doc_content then
ins(
"<dd><i style=\"font-size: larger;\">Berikut merupakan " ..
"[[Bantuan:Mendokumenkan templat dan modul|pendokumenan]] yang " ..
"dijana oleh [[" .. doc_content_source .. "]]. <sup>[[" ..
new_title(doc_content_source):fullUrl { action = "edit" } ..
" sunting]]</sup> </i></dd>")
elseif not nodoc then
if doc_title then
ins(
"<dd><i style=\"font-size: larger;\">Laman " .. pagetype ..
" ini kekurangan [[Bantuan:Mendokumenkan templat dan modul|sublaman pendokumenan]]. " ..
(fallback_docs and "Anda boleh " or "Minta tolong ") ..
"[" .. doc_title:fullUrl { action = "edit", preload = preload }
.. " ciptakan laman pendokumenan tersebut].</i></dd>\n")
else
ins(
"<dd><i style=\"font-size: larger; color: var(--wikt-palette-red-9,#FF0000);\">Tidak dapat menjana secara automatik " ..
"pendokumenan untuk " .. pagetype .. " ini.</i></dd>\n")
end
end
end
if startswith(title.fullText, "MediaWiki:Gadget-") then
local is_gadget = false
for line in gline(new_title("MediaWiki:Gadgets-definition").content) do
local gadget, items = line:match("^%*%s*(%a[%w_-]*)%[.-%]|(.+)$")
if not gadget then
gadget, items = line:match("^%*%s*(%a[%w_-]*)|(.+)$")
end
if gadget then
items = split(items, "|")
for i, item in ipairs(items) do
if title.fullText == ("MediaWiki:Gadget-" .. item) then
is_gadget = true
ins("<dd> ''Skrip ini merupakan sebahagian daripada <code>")
ins(gadget)
ins("</code> gajet ([")
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(" sunting takrifan])'' <dl>")
ins("<dd> ''Huraian ([")
ins(tostring(full_url("MediaWiki:Gadget-" .. gadget, { action = "edit" })))
ins(" sunting])'': ")
ins(preprocess(new_message('Gadget-' .. gadget):plain()))
ins(" </dd>")
table.remove(items, i)
if #items > 0 then
for j, item in ipairs(items) do
items[j] = '[[MediaWiki:Gadget-' .. item .. '|' .. item .. ']]'
end
ins("<dd> ''Bahagian lain'': ")
ins(list_to_text(items))
ins("</dd>")
end
ins("</dl></dd>")
break
end
end
end
end
if not is_gadget then
ins("<dd> ''Skrip ini bukanlah sebahagian daripada mana-mana [")
ins(tostring(full_url("Khas:Gadgets", { uselang = "ms" })))
ins(' gajet] ([')
ins(tostring(full_url("MediaWiki:Gadgets-definition", { action = "edit" })))
ins(' sunting takrifan]).</dd>')
-- else
-- inscat("Wiktionary gadgets")
end
end
if old_doc_title then
ins("<dd> ''Dilencong daripada'' [")
ins(old_doc_title:fullUrl { redirect = "no" })
ins(" ")
ins(old_doc_title.fullText)
ins("] ([")
ins(old_doc_title:fullUrl { action = "edit" })
ins(" sunting]).</dd>\n")
end
if not args.nolinks then
local links = {}
local function inslinks(txt)
insert(links, txt)
end
if title.isSubpage and not args.notsubpage then
inslinks("[[:" .. title.nsText .. ":" .. title.rootText .. "|laman akar]]")
inslinks("[[Khas:PrefixIndex/" .. title.nsText .. ":" .. title.rootText .. "/|sublaman laman akar]]")
else
inslinks("[[Khas:PrefixIndex/" .. title.fullText .. "/|senarai sublaman]]")
end
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidetrans = true, hideredirs = true })) ..
" pautan]")
if contentModel ~= "Scribunto" then
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hidetrans = true })) ..
" lencongan]")
end
if is_script_or_stylesheet then
if user_name then
inslinks("[[Khas:MyPage" .. title.text:sub(#title.rootText + 1) .. "|milik anda]]")
end
else
inslinks(
"[" ..
tostring(full_url("Khas:WhatLinksHere/" .. title.fullText, { hidelinks = true, hideredirs = true })) ..
" transklusi]")
end
if contentModel == "Scribunto" then
local is_testcases = title.isSubpage and title.subpageText == "testcases"
local without_subpage = title.nsText .. ":" .. title.baseText
if is_testcases then
inslinks("[[:" .. without_subpage .. "|modul ujian]]")
else
inslinks("[[" .. title.fullText .. "/testcases|kes ujian]]")
end
if user_name then
inslinks("[[Pengguna:" .. user_name .. "|laman pengguna]]")
inslinks("[[Perbincangan pengguna:" .. user_name .. "|laman perbincangan pengguna]]")
inslinks("[[Khas:PrefixIndex/Pengguna:" .. user_name .. "/|ruang pengguna]]")
-- If sandbox module, add a link to the module that this is a sandbox of.
-- Exclude user sandbox modules like [[User:Dine2016/sandbox]].
elseif title.text:find("^sandbox%d*/") or title.text:find("/sandbox%d*%f[/%z]") then
inscat("Modul kotak pasir")
-- Sandbox modules don’t really need documentation.
needs_doc = false
-- Don't track user sandbox modules.
local text_title = new_title(title.text)
if not (text_title and text_title.nsText == "Pengguna") then
local diff
local sandbox_of = title.text:match("^(.*)/sandbox%d*%f[/%z]")
if sandbox_of then
track("sandbox to be moved")
else
sandbox_of = title.text:match("^sandbox%d*/(.*)$")
end
if not sandbox_of then
error(("Internal error: Something wrong, couldn't extract sandbox-of module from title '%s'")
:format(title.text))
end
sandbox_of = title.nsText .. ":" .. sandbox_of
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
end
-- If not a sandbox module, add link to sandbox module.
-- Sometimes there are multiple sandboxes for a single Modul:
-- [[Modul:sandbox/sa-pronunc]], [[Modul:sandbox2/sa-pronunc]].
else
local sandbox_title
local user_prefix, user_rest = title.text:match("^(Pengguna:.-/)(.*)$")
if not user_prefix then
user_prefix = ""
user_rest = title.text
end
sandbox_title = title.nsText .. ":" .. user_prefix .. "sandbox/" .. user_rest
local sandbox_link = "[[:" .. sandbox_title .. "|kotak pasir]]"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks(sandbox_link .. (diff or ""))
end
end
if title.nsText == "Templat" then
-- Error search: all(any namespace), hastemplate (show pages using the template), insource (show source code), incategory (any/specific error) -- [[mw:Help:CirrusSearch]], [[w:Help:Searching/Regex]]
-- apparently same with/without: &profile=advanced&fulltext=1
local errorq = 'searchengineselect=mediawiki&search=all: hastemplate:\"' ..
title.rootText .. '\" insource:\"' .. title.rootText .. '\" incategory:'
local eincategory =
"Laman_yang_ada_ralat_skrip|ralat_ParserFunction|DisplayTitle_errors|Pages_with_ISBN_errors|Pages_with_ISSN_errors|Pages_with_reference_errors|Pages_with_syntax_highlighting_errors|Pages_with_TemplateStyles_errors"
inslinks(
'[' .. tostring(full_url('Khas:Search', errorq .. eincategory)) .. ' ralat]'
.. ' (' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'ralat_ParserFunction')) .. ' penghurai]'
.. '/' ..
'[' .. tostring(full_url('Khas:Search', errorq .. 'Laman_yang_ada_ralat_skrip')) .. ' modul]'
.. ')'
)
if title.isSubpage and title.text:find("/sandbox%d*%f[/%z]") then -- This is a sandbox template.
-- At the moment there are no user sandbox templates with subpage
-- “/sandbox”.
inscat("Templat kotak pasir")
-- Sandbox templates don’t really need documentation.
needs_doc = false
-- Will behave badly if “/sandbox” occurs twice in title!
local sandbox_of = title.fullText:gsub("/sandbox%d*%f[/%z]", "")
local diff
if title_exists(sandbox_of) then
diff = " (" .. compare_pages(title.fullText, sandbox_of, "diff") .. ")"
else
track("no sandbox of")
end
inslinks("[[:" .. sandbox_of .. "|kotak pasir bagi]]" .. (diff or ""))
-- This is a template that can have a sandbox.
elseif not args.nosandbox then -- unless we tell it not to
local sandbox_title = title.fullText .. "/sandbox"
local diff
if title_exists(sandbox_title) then
diff = " (" .. compare_pages(title.fullText, sandbox_title, "diff") .. ")"
end
inslinks("[[:" .. sandbox_title .. "|kotak pasir]]" .. (diff or ""))
end
end
if #links > 0 then
ins("<dd> ''Pautan berguna'': " .. concat(links, " • ") .. "</dd>")
end
end
ins("</dl>\n")
-- Show error from [[Modul:category tree/topic cat/data]] on its submodules'
-- documentation to, for instance, warn about duplicate labels.
if startswith(title.fullText, "Modul:category tree/topic/") then
local ok, err = pcall(require, "Modul:category tree/topic/data")
if not ok then
ins('<span class="error">' .. err .. '</span>\n\n')
end
end
if doc_title and doc_title.exists then
-- Override automatic documentation, if present.
doc_content = expand_template { title = doc_title.fullText }
elseif not doc_content and fallback_docs then
doc_content = expand_template {
title = fallback_docs,
args = {
['user'] = user_name,
['page'] = title.fullText,
['skin name'] = skin_name,
},
}
end
if doc_content then
ins(doc_content)
end
ins(('\n<%s style="clear: both;" />'):format(args.hr == "below" and "hr" or "br"))
if cats_auto_generated and not cats[1] and (not doc_content or not doc_content:find("%[%[Kategori:")) then
if contentModel == "Scribunto" then
inscat("Modul belum dikategorikan")
-- elseif title.nsText == "Templat" then
-- inscat("Templat belum dikategorikan")
end
end
if needs_doc then
inscat("Templat dan modul yang memerlukan pendokumenan")
end
for _, cat in ipairs(cats) do
ins("[[Kategori:" .. cat .. "]]")
end
ins("</div>\n")
return concat(output)
end
function export.module_auto_doc_table()
local parts = {}
local function ins(text)
insert(parts, text)
end
ins('{|class="wikitable"')
ins("! Regex !! Kategori !! Modul yang dikendalikan")
for _, spec in ipairs(module_regex) do
local cat_text
local cats = spec.cat
if cats then
local cat_parts = {}
if type(cats) == "string" then
cats = { cats }
end
for _, cat in ipairs(cats) do
insert(cat_parts, ("<code>%s</code>"):format((cat:gsub("|", "|"))))
end
cat_text = concat(cat_parts, ", ")
else
cat_text = "''(tidak dinyatakan secara khusus)''"
end
ins("|-")
ins(("| <code>%s</code> || %s || %s"):format(spec.regex, cat_text,
is_callable(spec.process) and "''(dikendali secara dalaman)''" or
type(spec.process) == "string" and ("[[Modul:pendokumenan/functions/%s]]"):format(spec.process) or
"''(tiada penjana pendokumenan)''"))
end
ins("|}")
return concat(parts, "\n")
end
return export
iqaaks3ujjv11d34e9oefx3iq8b1c5i
Modul:names
828
27469
375348
244788
2026-09-22T03:12:17Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708781|92708781]])
375348
Scribunto
text/plain
local export = {}
local m_languages = require("Module:languages")
local m_links = require("Module:links")
local m_utilities = require("Module:utilities")
local m_str_utils = require("Module:string utilities")
local m_table = require("Module:table")
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local parameter_utilities_module = "Module:parameter utilities"
local parse_interface_module = "Module:parse interface"
local parse_utilities_module = "Module:parse utilities"
local enlang = m_languages.getByCode("ms")
local rsubn = m_str_utils.gsub
local rsplit = m_str_utils.split
local u = m_str_utils.char
local function rsub(str, from, to)
return (rsubn(str, from, to))
end
local TEMP_LESS_THAN = u(0xFFF2)
local force_cat = false -- for testing
--[=[
FIXME:
1. from=the Bible (DONE)
2. origin=18th century [DONE]
3. popular= (DONE)
4. varoftype= (DONE)
5. eqtype= [DONE]
6. dimoftype= [DONE]
7. from=de:Elisabeth (same language) (DONE)
8. blendof=, blendof2= [DONE]
9. varform, dimform [DONE]
10. from=English < Latin [DONE]
11. usage=rare -> categorize as rare?
12. dimeq= (also vareq=?) [DONE]
13. fromtype= [DONE]
14. <tr:...> and similar params [DONE]
]=]
-- Used in category code; name types which are full-word end-matching substrings of longer name types (e.g. "surnames"
-- of "male surnames", but not "male surnames" of "female surnames" because "male" only matches a part of the word
-- "female") should follow the longer name.
export.personal_name_types = {
"nama keluarga lelaki", "nama keluarga perempuan", "nama keluarga bebas jantina", "nama keluarga",
"patronimik", "matronimik",
}
export.personal_name_type_set = m_table.listToSet(export.personal_name_types)
export.given_name_genders = {
lelaki = {type = "human"},
perempuan = {type = "human"},
uniseks = {type = "human", cat = {"nama diri lelaki", "nama diri perempuan", "nama diri uniseks"}, article = ""},
["jantina tidak diketahui"] = {type = "human", cat = {}, track = true},
haiwan = {type = "animal", track = true},
kucing = {type = "animal"},
lembu = {type = "animal"},
anjing = {type = "animal"},
kuda = {type = "animal"},
khinzir = {type = "animal"},
}
local function get_given_name_cats(gender, props)
local cats = props.cat
if not cats then
if props.type == "animal" then
cats = {"nama " .. gender}
else
cats = {"nama diri " .. gender}
end
end
return cats
end
do
local function do_cat(cat)
if not export.personal_name_type_set[cat] then
export.personal_name_type_set[cat] = true
table.insert(export.personal_name_types, cat)
end
end
for gender, props in pairs(export.given_name_genders) do
local cats = get_given_name_cats(gender, props)
for _, cat in ipairs(cats) do
do_cat("bentuk singkat " .. cat)
do_cat("bentuk agam " .. cat)
do_cat(cat)
end
end
do_cat("nama diri")
end
local translit_name_type_list = {
"nama keluarga", "nama diri lelaki", "nama diri perempuan", "nama diri uniseks",
"patronimik"
}
local function track(page)
require("Module:debug").track("names/" .. page)
end
-- Get raw text, for use in computing the indefinite article. Use get_plaintext() in [[Module:utilities]] and also
-- remove parens that may surround qualifier or label text preceding a term.
local function get_rawtext(text)
text = m_utilities.get_plaintext(text)
text = text:gsub("[()%[%]]", "")
return text
end
--[=[
Parse a term and associated properties. This works with parameters of the form 'Karlheinz' or
'Kunigunde<q:medieval, now rare>' or 'non:Óláfr' or 'ru:Фру́нзе<tr:Frúnzɛ><q:rare>' where the modifying properties
are contained in <...> specifications after the term. `term` is the full parameter value including any angle brackets
and colons; `paramname` is the name of the parameter that this value comes from, for error purposes; `deflang` is a
language object used in the return value when the language isn't specified (e.g. in the examples 'Karlheinz' and
'Kunigunde<q:medieval, now rare>' above); `allow_explicit_lang` indicates whether the language can be explicitly given
(e.g. in the examples 'non:Óláfr' or 'ru:Фру́нзе<tr:Frúnzɛ><q:rare>' above).
Normally the return value is a terminfo object that can be passed to full_link() in [[Module:links]]), additionally
with optional fields `.q`, `.qq`, `.l`, `.ll`, `.refs` and `.eq` (a list of objects of the same form as the returned
terminfo object. However, if `allow_multiple_terms` is given, multiple comma-separated names can be given in `term`,
and the return value is a list of objects of the form described just above.
]=]
local function parse_term_with_annotations(term, paramname, deflang, allow_explicit_lang, allow_multiple_terms)
local param_mods = require(parameter_utilities_module).construct_param_mods {
{group = {"link", "l", "q", "ref"}},
{param = "eq", convert = function(eqval, parse_err)
return parse_term_with_annotations(eqval, paramname .. ".eq", enlang, false, "allow multiple terms")
end},
}
local function generate_obj(term, parse_err)
local termlang
if allow_explicit_lang then
local actual_term
actual_term, termlang = require(parse_interface_module).parse_term_with_lang {
term = term,
parse_err = parse_err,
paramname = paramname,
}
term = actual_term or term
end
return {
term = term,
lang = termlang or deflang,
}
end
return require(parse_interface_module).parse_inline_modifiers(term, {
param_mods = param_mods,
paramname = paramname,
generate_obj = generate_obj,
splitchar = allow_multiple_terms and "," or nil,
})
end
--[=[
Link a single term. If `do_language_link` is given and a given term's language is English, the link will be constructed
using language_link() in [[Module:links]]; otherwise, with full_link(). `termobj` is an object as returned by
parse_term_with_annotations(), i.e. it is suitable for passing to [[Module:links]] and additionally contains optional
fields `.q`, `.qq`, `.l`, `.ll`, `.refs` and `.eq` (a list of objects of the same form as `termobj`).
]=]
local function link_one_term(termobj, do_language_link)
local link
if do_language_link and termobj.lang:getCode() == "ms" then
link = m_links.language_link(termobj)
else
link = m_links.full_link(termobj)
end
if termobj.q and termobj.q[1] or termobj.qq and termobj.qq[1] or
termobj.l and termobj.l[1] or termobj.ll and termobj.ll[1] or termobj.refs and termobj.refs[1] then
link = require(decorations_module).format_decorations {
lang = termobj.lang,
text = link,
q = termobj.q,
qq = termobj.qq,
l = termobj.l,
ll = termobj.ll,
refs = termobj.refs,
}
end
if termobj.eq then
local eqtext = {}
for _, eqobj in ipairs(termobj.eq) do
table.insert(eqtext, link_one_term(eqobj, true))
end
link = link .. " [=" .. m_table.serialCommaJoin(eqtext, {conj = "atau"}) .. "]"
end
return link
end
--[=[
Link the terms in `terms`, and join them using the conjunction in `conj` (defaulting to "or"). Joining is done using
serialCommaJoin() in [[Module:table]], so that e.g. two terms are joined as "TERM or TERM" while three terms are joined
as "TERM, TERM or TERM" with special CSS spans before the final "or" to allow an "Oxford comma" to appear if configured
appropriately. (However, if `conj` is the special value ", ", joining is done directly using that value.)
If `include_langname` is given, the language of the first term will be prepended to the joined terms. If
`do_language_link` is given and a given term's language is English, the link will be constructed using language_link()
in [[Module:links]]; otherwise, with full_link(). Each term in `terms` is an object as returned by
parse_term_with_annotations().
]=]
local function join_terms(terms, include_langname, do_language_link, conj)
local links = {}
local langnametext
for _, termobj in ipairs(terms) do
if include_langname and not langnametext then
langnametext = termobj.lang:getCanonicalName() .. " "
end
table.insert(links, link_one_term(termobj, do_language_link))
end
local joined_terms
if conj == ", " then
joined_terms = table.concat(links, conj)
else
joined_terms = m_table.serialCommaJoin(links, {conj = conj or "atau"})
end
return (langnametext or "") .. joined_terms
end
--[=[
Gather the parameters for multiple names and link each name using full_link() (for foreign names) or language_link()
(for English names), joining the names using serialCommaJoin() in [[Module:table]] with the conjunction `conj`
(defaulting to "or"). (However, if `conj` is the special value ", ", joining is done directly using that value.)
This can be used, for example, to fetch and join all the masculine equivalent names for a feminine given name. Each
name is specified using parameters beginning with `pname` in `args`, e.g. "m", "m2", "m3", etc. `lang` is a language
object specifying the language of the names (defaulting to English), for use in linking them. If `allow_explicit_lang`
is given, the language of the terms can be specified explicitly by prefixing a term with a language code, e.g.
'sv:Björn' or 'la:[[Nicolaus|Nīcolāī]]'. This function assumes that the parameters have already been parsed by
[[Module:parameters]] and gathered into lists, so that e.g. all "mN" parameters are in a list in args["m"].
]=]
local function join_names(lang, args, pname, conj, allow_explicit_lang)
local termobjs = {}
local do_language_link = false
if not lang then
lang = enlang
do_language_link = true
end
local function process_one_term(term, i)
for _, termobj in ipairs(parse_term_with_annotations(term, pname .. (i == 1 and "" or i), lang,
allow_explicit_lang, "allow multiple terms")) do
table.insert(termobjs, termobj)
end
end
if not args[pname] then
return "", 0
elseif type(args[pname]) == "table" then
for i, term in ipairs(args[pname]) do
process_one_term(term, i)
end
else
process_one_term(args[pname], 1)
end
return join_terms(termobjs, nil, do_language_link, conj or "dan"), #termobjs
end
local function get_eqtext(args)
local eqsegs = {}
local lastlang = nil
local last_eqseg = {}
local function process_one_term(term, i)
for _, termobj in ipairs(parse_term_with_annotations(term, "eq" .. (i == 1 and "" or i), enlang,
"allow explicit lang", "allow multiple terms")) do
local termlang = termobj.lang:getCode()
if lastlang and lastlang ~= termlang then
if #last_eqseg > 0 then
table.insert(eqsegs, last_eqseg)
end
last_eqseg = {}
end
lastlang = termlang
table.insert(last_eqseg, termobj)
end
end
if type(args.eq) == "table" then
for i, term in ipairs(args.eq) do
process_one_term(term, i)
end
elseif type(args.eq) == "string" then
process_one_term(args.eq, 1)
end
if #last_eqseg > 0 then
table.insert(eqsegs, last_eqseg)
end
local eqtextsegs = {}
for _, eqseg in ipairs(eqsegs) do
table.insert(eqtextsegs, join_terms(eqseg, "include langname"))
end
return m_table.serialCommaJoin(eqtextsegs, {conj = "atau"})
end
local function get_fromtext(lang, args)
local catparts = {}
local fromsegs = {}
local i = 1
local function parse_from(from)
local unrecognized = false
local prefix, suffix
if from == "nama keluarga" or from == "nama diri" or from == "nama panggilan" or from == "nama tempat" or from == "kata nama am" or from == "nama bulan" then
prefix = "dipindahkan daripada "
suffix = from
table.insert(catparts, from)
elseif from == "patronimik" or from == "matronimik" or from == "ciptaan baharu" then
prefix = "berasal "
suffix = "sebagai " .. from
table.insert(catparts, from)
elseif from == "pekerjaan" or from == "etnonim" then
prefix = "berasal "
suffix = "sebagai " .. from
table.insert(catparts, from)
elseif from == "Alkitab" then
prefix = "berasal "
suffix = "daripada Alkitab"
table.insert(catparts, from)
else
prefix = "daripada "
if from:find(":") then
local termobj = parse_term_with_annotations(from, "from" .. (i == 1 and "" or i), lang,
"allow explicit lang")
local fromlangname = ""
if termobj.lang:getCode() ~= lang:getCode() then
-- If name is derived from another name in the same language, don't include lang name after text
-- "from " or create a category like "German male given names derived from German".
local canonical_name = termobj.lang:getCanonicalName()
fromlangname = "bahasa " .. canonical_name .. " "
table.insert(catparts, canonical_name)
end
suffix = fromlangname .. link_one_term(termobj)
else
local family = from:match("^[Bb]ahasa%-bahasa (.+)$")
if family then
if require("Module:families").getByCanonicalName(family) then
table.insert(catparts, from)
else
unrecognized = true
end
suffix = from
else
if m_languages.getByCanonicalName(from, nil, "allow etym") then
table.insert(catparts, from)
else
unrecognized = true
end
suffix = from
end
end
end
if unrecognized then
track("unrecognized from")
track("unrecognized from/" .. from)
end
return prefix, suffix
end
local last_fromseg = nil
local put = require(parse_utilities_module)
local from_args = args.from or {}
if type(from_args) == "string" then
from_args = {from_args}
end
while from_args[i] do
-- We may have multiple comma-separated items, each of which may have multiple items separated by a
-- space-delimited < sign, each of which may have inline modifiers with embedded commas in them. To handle
-- this correctly, first replace space-delimited < signs with a special character, then split on balanced
-- <...> and [...] signs, then split on comma, then rejoin the stuff between commas. We will then split on
-- TEMP_LESS_THAN (the replacement for space-delimited < signs) and reparse.
local rawfroms = rsub(from_args[i], "%s+<%s+", TEMP_LESS_THAN)
local segments = put.parse_multi_delimiter_balanced_segment_run(rawfroms, {{"<", ">"}, {"[", "]"}})
local comma_separated_groups = put.split_alternating_runs_on_comma(segments)
for j, comma_separated_group in ipairs(comma_separated_groups) do
comma_separated_groups[j] = table.concat(comma_separated_group)
end
for _, rawfrom in ipairs(comma_separated_groups) do
local froms = rsplit(rawfrom, TEMP_LESS_THAN)
if #froms == 1 then
local prefix, suffix = parse_from(froms[1])
if last_fromseg and (last_fromseg.has_multiple_froms or last_fromseg.prefix ~= prefix) then
table.insert(fromsegs, last_fromseg)
last_fromseg = nil
end
if not last_fromseg then
last_fromseg = {prefix = prefix, suffixes = {}}
end
table.insert(last_fromseg.suffixes, suffix)
else
if last_fromseg then
table.insert(fromsegs, last_fromseg)
last_fromseg = nil
end
local first_suffixpart = ""
local rest_suffixparts = {}
for j, from in ipairs(froms) do
local prefix, suffix = parse_from(from)
if j == 1 then
first_suffixpart = prefix .. suffix
else
table.insert(rest_suffixparts, prefix .. suffix)
end
end
local full_suffix = first_suffixpart .. " [yang seterusnya " .. table.concat(rest_suffixparts, ", yang seterusnya ") .. "]"
last_fromseg = {prefix = "", has_multiple_froms = true, suffixes = {full_suffix}}
end
end
i = i + 1
end
table.insert(fromsegs, last_fromseg)
local fromtextsegs = {}
for _, fromseg in ipairs(fromsegs) do
table.insert(fromtextsegs, fromseg.prefix .. m_table.serialCommaJoin(fromseg.suffixes, {conj = "atau"}))
end
return m_table.serialCommaJoin(fromtextsegs, {conj = "atau"}), catparts
end
local function parse_given_name_genders(genderspec)
if export.given_name_genders[genderspec] then -- optimization
return {{
type = genderspec,
props = export.given_name_genders[genderspec],
}}, export.given_name_genders[genderspec].type == "animal"
end
local is_animal = nil
local param_mods = require(parameter_utilities_module).construct_param_mods {
{group = {"l", "q", "ref"}},
{param = {"text", "article"}},
}
local function generate_obj(term, parse_err)
if not export.given_name_genders[term] then
local valid_genders = {}
for k, _ in pairs(export.given_name_genders) do
table.insert(valid_genders, k)
end
table.sort(valid_genders)
parse_err(("Jantina '%s' tidak dikenali: jantina yang sah ialah %s"):format(
term, table.concat(valid_genders, ", ")))
end
return {
type = term,
props = export.given_name_genders[term],
}
end
local retval = require(parse_interface_module).parse_inline_modifiers(genderspec, {
param_mods = param_mods,
paramname = "2",
generate_obj = generate_obj,
splitchar = ",",
})
for _, spec in ipairs(retval) do
local this_is_animal = spec.props.type == "animal"
if is_animal == nil then
is_animal = this_is_animal
elseif is_animal ~= this_is_animal then
error("Jenis nama haiwan dan jantina manusia tidak boleh dicampurkan")
end
end
return retval, is_animal
end
local function generate_given_name_genders(lang, genders)
local parts = {}
for _, spec in ipairs(genders) do
local text
if spec.text then
-- NOTE: This assumes no % sign in the gender type, which seems safe.
text = spec.text:gsub("%+", spec.type)
else
if spec.props.type == "animal" then
text = "[[" .. spec.type .. "]]"
else
text = spec.type
end
end
if spec.q and spec.q[1] or spec.qq and spec.qq[1] or spec.l and spec.l[1] or spec.ll and spec.ll[1] or
spec.refs and spec.refs[1] then
text = require(decorations_module).format_decorations {
lang = lang,
text = text,
q = spec.q,
qq = spec.qq,
l = spec.l,
ll = spec.ll,
refs = spec.refs,
raw = true,
}
end
table.insert(parts, text)
end
local retval = m_table.serialCommaJoin(parts, {conj = "atau"})
local article = genders[1].article
if not article and not genders[1].text and not genders[1].q and not genders[1].l then
article = genders[1].props.article
end
if not article then
article = ""
end
return retval, article
end
-- The entry point for {{given name}}.
function export.given_name(frame)
local parent_args = frame:getParent().args
local compat = parent_args.lang
local offset = compat and 0 or 1
local lang_index = compat and "lang" or 1
local boolean = {type = "boolean"}
local list = {list = true}
local args = require("Module:parameters").process(parent_args, {
[lang_index] = {required = true, type = "language", default = "und"},
["gender"] = {default = "jantina tidak diketahui"},
[1 + offset] = {alias_of = "gender"},
["usage"] = true,
["origin"] = true,
["popular"] = true,
["populartype"] = true,
["meaning"] = list,
["meaningtype"] = true,
["addl"] = true,
["nocap"] = boolean,
-- initial article: A or An
["A"] = true,
["sort"] = true,
["from"] = true,
[2 + offset] = {alias_of = "from"},
["fromtype"] = true,
["xlit"] = true,
["eq"] = true,
["eqtype"] = true,
["varof"] = true,
["varoftype"] = true,
["var"] = {alias_of = "varof"},
["vartype"] = {alias_of = "varoftype"},
["varform"] = true,
["varformtype"] = true,
["dimof"] = true,
["dimoftype"] = true,
["dim"] = {alias_of = "dimof"},
["dimtype"] = {alias_of = "dimoftype"},
["dimform"] = true,
["dimformtype"] = true,
["augof"] = true,
["augoftype"] = true,
["aug"] = {alias_of = "augof"},
["augtype"] = {alias_of = "augoftype"},
["augform"] = true,
["augformtype"] = true,
["clipof"] = true,
["clipoftype"] = true,
["blend"] = true,
["blendtype"] = true,
["m"] = true,
["mtype"] = true,
["f"] = true,
["ftype"] = true,
["nocat"] = boolean,
})
local textsegs = {}
local lang = args[lang_index]
local langcode = lang:getCode()
local function fetch_typetext(param)
return args[param] and args[param] .. " " or ""
end
local genders, is_animal = parse_given_name_genders(args.gender)
local dimoftext, numdimofs = join_names(lang, args, "dimof")
local augoftext, numaugofs = join_names(lang, args, "augof")
local xlittext = join_names(nil, args, "xlit")
local blendtext = join_names(lang, args, "blend") -- formerly the only one that used "and" instead of "or"
local varoftext = join_names(lang, args, "varof")
local clipoftext = join_names(lang, args, "clipof")
local mtext = join_names(lang, args, "m")
local ftext = join_names(lang, args, "f")
local varformtext, numvarforms = join_names(lang, args, "varform", ", ")
local dimformtext, numdimforms = join_names(lang, args, "dimform", ", ")
local augformtext, numaugforms = join_names(lang, args, "augform", ", ")
local meaningsegs = {}
for _, meaning in ipairs(args.meaning) do
table.insert(meaningsegs, '“' .. meaning .. '”')
end
local meaningtext = m_table.serialCommaJoin(meaningsegs, {conj = "atau"})
local eqtext = get_eqtext(args)
local function ins(txt)
table.insert(textsegs, txt)
end
local dimoftype = args.dimoftype
local augoftype = args.augoftype
local added_text = nil
if numdimofs > 0 then
added_text = (dimoftype and dimoftype .. " " or "") .. "[[bentuk singkat]]" ..
(xlittext ~= "" and ", " .. xlittext .. "," or "") .. " daripada "
elseif numaugofs > 0 then
added_text = (augoftype and augoftype .. " " or "") .. "[[bentuk agam]]" ..
(xlittext ~= "" and ", " .. xlittext .. "," or "") .. " daripada "
end
local force_plural = false
if added_text ~= nil then
if args.dimof == "-" then
dimoftext = ""
force_plural = true
else
added_text = added_text .. ""
end
ins(added_text)
end
local article = args.A
if not article and textsegs[1] then
article = ""
end
ins("[[nama diri]]")
if not is_animal then
local gendertext, gender_article = generate_given_name_genders(lang, genders)
article = article or gender_article
ins(" ")
ins(gendertext)
end
article = article or "" -- if no article set yet, it's "a" based on "given name"
if langcode == "ms" and not args.nocap then
article = mw.getContentLanguage():ucfirst(article)
end
local need_comma = false
if numdimofs > 0 then
ins(" " .. dimoftext)
need_comma = not is_animal
elseif numaugofs > 0 then
ins(" " .. augoftext)
need_comma = not is_animal
elseif xlittext ~= "" then
ins(", " .. xlittext)
need_comma = true
end
if is_animal then
if need_comma then
ins(",")
end
need_comma = true
ins(" untuk ")
local gendertext, gender_article = generate_given_name_genders(lang, genders)
ins(gender_article ~= "" and gender_article .. " " or "")
ins(gendertext)
end
local from_catparts = {}
if args.from then
if need_comma then
ins(",")
end
need_comma = true
ins(" " .. fetch_typetext("fromtype"))
local textseg, this_catparts = get_fromtext(lang, args)
for _, catpart in ipairs(this_catparts) do
m_table.insertIfNot(from_catparts, catpart)
end
ins(textseg)
end
if meaningtext ~= "" then
if need_comma then
ins(",")
end
need_comma = true
ins(" " .. fetch_typetext("meaningtype") .. "bermaksud " .. meaningtext)
end
if args.origin then
if need_comma then
ins(",")
end
need_comma = true
ins(" berasal daripada " .. args.origin)
end
if args.usage then
if need_comma then
ins(",")
end
ins(" dengan penggunaan " .. args.usage)
end
if varoftext ~= "" then
ins(", " ..fetch_typetext("varoftype") .. "bentuk variasi daripada " .. varoftext)
end
if clipoftext ~= "" then
ins(", " .. fetch_typetext("clipoftype") .. "kependekan daripada " .. clipoftext)
end
if blendtext ~= "" then
ins(", " .. fetch_typetext("blendtype") .. "lakuran daripada " .. blendtext)
end
if args.popular then
ins(", " .. fetch_typetext("populartype") .. "popular " .. args.popular)
end
if mtext ~= "" then
ins(", " .. fetch_typetext("mtype") .. "padanan maskulin " .. mtext)
end
if ftext ~= "" then
ins(", " .. fetch_typetext("ftype") .. "padanan feminin " .. ftext)
end
if eqtext ~= "" then
ins(", " .. fetch_typetext("eqtype") .. "berpadanan dengan " .. eqtext)
end
if args.addl then
if args.addl:find("^;") then
ins(args.addl)
elseif args.addl:find("^_") then
ins(" " .. args.addl:sub(2))
else
ins(", " .. args.addl)
end
end
if varformtext ~= "" then
ins("; " .. fetch_typetext("varformtype") .. "bentuk variasi" .. " " ..
varformtext)
end
if dimformtext ~= "" then
ins("; " .. fetch_typetext("dimformtype") .. "bentuk singkat" .. " " ..
dimformtext)
end
if augformtext ~= "" then
ins("; " .. fetch_typetext("augformtype") .. "bentuk agam" .. " " ..
augformtext)
end
local text = "<span class='use-with-mention'>" .. (article ~= "" and article .. " " or "") .. table.concat(textsegs) .. "</span>"
if args.nocat then
return text
end
local categories = {}
local langname = " bahasa " .. lang:getCanonicalName()
local function insert_cats(dimaugof)
if dimaugof == "" and genders[1].props.type == "human" then
-- No category such as "English diminutives of given names"
table.insert(categories, "nama diri" .. langname)
end
local function insert_cat(cat)
table.insert(categories, dimaugof .. cat .. langname)
for _, catpart in ipairs(from_catparts) do
table.insert(categories, dimaugof .. cat .. langname .. " daripada " .. catpart)
end
end
for _, spec in ipairs(genders) do
local typ = spec.type
if spec.props.track then
track(typ)
end
local cats = get_given_name_cats(spec.type, spec.props)
for _, cat in ipairs(cats) do
insert_cat(cat)
end
end
end
insert_cats("")
if numdimofs > 0 then
insert_cats("bentuk singkat ")
elseif numaugofs > 0 then
insert_cats("bentuk agam ")
end
return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat)
end
-- The entry point for {{surname}}, {{patronymic}} and {{matronymic}}.
function export.surname(frame)
local iargs = require("Module:parameters").process(frame.args, {
["type"] = {required = true, set = {"nama keluarga", "patronimik", "matronimik"}},
})
local parent_args = frame:getParent().args
local compat = parent_args.lang
local offset = compat and 0 or 1
if parent_args.dot or parent_args.nodot then
error("dot= dan nodot= tidak lagi disokong dalam [[Template:" .. iargs.type .. "]] kerana tanda " ..
"noktah tidak lagi ditambah secara lalai; tambahkannya sendiri selepas templat jika diperlukan")
end
local lang_index = compat and "lang" or 1
local boolean = {type = "boolean"}
local list = {list = true}
local gender_arg = iargs.type == "nama keluarga" and "g" or 1 + offset
local adj_arg = iargs.type == "nama keluarga" and 1 + offset or 2 + offset
local args = require("Module:parameters").process(parent_args, {
[lang_index] = {required = true, type = "language", template_default = "und"},
[gender_arg] = iargs.type == "nama keluarga" and true or {required = true, template_default = "tidak diketahui"}, -- gender(s)
[adj_arg] = true, -- adjective/qualifier
["usage"] = true,
["origin"] = true,
["popular"] = true,
["populartype"] = true,
["meaning"] = list,
["meaningtype"] = true,
["parent"] = true,
["addl"] = true,
["nocap"] = boolean,
-- initial article: by default A or An (English), a or an (otherwise)
["A"] = true,
["sort"] = true,
["from"] = true,
["fromtype"] = true,
["xlit"] = true,
["eq"] = true,
["eqtype"] = true,
["varof"] = true,
["varoftype"] = true,
["var"] = {alias_of = "varof"},
["vartype"] = {alias_of = "varoftype"},
["varform"] = true,
["varformtype"] = true,
["clipof"] = true,
["clipoftype"] = true,
["blend"] = true,
["blendtype"] = true,
["m"] = true,
["mtype"] = true,
["f"] = true,
["ftype"] = true,
["nocat"] = boolean,
})
local textsegs = {}
local lang = args[lang_index]
local langcode = lang:getCode()
local function fetch_typetext(param)
return args[param] and args[param] .. " " or ""
end
local saw_male = false
local saw_female = false
local genders = {}
if args[gender_arg] then
for _, g in ipairs(require(parse_interface_module).split_on_comma(args[gender_arg])) do
if g == "tidak diketahui" or g == "jantina tidak diketahui" or g == "?" then
g = "jantina tidak diketahui"
track("jantina tidak diketahui")
elseif g == "uniseks" or g == "bebas jantina" or g == "c" then
g = "bebas jantina"
saw_male = true
saw_female = true
elseif g == "m" or g == "lelaki" then
g = "lelaki"
saw_male = true
elseif g == "f" or g == "perempuan" then
g = "perempuan"
saw_female = true
else
error("Jantina tidak dikenali: " .. g)
end
table.insert(genders, g)
end
end
local adj = args[adj_arg]
local xlittext = join_names(nil, args, "xlit")
local blendtext = join_names(lang, args, "blend", "dan")
local varoftext = join_names(lang, args, "varof")
local clipoftext = join_names(lang, args, "clipof")
local mtext = join_names(lang, args, "m")
local ftext = join_names(lang, args, "f")
local parenttext = join_names(lang, args, "parent", nil, "allow explicit lang")
local varformtext, numvarforms = join_names(lang, args, "varform", ", ")
local meaningsegs = {}
for _, meaning in ipairs(args.meaning) do
table.insert(meaningsegs, '“' .. meaning .. '”')
end
if parenttext ~= "" then
local child = saw_male and not saw_female and "anak lelaki" or saw_female and not saw_male and "anak perempuan" or
"anak lelaki/perempuan"
table.insert(meaningsegs, ("“%s kepada %s”"):format(child, parenttext))
end
local meaningtext = m_table.serialCommaJoin(meaningsegs, {conj = "atau"})
local eqtext = get_eqtext(args)
local function ins(txt)
table.insert(textsegs, txt)
end
ins("<span class='use-with-mention'>")
-- If gender is supplied, it goes before the specified adjective in adj=. The only value of gender that uses "an" is
-- "unknown-gender" (note that "unisex" wouldn't use it but in any case we map "unisex" to "common-gender"). If gender
-- isn't supplied, look at the first letter of the value of adj= if supplied; otherwise, the article is always "a"
-- because the word "surname", "patronymic" or "matronymic" follows. Capitalize "A"/"An" if English.
local article
if args.A then
article = args.A
else
article = #genders > 0 and genders[1] == "jantina tidak diketahui" and "" or
#genders == 0 and adj and "" or
""
if langcode == "ms" and not args.nocap then
article = mw.getContentLanguage():ucfirst(article)
end
end
ins(article ~= "" and article .. " " or "")
ins("[[" .. iargs.type .. "]]")
if #genders > 0 then
ins(" " .. table.concat(genders, " atau "))
end
if adj then
ins(" " .. adj)
end
local need_comma = false
if xlittext ~= "" then
ins(", " .. xlittext)
need_comma = true
end
local from_catparts = {}
if args.from then
if need_comma then
ins(",")
end
need_comma = true
ins(" " .. fetch_typetext("fromtype"))
local textseg, this_catparts = get_fromtext(lang, args)
for _, catpart in ipairs(this_catparts) do
m_table.insertIfNot(from_catparts, catpart)
end
ins(textseg)
end
if meaningtext ~= "" then
if need_comma then
ins(",")
end
need_comma = true
ins(" " .. fetch_typetext("meaningtype") .. "bermaksud " .. meaningtext)
end
if args.origin then
if need_comma then
ins(",")
end
need_comma = true
ins(" berasal daripada " .. args.origin)
end
if args.usage then
if need_comma then
ins(",")
end
need_comma = true
ins(" dengan penggunaan " .. args.usage)
end
if varoftext ~= "" then
ins(", " ..fetch_typetext("varoftype") .. "bentuk variasi daripada " .. varoftext)
end
if clipoftext ~= "" then
ins(", " .. fetch_typetext("clipoftype") .. "kependekan daripada " .. clipoftext)
end
if blendtext ~= "" then
ins(", " .. fetch_typetext("blendtype") .. "lakuran daripada " .. blendtext)
end
if args.popular then
ins(", " .. fetch_typetext("populartype") .. "popular " .. args.popular)
end
if mtext ~= "" then
ins(", " .. fetch_typetext("mtype") .. "padanan maskulin " .. mtext)
end
if ftext ~= "" then
ins(", " .. fetch_typetext("ftype") .. "padanan feminin " .. ftext)
end
if eqtext ~= "" then
ins(", " .. fetch_typetext("eqtype") .. "berpadanan dengan " .. eqtext)
end
if args.addl then
if args.addl:find("^;") then
ins(args.addl)
elseif args.addl:find("^_") then
ins(" " .. args.addl:sub(2))
else
ins(", " .. args.addl)
end
end
if varformtext ~= "" then
ins("; " .. fetch_typetext("varformtype") .. "bentuk variasi" ..
" " .. varformtext)
end
ins("</span>")
local text = table.concat(textsegs)
if args.nocat then
return text
end
local categories = {}
local langname = " bahasa " .. lang:getCanonicalName()
local function insert_cats(g)
g = g and " " .. g or ""
table.insert(categories, iargs.type .. g .. langname)
for _, catpart in ipairs(from_catparts) do
table.insert(categories, iargs.type .. g .. langname .. " daripada " .. catpart)
end
end
insert_cats(nil)
local function insert_cats_gender(g)
if g == "jantina tidak diketahui" then
return
end
if g == "bebas jantina" then
insert_cats_gender("lelaki")
insert_cats_gender("perempuan")
end
insert_cats(g)
end
for _, g in ipairs(genders) do
insert_cats_gender(g)
end
return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat)
end
-- The entry point for {{name translit}}, {{name respelling}}, {{name obor}} and {{foreign name}}.
function export.name_translit(frame)
local boolean = {type = "boolean"}
local iargs = require("Module:parameters").process(frame.args, {
["desctext"] = {required = true},
["obor"] = boolean,
["foreign_name"] = boolean,
})
local parent_args = frame:getParent().args
local params = {
[1] = {required = true, type = "language", template_default = "ms"},
[2] = {required = true, type = "language", sublist = true, template_default = "ru"},
[3] = {list = true, allow_holes = true},
["type"] = {required = true, set = translit_name_type_list, sublist = true, default = "patronimik"},
["dim"] = boolean,
["aug"] = boolean,
["nocap"] = boolean,
["addl"] = true,
["sort"] = true,
["pagename"] = true,
["nocat"] = boolean,
}
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{group = {"link", "q", "l", "ref"}},
{param = {"xlit", "eq"}},
}
local names, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 3,
track_module = "names/name translit",
disallow_custom_separators = true,
-- Use the first source language as the language of the specified names.
lang = function(args) return args[2][1] end,
sc = "sc.default",
-- We display decorations ourselves so we can display them "raw", as we surround everything with
-- 'use-with-mention' and want to avoid double-tagging.
no_show_decorations = true,
}
local lang = args[1]
local langcode = lang:getCode()
local sources = args[2]
local pagename = args.pagename or mw.loadData("Module:headword/data").pagename
local textsegs = {}
local function ins(txt)
table.insert(textsegs, txt)
end
ins("<span class='use-with-mention'>")
local desctext = iargs.desctext
if langcode == "ms" and not args.nocap then
desctext = mw.getContentLanguage():ucfirst(desctext)
end
ins(desctext .. " ")
if not iargs.foreign_name then
ins("daripada ")
end
local langsegs = {}
for i, source in ipairs(sources) do
local sourcename = source:getCanonicalName()
local function get_source_link()
local term_to_link = names[1] and names[1].term or pagename
-- We link the language name to either the first specified name or the pagename, in the following
-- circumstances:
-- (1) More than one language was given along with at least one name; or
-- (2) We're handling {{foreign name}} or {{name obor}}, and no name was given.
-- The reason for (1) is that if more than one language was given, we want a link to the name
-- in each language, as the name that's displayed is linked only to the first specified language.
-- However, if only one language was given, linking the language to the name is redundant.
-- The reason for (2) is that {{foreign name}} is often used when the name in the destination language
-- is spelled the same as the name in the source language (e.g. [[Clinton]] or [[Obama]] in Italian),
-- and in that case no name will be explicitly specified but we still want a link to the name in the
-- source language. The reason we restrict this to {{foreign name}} or {{name obor}}, not to
-- {{name translit}} or {{name respelling}}, is that {{name translit}} and {{name respelling}} ought to be
-- used for names spelled differently in the destination language (either transliterated or respelled), so
-- assuming the pagename is the name in the source language is wrong.
if names[1] and #sources > 1 or (iargs.foreign_name or iargs.obor) and not names[1] then
return m_links.language_link{
lang = sources[i], term = term_to_link, alt = sourcename, tr = "-"
}
else
return sourcename
end
end
if i == 1 and not iargs.foreign_name then
-- If at least one name is given, we say "A transliteration of the LANG surname FOO", linking LANG to FOO.
-- Otherwise we say "A transliteration of a LANG surname".
if names[1] then
table.insert(langsegs, get_source_link())
else
table.insert(langsegs, sourcename)
end
else
table.insert(langsegs, get_source_link())
end
end
local langseg_text = m_table.serialCommaJoin(langsegs, {conj = "atau"})
local augdim_text
if args.dim then
augdim_text = " [[bentuk singkat]]"
elseif args.aug then
augdim_text = " [[bentuk agam]]"
else
augdim_text = ""
end
local nametype_linked = {}
for _, nametype in ipairs(args["type"]) do
if nametype == "nama keluarga" or nametype == "patronimik" then
table.insert(nametype_linked, "[[" .. nametype .. "]]")
elseif nametype == "nama diri lelaki" then
table.insert(nametype_linked, "[[nama diri]] lelaki")
elseif nametype == "nama diri perempuan" then
table.insert(nametype_linked, "[[nama diri]] perempuan")
elseif nametype == "nama diri uniseks" then
table.insert(nametype_linked, "[[nama diri]] uniseks")
else
table.insert(nametype_linked, nametype)
end
end
local nametype_text = m_table.serialCommaJoin(nametype_linked, {conj = "atau"}) .. augdim_text
if not iargs.foreign_name then
ins(nametype_text)
ins(" bahasa " .. langseg_text)
if names[1] then
ins(" ")
end
else
ins(nametype_text)
ins(" dalam bahasa " .. langseg_text)
if names[1] then
ins(", ")
end
end
local linked_names = {}
local embedded_comma = false
for _, name in ipairs(names) do
local linked_name = m_links.full_link(name, "term")
if name.q and name.q[1] or name.qq and name.qq[1] or name.l and name.l[1] or name.ll and name.ll[1] or
name.refs and name.refs[1] then
linked_name = require(decorations_module).format_decorations {
lang = name.lang,
text = linked_name,
q = name.q,
qq = name.qq,
l = name.l,
ll = name.ll,
refs = name.refs,
raw = true, -- since we put 'use-with-mention' around the entire text
}
end
if name.xlit then
embedded_comma = true
linked_name = linked_name .. ", " .. m_links.language_link { lang = enlang, term = name.xlit }
end
if name.eq then
embedded_comma = true
linked_name = linked_name .. ", berpadanan dengan " .. m_links.language_link { lang = enlang, term = name.eq }
end
table.insert(linked_names, linked_name)
end
if embedded_comma then
ins(table.concat(linked_names, "; atau daripada "))
else
ins(m_table.serialCommaJoin(linked_names, {conj = "atau"}))
end
if args.addl then
if args.addl:find("^;") then
ins(args.addl)
elseif args.addl:find("^_") then
ins(" " .. args.addl:sub(2))
else
ins(", " .. args.addl)
end
end
ins("</span>")
local text = table.concat(textsegs)
if args.nocat then
return text
end
local categories = {}
local function inscat(cat)
table.insert(categories, cat)
end
for _, nametype in ipairs(args.type) do
local function insert_cats(dimaugof)
local function insert_cats_type(ty)
if ty == "nama diri uniseks" then
insert_cats_type("nama diri lelaki")
insert_cats_type("nama diri perempuan")
end
for _, source in ipairs(sources) do
inscat("Kemasan bahasa " .. lang:getFullName() .. " daripada " .. dimaugof .. ty .. " bahasa " .. source:getCanonicalName())
inscat("Perkataan bahasa " .. lang:getFullName() .. " diterbitkan daripada bahasa " .. source:getCanonicalName())
inscat("Perkataan bahasa " .. lang:getFullName() .. " dipinjam daripada bahasa " .. source:getCanonicalName())
if iargs.obor then
inscat("Pinjaman ortografi bahasa " .. lang:getFullName() .. " daripada bahasa " .. source:getCanonicalName())
end
if source:getCode() ~= source:getFullCode() then
-- etymology language
inscat("Kemasan bahasa " .. lang:getFullName() .. " daripada " .. dimaugof .. ty .. " bahasa " .. source:getFullName())
end
end
end
insert_cats_type(nametype)
end
insert_cats("")
if args.dim then
insert_cats("bentuk singkat ")
end
if args.aug then
insert_cats("bentuk agam ")
end
end
return text .. m_utilities.format_categories(categories, lang, args.sort, nil, force_cat)
end
return export
b2vgwn57cpwfm78qqio840nz6cbmyyt
Modul:etymology/templates/descendant
828
27509
375370
365964
2026-09-22T05:01:25Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708899|92708899]])
375370
Scribunto
text/plain
local export = {}
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local descendants_tree_module = "Module:descendants tree"
local etymology_style_css = "Module:etymology/style.css"
local labels_module = "Module:labels"
local languages_module = "Module:languages"
local links_module = "Module:links"
local parameter_utilities_module = "Module:parameter utilities"
local scripts_module = "Module:scripts"
local table_module = "Module:table"
local table_module_list_to_set = "Module:table/listToSet"
local template_styles_module = "Module:TemplateStyles"
local concat = table.concat
local insert = table.insert
local list_to_set = require(table_module_list_to_set)
local error_on_no_descendants = false
local function track(page)
return require(debug_track_module)("descendant/" .. page)
end
local function ine(arg)
if arg == "" then
return nil
else
return arg
end
end
local function add_tooltip(text, tooltip)
return '<span class="desc-arr" title="' .. tooltip .. '">' .. text .. '</span>'
end
-- Boolean params indicating whether a descendant term (or all terms) are particular sorts of borrowings.
local bortypes = {"inh", "bor", "lbor", "slb", "obor", "translit", "der", "clq", "pclq", "sml", "unc"}
local bortype_set = list_to_set(bortypes)
-- Aliases of clq=.
local calque_aliases = {"cal", "calq", "calque"}
local calque_alias_set = list_to_set(calque_aliases)
-- Aliases of pclq=.
local partial_calque_aliases = {"pcal", "pcalq", "pcalque"}
local partial_calque_alias_set = list_to_set(partial_calque_aliases)
local semi_learned_borrowing_aliases = {"slbor"}
local semi_learned_borrowing_alias_set = list_to_set(semi_learned_borrowing_aliases)
--- Return a function of one argument `field` (a param name), which fetches `args`[`field`].default if index == 0, else
--- `container`[`field`].
local function get_val(container, args, index)
return function(field)
if index == 0 then
return args[field].default
else
return container[field]
end
end
end
local function get_arrow(container, args, index)
local val = get_val(container, args, index)
local arrow
if val("bor") then
arrow = add_tooltip("→", "pinjaman")
elseif val("lbor") then
arrow = add_tooltip("→", "pinjaman terpelajar")
elseif val("slb") then
arrow = add_tooltip("→", "pinjaman terpelajar separa")
elseif val("obor") then
arrow = add_tooltip("→", "pinjaman ortografi")
elseif val("translit") then
arrow = add_tooltip("→", "transliterasi")
elseif val("clq") then
arrow = add_tooltip("→", "pinjaman terjemahan")
elseif val("pclq") then
arrow = add_tooltip("→", "pinjaman terjemahan separa")
elseif val("sml") then
arrow = add_tooltip("→", "pinjaman semantik")
elseif val("inh") or (val("unc") and not val("der")) then
arrow = add_tooltip(">", "diwariskan")
else
arrow = ""
end
-- allow der=1 in conjunction with bor=1 to indicate e.g. English "pars recta"
-- derived and borrowed from Latin "pars".
if val("der") then
arrow = arrow .. add_tooltip("⇒", "reshaped by analogy or addition of morphemes")
end
if val("unc") then
arrow = arrow .. add_tooltip("?", "uncertain")
end
if arrow ~= "" then
arrow = arrow .. " "
end
return arrow
end
-- Return the pre-decoration text for the `index`th term, or the overall pre-decoration text if index == 0.
local function get_pre_decorations(container, args, index)
if index > 0 then
-- per term decorations are handled at the subitem level, by full_link().
return nil, nil
end
local val = get_val(container, args, index)
return val("l"), val("q")
end
-- Return the post-decoration text for the `index`th term, or the overall post-decoration text if index == 0.
local function get_post_decorations(container, args, index, lang)
local val = get_val(container, args, index)
local boolean_labels = {}
if val("inh") then
insert(boolean_labels, "diwariskan")
end
if val("lbor") then
insert(boolean_labels, "terpelajar")
end
if val("slb") then
insert(boolean_labels, "terpelajar separa")
end
if val("translit") then
insert(boolean_labels, "transliterasi")
end
if val("clq") then
insert(boolean_labels, "pinjaman terjemahan")
end
if val("pclq") then
insert(boolean_labels, "pinjaman terjemahan separa")
end
if val("sml") then
insert(boolean_labels, "pinjaman semantik")
end
if index > 0 then
-- per term decorations are handled at the subitem level, by full_link().
return boolean_labels
else
local quals, dash_labels
quals = val("qq")
if val("ll") then
local labels = require(labels_module).show_labels {
lang = lang,
labels = val("ll"),
nocat = true,
open = false,
close = false,
no_track_already_seen = true,
ok_to_destructively_modify = true, -- doesn't apply to `labels`
}
if labels ~= "" then
dash_labels = " — " .. labels
end
end
return boolean_labels, quals, dash_labels
end
end
local function desc_or_desc_tree(frame, desc_tree)
local params
local boolean = {type = "boolean"}
if desc_tree then
params = {
[1] = {required = true, type = "language", family = true, default = "gem-pro"},
[2] = {required = true, list = true, allow_holes = true, default = "*fuhsaz"},
notext = boolean,
noalts = boolean,
noparent = boolean,
}
else
params = {
[1] = {required = true, type = "language", family = true, default = "en"},
[2] = {list = true, allow_holes = true, template_default = "word"},
alts = boolean,
}
end
-- Add other single params.
params.sclang = boolean
params.sclb = {replaced_by = "sclang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"}
params.nolang = boolean
params.nolb = {replaced_by = "nolang", reason = "to avoid confusion with 'labels' as in [[Template:lb]]"}
local parent_args
if frame.args[1] then
parent_args = frame.args
else
parent_args = frame:getParent().args
end
-- Error to catch most uses of old-style parameters.
if ine(parent_args[4]) and not ine(parent_args[3]) and not ine(parent_args.tr2) and not ine(parent_args.ts2)
and not ine(parent_args.t2) and not ine(parent_args.gloss2) and not ine(parent_args.g2)
and not ine(parent_args.alt2) then
error("You specified a term in 4= and not one in 3=. You probably meant to use t= to specify a gloss instead. "
.. "If you intended to specify two terms, put the second term in 3=.")
end
if not ine(parent_args[3]) and not ine(parent_args.alt2) and not ine(parent_args.tr2) and not ine(parent_args.ts2)
and ine(parent_args.g2) then
error("You specified a gender in g2= but no term in 3=. You were probably trying to specify two genders for "
.. "a single term. To do that, put both genders in g=, comma-separated.")
end
local m_param_utils = require(parameter_utilities_module)
local param_mods = m_param_utils.construct_param_mods {
{group = {"link", "ref", "l", "q"}},
{param = "lb", replaced_by = false, instead = "use 'l' for left labels or 'll' for right labels"},
{param = bortypes, type = "boolean", overall = true, separate_no_index = true},
{param = calque_aliases, alias_of = "clq"},
{param = partial_calque_aliases, alias_of = "pclq"},
{param = semi_learned_borrowing_aliases, alias_of = "slb"},
}
local groups, args, globalprops = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 2,
-- Need some work to support this.
-- parse_lang_prefix = true,
track_module = "descendant",
-- Due to allowing families as langs and substituting 'und', it's easier to do this later.
-- lang = function() ... end
sc = "sc.default",
splitchar = "[,~]",
subitem_separator_map = {[","] = " / ", ["~"] = " ~ "},
pre_normalize_modifiers = function(data)
local modtext = data.modtext
modtext = modtext:match("^<(.*)>$")
if not modtext then
error(("Internal error: Passed-in modifier isn't surrounded by angle brackets: %s"):format(
data.modtext))
end
if bortype_set[modtext] or calque_alias_set[modtext] or partial_calque_alias_set[modtext] or semi_learned_borrowing_alias_set[modtext] then
modtext = modtext .. ":1"
end
return "<" .. modtext .. ">"
end,
}
local lang = args[1]
local namespace = mw.title.getCurrentTitle().nsText
if (namespace == "" or namespace == "Rekonstruksi") and (
lang:hasType("appendix-constructed") and not lang:hasType("regular")) then
error("Istilah dalam bahasa binaan lampiran-sahaja tidak boleh diberikan sebagai keturunan.")
end
local fetch_alt_forms = desc_tree and not args.noalts or not desc_tree and args.alts
local m_desctree
if desc_tree or fetch_alt_forms then
m_desctree = require(descendants_tree_module)
end
if lang:getCode() ~= lang:getFullCode() then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/etymological]]
track("etymological")
track("etymological/" .. lang:getCode())
end
local is_family = lang:hasType("family")
local proxy_lang
if is_family then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/family]]
track("family")
track("family/" .. lang:getCode())
proxy_lang = require(languages_module).getByCode("und")
else
proxy_lang = lang
end
local langname
if is_family then
-- The display form for families includes the word "languages", which we probably don't want to
-- display.
langname = lang:getCanonicalName()
else
langname = lang:getDisplayForm()
end
local langtag
if args.sclang then
local sc_to_use = args.sc.default
if not sc_to_use then
local first_termobj = groups[1] and groups[1].terms[1]
if not first_termobj then
error("sclang= diberikan tetapi tiada istilah untuk memaparkan nama skrip")
end
sc_to_use = first_termobj.sc
if not sc_to_use then
local first_term = first_termobj.term or first_termobj.alt
if not first_term then
error("sclang= diberikan tetapi item pertama tiada istilah atau bentuk paparan untuk memaparkan nama skrip")
end
if first_termobj.lang then
sc_to_use = first_termobj.lang:findBestScript(first_term)
elseif is_family then
sc_to_use = require(scripts_module).findBestScriptWithoutLang(first_term, "none is last resort")
else
sc_to_use = lang:findBestScript(first_term)
end
end
end
langtag = sc_to_use:getDisplayForm(lang)
else
langtag = langname
end
local terms_for_descendant_trees = {}
-- Keep track of descendants whose descendant tree we fetch. Don't fetch the same descendant tree twice (which
-- can happen especially with Arabic-script terms with the same unvocalized spelling but differing vocalization).
-- This happens e.g. with Ottoman Turkish [[پورتقال]], which has {{desctree|fa-cls|پُرْتُقَال|پُرْتِقَال|bor=1}}, with
-- two terms that have the same unvocalized spelling.
local terms_and_ids_fetched = {}
local descendant_terms_seen = {}
local parts = {}
for i, group in ipairs(groups) do
local group_parts = {}
local terms_for_alt_forms = {}
for _, item in ipairs(group.terms) do
local link = ""
item.lang = item.lang or proxy_lang
item.track_sc = true
-- Construct a link out of `item`. Also add the term to the list of descendant trees and/or alternative
-- forms to fetch, if the page+ID combination hasn't already been seen.
if item.term ~= "-" then -- including term == nil
link = require(links_module).full_link(item, nil, true)
if item.term and (desc_tree or fetch_alt_forms) then
local m_links = require(links_module)
-- Fetches information under entry. If term is of type A//B, it checks A.
local entry_name = m_links.get_link_page(m_links.remove_links(mw.ustring.gsub(item.term, "//.+$", "")), lang, item.sc)
-- NOTE: We use the term and ID as the key, but not the language. This is OK currently because
-- all terms have the same language; but if we ever add support for a term-specific language,
-- we need to fix this.
local term_and_id = item.id and entry_name .. "!!!" .. item.id or entry_name
if not terms_and_ids_fetched[term_and_id] then
terms_and_ids_fetched[term_and_id] = true
local term_for_fetching = {
lang = lang, entry_name = entry_name, id = item.id
}
if desc_tree then
if is_family then
error("Tiada sokongan pada masa ini (dan mungkin tidak akan ada) untuk mengambil pokok keturunan apabila kod keluarga diberikan sebagai ganti kod bahasa")
end
if error_on_no_descendants then
require(table_module).insertIfNot(descendant_terms_seen,
{ term = item.term, id = item.id })
end
table.insert(terms_for_descendant_trees, term_for_fetching)
end
if fetch_alt_forms then
if is_family then
error("Tiada sokongan pada masa ini (dan mungkin tidak akan ada) untuk mengambil bentuk alternatif apabila kod keluarga diberikan sebagai ganti kod bahasa")
end
-- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/alts]]
track("alts")
table.insert(terms_for_alt_forms, term_for_fetching)
end
end
end
elseif item.tr or item.ts or item.gloss or item.genders then
-- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term]]
track("no term")
item.term = nil
item.show_decorations = true
link = require(links_module).full_link(item, nil, true)
link = link
:gsub("<small>%[Istilah%?%]</small> ", "")
:gsub("<small>%[Istilah%?%]</small> ", "")
:gsub("%[%[Category:[^%[%]]+ term requests%]%]", "")
else -- display no link at all
-- [[Special:WhatLinksHere/Wiktionary:Tracking/descendant/no term or annotations]]
track("no term or annotations")
end
if link ~= "" then
insert(group_parts, item.separator)
insert(group_parts, link)
end
end
if group_parts[1] then
for _, altterm in ipairs(terms_for_alt_forms) do
local altform = m_desctree.get_alternative_forms(altterm.lang, altterm.entry_name, altterm.id,
globalprops.use_semicolon and "; " or ", ")
if altform ~= "" then
insert(group_parts, globalprops.use_semicolon and "; " or ", ")
insert(group_parts, altform)
end
end
local group_link = concat(group_parts)
insert(parts, group.separator)
if not args.notext then
insert(parts, get_arrow(group, args, i))
end
-- no pre-qualifiers/labels and no post-qualifiers/dash-labels
local post_boolean_labels = get_post_decorations(group, args, i, proxy_lang)
if post_boolean_labels and post_boolean_labels[1] then
group_link = require(decorations_module).format_decorations {
lang = proxy_lang,
text = group_link,
ll = post_boolean_labels,
}
end
insert(parts, group_link)
end
end
local descendant_trees = {}
for _, descterm in ipairs(terms_for_descendant_trees) do
-- When I ([[User:Benwing2]]) first implemented this in Nov 2020, I had `maxmaxindex > 1` as the last argument.
-- Since then, [[User:Fytcha]] changed the last param to `true`.
local descendant_tree = m_desctree.get_descendants(descterm.lang, descterm.entry_name, descterm.id, true)
if descendant_tree and descendant_tree ~= "" then
insert(descendant_trees, descendant_tree)
end
end
if error_on_no_descendants and desc_tree and not descendant_trees[1] then
local function format_term_seen(term_seen)
if term_seen.id then
return ("[[%s]] dengan ID '%s'"):format(term_seen.term, term_seen.id)
else
return ("[[%s]]"):format(term_seen.term)
end
end
if #descendant_terms_seen == 0 then
error("[[Template:desctree]] dipanggil tetapi tiada istilah untuk mendapatkan keturunan")
elseif #descendant_terms_seen == 1 then
error(("Tiada bahagian Keturunan ditemui dalam entri %s di bawah pengepala untuk %s"):format(
format_term_seen(descendant_terms_seen[1]), lang:getFullName()))
else
for i, term_seen in ipairs(descendant_terms_seen) do
descendant_terms_seen[i] = format_term_seen(term_seen)
end
error(("Tiada bahagian Keturunan ditemui dalam mana-mana entri %s di bawah pengepala untuk %s"):format(
concat(descendant_terms_seen, ", "), lang:getFullName()))
end
end
local descendants = concat(descendant_trees)
if args.noparent then
return descendants
end
local initial_labels, initial_quals = get_pre_decorations(nil, args, 0)
local final_boolean_labels, final_quals, final_dash_labels = get_post_decorations(nil, args, 0, proxy_lang)
local all_linktext = concat(parts)
if initial_labels and initial_labels[1] or initial_quals and initial_quals[1] or
final_boolean_labels and final_boolean_labels[1] or final_quals and final_quals[1] then
all_linktext = require(decorations_module).format_decorations {
lang = proxy_lang,
text = all_linktext,
l = initial_labels,
q = initial_quals,
ll = final_boolean_labels,
qq = final_quals,
}
end
if final_dash_labels then
all_linktext = all_linktext .. final_dash_labels
end
all_linktext = all_linktext .. descendants
if args.notext then
return all_linktext
end
local initial_arrow = get_arrow(nil, args, 0)
if args.nolang then
return initial_arrow .. all_linktext
else
return concat { initial_arrow, langtag, ":", all_linktext ~= "" and " " or "", all_linktext }
end
end
function export.descendant(frame)
return desc_or_desc_tree(frame, false) .. require(template_styles_module)(etymology_style_css)
end
function export.descendants_tree(frame)
return desc_or_desc_tree(frame, true)
end
return export
4kh1nl61q5k5lsm5uswzc33ooz8wsiy
Modul:languages/data/exceptional
828
33718
375374
375287
2026-09-22T05:25:54Z
Hakimi97
2668
Betulkan keluarga bahasa bahasa Temuan terus ke bahasa Melayik Purba
375374
Scribunto
text/plain
local m_langdata = require("Module:languages/data")
-- Loaded on demand, as it may not be needed (depending on the data).
local function u(...)
u = require("Module:string utilities").char
return u(...)
end
local c = m_langdata.chars
local p = m_langdata.puaChars
local s = m_langdata.shared
local m = {}
m["aav-khs-pro"] = {
"Khasi Purba",
116773216,
"aav-khs",
"Latn",
type = "reconstructed",
}
m["aav-nic-pro"] = {
"Nicobar Purba",
116773793,
"aav-nic",
"Latn",
type = "reconstructed",
}
m["aav-pkl-pro"] = {
"Pnar-Khasi-Lyngngam Purba",
116773259,
"aav-pkl",
"Latn",
type = "reconstructed",
}
m["aav-pro"] = { -- mkh-pro akan digabungkan ke dalam ini
"Austroasia Purba",
116773186,
"aav",
"Latn",
type = "reconstructed",
}
m["afa-pro"] = {
"Afroasia Purba",
269125,
"afa",
"Latn",
type = "reconstructed",
}
m["alg-aga"] = {
"Agawam",
nil,
"alg-eas",
"Latn",
}
m["alg-pro"] = {
"Algonquian Purba",
7251834,
"alg",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = "·"},
}
m["alv-ama"] = {
"Amasi",
4740400,
"nic-grs",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron},
}
m["alv-bgu"] = {
"Bainouk Gubeeher",
17002646,
"alv-bny",
"Latn",
}
m["alv-bua-pro"] = {
"Bua Purba",
116773723,
"alv-bua",
"Latn",
type = "reconstructed",
}
m["alv-cng-pro"] = {
"Cangin Purba",
116773726,
"alv-cng",
"Latn",
type = "reconstructed",
}
m["alv-edo-pro"] = {
"Edoid Purba",
116773206,
"alv-edo",
"Latn",
type = "reconstructed",
}
m["alv-fli-pro"] = {
"Fali Purba",
116773754,
"alv-fli",
"Latn",
type = "reconstructed",
}
m["alv-gbe-pro"] = {
"Gbe Purba",
116773208,
"alv-gbe",
"Latn",
type = "reconstructed",
}
m["alv-gng-pro"] = {
"Guang Purba",
116773757,
"alv-gng",
"Latn",
type = "reconstructed",
}
m["alv-gtm-pro"] = {
"Togo Tengah Purba",
116773732,
"alv-gtm",
"Latn",
type = "reconstructed",
}
m["alv-gwa"] = {
"Gwara",
16945580,
"nic-pla",
"Latn",
}
m["alv-hei-pro"] = {
"Heiban Purba",
116773760,
"alv-hei",
"Latn",
type = "reconstructed",
}
m["alv-ido-pro"] = {
"Idomoid Purba",
116773764,
"alv-ido",
"Latn",
type = "reconstructed",
}
m["alv-igb-pro"] = {
"Igboid Purba",
116773765,
"alv-igb",
"Latn",
type = "reconstructed",
}
m["alv-kwa-pro"] = {
"Kwa Purba",
116773780,
"alv-kwa",
"Latn",
type = "reconstructed",
}
m["alv-mum-pro"] = {
"Mumuye Purba",
116773791,
"alv-mum",
"Latn",
type = "reconstructed",
}
m["alv-nup-pro"] = {
"Nupoid Purba",
116773795,
"alv-nup",
"Latn",
type = "reconstructed",
}
m["alv-pro"] = {
"Atlantik-Congo Purba",
116732838,
"alv",
"Latn",
type = "reconstructed",
}
m["alv-edk-pro"] = {
"Edekiri Purba",
nil,
"alv-edk",
"Latn",
type = "reconstructed",
}
m["alv-yor-pro"] = {
"Yoruba Purba",
nil,
"alv-yor",
"Latn",
type = "reconstructed",
}
m["alv-yrd-pro"] = {
"Yoruboid Purba",
116773824,
"alv-yrd",
"Latn",
type = "reconstructed",
}
m["alv-von-pro"] = {
"Volta-Niger Purba",
116773820,
"alv-von",
"Latn",
type = "reconstructed",
}
m["apa-pro"] = {
"Apache Purba",
116773135,
"apa",
"Latn",
type = "reconstructed",
}
m["aql-pro"] = {
"Algik Purba",
18389588,
"aql",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = "·"},
}
m["art-adu"] = {
"Adûni",
1232159,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-bel"] = {
"Kreol Belter",
108055510,
"art",
"Latn",
type = "appendix-constructed",
sort_key = {
remove_diacritics = c.acute,
from = {"ɒ"},
to = {"a"},
},
}
m["art-blk"] = {
"Bolak",
2909283,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-bsp"] = {
"Bahasa Hitam",
686210,
"art",
"Latn, Teng",
type = "appendix-constructed",
}
m["art-com"] = {
"Communicationssprache",
35227,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-dtk"] = {
"Dothraki",
2914733,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-elo"] = {
"Eloi",
nil,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-gld"] = {
"Goa'uld",
19823,
"art",
"Latn, Egyp, Mero",
type = "appendix-constructed",
}
m["art-lap"] = {
"Lapine",
6488195,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-man"] = {
"Mandalorian",
54289,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-mun"] = {
"Mundolinco",
851355,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-nav"] = {
"Naʼvi",
316939,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-vlh"] = {
"Valyria Tinggi",
64483808,
"art",
"Latn",
type = "appendix-constructed",
}
m["ath-nic"] = {
"Nicola",
20609,
"ath-nor",
"Latn",
}
m["ath-pro"] = {
"Athabaska Purba",
104841722,
"ath",
"Latn",
type = "reconstructed",
}
m["auf-pro"] = {
"Arawa Purba",
116773706,
"auf",
"Latn",
type = "reconstructed",
}
m["aus-alu"] = {
"Alungul",
16827670,
"aus-pmn",
"Latn",
}
m["aus-and"] = {
"Andjingith",
4754509,
"aus-pmn",
"Latn",
}
m["aus-ang"] = {
"Angkula",
16828520,
"aus-pmn",
"Latn",
}
m["aus-arn-pro"] = {
"Arnhem Purba",
116773720,
"aus-arn",
"Latn",
type = "reconstructed",
}
m["aus-bra"] = {
"Barranbinya",
4863220,
"aus-pmn",
"Latn",
}
m["aus-brm"] = {
"Barunggam",
4865914,
"aus-pmn",
"Latn",
}
m["aus-cww-pro"] = {
"New South Wales Tengah Purba",
116773199,
"aus-cww",
"Latn",
type = "reconstructed",
}
m["aus-dal-pro"] = {
"Daly Purba",
116773743,
"aus-dal",
"Latn",
type = "reconstructed",
}
m["aus-guw"] = {
"Guwar",
6652138,
"aus-pam",
"Latn",
}
m["aus-lsw"] = {
"Little Swanport",
6652138,
"qfa-unc",
"Latn",
}
m["aus-mbi"] = {
"Mbiywom",
6799701,
"aus-pmn",
"Latn",
}
m["aus-ngk"] = {
"Ngkoth",
7022405,
"aus-pmn",
"Latn",
}
m["aus-nyu-pro"] = {
"Nyulnyulan Purba",
116773797,
"aus-nyu",
"Latn",
type = "reconstructed",
}
m["aus-pam-pro"] = {
"Pama-Nyunga Purba",
33942,
"aus-pam",
"Latn",
type = "reconstructed",
}
m["aus-tul"] = {
"Tulua",
16938541,
"aus-pam",
"Latn",
}
m["aus-uwi"] = {
"Uwinymil",
7903995,
"aus-arn",
"Latn",
}
m["aus-wdj-pro"] = {
"Iwaidjan Purba",
116773767,
"aus-wdj",
"Latn",
type = "reconstructed",
}
m["aus-won"] = {
"Wong-gie",
nil,
"aus-pam",
"Latn",
}
m["aus-wul"] = {
"Wulguru",
8039196,
"aus-dyb",
"Latn",
}
m["aus-ynk"] = { -- kontras nny
"Yangkaal",
3913770,
"aus-tnk",
"Latn",
}
m["awd-amc-pro"] = {
"Amuesha-Chamicuro Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-kmp-pro"] = {
"Kampa Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-prw-pro"] = {
"Paresi-Waura Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-ama"] = {
"Amarizana",
16827787,
"awd",
"Latn",
}
m["awd-ana"] = {
"Anauyá",
16828252,
"awd",
"Latn",
}
m["awd-apo"] = {
"Apolista",
16916645,
"awd",
"Latn",
}
m["awd-cab"] = {
"Cabre",
16850160,
"awd",
"Latn",
}
m["awd-gnu"] = {
"Guinau",
3504087,
"awd",
"Latn",
}
m["awd-kar"] = {
"Cariay",
16920253,
"awd",
"Latn",
}
m["awd-kaw"] = {
"Kawishana",
6379993,
"awd-nwk",
"Latn",
}
m["awd-kus"] = {
"Kustenau",
5196293,
"awd",
"Latn",
}
m["awd-man"] = {
"Manao",
6746920,
"awd",
"Latn",
}
m["awd-mar"] = {
"Marawan",
6755108,
"awd",
"Latn",
}
m["awd-mpr"] = {
"Maipure",
6736872,
"awd",
"Latn",
}
m["awd-mrt"] = {
"Mariaté",
16910017,
"awd-nwk",
"Latn",
}
m["awd-nwk-pro"] = {
"Nawiki Purba",
116773234,
"awd-nwk",
"Latn",
type = "reconstructed",
}
m["awd-pai"] = {
"Paikoneka",
128807835,
"awd",
"Latn",
}
m["awd-pas"] = {
"Pasé",
7143168,
"awd-nwk",
"Latn",
}
m["awd-pro"] = {
"Arawak Purba",
97573478,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-she"] = {
"Shebayo",
7492248,
"awd",
"Latn",
}
m["awd-taa-pro"] = {
"Ta-Arawak Purba",
116773282,
"awd-taa",
"Latn",
type = "reconstructed",
}
m["awd-wai"] = {
"Wainumá",
16910017,
"awd-nwk",
"Latn",
}
m["awd-war"] = {
"Warekena Kuno",
105320180,
"awd-nwk",
"Latn",
}
m["awd-yum"] = {
"Yumana",
8061062,
"awd-nwk",
"Latn",
}
m["azc-caz"] = {
"Cazcan",
5055514,
"azc",
"Latn",
}
m["azc-cup-pro"] = {
"Cupan Purba",
116773738,
"azc-cup",
"Latn",
type = "reconstructed",
}
m["azc-ktn"] = {
"Kitanemuk",
3197558,
"azc-tak",
"Latn",
}
m["azc-nah-pro"] = {
"Nahua Purba",
7251860,
"azc-nah",
"Latn",
type = "reconstructed",
}
m["azc-nic"] = {
"Nicoleño",
50241488,
"azc",
"Latn",
}
m["azc-num-pro"] = {
"Numik Purba",
116773247,
"azc-num",
"Latn",
type = "reconstructed",
}
m["azc-pro"] = {
"Uto-Aztek Purba",
96400333,
"azc",
"Latn",
type = "reconstructed",
}
m["azc-tak-pro"] = {
"Takik Purba",
116773283,
"azc-tak",
"Latn",
type = "reconstructed",
}
m["azc-tat"] = {
"Tataviam",
743736,
"azc",
"Latn",
}
m["ber-pro"] = {
"Berber Purba",
2855698,
"ber",
"Latn",
type = "reconstructed",
}
m["ber-fog"] = {
"Fogaha",
107610173,
"ber",
"Latn",
}
m["ber-zuw"] = {
"Zuwara",
4117169,
"ber",
"Latn",
}
m["bnt-bal"] = {
"Balong",
93935237,
"bnt-bbo",
"Latn",
}
m["bnt-bon"] = {
"Boma Nkuu",
nil,
"bnt",
"Latn",
}
m["bnt-boy"] = {
"Boma Yumu",
nil,
"bnt",
"Latn",
}
m["bnt-bwa"] = {
"Bwala",
128810345,
"bnt-tek",
"Latn",
}
m["bnt-cmw"] = {
"Chimwiini",
4958328,
"bnt-swh",
"Latn",
}
m["bnt-ind"] = {
"Indanga",
51412803,
"bnt",
"Latn",
}
m["bnt-lal"] = {
"Lala (Afrika Selatan)",
6480154,
"bnt-ngu",
"Latn",
}
m["bnt-mpi"] = {
"Mpiin",
93937013,
"bnt-bdz",
"Latn",
}
m["bnt-mpu"] = {
"Mpuono", -- jangan dikelirukan dengan Mbuun zmp
36056,
"bnt",
"Latn",
}
m["bnt-ngu-pro"] = {
"Nguni Purba",
961559,
"bnt-ngu",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.caron},
}
m["bnt-phu"] = {
"Phuthi",
33796,
"bnt-ngu",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute},
}
m["bnt-pro"] = {
"Bantu Purba",
3408025,
"bnt",
"Latn",
type = "reconstructed",
sort_key = "bnt-pro-sortkey",
}
m["bnt-sab-pro"] = {
"Sabaki Purba",
nil, -- Q2209395 ialah kod untuk keluarga Sabaki
"bnt-sab",
"Latn",
type = "reconstructed",
}
m["bnt-sbo"] = {
"Boma Selatan",
nil,
"bnt",
"Latn",
}
m["bnt-sts-pro"] = {
"Sotho-Tswana Purba",
116773278,
"bnt-sts",
"Latn",
type = "reconstructed",
}
m["btk-pro"] = {
"Batak Purba",
116773191,
"btk",
"Latn",
type = "reconstructed",
}
m["cau-abz-pro"] = {
"Abkhaz-Abaza Purba",
7251831,
"cau-abz",
"Latn",
type = "reconstructed",
}
m["cau-and-pro"] = {
"Andi Purba",
nil,
"cau-and",
"Latn",
type = "reconstructed",
}
m["cau-ava-pro"] = {
"Avar-Andi Purba",
116773187,
"cau-ava",
"Latn",
type = "reconstructed",
}
m["cau-cir-pro"] = {
"Circassia Purba",
7251838,
"cau-cir",
"Latn",
type = "reconstructed",
}
m["cau-drg-pro"] = {
"Dargwa Purba",
116773205,
"cau-drg",
"Latn",
type = "reconstructed",
}
m["cau-lzg-pro"] = {
"Lezghi Purba",
116773223,
"cau-lzg",
"Latn",
type = "reconstructed",
}
m["cau-nec-pro"] = {
"Kaukasia Timur Laut Purba",
116773244,
"cau-nec",
"Latn",
type = "reconstructed",
}
m["cau-nkh-pro"] = {
"Nakh Purba",
108032840,
"cau-nkh",
"Latn",
type = "reconstructed",
}
m["cau-nwc-pro"] = {
"Kaukasia Barat Laut Purba",
7251861,
"cau-nwc",
"Latn",
type = "reconstructed",
}
m["cau-tsz-pro"] = {
"Tsez Purba",
116773287,
"cau-tsz",
"Latn",
type = "reconstructed",
}
m["cba-ata"] = {
"Atanques",
4812783,
"cba",
"Latn",
}
m["cba-cat"] = {
"Catío Chibcha",
7083619,
"cba",
"Latn",
}
m["cba-dor"] = {
"Dorasque",
5297532,
"cba",
"Latn",
}
m["cba-dui"] = {
"Duit",
3041061,
"cba",
"Latn",
}
m["cba-hue"] = {
"Huetar",
35514,
"cba",
"Latn",
}
m["cba-nut"] = {
"Nutabe",
7070405,
"cba",
"Latn",
}
m["cba-pro"] = {
"Chibchan Purba",
116773203,
"cba",
"Latn",
type = "reconstructed",
}
m["ccs-pro"] = {
"Kartvelia Purba",
2608203,
"ccs",
"Latn",
type = "reconstructed",
strip_diacritics = {
from = {"q̣", "p̣", "ʓ", "ċ"},
to = {"q̇", "ṗ", "ʒ", "c̣"}
},
}
m["ccs-gzn-pro"] = {
"Georgia-Zan Purba",
23808119,
"ccs-gzn",
"Latn",
type = "reconstructed",
strip_diacritics = {
from = {"q̣", "p̣", "ʓ", "ċ"},
to = {"q̇", "ṗ", "ʒ", "c̣"}
},
}
m["cdc-cbm-pro"] = {
"Chadik Tengah Purba",
116773197,
"cdc-cbm",
"Latn",
type = "reconstructed",
}
m["cdc-mas-pro"] = {
"Masa Purba",
116773789,
"cdc-mas",
"Latn",
type = "reconstructed",
}
m["cdc-pro"] = {
"Chadik Purba",
116773201,
"cdc",
"Latn",
type = "reconstructed",
}
m["cdd-pro"] = {
"Caddoan Purba",
116773725,
"cdd",
"Latn",
type = "reconstructed",
}
m["cel-bry-pro"] = {
"Britonik Purba",
1248800,
"cel-bry",
"Latn, Polyt",
sort_key = {
Latn = "cel-bry-pro-sortkey",
},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["cel-gal"] = {
"Gallaecia",
3094789,
"cel-his",
}
m["cel-gau"] = {
"Gaul",
29977,
"cel",
"Latn, Polyt, Ital",
strip_diacritics = {
Latn = {remove_diacritics = c.macron .. c.breve .. c.diaer},
},
sort_key = {
Latn = "cel-bry-pro-sortkey",
},
-- translit Ital dalam [[Module:scripts/data]]
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["cel-pro"] = {
"Keltik Purba",
653649,
"cel",
"Latn",
type = "reconstructed",
sort_key = "cel-pro-sortkey",
}
m["chi-pro"] = {
"Chimakuan Purba",
116773734,
"chi",
"Latn",
type = "reconstructed",
}
m["chm-pro"] = {
"Mari Purba",
116773788,
"chm",
"Latn",
type = "reconstructed",
}
m["cmc-pro"] = {
"Chamik Purba",
114793834,
"cmc",
"Latn",
type = "reconstructed",
}
m["crp-bip"] = {
"Pijin Basque-Iceland",
810378,
"crp",
"Latn",
ancestors = "eu",
}
m["crp-cpr"] = {
"Pijin Rusia-China",
nil,
"crp",
"Hani, Cyrl, Latn",
ancestors = "ru, zh",
translit = {Cyrl = "ru-translit"},
strip_diacritics = {
Cyrl = {remove_diacritics = c.acute .. c.grave .. c.macron},
},
}
m["crp-gep"] = {
"Pijin Greenland Barat",
17036301,
"crp",
"Latn",
ancestors = "kl",
}
m["crp-kia"] = {
"Pijin Jerman Kiautschou",
108314615,
"crp",
"Latn",
ancestors = "de",
}
m["crp-mar"] = {
"Bahasa Roh Maroon",
1093206,
"crp",
"Latn",
ancestors = "en",
}
m["crp-mpp"] = {
"Pijin Portugis Macau",
128804537,
"crp",
"Hant, Latn",
ancestors = "pt",
sort_key = {Hant = "Hani-sortkey"},
}
m["crp-rsn"] = {
"Russenorsk",
505125,
"crp",
"Cyrl, Latn",
ancestors = "nn, ru",
translit = {Cyrl = "ru-translit"},
}
m["crp-spp"] = {
"Pijin Ladang Samoa",
7409948,
"crp",
"Latn",
ancestors = "en",
}
m["crp-slb"] = {
"Inggeris Solombala",
7558525,
"crp",
"Cyrl, Latn",
ancestors = "en, ru",
translit = {Cyrl = "ru-translit"},
}
m["crp-tpr"] = {
"Pijin Rusia Taimyr",
16930506,
"crp",
"Cyrl",
ancestors = "ru",
translit = "ru-translit",
}
m["csu-bba-pro"] = {
"Bongo-Bagirmi Purba",
116773722,
"csu-bba",
"Latn",
type = "reconstructed",
}
m["csu-maa-pro"] = {
"Mangbetu Purba",
116773786,
"csu-maa",
"Latn",
type = "reconstructed",
}
m["csu-pro"] = {
"Sudan Tengah Purba",
116773730,
"csu",
"Latn",
type = "reconstructed",
}
m["csu-sar-pro"] = {
"Sara Purba",
116773809,
"csu-sar",
"Latn",
type = "reconstructed",
}
m["cus-ash"] = {
"Ashraaf",
4805855,
"cus-som",
"Latn",
}
m["cus-hec-pro"] = {
"Kusyi Timur Tanah Tinggi Purba",
116773761,
"cus-hec",
"Latn",
type = "reconstructed",
}
m["cus-som-pro"] = {
"Somaloid Purba",
nil,
"cus-som",
"Latn",
type = "reconstructed",
}
m["cus-sou-pro"] = {
"Kusyi Selatan Purba",
126081567,
"cus-sou",
"Latn",
type = "reconstructed",
}
m["cus-pro"] = {
"Kusyi Purba",
116773204,
"cus",
"Latn",
type = "reconstructed",
}
m["dmn-dam"] = {
"Dama (Sierra Leone)",
19601574,
"dmn",
"Latn",
}
m["dra-bry"] = {
"Beary",
1089116,
"qfa-mix",
"Mlym, Knda",
ancestors = "ml, tcy",
-- translit Knda dalam [[Module:scripts/data]]
-- translit Mlym dalam [[Module:scripts/data]]
}
m["dra-cen-pro"] = {
"Dravidia Tengah Purba",
nil,
"dra-cen",
"Latn",
type = "reconstructed",
}
m["dra-mkn"] = {
"Kannada Pertengahan",
128810572,
"dra-kan",
"Knda",
-- translit Knda dalam [[Module:scripts/data]]
}
m["dra-nor-pro"] = {
"Dravidia Utara Purba",
124433593,
"dra-nor",
"Latn",
type = "reconstructed",
}
m["dra-okn"] = {
"Kannada Kuno",
15723156,
"dra-kan",
"Knda",
-- translit Knda dalam [[Module:scripts/data]]
}
m["dra-ote"] = {
"Telugu Kuno",
126720868,
"dra-tel",
"Telu",
translit = "te-translit",
}
m["dra-pro"] = {
"Dravidia Purba",
1702853,
"dra",
"Latn",
type = "reconstructed",
}
m["dra-sdo-pro"] = {
"Dravidia Selatan I Purba",
104847952, -- "Proto-Dravidia Selatan" Wikipedia ialah Proto-Dravidia Selatan I dalam skema ini.
"dra-sdo",
"Latn",
type = "reconstructed",
}
m["dra-sdt-pro"] = {
"Dravidia Selatan II Purba",
128885257,
"dra-sdt",
"Latn",
type = "reconstructed",
}
m["dra-sou-pro"] = {
"Dravidia Selatan Purba",
128886121,
"dra-sou",
"Latn",
type = "reconstructed",
}
m["egx-dem"] = {
"Mesir Demotik",
36765,
"egx",
"Latn, Egyd, Polyt",
sort_key = {
Latn = {
remove_diacritics = "'%-%s",
from = {"ꜣ", "j", "e", "ꜥ", "y", "w", "b", "p", "f", "m", "n", "r", "l", "ḥ", "ḫ", "h̭", "ẖ", "h", "š", "s", "q", "k", "g", "ṱ", "ṯ", "t", "ḏ", "%.", "⸗"},
to = {p[1], p[2], p[3], p[4], p[5], p[6], p[7], p[8], p[9], p[10], p[11], p[12], p[13], p[15], p[16], p[16], p[17], p[14], p[19], p[18], p[20], p[21], p[22], p[23], p[24], p[23], p[25], p[26], p[26]}
},
},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["dmn-pro"] = {
"Mande Purba",
116773785,
"dmn",
"Latn",
type = "reconstructed",
}
m["dmn-mdw-pro"] = {
"Mande Barat Purba",
116773822,
"dmn-mdw",
"Latn",
type = "reconstructed",
}
m["dru-pro"] = {
"Rukai Purba",
116773807,
"map",
"Latn",
type = "reconstructed",
}
m["ero-gsz"] = {
"Geshiza",
nil,
"ero",
"Latn",
}
m["ero-nya"] = {
"Nyagrong Minyag",
nil,
"ero",
"Latn",
}
m["ero-tau"] = {
"Stau",
nil,
"ero",
"Latn",
}
m["esx-esk-pro"] = {
"Eskimo Purba",
7251842,
"esx-esk",
"Latn",
type = "reconstructed",
}
m["esx-ink"] = {
"Inuktun",
1671647,
"esx-inu",
"Latn",
}
m["esx-inq"] = {
"Inuinnaqtun",
28070,
"esx-inu",
"Latn",
}
m["esx-inu-pro"] = {
"Inuit Purba",
60785588,
"esx-inu",
"Latn",
type = "reconstructed",
}
m["esx-pro"] = {
"Eskimo-Aleut Purba",
7251843,
"esx",
"Latn",
type = "reconstructed",
}
m["esx-tut"] = {
"Tunumiisut",
15665389,
"esx-inu",
"Latn",
}
m["euq-pro"] = {
"Basque Purba",
938011,
"euq",
"Latn",
type = "reconstructed",
}
m["gba-pro"] = {
"Gbaya Purba",
nil,
"gba",
"Latn",
type = "reconstructed",
}
m["gem-pro"] = {
"Jermanik Purba",
669623,
"gem",
"Latn",
type = "reconstructed",
sort_key = "gem-pro-sortkey",
}
m["gme-bur"] = {
"Burgundia",
47625,
"gme",
"Latn",
}
m["gme-cgo"] = {
"Goth Crimea",
36211,
"gme",
"Latn",
}
m["gmq-gut"] = {
"Gutnish",
1256646,
"gmq",
"Latn",
ancestors = "gmq-ogt",
}
m["gmq-jmk"] = {
"Jamtish",
35512,
"gmq-eas",
"Latn",
}
m["gmq-mno"] = {
"Norway Pertengahan",
3417070,
"gmq-wes",
"Latn",
}
m["gmq-oda"] = {
"Denmark Kuno",
12330003,
"gmq-eas",
"Latn, Runr",
strip_diacritics = {remove_diacritics = c.macron},
}
m["gmq-ogt"] = {
"Gutnish Kuno",
1133488,
"gmq",
"Latn, Runr",
ancestors = "non",
}
m["gmq-osw"] = {
"Sweden Kuno",
2417210,
"gmq-eas",
"Latn, Runr",
strip_diacritics = {remove_diacritics = c.macron},
}
m["gmq-pro"] = {
"Norse Purba",
1671294,
"gmq",
"Runr",
translit = "Runr-translit",
}
m["gmq-scy"] = {
"Scanian",
768017,
"gmq-eas",
"Latn",
}
m["gmw-bgh"] = {
"Bergish",
329030,
"gmw-frk",
"Latn",
}
m["gmw-cfr"] = {
"Franconia Tengah",
572197,
"gmw-hgm",
"Latn",
ancestors = "gmh",
wikimedia_codes = "ksh",
}
m["gmw-ecg"] = {
"Jerman Tengah Timur",
499344, -- merangkumi Q699284, Q152965
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gmw-fin"] = {
"Fingallian",
3072588,
"gmw-ian",
"Latn",
}
m["gmw-gts"] = {
"Gottscheerish",
533109,
"gmw-hgm",
"Latn",
ancestors = "bar",
}
m["gmw-jdt"] = {
"Belanda Jersey",
1687911,
"gmw-frk",
"Latn",
ancestors = "nl",
}
m["gmw-msc"] = {
"Scots Pertengahan",
3327000,
"gmw-ang",
"Latn",
ancestors = "enm-esc",
}
m["gmw-pro"] = {
"Jermanik Barat Purba",
78079021,
"gmw",
"Latn, Runr",
-- type = "reconstructed",
-- sebahagian besarnya tetapi tidak sepenuhnya direkonstruksi (seperti Proto-Norse); lihat BP Apr '24, tetapkan kembali kepada direkonstruksi (?) jika 'anti-asterisk' ditambah
sort_key = "gmw-pro-sortkey",
}
m["gmw-rfr"] = {
"Franconia Rhine",
707007,
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gmw-stm"] = {
"Schwaben Szatmár",
2223059,
"gmw-hgm",
"Latn",
ancestors = "swg",
}
m["gmw-tsx"] = {
"Saxon Transylvania",
260942,
"gmw-hgm",
"Latn",
ancestors = "gmw-cfr",
}
m["gmw-vog"] = {
"Jerman Volga",
312574,
"gmw-hgm",
"Latn",
ancestors = "gmw-rfr",
}
m["gmw-zps"] = {
"Jerman Zipser",
205548,
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gn-cls"] = {
"Guarani Klasik",
17478065,
"gn",
"Latn",
}
m["grk-cal"] = {
"Yunani Calabria",
1146398,
"grk",
"Latn, Grek",
ancestors = "grk-ita",
translit = {
Grek = "el-translit",
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-ita"] = {
"Yunani Italiot",
19720507,
"grk",
"Latn, Grek",
ancestors = "gkm",
translit = {
Grek = "el-translit",
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-mar"] = {
"Yunani Mariupol",
4400023,
"grk",
"Cyrl, Latn, Grek",
ancestors = "gkm",
translit = {
Cyrl = "grk-mar-translit",
Grek = "grk-mar-translit",
},
override_translit = true,
strip_diacritics = {
Cyrl = {remove_diacritics = c.acute},
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-pro"] = {
"Hellenik Purba",
1231805,
"grk",
"Latn, Polyt",
type = "reconstructed",
sort_key = {Latn = {
from = {"ʰ", "ʷ"},
to = {"h", "w"},
remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.caron .. c.CGJ
}},
display_text = {Latn = {
from = {"([dlLt])" .. c.caron},
to = {"%1" .. c.CGJ .. c.caron},
}},
strip_diacritics = {Latn = {
from = {"([dlLt])" .. c.caron},
to = {"%1" .. c.CGJ .. c.caron},
}},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
-- NOTA: dahulunya tiada translit ditentukan untuk Polyt; mungkin ketinggalan secara tidak sengaja; jika tidak, tetapkan Polyt = false dalam
-- bahagian translit
}
m["hmn-pro"] = {
"Hmongik Purba",
116773210,
"hmn",
"Latn",
type = "reconstructed",
}
m["hmx-mie-pro"] = {
"Mienik Purba",
116773229,
"hmx-mie",
"Latn",
type = "reconstructed",
}
m["hmx-pro"] = {
"Hmong-Mien Purba",
7251846,
"hmx",
"Latn",
type = "reconstructed",
}
m["hyx-pro"] = {
"Armenia Purba",
3848498,
"hyx",
"Latn",
type = "reconstructed",
}
m["iir-nur-pro"] = {
"Nuristani Purba",
116773248,
"iir-nur",
"Latn",
type = "reconstructed",
}
m["iir-pro"] = {
"Indo-Iran Purba",
966439,
"iir",
"Latn",
type = "reconstructed",
}
m["ijo-pro"] = {
"Ijoid Purba",
116773766,
"ijo",
"Latn",
type = "reconstructed",
}
m["inc-apa"] = {
"Apabhramsa",
616419,
"inc-mid",
"Deva, Shrd, Sidd",
ancestors = "pra",
translit = {
Deva = "sa-translit",
-- translit Shrd dalam [[Module:scripts/data]]
-- translit Sidd dalam [[Module:scripts/data]]
},
}
m["inc-ash"] = {
"Prakrit Ashoka",
104854379,
"inc-mid",
"Brah, Khar",
ancestors = "sa",
translit = {
-- translit Brah dalam [[Module:scripts/data]]
Khar = "Khar-translit",
},
}
m["inc-dng-pro"] = {
"Dangari Purba",
nil,
"inc-dng",
"Latn",
type = "reconstructed",
}
m["inc-kam"] = {
"Prakrit Kamarupi",
6356097,
"inc-bas",
"Brah, Sidd",
-- translit Brah, Sidd dalam [[Module:scripts/data]]
}
m["inc-kho"] = {
"Kholosi",
24952008,
"inc-snd",
"Latn",
}
m["inc-khr"] = {
"Khortha",
13406670,
"inc-sad",
"Deva, Kthi",
translit = {
Deva = "bho-translit",
Kthi = "bho-Kthi-translit",
},
}
m["inc-krd-pro"] = {
"Kamta Purba",
128816843,
"inc-bas",
"Latn",
ancestors = "inc-kam",
type = "reconstructed",
}
m["inc-mas"] = {
"Assam Pertengahan",
128806836,
"inc-bas",
"as-Beng",
ancestors = "inc-oas",
translit = "inc-mas-translit",
}
m["inc-mbn"] = {
"Benggali Pertengahan",
113559927,
"inc-bas",
"Beng",
ancestors = "inc-obn",
translit = "inc-mbn-translit",
}
m["inc-mgu"] = {
"Gujarati Pertengahan",
24907429,
"inc-wes",
"Deva",
ancestors = "inc-ogu",
}
m["inc-mor"] = {
"Odia Pertengahan",
128810882,
"inc-eas",
"Orya",
ancestors = "inc-oor",
}
m["inc-oas"] = {
"Assam Awal",
85758237,
"inc-bas",
"as-Beng",
ancestors = "inc-kam",
translit = "inc-oas-translit",
}
m["inc-oaw"] = {
"Awadhi Kuno",
nil,
"inc-hie",
"Deva, Kthi, Aran",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Deva = "sa-translit",
Kthi = "sa-Kthi-translit",
Aran = "inc-ohi-translit",
},
}
m["inc-obn"] = {
"Benggali Kuno",
113559926,
"inc-bas",
"Beng",
}
m["inc-ogu"] = {
"Gujarati Kuno",
24907427,
"inc-wes",
"Deva",
translit = "sa-translit",
}
m["inc-ohi"] = {
"Hindi Kuno",
48767781,
"inc-hiw",
"Deva, Aran",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Deva = "sa-translit",
Aran = "inc-ohi-translit",
},
}
m["inc-oor"] = {
"Odia Kuno",
128807801,
"inc-eas",
"Orya",
}
m["inc-opa"] = {
"Punjabi Kuno",
115270971,
"inc-pan",
"Guru, Aran",
translit = {
Guru = "inc-opa-Guru-translit",
Aran = "pa-Aran-translit",
},
strip_diacritics = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun},
}
m["inc-pro"] = {
"Indo-Arya Purba",
23808344,
"inc",
"Latn",
type = "reconstructed",
}
m["inc-sar"] = {
"Sarazi",
85799728,
"him",
"Aran, Deva, Takr",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Aran = "ur-translit",
Deva = "hi-translit",
-- Takr = "Takr-translit",
},
}
m["ine-ana-pro"] = {
"Anatolia Purba",
7251833,
"ine-ana",
"Latn",
type = "reconstructed",
}
m["ine-bsl-pro"] = {
"Balto-Slavik Purba",
1703347,
"ine-bsl",
"Latn",
type = "reconstructed",
sort_key = {
from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", c.acute, c.macron, "ˀ"},
to = {"a", "e", "i", "o", "u"}
},
}
m["ine-kal"] = {
"Kalašma",
122770439,
"ine-ana",
"Xsux",
}
m["ine-pae"] = {
"Paeonia",
2705672,
"ine",
"Polyt",
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["ine-pro"] = {
"Indo-Eropah Purba",
37178,
"ine",
"Latn",
type = "reconstructed",
sort_key = {
from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", "ĺ", "ḿ", "ń", "ŕ", "ǵ", "ḱ", "ʰ", "ʷ", "₁", "₂", "₃", c.ringbelow, c.acute, c.macron},
to = {"a", "e", "i", "o", "u", "l", "m", "n", "r", "g'", "k'", "¯h", "¯w", "1", "2", "3"}
},
}
m["ine-toc-pro"] = {
"Tocharia Purba",
104841462,
"ine-toc",
"Latn",
type = "reconstructed",
}
m["xme-old"] = {
"Median Kuno",
36461,
"xme",
"Polyt, Latn",
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["xme-mid"] = {
"Median Pertengahan",
12836150,
"xme",
"Latn",
}
m["xme-ker"] = {
"Kerman",
129850,
"xme",
"Arab, Latn, Hebr",
ancestors = "xme-mid",
-- display_text, strip_diacritics, sort_key Hebr dalam [[Module:scripts/data]]
}
m["xme-taf"] = {
"Tafreshi",
nil,
"xme",
"Arab, Latn",
ancestors = "xme-mid",
}
m["xme-ttc-pro"] = {
"Tatik Purba",
122973870,
"xme-ttc",
"Latn",
ancestors = "xme-mid",
}
m["xme-kls"] = {
"Kalasuri",
nil,
"xme-ttc",
ancestors = "xme-ttc-nor",
}
m["xme-klt"] = {
"Kilit",
3612452,
"xme-ttc",
"Cyrl", -- dan Arab?
}
m["xme-ott"] = {
"Tati Kuno",
434697,
"xme-ttc",
"Arab, Latn",
}
m["ira-kms-pro"] = {
"Komisenia Purba",
116773777,
"ira-kms",
"Latn",
type = "reconstructed",
}
m["ira-mpr-pro"] = {
"Medo-Parthia Purba",
116773227,
"ira-mpr",
"Latn",
type = "reconstructed",
}
m["ira-pat-pro"] = {
"Pathan Purba",
116773255,
"ira-pat",
"Latn",
type = "reconstructed",
}
m["ira-pro"] = {
"Iran Purba",
4167865,
"ira",
"Latn",
type = "reconstructed",
}
m["ira-zgr-pro"] = {
"Zaza-Gorani Purba",
116775031,
"ira-zgr",
"Latn",
type = "reconstructed",
}
m["xsc-pro"] = {
"Scythia Purba",
116773273,
"xsc",
"Latn",
type = "reconstructed",
}
m["xsc-sar-pro"] = {
"Sarmatia Purba",
116773249,
"xsc-sar",
"Latn",
type = "reconstructed",
}
m["xsc-skw-pro"] = {
"Saka-Wakhi Purba",
116773267,
"xsc-skw",
"Latn",
type = "reconstructed",
}
m["xsc-sak-pro"] = {
"Saka Purba",
116773264,
"xsc-sak",
"Latn",
type = "reconstructed",
}
m["ira-sym-pro"] = {
"Shughni-Yazghulami-Munji Purba",
116773813,
"ira-sym",
"Latn",
type = "reconstructed",
}
m["ira-sgi-pro"] = {
"Sanglechi-Ishkashimi Purba",
116773808,
"ira-sgi",
"Latn",
type = "reconstructed",
}
m["ira-mny-pro"] = {
"Munji-Yidgha Purba",
116773792,
"ira-mny",
"Latn",
type = "reconstructed",
}
m["ira-shy-pro"] = {
"Shughni-Yazghulami Purba",
116773812,
"ira-shy",
"Latn",
type = "reconstructed",
}
m["ira-shr-pro"] = {
"Shughni-Roshani Purba",
116773811,
"ira-shr",
"Latn",
type = "reconstructed",
}
m["ira-sgc-pro"] = {
"Sogdia Purba",
116773276,
"ira-sgc",
"Latn",
type = "reconstructed",
}
m["ira-wnj"] = {
"Vanji",
3398419,
"ira-shy",
"Latn",
}
m["iro-ere"] = {
"Erie",
5388365,
"iro-nor",
"Latn",
}
m["iro-min"] = {
"Mingo",
128531,
"iro-nor",
"Latn",
ietf_subtag = "i-mingo", -- tag IETF yang diwarisi
}
m["iro-nor-pro"] = {
"Iroquois Utara Purba",
116773242,
"iro-nor",
"Latn",
type = "reconstructed",
}
m["iro-pro"] = {
"Iroquois Purba",
7251852,
"iro",
"Latn",
type = "reconstructed",
}
m["itc-pro"] = {
"Italik Purba",
17102720,
"itc",
"Latn",
type = "reconstructed",
}
m["itc-psa"] = {
"Pra-Samnit",
7239186,
"itc-sbl",
"Ital, Polyt, Latn",
-- translit Ital dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja)
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["jpx-hcj"] = {
"Hachijō",
5637049,
"jpx",
"Jpan",
ancestors = "ojp-eas",
translit = s["jpx-translit"],
display_text = s["jpx-displaytext"],
strip_diacritics = s["jpx-stripdiacritics"],
sort_key = s["jpx-sortkey"],
}
m["jpx-pro"] = {
"Jepunik Purba",
3924309,
"jpx",
"Latn",
type = "reconstructed",
}
m["jpx-ryu-pro"] = {
"Ryukyu Purba",
56349069,
"jpx-ryu",
"Latn",
type = "reconstructed",
}
m["kar-pro"] = {
"Karen Purba",
85794783,
"kar",
"Latn",
type = "reconstructed",
}
m["kca-eas"] = {
"Khanty Timur",
30304622,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
-- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka)
sort_key = { Cyrl = { from = {""}, to = {""} } },
}
m["kca-nor"] = {
"Khanty Utara",
30304527,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
-- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka)
sort_key = { Cyrl = { from = {""}, to = {""} } },
}
m["kca-pro"] = {
"Khanty Purba",
127505171,
"kca",
"Latn",
type = "reconstructed",
}
m["kca-sou"] = {
"Khanty Selatan",
30304618,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
}
m["khi-kho-pro"] = {
"Khoe Purba",
116773218,
"khi-kho",
"Latn",
type = "reconstructed",
}
m["khi-kun"] = {
"ǃKung",
32904,
"khi-kxa",
"Latn",
}
m["ko-ear"] = {
"Korea Moden Awal",
756014,
"qfa-kor",
"Kore",
ancestors = "okm",
translit = "okm-translit",
-- strip_diacritics Kore dalam [[Module:scripts/data]]
}
m["kro-pro"] = {
"Kru Purba",
116773778,
"kro",
"Latn",
type = "reconstructed",
}
m["ku-pro"] = {
"Kurdi Purba",
116773221,
"ku",
"Latn",
type = "reconstructed",
}
m["map-ata-pro"] = {
"Atayalik Purba",
116773151,
"map-ata",
"Latn",
type = "reconstructed",
}
m["map-bms"] = {
"Banyumasan",
33219,
"map",
"Latn, Java",
}
m["map-pro"] = {
"Austronesia Purba",
49230,
"map",
"Latn",
type = "reconstructed",
}
m["mis-hkl"] = {
"Hokkien Peranakan Kelantan",
108794818,
"qfa-mix",
ancestors = "nan-hbl, sou, mfa",
}
m["mis-idn"] = {
"Idiom Neutral",
35847,
"art",
"Latn",
type = "appendix-constructed",
}
m["mis-isa"] = {
"Isauria",
16956868,
nil,
-- "Xsux, Hluw, Latn",
}
m["mis-jie"] = {
"Jie",
124424186,
nil,
"Hani",
sort_key = "Hani-sortkey",
}
m["mis-jzh"] = {
"Jizhao",
45242758,
"qfa-bej",
"Latn",
}
m["mis-kas"] = {
"Kassite",
35612,
nil,
"Xsux",
}
m["mis-mmd"] = {
"Mimi Decorse",
6862206,
nil,
"Latn",
}
m["mis-mmn"] = {
"Mimi Nachtigal",
6862207,
nil,
"Latn",
}
m["mis-phi"] = {
"Filistin",
2230924,
nil,
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja)
}
m["mis-rou"] = {
"Rouran",
48816637,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-tdl"] = {
"Turdulia",
133176492,
}
m["mis-tdt"] = {
"Turdetania",
133176461,
}
m["mis-tnw"] = {
"Tangwang",
7683179,
"qfa-mix",
"Latn",
ancestors = "cmn, sce",
}
m["mis-tuh"] = {
"Tuyuhun",
48816625,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-tuo"] = {
"Tuoba",
48816629,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-wuh"] = {
"Wuhuan",
118976867,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-xbi"] = {
"Xianbei",
4448647,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-xnu"] = {
"Xiongnu",
10901674,
nil,
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mjg-mgl"] = {
"Mongghul",
53765528,
"mjg",
"Latn", -- juga Mong, Cyrl?
}
m["mjg-mgr"] = {
"Mangghuer",
56285392,
"mjg",
"Latn", -- juga Mong, Cyrl?
}
m["mkh-asl-pro"] = {
"Asli Purba",
55630680,
"mkh-asl",
"Latn",
type = "reconstructed",
}
m["mkh-ban-pro"] = {
"Bahnar Purba",
116773189,
"mkh-ban",
"Latn",
type = "reconstructed",
}
m["mkh-kat-pro"] = {
"Katuik Purba",
116773772,
"mkh-kat",
"Latn",
type = "reconstructed",
}
m["mkh-khm-pro"] = {
"Khmuik Purba",
116773774,
"mkh-khm",
"Latn",
type = "reconstructed",
}
m["mkh-kmr-pro"] = {
"Khmer Purba",
55630684,
"mkh-kmr",
"Latn",
type = "reconstructed",
}
m["mkh-mmn"] = {
"Mon Pertengahan",
121337926,
"mkh-mnc",
"Latn, Mymr", --dan juga Pallava
ancestors = "omx",
}
m["mkh-mnc-pro"] = {
"Monik Purba",
116773231,
"mkh-mnc",
"Latn",
type = "reconstructed",
}
m["mkh-mvi"] = {
"Vietnam Pertengahan",
9199,
"mkh-vie",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mkh-pal-pro"] = {
"Palaungik Purba",
104847372,
"mkh-pal",
"Latn",
type = "reconstructed",
}
m["mkh-pea-pro"] = {
"Pearik Purba",
116773804,
"mkh-pea",
"Latn",
type = "reconstructed",
}
m["mkh-pkn-pro"] = {
"Pakanik Purba",
116773803,
"mkh-pkn",
"Latn",
type = "reconstructed",
}
m["mkh-pro"] = { --Ini akan digabungkan ke dalam aav-pro 2015.
"Mon-Khmer Purba",
7251859,
"mkh",
"Latn",
type = "reconstructed",
}
m["mnw-tha"] = { -- Untuk dibuang.
"Mon Thailand",
nil,
"mkh-mnc",
"Mymr, Thai",
ancestors = "mkh-mmn",
sort_key = {
from = {"[%p]", "ျ", "ြ", "ွ", "ှ", "ၞ", "ၟ", "ၠ", "ၚ", "ဿ", "[็-๎]", "([เแโใไ])([ก-ฮ])ฺ?"},
to = {"", "္ယ", "္ရ", "္ဝ", "္ဟ", "္န", "္မ", "္လ", "င", "သ္သ", "", "%2%1"}
},
}
m["mkh-vie-pro"] = {
"Vietik Purba",
109432616,
"mkh-vie",
"Latn",
type = "reconstructed",
}
m["mns-cen"] = {
"Mansi Tengah",
128810384,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mns-nor"] = {
"Mansi Utara",
30304537,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mns-pro"] = {
"Mansi Purba",
128883093,
"mns",
"Latn",
type = "reconstructed",
}
m["mns-sou"] = {
"Mansi Selatan",
30304629,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mun-pro"] = {
"Munda Purba",
105102373,
"mun",
"Latn",
type = "reconstructed",
}
m["myn-chl"] = { -- peringkat selepas ''emy''
"Ch'olti'",
873995,
"myn",
"Latn",
}
m["myn-pro"] = {
"Maya Purba",
3321532,
"myn",
"Latn",
type = "reconstructed",
}
m["nai-ala"] = {
"Alazapa",
128810233,
nil,
"Latn",
}
m["nai-bay"] = {
"Bayogoula",
1563704,
nil,
"Latn",
}
m["nai-cal"] = {
"Calusa",
51782,
nil,
"Latn",
}
m["nai-chi"] = {
"Chiquimulilla",
25339627,
"nai-xin",
"Latn",
}
m["nai-chu-pro"] = {
"Chumash Purba",
116773736,
"nai-chu",
"Latn",
type = "reconstructed",
}
m["nai-cig"] = {
"Ciguayo",
20741700,
nil,
"Latn",
}
m["nai-ckn-pro"] = {
"Chinook Purba",
116773735,
"nai-ckn",
"Latn",
type = "reconstructed",
}
m["nai-guz"] = {
"Guazacapán",
19572028,
"nai-xin",
"Latn",
}
m["nai-hit"] = {
"Hitchiti",
1542882,
"nai-mus",
"Latn",
}
m["nai-ipa"] = {
"Ipai",
3027474,
"nai-yuc",
"Latn",
}
m["nai-jtp"] = {
"Jutiapa",
nil,
"nai-xin",
"Latn",
}
m["nai-jum"] = {
"Jumaytepeque",
25339626,
"nai-xin",
"Latn",
}
m["nai-kat"] = {
"Kathlamet",
6376639,
"nai-ckn",
"Latn",
}
m["nai-klp-pro"] = {
"Kalapuya Purba",
116773771,
"nai-klp",
"Latn",
type = "reconstructed",
}
m["nai-knm"] = {
"Konomihu",
3198734,
"nai-shs",
"Latn",
}
m["nai-kum"] = {
"Kumeyaay",
4910139,
"nai-yuc",
"Latn",
}
m["nai-mac"] = {
"Macoris",
21070851,
nil,
"Latn",
}
m["nai-mdu-pro"] = {
"Maidu Purba",
116773784,
"nai-mdu",
"Latn",
type = "reconstructed",
}
m["nai-miz-pro"] = {
"Mixe-Zoque Purba",
7251858,
"nai-miz",
"Latn",
type = "reconstructed",
}
m["nai-mus-pro"] = {
"Muskogi Purba",
116775368,
"nai-mus",
"Latn",
type = "reconstructed",
}
m["nai-nao"] = {
"Naolan",
6964594,
nil,
"Latn",
}
m["nai-nrs"] = {
"Shasta Sungai Baru",
7011254,
"nai-shs",
"Latn",
}
m["nai-okw"] = {
"Okwanuchu",
3350126,
"nai-shs",
"Latn",
}
m["nai-per"] = {
"Pericú",
3375369,
nil,
"Latn",
}
m["nai-pic"] = {
"Picuris",
7191257,
"nai-kta",
"Latn",
}
m["nai-plp-pro"] = {
"Penuti Penara Purba",
116773806,
"nai-plp",
"Latn",
type = "reconstructed",
}
m["nai-pom-pro"] = {
"Pomo Purba",
116773262,
"nai-pom",
"Latn",
type = "reconstructed",
}
m["nai-qng"] = {
"Quinigua",
36360,
nil,
"Latn",
}
m["nai-sca-pro"] = { -- PERHATIAN 'sio-pro' "Proto-Siouan" iaitu Proto-Sioux Barat
"Siouan-Catawba Purba",
116773275,
"nai-sca",
"Latn",
type = "reconstructed",
}
m["nai-sin"] = {
"Sinacantán",
24190249,
"nai-xin",
"Latn",
}
m["nai-sln"] = {
"Lenca Salvador",
3229434,
"nai-len",
"Latn",
}
m["nai-spt"] = {
"Sahaptin",
3833015,
"nai-shp",
"Latn",
}
m["nai-tap"] = {
"Tapachultec",
7684401,
"nai-miz",
"Latn",
}
m["nai-taw"] = {
"Tawasa",
7689233,
nil,
"Latn",
}
m["nai-teq"] = {
"Tequistlatec",
2964454,
"nai-tqn",
"Latn",
}
m["nai-tip"] = {
"Tipai",
3027471,
"nai-yuc",
"Latn",
}
m["nai-tot-pro"] = {
"Totozoquean Purba",
116773285,
"nai-tot",
"Latn",
type = "reconstructed",
}
m["nai-tsi-pro"] = {
"Tsimshianik Purba",
nil,
"nai-tsi",
"Latn",
type = "reconstructed",
}
m["nai-utn-pro"] = {
"Utik Purba",
116773290,
"nai-utn",
"Latn",
type = "reconstructed",
}
m["nai-wai"] = {
"Waikuri",
3118702,
nil,
"Latn",
}
m["nai-wji"] = {
"Jicaque Barat",
3178610,
"nai-jcq",
"Latn",
}
m["nai-yup"] = {
"Yupiltepeque",
25339628,
"nai-xin",
"Latn",
}
m["nan-dat"] = {
"Min Datian",
19855572,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-hbl"] = {
"Hokkien",
1624231,
"zhx-nan",
"Hants, Latn, Bopo, Kana",
wikimedia_codes = "zh-min-nan",
generate_forms = "zh-generateforms",
sort_key = {
Hani = "Hani-sortkey",
Kana = "Kana-sortkey"
},
}
m["nan-hlh"] = {
"Min Hailufeng",
120755728,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-lnx"] = {
"Min Longyan",
6674568,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-tws"] = {
"Teochew",
36759,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["nan-zhe"] = {
"Min Zhenan",
3846710,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-zsh"] = {
"Min Sanxiang",
7420769,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["ngf-bin-pro"] = {
"Binandere Purba",
137881672,
"ngf-bin",
"Latn",
type = "reconstructed",
}
m["ngf-pro"] = {
"Trans-New Guinea Purba",
85794785,
"ngf",
"Latn",
type = "reconstructed",
}
m["nic-bco-pro"] = {
"Benue-Congo Purba",
116773194,
"nic-bco",
"Latn",
type = "reconstructed",
}
m["nic-bod-pro"] = {
"Bantoid Purba",
116773190,
"nic-bod",
"Latn",
type = "reconstructed",
}
m["nic-eov-pro"] = {
"Oti-Volta Timur Purba",
116773753,
"nic-eov",
"Latn",
type = "reconstructed",
}
m["nic-gns-pro"] = {
"Gurunsi Purba",
116773759,
"nic-gns",
"Latn",
type = "reconstructed",
}
m["nic-grf-pro"] = {
"Grassfields Purba",
116773755,
"nic-grf",
"Latn",
type = "reconstructed",
}
m["nic-gur-pro"] = {
"Gur Purba",
116773758,
"nic-gur",
"Latn",
type = "reconstructed",
}
m["nic-jkn-pro"] = {
"Jukunoid Purba",
116773769,
"nic-jkn",
"Latn",
type = "reconstructed",
}
m["nic-lcr-pro"] = {
"Cross River Hilir Purba",
116773782,
"nic-lcr",
"Latn",
type = "reconstructed",
}
m["nic-ogo-pro"] = {
"Ogoni Purba",
116773799,
"nic-ogo",
"Latn",
type = "reconstructed",
}
m["nic-ovo-pro"] = {
"Oti-Volta Purba",
116773802,
"nic-ovo",
"Latn",
type = "reconstructed",
}
m["nic-plt-pro"] = {
"Plateau Purba",
116773805,
"nic-plt",
"Latn",
type = "reconstructed",
}
m["nic-pro"] = {
"Niger-Congo Purba",
108000748,
"nic",
"Latn",
type = "reconstructed",
}
m["nic-ubg-pro"] = {
"Ubangi Purba",
116773818,
"nic-ubg",
"Latn",
type = "reconstructed",
}
m["nic-ucr-pro"] = {
"Cross River Hulu Purba",
116773819,
"nic-ucr",
"Latn",
type = "reconstructed",
}
m["nic-vco-pro"] = {
"Volta-Congo Purba",
116773293,
"nic-vco",
"Latn",
type = "reconstructed",
}
m["njo-jgl"] = {
"Ao Chungli",
55607615,
"njo",
"Latn",
}
m["njo-mng"] = {
"Ao Mongsen",
85383221,
"njo",
"Latn",
}
m["nub-har"] = {
"Haraza",
19572059,
"nub",
"Arab, Latn",
}
m["nub-pro"] = {
"Nubia Purba",
116773246,
"nub",
"Latn",
type = "reconstructed",
}
m["omq-cha-pro"] = {
"Chatino Purba",
116773202,
"omq-cha",
"Latn",
type = "reconstructed",
}
m["omq-maz-pro"] = {
"Mazatec Purba",
116773790,
"omq-maz",
"Latn",
type = "reconstructed",
}
m["omq-mix-pro"] = {
"Mixtecan Purba",
21573423,
"omq-mix",
"Latn",
type = "reconstructed",
}
m["omq-mxt-pro"] = {
"Mixtec Purba",
21573424,
"omq-mxt",
"Latn",
type = "reconstructed",
}
m["omq-otp-pro"] = {
"Oto-Pamean Purba",
116773251,
"omq-otp",
"Latn",
type = "reconstructed",
}
m["omq-pro"] = {
"Oto-Manguean Purba",
33669,
"omq",
"Latn",
type = "reconstructed",
}
m["omq-sjq"] = {
"Chatino San Juan Quiahije",
138330751,
"omq-cha",
"Latn",
}
m["omq-tel"] = {
"Mixtec Teposcolula",
nil,
"omq-mxt",
"Latn",
}
m["omq-teo"] = {
"Chatino Teojomulco",
25340451,
"omq-cha",
"Latn",
}
m["omq-tri-pro"] = {
"Triqui Purba",
116773817,
"omq-tri",
"Latn",
type = "reconstructed",
}
m["omq-zap-pro"] = {
"Zapotecan Purba",
116773297,
"omq-zap",
"Latn",
type = "reconstructed",
}
m["omq-zpc-pro"] = {
"Zapotec Purba",
116773296,
"omq-zpc",
"Latn",
type = "reconstructed",
}
m["omv-aro-pro"] = {
"Aroid Purba",
116773721,
"omv-aro",
"Latn",
type = "reconstructed",
}
m["omv-diz-pro"] = {
"Dizoid Purba",
116773750,
"omv-diz",
"Latn",
type = "reconstructed",
}
m["omv-pro"] = {
"Omotik Purba",
116773800,
"omv",
"Latn",
type = "reconstructed",
}
m["oto-otm-pro"] = {
"Otomi Purba",
5908710,
"oto-otm",
"Latn",
type = "reconstructed",
}
m["oto-pro"] = {
"Otomian Purba",
116773252,
"oto",
"Latn",
type = "reconstructed",
}
m["paa-kmn"] = {
"Kómnzo",
18344310,
"paa-wko",
"Latn",
}
m["paa-kwn"] = {
"Kuwani",
6449056,
"qfa-unc", -- kurang dibuktikan, mungkin sama dengan atau berkaitan dengan Kalabra
"Latn",
}
m["paa-lei"] = {
"Leitre",
85776228,
"paa-isk",
}
m["paa-nha-pro"] = {
"Halmahera Utara Purba",
116773241,
"paa-nha",
"Latn",
type = "reconstructed"
}
m["paa-nun"] = {
"Nungon",
128807788,
"ngf-ynu",
"Latn",
}
m["phi-din"] = {
"Agta Dinapigue",
16945774,
"phi",
"Latn",
}
m["phi-kal-pro"] = {
"Kalamian Purba",
116773213,
"phi-kal",
"Latn",
type = "reconstructed",
}
m["phi-nag"] = {
"Agta Nagtipunan",
16966111,
"phi",
"Latn",
}
m["phi-pro"] = {
"Filipina Purba",
18204898,
"phi",
"Latn",
type = "reconstructed",
}
m["poz-abi"] = {
"Abai",
19570729,
"poz-san",
"Latn",
}
m["poz-bal"] = {
"Baliledo",
4850912,
"poz",
"Latn",
}
m["poz-btk-pro"] = {
"Bungku-Tolaki Purba",
116773724,
"poz-btk",
"Latn",
type = "reconstructed",
}
m["poz-cet-pro"] = {
"Melayu-Polinesia Tengah-Timur Purba",
2269883,
"poz-cet",
"Latn",
type = "reconstructed",
}
m["poz-hce-pro"] = {
"Halmahera-Cenderawasih Purba",
116773209,
"poz-hce",
"Latn",
type = "reconstructed",
}
m["poz-lgx-pro"] = {
"Lampung Purba",
116773222,
"poz-lgx",
"Latn",
type = "reconstructed",
}
m["poz-mcm-pro"] = {
"Melayu-Chamik Purba",
116773225,
"poz-mcm",
"Latn",
type = "reconstructed",
}
m["poz-mic-pro"] = {
"Mikronesia Purba",
111939079,
"poz-mic",
"Latn",
type = "reconstructed",
}
m["poz-mly-pro"] = {
"Melayik Purba",
98057728,
"poz-mly",
"Latn",
type = "reconstructed",
}
m["poz-msa-pro"] = {
"Melayu-Sumbawa Purba",
116773226,
"poz-msa",
"Latn",
type = "reconstructed",
}
m["poz-nes"] = {
"Nese",
2157412,
"poz-vnc",
"Latn",
}
m["poz-oce-pro"] = {
"Oceania Purba",
141741,
"poz-oce",
"Latn",
type = "reconstructed",
}
m["poz-pcc-pro"] = {
"Pasifik Tengah Purba",
111962726,
"poz-pcc",
"Latn",
type = "reconstructed",
}
m["poz-pep-pro"] = {
"Polinesia Timur Purba",
113988745,
"poz-pep",
"Latn",
type = "reconstructed",
}
m["poz-pnp-pro"] = {
"Polinesia Teras Purba",
113988746,
"poz-pnp",
"Latn",
type = "reconstructed",
}
m["poz-pol-pro"] = {
"Polinesia Purba",
1658709,
"poz-pol",
"Latn",
type = "reconstructed",
}
m["poz-pro"] = {
"Melayu-Polinesia Purba",
3832960,
"poz",
"Latn",
type = "reconstructed",
}
m["poz-sml"] = {
"Melayu Sarawak",
4251702,
"poz-mly",
"Latn, Arab",
}
m["poz-ssw-pro"] = {
"Sulawesi Selatan Purba",
116773279,
"poz-ssw",
"Latn",
type = "reconstructed",
}
m["poz-swa-pro"] = {
"Sarawak Utara Purba",
116773243,
"poz-swa",
"Latn",
type = "reconstructed",
}
m["poz-ter"] = {
"Melayu Terengganu",
4207412,
"poz-mly",
"Latn, Arab",
}
m["pqe-pro"] = {
"Melayu-Polinesia Timur Purba",
2269883,
"pqe",
"Latn",
type = "reconstructed",
}
m["pra-niy"] = {
"Prakrit Niya",
11991601,
"inc-mid",
"Khar",
ancestors = "inc-ash",
translit = "Khar-translit",
}
m["qfa-adm-pro"] = {
"Andaman Besar Purba",
116773756,
"qfa-adm",
"Latn",
type = "reconstructed",
}
m["qfa-bet-pro"] = {
"Be-Tai Purba",
116773193,
"qfa-bet",
"Latn",
type = "reconstructed",
}
m["qfa-cka-pro"] = {
"Chukotko-Kamchatka Purba",
7251837,
"qfa-cka",
"Latn",
type = "reconstructed",
}
m["qfa-hur-pro"] = {
"Hurro-Urartia Purba",
116773211,
"qfa-hur",
"Latn",
type = "reconstructed",
}
m["qfa-kad-pro"] = {
"Kadu Purba",
116773770,
"qfa-kad",
"Latn",
type = "reconstructed",
}
m["qfa-kms-pro"] = {
"Kam-Sui Purba",
55630682,
"qfa-kms",
"Latn",
type = "reconstructed",
}
m["qfa-kor-pro"] = {
"Korea Purba",
467883,
"qfa-kor",
"Latn",
type = "reconstructed",
}
m["qfa-kra-pro"] = {
"Kra Purba",
7251854,
"qfa-kra",
"Latn",
type = "reconstructed",
}
m["qfa-lic-pro"] = {
"Hlai Purba",
7251845,
"qfa-lic",
"Latn",
type = "reconstructed",
}
m["qfa-onb-pro"] = {
"Be Purba",
116773192,
"qfa-onb",
"Latn",
type = "reconstructed",
}
m["qfa-ong-pro"] = {
"Onga Purba",
116773801,
"qfa-ong",
"Latn",
type = "reconstructed",
}
m["qfa-tak-pro"] = {
"Kra-Dai Purba",
104901616,
"qfa-tak",
"Latn",
type = "reconstructed",
}
m["qfa-yen-pro"] = {
"Yenisei Purba",
27639,
"qfa-yen",
"Latn",
type = "reconstructed",
}
m["qfa-yuk-pro"] = {
"Yukaghir Purba",
116773294,
"qfa-yuk",
"Latn",
type = "reconstructed",
}
m["qwe-kch"] = {
"Kichwa",
1740805,
"qwe",
"Latn",
ancestors = "qu",
}
m["qwe-pro"] = {
"Quechua Purba",
5575757,
"qwe",
"Latn",
type = "reconstructed",
}
m["roa-ang"] = {
"Angevin",
56782,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-bbn"] = {
"Bourbonnais-Berrichon",
2899128,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-brg"] = {
"Bourguignon",
508332,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-can"] = {
"Cantabrian",
917021,
"roa-asl",
"Latn",
}
m["roa-cha"] = {
"Champenois",
430018,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-fcm"] = {
"Franc-Comtois",
510561,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-gal"] = {
"Gallo",
37300,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-gib"] = {
"Gallo-Italik Basilicata",
3094838,
"roa-git",
ancestors = "pms-old",
"Latn",
}
m["roa-gis"] = {
"Gallo-Italik Sicily",
2629019,
"roa-git",
"Latn",
ancestors = "pms-old",
}
m["roa-leo"] = {
"Leon",
34108,
"roa-asl",
"Latn",
}
m["roa-lor"] = {
"Lorrain",
671198,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-oca"] = {
"Catalonia Kuno",
15478520,
"roa-ocr",
"Latn",
sort_key = {remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla .. "·"},
}
m["roa-ole"] = {
"Leon Kuno",
125977465,
"roa-asl",
"Latn",
}
m["roa-ona"] = {
"Navarro-Aragon Kuno",
2736184,
"roa-nar",
"Latn",
}
m["roa-opt"] = {
"Galicia-Portugis Kuno",
1072111,
"roa-gap",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ},
}
m["roa-orl"] = {
"Orléanais",
28497058,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-poi"] = {
"Poitevin-Saintongeais",
514123,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-tar"] = {
"Tarantino",
695526,
"roa-itr",
"Latn",
wikimedia_codes = "roa-tara",
}
m["sai-all"] = {
"Allentiac",
19570789,
"sai-hrp",
"Latn",
}
m["sai-and"] = {
"Andoquero",
16828359,
"sai-wit",
"Latn",
}
m["sai-ayo"] = {
"Ayomán",
16937754,
"sai-jir",
"Latn",
}
m["sai-bae"] = {
"Baenan",
3401998,
"qfa-unc", -- pupus, kurang dibuktikan; hanya dikenali melalui 9 perkataan
"Latn",
}
m["sai-bag"] = {
"Bagua",
5390321,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin bahasa Carib
"Latn",
}
m["sai-bet"] = {
"Betoi",
926551,
"qfa-iso",
"Latn",
}
m["sai-bor-pro"] = {
"Bora Purba",
nil,
"sai-bor",
"Latn",
}
m["sai-cac"] = {
"Cacán",
945482,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-caq"] = {
"Caranqui",
2937753,
"sai-bar",
"Latn",
}
m["sai-car-pro"] = {
"Carib Purba",
116773196,
"sai-car",
"Latn",
type = "reconstructed",
}
m["sai-cat"] = {
"Catacao",
5051136,
"sai-ctc",
"Latn",
}
m["sai-cer-pro"] = {
"Cerrado Purba",
116773200,
"sai-cer",
"Latn",
type = "reconstructed",
}
m["sai-chi"] = {
"Chirino",
5390321,
"qfa-unc", -- pupus, hanya empat perkataan diketahui; mungkin berkaitan dengan Candoshi-Shapra (cbu)
"Latn",
}
m["sai-chn"] = {
"Chaná",
5072718,
"sai-crn",
"Latn",
}
m["sai-chp"] = {
"Chapacura",
5072884,
"sai-cpc",
"Latn",
}
m["sai-chr"] = {
"Charrua",
5086680,
"sai-crn",
"Latn",
}
m["sai-chu"] = {
"Churuya",
5118339,
"sai-guh",
"Latn",
}
m["sai-cje-pro"] = {
"Jê Tengah Purba",
116773198,
"sai-cje",
"Latn",
type = "reconstructed",
}
m["sai-cmg"] = {
"Comechingon",
6644203,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-cno"] = {
"Chono",
5104704,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan, mungkin palsu
"Latn",
}
m["sai-cnr"] = {
"Cañari",
5055572,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Chimuan atau Barbacoan
"Latn",
}
m["sai-coe"] = {
"Coeruna",
6425639,
"sai-wit",
"Latn",
}
m["sai-col"] = {
"Colán",
5141893,
"sai-ctc",
"Latn",
}
m["sai-cop"] = {
"Copallén",
5390321,
"qfa-unc", -- pupus, hanya empat perkataan dibuktikan; mungkin Cholonan
"Latn",
}
m["sai-crd"] = {
"Coroado Puri",
24191321,
"sai-mje",
"Latn",
}
m["sai-ctq"] = {
"Catuquinaru",
16858455,
"qfa-unc", -- pupus, kurang dibuktikan; kosa kata tidak menyerupai bahasa lain
"Latn",
}
m["sai-cul"] = {
"Culli",
2879660,
"qfa-unc", -- pupus, kurang dibuktikan; sering dianggap sebagai pencilan
"Latn",
}
m["sai-cva"] = {
"Cueva",
5192644,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Chocoan
"Latn",
}
m["sai-esm"] = {
"Esmeralda",
3058083,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin berkaitan dengan Yaruro
"Latn",
}
m["sai-ewa"] = {
"Ewarhuyana",
16898104,
nil,
"Latn",
}
m["sai-gam"] = {
"Gamela",
5403661,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin pencilan
"Latn",
}
m["sai-gay"] = {
"Gayón",
5528902,
"sai-jir",
"Latn",
}
m["sai-gmo"] = {
"Guamo",
5613495,
"qfa-unc", -- pupus; "Kaufman (1990) mendapati hubungan dengan bahasa-bahasa Chapacuran meyakinkan." [Wikipedia] Dianggap sebagai pencilan oleh Campbell (2024).
"Latn",
}
m["sai-gua"] = {
"Guachí",
5613172,
"sai-guc",
"Latn",
}
m["sai-gue"] = {
"Güenoa",
5626799,
"sai-crn",
"Latn",
}
m["sai-hau"] = {
"Haush",
3128376,
"sai-cho",
"Latn",
}
m["sai-jee-pro"] = {
"Jê Purba",
116773212,
"sai-jee",
"Latn",
type = "reconstructed",
}
m["sai-jko"] = {
"Jeikó",
6176527,
"sai-mje",
"Latn",
}
m["sai-jrj"] = {
"Jirajara",
6202966,
"sai-jir",
"Latn",
}
m["sai-kat"] = { -- kontras xoo, kzw, sai-xoc
"Katembri",
6375925,
"qfa-unc", -- pupus, kurang dibuktikan; "Kaufman (1990) telah menghubungkannya dengan bahasa Taruma yang hampir pupus, walaupun ini tidak diterima oleh sarjana lain." [Wikipedia]
"Latn",
}
m["sai-mal"] = {
"Malalí",
6741212,
"sai-mje", -- dianggap sebagai bahasa Maxakalían yang paling divergen (subbahagian kepada Macro-Jê), yang mana kami tiada entri
"Latn",
}
m["sai-mar"] = {
"Maratino",
6755055,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Uto-Aztecan
"Latn",
}
m["sai-mat"] = {
"Matanawi",
6786047,
"qfa-unc", -- pupus; sama ada pencilan atau berkait jauh dengan bahasa-bahasa Muran; Campbell (2024) menyenarainya sebagai pencilan, Glottolog memberikannya sebagai tidak terkelas
"Latn",
}
m["sai-mcn"] = {
"Mocana",
3402048,
"qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad)
"Latn",
}
m["sai-men"] = {
"Menien",
16890110,
"sai-mje",
"Latn",
}
m["sai-mil"] = {
"Millcayac",
19573012,
"sai-hrp",
"Latn",
}
m["sai-mlb"] = {
"Malibu",
134374036,
"qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad)
"Latn",
}
m["sai-msk"] = {
"Masakará",
6782426,
"sai-mje",
"Latn",
}
m["sai-muc"] = {
"Mucuchí",
6931290,
nil, -- lazimnya dianggap sebagai Timotean, yang mana kami tiada entri
"Latn",
}
m["sai-mue"] = {
"Muellama",
16886936,
"sai-bar",
"Latn",
}
m["sai-muz"] = {
"Muzo",
6644203,
"qfa-unc", -- bahasa pupus di Colombia, kurang dibuktikan; mungkin Pijao (Cariban)
"Latn",
}
m["sai-mys"] = {
"Maynas",
16919393,
"sai-cah", -- mengikut Campbell (2024); dahulu dianggap tidak terkelas
"Latn",
}
m["sai-nat"] = {
"Natú",
9006749,
"qfa-unc", -- pupus, kurang dibuktikan; "hanya Greenberg yang berani mengelaskannya".[Wikipedia, memetik Moseley, Christopher; Asher, R. E.; Tait, Mary (1994), Atlas of the world's languages]
"Latn",
}
m["sai-nje-pro"] = {
"Jê Utara Purba",
116773245,
"sai-nje",
"Latn",
type = "reconstructed",
}
m["sai-opo"] = {
"Opón",
7099152,
"sai-car",
"Latn",
}
m["sai-oto"] = {
"Otomaco",
16879234,
"sai-otm",
"Latn",
}
m["sai-pal"] = {
"Palta",
3042978,
"qfa-unc", -- pupus, tidak terkelas; mungkin Chicham
"Latn",
}
m["sai-pam"] = {
"Pamigua",
5908689,
"sai-tin",
"Latn",
}
m["sai-par"] = {
"Paratió",
16890038,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Xukuruan
"Latn",
}
m["sai-peb"] = {
"Peba",
3373890,
"sai-pey",
"Latn",
}
m["sai-pnz"] = {
"Panzaleo",
3123275,
"qfa-unc", -- pupus, tidak terkelas; mungkin Paezan
"Latn",
}
m["sai-prh"] = {
"Puruhá",
3410994,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin dalam keluarga dengan Cañari
"Latn",
}
m["sai-ptg"] = {
"Patagón",
128807870,
"sai-tar", -- pupus, hanya diketahui daripada 4 perkataan, yang mencadangkan susur galur Cariban (Campbell 2024)
"Latn",
}
m["sai-pur"] = {
"Purukotó",
7261622,
"sai-pem",
"Latn",
}
m["sai-pyg"] = {
"Payaguá",
7156643,
"sai-guc",
"Latn",
}
m["sai-pyk"] = {
"Pykobjê",
98113977,
"sai-nje",
"Latn",
}
m["sai-qmb"] = {
"Quimbaya",
7272043,
"qfa-unc", -- pupus, mungkin tidak wujud; sedikit perkataan yang diketahui
"Latn",
}
m["sai-qtm"] = {
"Quitemo",
7272651,
"sai-cpc",
"Latn",
}
m["sai-rab"] = {
"Rabona",
6644203,
"qfa-unc", -- pupus, kurang dibuktikan, kebanyakan nama tumbuhan; mungkin Candoshi-Shapra
"Latn",
}
m["sai-ram"] = {
"Ramanos",
16902824,
"qfa-unc", -- pupus, kurang dibuktikan, mungkin pencilan; mengikut Glottolog: "senarai perkataan yang kerdil ... tidak menunjukkan persamaan yang meyakinkan dengan bahasa sekeliling"
"Latn",
}
m["sai-sac"] = {
"Sácata",
5390321,
"qfa-unc", -- pupus, hanya 3 perkataan diketahui; mungkin Candoshí atau Arawak
"Latn",
}
m["sai-san"] = {
"Sanaviron",
16895999,
"qfa-unc", -- pupus, tidak terkelas; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-sap"] = {
"Sapará",
7420922,
"sai-car",
"Latn",
}
m["sai-sec"] = {
"Sechura",
7442912,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan
"Latn",
}
m["sai-sin"] = {
"Sinúfana",
7525275,
"qfa-unc", -- hampir pupus, kurang dibuktikan; mungkin Chocoan
"Latn",
}
m["sai-sje-pro"] = {
"Jê Selatan Purba",
116773814,
"sai-sje",
"Latn",
type = "reconstructed",
}
m["sai-tab"] = {
"Tabancale",
5390321,
"qfa-unc", -- pupus, hanya 5 perkataan diketahui; tiada kaitan yang jelas, mungkin pencilan
"Latn",
}
m["sai-tal"] = {
"Tallán",
16910468,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan
"Latn",
}
m["sai-tap"] = {
"Tapayuna",
30719984,
"sai-nje",
"Latn",
}
m["sai-tar-pro"] = {
"Taranoan Purba",
116773816,
"sai-tar",
"Latn",
type = "reconstructed",
}
m["sai-teu"] = {
"Teushen",
3519243,
"qfa-unc", -- mungkin pupus menjelang 1950-an; mungkin Chonan
"Latn",
}
m["sai-tim"] = {
"Timote",
7806995,
nil, -- mungkin dalam keluarga Timote kecil
"Latn",
}
m["sai-tpr"] = {
"Taparita",
7684460,
"sai-otm",
"Latn",
}
m["sai-trr"] = {
"Tarairiú",
7685313,
"qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan
"Latn",
}
m["sai-wai"] = {
"Waitaká",
16918610,
"qfa-unc", -- pupus, mungkin Purian
"Latn",
}
m["sai-way"] = {
"Wayumara",
7960726,
"sai-car",
"Latn",
}
m["sai-wit-pro"] = {
"Witotoan Purba",
116773823,
"sai-wit",
"Latn",
type = "reconstructed",
}
m["sai-wnm"] = {
"Wanham",
16879440,
"sai-cpc",
"Latn",
}
m["sai-xoc"] = { -- kontras xoo, kzw, sai-kat
"Xocó",
12953620,
"qfa-unc", -- pupus dan kurang dibuktikan; tidak jelas sama ada satu atau tiga bahasa
"Latn",
}
m["sai-yao"] = {
"Yao (Amerika Selatan)",
16979655,
"sai-ven",
"Latn",
}
m["sai-yar"] = { -- bukan keluarga yang sama dengan 'suy'
"Yarumá",
3505859,
"sai-pek",
"Latn",
}
m["sai-yri"] = {
"Yuri",
2669157,
"sai-tyu",
"Latn",
}
m["sai-yup"] = {
"Yupua",
8061430,
"sai-tuc",
"Latn",
}
m["sai-yur"] = {
"Yurumanguí",
1281291,
"qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan
"Latn",
}
m["sal-pro"] = {
"Salish Purba",
116773269,
"sal",
"Latn",
type = "reconstructed",
}
m["sdv-daj-pro"] = {
"Daju Purba",
116773739,
"sdv-daj",
"Latn",
type = "reconstructed",
}
m["sdv-eje-pro"] = {
"Jebel Timur Purba",
116773751,
"sdv-eje",
"Latn",
type = "reconstructed",
}
m["sdv-nil-pro"] = {
"Nilotik Purba",
116773794,
"sdv-nil",
"Latn",
type = "reconstructed",
}
m["sdv-nyi-pro"] = {
"Nyima Purba",
116773796,
"sdv-nyi",
"Latn",
type = "reconstructed",
}
m["sdv-tmn-pro"] = {
"Taman Purba",
116773815,
"sdv-tmn",
"Latn",
type = "reconstructed",
}
m["sel-nor"] = {
"Selkup Utara",
30304565,
"sel",
"Cyrl",
translit = "sel-nor-translit",
}
m["sel-pro"] = {
"Selkup Purba",
128884235,
"sel",
"Latn",
type = "reconstructed",
}
m["sel-sou"] = {
"Selkup Selatan",
30304639,
"sel",
"Cyrl",
translit = "sel-sou-translit",
}
m["sem-amm"] = {
"Ammon",
279181,
"sem-can",
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]]
}
m["sem-amo"] = {
"Amor",
35941,
"sem-nwe",
"Xsux, Latn",
}
m["sem-cha"] = {
"Chaha",
35543,
"sem-eth",
"Ethi",
translit = "Ethi-translit",
}
m["sem-dad"] = {
"Dadan",
21838040,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-dum"] = {
"Dumait",
128810397,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-has"] = {
"Hasait",
3541433,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-his"] = {
"Hisma",
22948260,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-mhr"] = {
"Muher",
33743,
"sem-eth",
"Latn",
}
m["sem-pro"] = {
"Samiah Purba",
1658554,
"sem",
"Latn",
type = "reconstructed",
}
m["sem-saf"] = {
"Safait",
472586,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-sam"] = {
"Samal",
85847147,
"sem-nwe",
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]]
}
m["sem-srb"] = {
"Arab Selatan Kuno",
35025,
"sem-osa",
"Sarb",
-- translit Sarb dalam [[Module:scripts/data]]
}
m["sem-tay"] = {
"Tayman",
24912301,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-tha"] = {
"Thamud",
843030,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-wes-pro"] = {
"Samiah Barat Purba",
98021726,
"sem-wes",
"Latn",
type = "reconstructed",
}
m["sio-pro"] = { -- PERHATIAN ini bukan 'nai-sca-pro' "Proto-Siouan-Catawban" iaitu Proto-Sioux Barat
"Sioux Purba",
34181,
"sio",
"Latn",
type = "reconstructed",
}
m["sit-aao-pro"] = {
"Naga Tengah Purba",
nil,
"sit-aao",
"Latn",
type = "reconstructed",
}
m["sit-bai-pro"] = {
"Bai Purba",
nil,
"sit-bai",
"Latn",
type = "reconstructed",
}
m["sit-ban"] = {
"Bangru",
56071779,
"sit-hrs",
"Latn",
}
m["sit-bdi-pro"] = {
"Bodish Purba",
nil,
"sit-bdi",
"Latn",
type = "reconstructed",
}
m["sit-bok"] = {
"Bokar",
4938727,
"sit-tan",
"Latn, Tibt",
override_translit = true,
-- translit, display_text, strip_diacritics, sort_key Tibt dalam [[Module:scripts/data]]
}
m["sit-cai"] = {
"Caijia",
5017528,
"sit-cln",
"Latn"
}
m["sit-cha"] = {
"Chairel",
5068066,
"sit-luu",
"Latn",
}
m["sit-ers-pro"] = {
"Ersu Purba",
nil,
"sit-ers",
"Latn",
type = "reconstructed",
}
m["sit-hrs-pro"] = {
"Hrusish Purba",
116773762,
"sit-hrs",
"Latn",
type = "reconstructed",
}
m["sit-jap"] = {
"Japhug",
3162245,
"sit-egy",
"Latn",
}
m["sit-kha-pro"] = {
"Kham Purba",
116773773,
"sit-kha",
"Latn",
type = "reconstructed",
}
m["sit-khb-pro"] = {
"Kho-Bwa Purba",
nil,
"sit-khb",
"Latn",
type = "reconstructed",
}
m["sit-khp-pro"] = {
"Puroik Purba",
nil,
"sit-khb",
"Latn",
type = "reconstructed",
}
m["sit-khw-pro"] = {
"Kho-Bwa Barat Purba",
nil,
"sit-khw",
"Latn",
type = "reconstructed",
}
m["sit-kon-pro"] = {
"Naga Utara Purba",
nil,
"sit-kon",
"Latn",
type = "reconstructed",
}
m["sit-liz"] = {
"Lizu",
6660653,
"sit-ers",
"Latn", -- dan Ersu Shaba
}
m["sit-lnj"] = {
"Longjia",
17096251,
"sit-cln",
"Latn"
}
m["sit-lrn"] = {
"Luren",
16946370,
"sit-cln",
"Latn"
}
m["sit-luu-pro"] = {
"Luish Purba",
116773783,
"sit-luu",
"Latn",
type = "reconstructed",
}
m["sit-nas-pro"] = {
"Naish Purba",
nil,
"sit-nas",
"Latn",
type = "reconstructed",
}
m["sit-prn"] = {
"Puiron",
7259048,
"sit-zem",
}
m["sit-pro"] = {
"Sino-Tibet Purba",
24839178,
"sit",
"Latn",
type = "reconstructed",
}
m["sit-sit"] = {
"Situ",
19840830,
"sit-egy",
"Latn",
}
m["sit-tam-pro"] = {
"Tamang Purba",
117469295,
"sit-tam",
"Latn",
type = "reconstructed",
}
m["sit-tan-pro"] = {
"Tani Purba",
116773284,
"sit-tan",
"Latn", -- memerlukan pengesahan
type = "reconstructed",
}
m["sit-tgm"] = {
"Tangam",
17041370,
"sit-tan",
"Latn",
}
m["sit-tng-pro"] = {
"Tangkhul Purba",
nil,
"sit-tng",
"Latn",
type = "reconstructed",
}
m["sit-tos"] = {
"Tosu",
7827899,
"sit-ers",
"Latn", -- juga Ersu Shaba
}
m["sit-tsh"] = {
"Tshobdun",
19840950,
"sit-egy",
"Latn",
}
m["sit-zbu"] = {
"Zbu",
19841106,
"sit-egy",
"Latn",
}
m["sla-pro"] = {
"Slav Purba",
747537,
"sla",
"Latn",
type = "reconstructed",
strip_diacritics = {
remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve,
remove_exceptions = {'ś'},
},
sort_key = {
from = {"č", "ď", "ě", "ę", "ь", "ľ", "ň", "ǫ", "ř", "š", "ś", "ť", "ъ", "ž"},
to = {"c²", "d²", "e²", "e³", "i²", "l²", "nj", "o²", "r²", "s²", "s³", "t²", "u²", "z²"},
}
}
m["smi-pro"] = {
"Sami Purba",
7251862,
"smi",
"Latn",
type = "reconstructed",
sort_key = {
from = {"ā", "č", "δ", "[ëē]", "ŋ", "ń", "ō", "š", "θ", "%([^()]+%)"},
to = {"a", "c²", "d", "e", "n²", "n³", "o", "s²", "t²"}
},
}
m["son-pro"] = {
"Songhai Purba",
116773277,
"son",
"Latn",
type = "reconstructed",
}
m["sqj-pro"] = {
"Albania Purba",
18210846,
"sqj",
"Latn",
type = "reconstructed",
}
m["ssa-klk-pro"] = {
"Kuliak Purba",
116773779,
"ssa-klk",
"Latn",
type = "reconstructed",
}
m["ssa-kom-pro"] = {
"Koma Purba",
116773775,
"ssa-kom",
"Latn",
type = "reconstructed",
}
m["ssa-pro"] = {
"Nilo-Sahara Purba",
116773236,
"ssa",
"Latn",
type = "reconstructed",
}
m["syd-pro"] = {
"Samoyed Purba",
7251863,
"syd",
"Latn",
type = "reconstructed",
}
m["tai-pro"] = {
"Tai Purba",
6583709,
"tai",
"Latn",
type = "reconstructed",
}
m["tai-swe-pro"] = {
"Tai Barat Daya Purba",
116773280,
"tai-swe",
"Latn",
type = "reconstructed",
}
m["tbq-bdg-pro"] = {
"Bodo-Garo Purba",
116773195,
"tbq-bdg",
"Latn",
type = "reconstructed",
}
m["tbq-blg"] = {
"Bailang",
2879843,
"tbq-lob",
"Hani",
sort_key = "Hani-sortkey",
}
m["tbq-brm-pro"] = {
"Burma Purba",
nil,
"tbq-brm",
"Latn",
type = "reconstructed",
}
m["tbq-gkh"] = {
"Gokhy",
5578069,
"tbq-sil",
"Latn",
}
m["tbq-kuk-pro"] = {
"Kuki-Chin Purba",
116773220,
"tbq-kuk",
"Latn",
type = "reconstructed",
}
m["tbq-lal-pro"] = {
"Lalo Purba",
116773781,
"tbq-lal",
"Latn",
type = "reconstructed",
}
m["tbq-laz"] = {
"Laze",
17007626,
"sit-nas",
"Latn",
}
m["tbq-lob-pro"] = {
"Lolo-Burma Purba",
116773224,
"tbq-lob",
"Latn",
type = "reconstructed",
}
m["tbq-lol-pro"] = {
"Lolo Purba",
7251855,
"tbq-lol",
"Latn",
type = "reconstructed",
}
m["tbq-mil"] = {
"Milang",
6850761,
"sit-gsi",
"Deva, Latn",
}
m["tbq-mor"] = {
"Moran",
6909216,
"tbq-bdg",
"Latn",
}
m["tbq-ngo"] = {
"Ngochang",
56582,
"tbq-brm",
"Latn",
}
-- tbq-pro kini khusus etimologi
m["trk-dkh"] = {
"Dukhan",
12809273,
"trk-ssb",
"Latn, Cyrl, Mong",
-- translit, display_text dan strip_diacritics Mong dalam [[Module:scripts/data]]
}
-- Seperti yang diuraikan dalam ''Dīwān Lughāt al-Turk'' karya Mahmud al-Kashgari abad ke-11.
m["trk-eog"] = {
"Oghuz Kuno Awal",
nil,
"trk-ogz",
"Arab",
strip_diacritics = {Arab = "ar-stripdiacritics"},
}
m["trk-oat"] = {
"Turki Anatolia Kuno",
7083390,
"trk-ogz",
"Arab",
strip_diacritics = {Arab = "ar-stripdiacritics"},
ancestors = "trk-eog",
}
m["trk-pro"] = {
"Turkik Purba",
3657773,
"trk",
"Latn",
type = "reconstructed",
standard_chars = {
Latn = " ()-abdegiklmnoprstuxyzïöüāčēīĺŋōŕšūǖȫẹ" .. c.macron,
}
}
m["tup-gua-pro"] = {
"Tupi-Guarani Purba",
116773288,
"tup-gua",
"Latn",
type = "reconstructed",
}
m["tup-kab"] = {
"Kabishiana",
15302988,
"tup",
"Latn",
}
m["tup-kaw"] = {
"Kawahiva",
6346712,
"tup-gua",
"Latn",
}
m["tup-pro"] = {
"Tupi Purba",
10354700,
"tup",
"Latn",
type = "reconstructed",
}
m["tuw-alk"] = {
"Alchuka",
113553616,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-bal"] = {
"Bala",
86730632,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-kkl"] = {
"Kyakala",
118875708,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-kli"] = {
"Kili",
6406892,
"tuw-ewe",
"Cyrl",
}
m["tuw-pro"] = {
"Tungus Purba",
85872335,
"tuw",
"Latn",
type = "reconstructed",
}
m["tuw-sol"] = {
"Solon",
30004,
"tuw-ewe",
}
m["urj-fin-pro"] = {
"Finnik Purba",
11883720,
"urj-fin",
"Latn",
type = "reconstructed",
}
m["urj-koo"] = {
"Komi Kuno",
86679962,
"kv",
"Perm, Cyrs",
translit = "urj-koo-translit",
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]; sebelum ini, strip_diacritics Cyrs tidak hadir
}
m["urj-kuk"] = {
"Kukkuzi",
107410460,
"urj-fin",
"Latn",
ancestors = "vot",
}
m["urj-kya"] = {
"Komi-Yazva",
2365210,
"kv",
"Cyrl",
translit = "kv-translit",
override_translit = true,
strip_diacritics = {remove_diacritics = c.acute},
}
m["urj-mdv-pro"] = {
"Mordvinik Purba",
116773232,
"urj-mdv",
"Latn",
type = "reconstructed",
}
m["urj-prm-pro"] = {
"Permik Purba",
116773257,
"urj-prm",
"Latn",
type = "reconstructed",
}
m["urj-pro"] = {
"Uralik Purba",
288765,
"urj",
"Latn",
type = "reconstructed",
}
m["urj-ugr-pro"] = {
"Ugrik Purba",
156631,
"urj-ugr",
"Latn",
type = "reconstructed",
}
m["xnd-pro"] = {
"Na-Dene Purba",
116773233,
"xnd",
"Latn",
type = "reconstructed",
}
m["xgn-pro"] = {
"Mongol Purba",
2493677,
"xgn",
"Latn",
type = "reconstructed",
sort_key = {
from = {"č", "i", "ï", "ǰ", "ŋ", "ö", "š", "ü"},
to = {"c", "i" .. p[1], "i", "j", "n" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1]},
},
}
m["yok-bvy"] = {
"Yokuts Buena Vista",
4985474,
"yok",
"Latn",
}
m["yok-dly"] = {
"Yokuts Delta",
70923266,
"yok",
"Latn",
}
m["yok-gsy"] = {
"Yokuts Gashowu",
3098708,
"yok",
"Latn",
}
m["yok-kry"] = {
"Yokuts Sungai Kings",
6413014,
"yok",
"Latn",
}
m["yok-nvy"] = {
"Yokuts Lembah Utara",
85789777,
"yok",
"Latn",
}
m["yok-ply"] = {
"Yokuts Palewyami",
2387391,
"yok",
"Latn",
}
m["yok-svy"] = {
"Yokuts Lembah Selatan",
12642473,
"yok",
"Latn",
}
m["yok-tky"] = {
"Yokuts Tule-Kaweah",
7851988,
"yok",
"Latn",
}
m["ypk-pro"] = {
"Yupik Purba",
116773295,
"ypk",
"Latn",
type = "reconstructed",
}
m["yrk-for"] = {
"Nenets Hutan",
1295107,
"yrk",
"Cyrl",
translit = "yrk-for-translit",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.dotabove},
}
m["yrk-tun"] = {
"Nenets Tundra",
36452,
"yrk",
"Cyrl",
strip_diacritics = {
from = {"ӑ", "а̄", "э̇", "ӣ", "ы̄", "ӯ", "ю̄", "я̆", "я̄"},
to = {"а", "а", "э", "и", "ы", "у", "ю", "я", "я"},
},
translit = "yrk-tun-translit",
}
m["zhx-min-pro"] = {
"Min Purba",
19646347,
"zhx-min",
"Latn",
type = "reconstructed",
}
m["zhx-sht"] = {
"Tuhua Shaozhou",
1920769,
"zhx",
"Nshu, Hants",
generate_forms = "zh-generateforms",
sort_key = {Hani = "Hani-sortkey"},
}
m["zhx-sic"] = {
"Sichuan",
2278732,
"zhx-man",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["zhx-tai"] = {
"Taishan",
2208940,
"zhx-yue",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["zle-ono"] = {
"Novgorod Kuno",
162013,
"zle",
"Cyrs, Glag",
translit = {Cyrs = "Cyrs-translit", Glag = "Glag-translit"},
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zle-ort"] = {
"Ruthenia Kuno",
13211,
"zle",
"Arab, Cyrs, Latn",
ancestors = "orv",
translit = {
Cyrs = "zle-ort-translit",
Arab = "zle-ort-Arab-translit",
},
strip_diacritics = {
Cyrs = {
remove_diacritics = m_langdata.chars_substitutions["Cyrs_remove_diacritics"],
remove_exceptions = {"Ї", "ї"},
},
Arab = "ar-stripdiacritics",
},
-- sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zls-chs"] = {
"Slav Gereja",
33251,
"zls",
"Cyrs, Glag, Latn",
ancestors = "cu",
translit = {
Cyrs = "Cyrs-translit",
Glag = "Glag-translit"
},
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zlw-ocs"] = {
"Czech Kuno",
593096,
"zlw",
"Latn",
}
m["zlw-opl"] = {
"Poland Kuno",
149838,
"zlw-lch",
"Latn",
strip_diacritics = {remove_diacritics = c.ringabove},
}
m["zlw-osk"] = {
"Slovak Kuno",
12776676,
"zlw",
"Latn",
}
m["zlw-slv"] = {
"Slovincia",
36822,
"zlw-pom",
"Latn",
strip_diacritics = {remove_diacritics = c.macron .. c.breve},
}
-- Kod tambahan untuk bahasa-bahasa yang digunakan di Malaysia, yang tidak wujud di Wikikamus Bahasa Inggeris
m["zlm-coa"] = {
"Melayu Terengganu Pesisir",
4207412,
"poz-mly",
"Latn, ms-Arab",
}
m["zlm-pah"] = {
"Melayu Pahang",
7310370,
"poz-mly",
"Latn",
}
m["tmw"] = {
"Temuan",
3025610,
"poz-mly",
"Latn",
}
m["kzt"] = {
"Dusun Tambunan",
12953514,
"poz-san",
"Latn",
}
return require("Module:languages").finalizeData(m, "language")
8hpheldymxlci4iezyguza9t4ttu1az
375376
375374
2026-09-22T05:35:06Z
Hakimi97
2668
Move "Temuan" back to Module:languages/data/3/t, and move "Dusun Tambunan" back to Module:languages/data/3/k
375376
Scribunto
text/plain
local m_langdata = require("Module:languages/data")
-- Loaded on demand, as it may not be needed (depending on the data).
local function u(...)
u = require("Module:string utilities").char
return u(...)
end
local c = m_langdata.chars
local p = m_langdata.puaChars
local s = m_langdata.shared
local m = {}
m["aav-khs-pro"] = {
"Khasi Purba",
116773216,
"aav-khs",
"Latn",
type = "reconstructed",
}
m["aav-nic-pro"] = {
"Nicobar Purba",
116773793,
"aav-nic",
"Latn",
type = "reconstructed",
}
m["aav-pkl-pro"] = {
"Pnar-Khasi-Lyngngam Purba",
116773259,
"aav-pkl",
"Latn",
type = "reconstructed",
}
m["aav-pro"] = { -- mkh-pro akan digabungkan ke dalam ini
"Austroasia Purba",
116773186,
"aav",
"Latn",
type = "reconstructed",
}
m["afa-pro"] = {
"Afroasia Purba",
269125,
"afa",
"Latn",
type = "reconstructed",
}
m["alg-aga"] = {
"Agawam",
nil,
"alg-eas",
"Latn",
}
m["alg-pro"] = {
"Algonquian Purba",
7251834,
"alg",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = "·"},
}
m["alv-ama"] = {
"Amasi",
4740400,
"nic-grs",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.tilde .. c.macron},
}
m["alv-bgu"] = {
"Bainouk Gubeeher",
17002646,
"alv-bny",
"Latn",
}
m["alv-bua-pro"] = {
"Bua Purba",
116773723,
"alv-bua",
"Latn",
type = "reconstructed",
}
m["alv-cng-pro"] = {
"Cangin Purba",
116773726,
"alv-cng",
"Latn",
type = "reconstructed",
}
m["alv-edo-pro"] = {
"Edoid Purba",
116773206,
"alv-edo",
"Latn",
type = "reconstructed",
}
m["alv-fli-pro"] = {
"Fali Purba",
116773754,
"alv-fli",
"Latn",
type = "reconstructed",
}
m["alv-gbe-pro"] = {
"Gbe Purba",
116773208,
"alv-gbe",
"Latn",
type = "reconstructed",
}
m["alv-gng-pro"] = {
"Guang Purba",
116773757,
"alv-gng",
"Latn",
type = "reconstructed",
}
m["alv-gtm-pro"] = {
"Togo Tengah Purba",
116773732,
"alv-gtm",
"Latn",
type = "reconstructed",
}
m["alv-gwa"] = {
"Gwara",
16945580,
"nic-pla",
"Latn",
}
m["alv-hei-pro"] = {
"Heiban Purba",
116773760,
"alv-hei",
"Latn",
type = "reconstructed",
}
m["alv-ido-pro"] = {
"Idomoid Purba",
116773764,
"alv-ido",
"Latn",
type = "reconstructed",
}
m["alv-igb-pro"] = {
"Igboid Purba",
116773765,
"alv-igb",
"Latn",
type = "reconstructed",
}
m["alv-kwa-pro"] = {
"Kwa Purba",
116773780,
"alv-kwa",
"Latn",
type = "reconstructed",
}
m["alv-mum-pro"] = {
"Mumuye Purba",
116773791,
"alv-mum",
"Latn",
type = "reconstructed",
}
m["alv-nup-pro"] = {
"Nupoid Purba",
116773795,
"alv-nup",
"Latn",
type = "reconstructed",
}
m["alv-pro"] = {
"Atlantik-Congo Purba",
116732838,
"alv",
"Latn",
type = "reconstructed",
}
m["alv-edk-pro"] = {
"Edekiri Purba",
nil,
"alv-edk",
"Latn",
type = "reconstructed",
}
m["alv-yor-pro"] = {
"Yoruba Purba",
nil,
"alv-yor",
"Latn",
type = "reconstructed",
}
m["alv-yrd-pro"] = {
"Yoruboid Purba",
116773824,
"alv-yrd",
"Latn",
type = "reconstructed",
}
m["alv-von-pro"] = {
"Volta-Niger Purba",
116773820,
"alv-von",
"Latn",
type = "reconstructed",
}
m["apa-pro"] = {
"Apache Purba",
116773135,
"apa",
"Latn",
type = "reconstructed",
}
m["aql-pro"] = {
"Algik Purba",
18389588,
"aql",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = "·"},
}
m["art-adu"] = {
"Adûni",
1232159,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-bel"] = {
"Kreol Belter",
108055510,
"art",
"Latn",
type = "appendix-constructed",
sort_key = {
remove_diacritics = c.acute,
from = {"ɒ"},
to = {"a"},
},
}
m["art-blk"] = {
"Bolak",
2909283,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-bsp"] = {
"Bahasa Hitam",
686210,
"art",
"Latn, Teng",
type = "appendix-constructed",
}
m["art-com"] = {
"Communicationssprache",
35227,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-dtk"] = {
"Dothraki",
2914733,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-elo"] = {
"Eloi",
nil,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-gld"] = {
"Goa'uld",
19823,
"art",
"Latn, Egyp, Mero",
type = "appendix-constructed",
}
m["art-lap"] = {
"Lapine",
6488195,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-man"] = {
"Mandalorian",
54289,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-mun"] = {
"Mundolinco",
851355,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-nav"] = {
"Naʼvi",
316939,
"art",
"Latn",
type = "appendix-constructed",
}
m["art-vlh"] = {
"Valyria Tinggi",
64483808,
"art",
"Latn",
type = "appendix-constructed",
}
m["ath-nic"] = {
"Nicola",
20609,
"ath-nor",
"Latn",
}
m["ath-pro"] = {
"Athabaska Purba",
104841722,
"ath",
"Latn",
type = "reconstructed",
}
m["auf-pro"] = {
"Arawa Purba",
116773706,
"auf",
"Latn",
type = "reconstructed",
}
m["aus-alu"] = {
"Alungul",
16827670,
"aus-pmn",
"Latn",
}
m["aus-and"] = {
"Andjingith",
4754509,
"aus-pmn",
"Latn",
}
m["aus-ang"] = {
"Angkula",
16828520,
"aus-pmn",
"Latn",
}
m["aus-arn-pro"] = {
"Arnhem Purba",
116773720,
"aus-arn",
"Latn",
type = "reconstructed",
}
m["aus-bra"] = {
"Barranbinya",
4863220,
"aus-pmn",
"Latn",
}
m["aus-brm"] = {
"Barunggam",
4865914,
"aus-pmn",
"Latn",
}
m["aus-cww-pro"] = {
"New South Wales Tengah Purba",
116773199,
"aus-cww",
"Latn",
type = "reconstructed",
}
m["aus-dal-pro"] = {
"Daly Purba",
116773743,
"aus-dal",
"Latn",
type = "reconstructed",
}
m["aus-guw"] = {
"Guwar",
6652138,
"aus-pam",
"Latn",
}
m["aus-lsw"] = {
"Little Swanport",
6652138,
"qfa-unc",
"Latn",
}
m["aus-mbi"] = {
"Mbiywom",
6799701,
"aus-pmn",
"Latn",
}
m["aus-ngk"] = {
"Ngkoth",
7022405,
"aus-pmn",
"Latn",
}
m["aus-nyu-pro"] = {
"Nyulnyulan Purba",
116773797,
"aus-nyu",
"Latn",
type = "reconstructed",
}
m["aus-pam-pro"] = {
"Pama-Nyunga Purba",
33942,
"aus-pam",
"Latn",
type = "reconstructed",
}
m["aus-tul"] = {
"Tulua",
16938541,
"aus-pam",
"Latn",
}
m["aus-uwi"] = {
"Uwinymil",
7903995,
"aus-arn",
"Latn",
}
m["aus-wdj-pro"] = {
"Iwaidjan Purba",
116773767,
"aus-wdj",
"Latn",
type = "reconstructed",
}
m["aus-won"] = {
"Wong-gie",
nil,
"aus-pam",
"Latn",
}
m["aus-wul"] = {
"Wulguru",
8039196,
"aus-dyb",
"Latn",
}
m["aus-ynk"] = { -- kontras nny
"Yangkaal",
3913770,
"aus-tnk",
"Latn",
}
m["awd-amc-pro"] = {
"Amuesha-Chamicuro Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-kmp-pro"] = {
"Kampa Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-prw-pro"] = {
"Paresi-Waura Purba",
nil,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-ama"] = {
"Amarizana",
16827787,
"awd",
"Latn",
}
m["awd-ana"] = {
"Anauyá",
16828252,
"awd",
"Latn",
}
m["awd-apo"] = {
"Apolista",
16916645,
"awd",
"Latn",
}
m["awd-cab"] = {
"Cabre",
16850160,
"awd",
"Latn",
}
m["awd-gnu"] = {
"Guinau",
3504087,
"awd",
"Latn",
}
m["awd-kar"] = {
"Cariay",
16920253,
"awd",
"Latn",
}
m["awd-kaw"] = {
"Kawishana",
6379993,
"awd-nwk",
"Latn",
}
m["awd-kus"] = {
"Kustenau",
5196293,
"awd",
"Latn",
}
m["awd-man"] = {
"Manao",
6746920,
"awd",
"Latn",
}
m["awd-mar"] = {
"Marawan",
6755108,
"awd",
"Latn",
}
m["awd-mpr"] = {
"Maipure",
6736872,
"awd",
"Latn",
}
m["awd-mrt"] = {
"Mariaté",
16910017,
"awd-nwk",
"Latn",
}
m["awd-nwk-pro"] = {
"Nawiki Purba",
116773234,
"awd-nwk",
"Latn",
type = "reconstructed",
}
m["awd-pai"] = {
"Paikoneka",
128807835,
"awd",
"Latn",
}
m["awd-pas"] = {
"Pasé",
7143168,
"awd-nwk",
"Latn",
}
m["awd-pro"] = {
"Arawak Purba",
97573478,
"awd",
"Latn",
type = "reconstructed",
}
m["awd-she"] = {
"Shebayo",
7492248,
"awd",
"Latn",
}
m["awd-taa-pro"] = {
"Ta-Arawak Purba",
116773282,
"awd-taa",
"Latn",
type = "reconstructed",
}
m["awd-wai"] = {
"Wainumá",
16910017,
"awd-nwk",
"Latn",
}
m["awd-war"] = {
"Warekena Kuno",
105320180,
"awd-nwk",
"Latn",
}
m["awd-yum"] = {
"Yumana",
8061062,
"awd-nwk",
"Latn",
}
m["azc-caz"] = {
"Cazcan",
5055514,
"azc",
"Latn",
}
m["azc-cup-pro"] = {
"Cupan Purba",
116773738,
"azc-cup",
"Latn",
type = "reconstructed",
}
m["azc-ktn"] = {
"Kitanemuk",
3197558,
"azc-tak",
"Latn",
}
m["azc-nah-pro"] = {
"Nahua Purba",
7251860,
"azc-nah",
"Latn",
type = "reconstructed",
}
m["azc-nic"] = {
"Nicoleño",
50241488,
"azc",
"Latn",
}
m["azc-num-pro"] = {
"Numik Purba",
116773247,
"azc-num",
"Latn",
type = "reconstructed",
}
m["azc-pro"] = {
"Uto-Aztek Purba",
96400333,
"azc",
"Latn",
type = "reconstructed",
}
m["azc-tak-pro"] = {
"Takik Purba",
116773283,
"azc-tak",
"Latn",
type = "reconstructed",
}
m["azc-tat"] = {
"Tataviam",
743736,
"azc",
"Latn",
}
m["ber-pro"] = {
"Berber Purba",
2855698,
"ber",
"Latn",
type = "reconstructed",
}
m["ber-fog"] = {
"Fogaha",
107610173,
"ber",
"Latn",
}
m["ber-zuw"] = {
"Zuwara",
4117169,
"ber",
"Latn",
}
m["bnt-bal"] = {
"Balong",
93935237,
"bnt-bbo",
"Latn",
}
m["bnt-bon"] = {
"Boma Nkuu",
nil,
"bnt",
"Latn",
}
m["bnt-boy"] = {
"Boma Yumu",
nil,
"bnt",
"Latn",
}
m["bnt-bwa"] = {
"Bwala",
128810345,
"bnt-tek",
"Latn",
}
m["bnt-cmw"] = {
"Chimwiini",
4958328,
"bnt-swh",
"Latn",
}
m["bnt-ind"] = {
"Indanga",
51412803,
"bnt",
"Latn",
}
m["bnt-lal"] = {
"Lala (Afrika Selatan)",
6480154,
"bnt-ngu",
"Latn",
}
m["bnt-mpi"] = {
"Mpiin",
93937013,
"bnt-bdz",
"Latn",
}
m["bnt-mpu"] = {
"Mpuono", -- jangan dikelirukan dengan Mbuun zmp
36056,
"bnt",
"Latn",
}
m["bnt-ngu-pro"] = {
"Nguni Purba",
961559,
"bnt-ngu",
"Latn",
type = "reconstructed",
sort_key = {remove_diacritics = c.grave .. c.acute .. c.circ .. c.caron},
}
m["bnt-phu"] = {
"Phuthi",
33796,
"bnt-ngu",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute},
}
m["bnt-pro"] = {
"Bantu Purba",
3408025,
"bnt",
"Latn",
type = "reconstructed",
sort_key = "bnt-pro-sortkey",
}
m["bnt-sab-pro"] = {
"Sabaki Purba",
nil, -- Q2209395 ialah kod untuk keluarga Sabaki
"bnt-sab",
"Latn",
type = "reconstructed",
}
m["bnt-sbo"] = {
"Boma Selatan",
nil,
"bnt",
"Latn",
}
m["bnt-sts-pro"] = {
"Sotho-Tswana Purba",
116773278,
"bnt-sts",
"Latn",
type = "reconstructed",
}
m["btk-pro"] = {
"Batak Purba",
116773191,
"btk",
"Latn",
type = "reconstructed",
}
m["cau-abz-pro"] = {
"Abkhaz-Abaza Purba",
7251831,
"cau-abz",
"Latn",
type = "reconstructed",
}
m["cau-and-pro"] = {
"Andi Purba",
nil,
"cau-and",
"Latn",
type = "reconstructed",
}
m["cau-ava-pro"] = {
"Avar-Andi Purba",
116773187,
"cau-ava",
"Latn",
type = "reconstructed",
}
m["cau-cir-pro"] = {
"Circassia Purba",
7251838,
"cau-cir",
"Latn",
type = "reconstructed",
}
m["cau-drg-pro"] = {
"Dargwa Purba",
116773205,
"cau-drg",
"Latn",
type = "reconstructed",
}
m["cau-lzg-pro"] = {
"Lezghi Purba",
116773223,
"cau-lzg",
"Latn",
type = "reconstructed",
}
m["cau-nec-pro"] = {
"Kaukasia Timur Laut Purba",
116773244,
"cau-nec",
"Latn",
type = "reconstructed",
}
m["cau-nkh-pro"] = {
"Nakh Purba",
108032840,
"cau-nkh",
"Latn",
type = "reconstructed",
}
m["cau-nwc-pro"] = {
"Kaukasia Barat Laut Purba",
7251861,
"cau-nwc",
"Latn",
type = "reconstructed",
}
m["cau-tsz-pro"] = {
"Tsez Purba",
116773287,
"cau-tsz",
"Latn",
type = "reconstructed",
}
m["cba-ata"] = {
"Atanques",
4812783,
"cba",
"Latn",
}
m["cba-cat"] = {
"Catío Chibcha",
7083619,
"cba",
"Latn",
}
m["cba-dor"] = {
"Dorasque",
5297532,
"cba",
"Latn",
}
m["cba-dui"] = {
"Duit",
3041061,
"cba",
"Latn",
}
m["cba-hue"] = {
"Huetar",
35514,
"cba",
"Latn",
}
m["cba-nut"] = {
"Nutabe",
7070405,
"cba",
"Latn",
}
m["cba-pro"] = {
"Chibchan Purba",
116773203,
"cba",
"Latn",
type = "reconstructed",
}
m["ccs-pro"] = {
"Kartvelia Purba",
2608203,
"ccs",
"Latn",
type = "reconstructed",
strip_diacritics = {
from = {"q̣", "p̣", "ʓ", "ċ"},
to = {"q̇", "ṗ", "ʒ", "c̣"}
},
}
m["ccs-gzn-pro"] = {
"Georgia-Zan Purba",
23808119,
"ccs-gzn",
"Latn",
type = "reconstructed",
strip_diacritics = {
from = {"q̣", "p̣", "ʓ", "ċ"},
to = {"q̇", "ṗ", "ʒ", "c̣"}
},
}
m["cdc-cbm-pro"] = {
"Chadik Tengah Purba",
116773197,
"cdc-cbm",
"Latn",
type = "reconstructed",
}
m["cdc-mas-pro"] = {
"Masa Purba",
116773789,
"cdc-mas",
"Latn",
type = "reconstructed",
}
m["cdc-pro"] = {
"Chadik Purba",
116773201,
"cdc",
"Latn",
type = "reconstructed",
}
m["cdd-pro"] = {
"Caddoan Purba",
116773725,
"cdd",
"Latn",
type = "reconstructed",
}
m["cel-bry-pro"] = {
"Britonik Purba",
1248800,
"cel-bry",
"Latn, Polyt",
sort_key = {
Latn = "cel-bry-pro-sortkey",
},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["cel-gal"] = {
"Gallaecia",
3094789,
"cel-his",
}
m["cel-gau"] = {
"Gaul",
29977,
"cel",
"Latn, Polyt, Ital",
strip_diacritics = {
Latn = {remove_diacritics = c.macron .. c.breve .. c.diaer},
},
sort_key = {
Latn = "cel-bry-pro-sortkey",
},
-- translit Ital dalam [[Module:scripts/data]]
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["cel-pro"] = {
"Keltik Purba",
653649,
"cel",
"Latn",
type = "reconstructed",
sort_key = "cel-pro-sortkey",
}
m["chi-pro"] = {
"Chimakuan Purba",
116773734,
"chi",
"Latn",
type = "reconstructed",
}
m["chm-pro"] = {
"Mari Purba",
116773788,
"chm",
"Latn",
type = "reconstructed",
}
m["cmc-pro"] = {
"Chamik Purba",
114793834,
"cmc",
"Latn",
type = "reconstructed",
}
m["crp-bip"] = {
"Pijin Basque-Iceland",
810378,
"crp",
"Latn",
ancestors = "eu",
}
m["crp-cpr"] = {
"Pijin Rusia-China",
nil,
"crp",
"Hani, Cyrl, Latn",
ancestors = "ru, zh",
translit = {Cyrl = "ru-translit"},
strip_diacritics = {
Cyrl = {remove_diacritics = c.acute .. c.grave .. c.macron},
},
}
m["crp-gep"] = {
"Pijin Greenland Barat",
17036301,
"crp",
"Latn",
ancestors = "kl",
}
m["crp-kia"] = {
"Pijin Jerman Kiautschou",
108314615,
"crp",
"Latn",
ancestors = "de",
}
m["crp-mar"] = {
"Bahasa Roh Maroon",
1093206,
"crp",
"Latn",
ancestors = "en",
}
m["crp-mpp"] = {
"Pijin Portugis Macau",
128804537,
"crp",
"Hant, Latn",
ancestors = "pt",
sort_key = {Hant = "Hani-sortkey"},
}
m["crp-rsn"] = {
"Russenorsk",
505125,
"crp",
"Cyrl, Latn",
ancestors = "nn, ru",
translit = {Cyrl = "ru-translit"},
}
m["crp-spp"] = {
"Pijin Ladang Samoa",
7409948,
"crp",
"Latn",
ancestors = "en",
}
m["crp-slb"] = {
"Inggeris Solombala",
7558525,
"crp",
"Cyrl, Latn",
ancestors = "en, ru",
translit = {Cyrl = "ru-translit"},
}
m["crp-tpr"] = {
"Pijin Rusia Taimyr",
16930506,
"crp",
"Cyrl",
ancestors = "ru",
translit = "ru-translit",
}
m["csu-bba-pro"] = {
"Bongo-Bagirmi Purba",
116773722,
"csu-bba",
"Latn",
type = "reconstructed",
}
m["csu-maa-pro"] = {
"Mangbetu Purba",
116773786,
"csu-maa",
"Latn",
type = "reconstructed",
}
m["csu-pro"] = {
"Sudan Tengah Purba",
116773730,
"csu",
"Latn",
type = "reconstructed",
}
m["csu-sar-pro"] = {
"Sara Purba",
116773809,
"csu-sar",
"Latn",
type = "reconstructed",
}
m["cus-ash"] = {
"Ashraaf",
4805855,
"cus-som",
"Latn",
}
m["cus-hec-pro"] = {
"Kusyi Timur Tanah Tinggi Purba",
116773761,
"cus-hec",
"Latn",
type = "reconstructed",
}
m["cus-som-pro"] = {
"Somaloid Purba",
nil,
"cus-som",
"Latn",
type = "reconstructed",
}
m["cus-sou-pro"] = {
"Kusyi Selatan Purba",
126081567,
"cus-sou",
"Latn",
type = "reconstructed",
}
m["cus-pro"] = {
"Kusyi Purba",
116773204,
"cus",
"Latn",
type = "reconstructed",
}
m["dmn-dam"] = {
"Dama (Sierra Leone)",
19601574,
"dmn",
"Latn",
}
m["dra-bry"] = {
"Beary",
1089116,
"qfa-mix",
"Mlym, Knda",
ancestors = "ml, tcy",
-- translit Knda dalam [[Module:scripts/data]]
-- translit Mlym dalam [[Module:scripts/data]]
}
m["dra-cen-pro"] = {
"Dravidia Tengah Purba",
nil,
"dra-cen",
"Latn",
type = "reconstructed",
}
m["dra-mkn"] = {
"Kannada Pertengahan",
128810572,
"dra-kan",
"Knda",
-- translit Knda dalam [[Module:scripts/data]]
}
m["dra-nor-pro"] = {
"Dravidia Utara Purba",
124433593,
"dra-nor",
"Latn",
type = "reconstructed",
}
m["dra-okn"] = {
"Kannada Kuno",
15723156,
"dra-kan",
"Knda",
-- translit Knda dalam [[Module:scripts/data]]
}
m["dra-ote"] = {
"Telugu Kuno",
126720868,
"dra-tel",
"Telu",
translit = "te-translit",
}
m["dra-pro"] = {
"Dravidia Purba",
1702853,
"dra",
"Latn",
type = "reconstructed",
}
m["dra-sdo-pro"] = {
"Dravidia Selatan I Purba",
104847952, -- "Proto-Dravidia Selatan" Wikipedia ialah Proto-Dravidia Selatan I dalam skema ini.
"dra-sdo",
"Latn",
type = "reconstructed",
}
m["dra-sdt-pro"] = {
"Dravidia Selatan II Purba",
128885257,
"dra-sdt",
"Latn",
type = "reconstructed",
}
m["dra-sou-pro"] = {
"Dravidia Selatan Purba",
128886121,
"dra-sou",
"Latn",
type = "reconstructed",
}
m["egx-dem"] = {
"Mesir Demotik",
36765,
"egx",
"Latn, Egyd, Polyt",
sort_key = {
Latn = {
remove_diacritics = "'%-%s",
from = {"ꜣ", "j", "e", "ꜥ", "y", "w", "b", "p", "f", "m", "n", "r", "l", "ḥ", "ḫ", "h̭", "ẖ", "h", "š", "s", "q", "k", "g", "ṱ", "ṯ", "t", "ḏ", "%.", "⸗"},
to = {p[1], p[2], p[3], p[4], p[5], p[6], p[7], p[8], p[9], p[10], p[11], p[12], p[13], p[15], p[16], p[16], p[17], p[14], p[19], p[18], p[20], p[21], p[22], p[23], p[24], p[23], p[25], p[26], p[26]}
},
},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["dmn-pro"] = {
"Mande Purba",
116773785,
"dmn",
"Latn",
type = "reconstructed",
}
m["dmn-mdw-pro"] = {
"Mande Barat Purba",
116773822,
"dmn-mdw",
"Latn",
type = "reconstructed",
}
m["dru-pro"] = {
"Rukai Purba",
116773807,
"map",
"Latn",
type = "reconstructed",
}
m["ero-gsz"] = {
"Geshiza",
nil,
"ero",
"Latn",
}
m["ero-nya"] = {
"Nyagrong Minyag",
nil,
"ero",
"Latn",
}
m["ero-tau"] = {
"Stau",
nil,
"ero",
"Latn",
}
m["esx-esk-pro"] = {
"Eskimo Purba",
7251842,
"esx-esk",
"Latn",
type = "reconstructed",
}
m["esx-ink"] = {
"Inuktun",
1671647,
"esx-inu",
"Latn",
}
m["esx-inq"] = {
"Inuinnaqtun",
28070,
"esx-inu",
"Latn",
}
m["esx-inu-pro"] = {
"Inuit Purba",
60785588,
"esx-inu",
"Latn",
type = "reconstructed",
}
m["esx-pro"] = {
"Eskimo-Aleut Purba",
7251843,
"esx",
"Latn",
type = "reconstructed",
}
m["esx-tut"] = {
"Tunumiisut",
15665389,
"esx-inu",
"Latn",
}
m["euq-pro"] = {
"Basque Purba",
938011,
"euq",
"Latn",
type = "reconstructed",
}
m["gba-pro"] = {
"Gbaya Purba",
nil,
"gba",
"Latn",
type = "reconstructed",
}
m["gem-pro"] = {
"Jermanik Purba",
669623,
"gem",
"Latn",
type = "reconstructed",
sort_key = "gem-pro-sortkey",
}
m["gme-bur"] = {
"Burgundia",
47625,
"gme",
"Latn",
}
m["gme-cgo"] = {
"Goth Crimea",
36211,
"gme",
"Latn",
}
m["gmq-gut"] = {
"Gutnish",
1256646,
"gmq",
"Latn",
ancestors = "gmq-ogt",
}
m["gmq-jmk"] = {
"Jamtish",
35512,
"gmq-eas",
"Latn",
}
m["gmq-mno"] = {
"Norway Pertengahan",
3417070,
"gmq-wes",
"Latn",
}
m["gmq-oda"] = {
"Denmark Kuno",
12330003,
"gmq-eas",
"Latn, Runr",
strip_diacritics = {remove_diacritics = c.macron},
}
m["gmq-ogt"] = {
"Gutnish Kuno",
1133488,
"gmq",
"Latn, Runr",
ancestors = "non",
}
m["gmq-osw"] = {
"Sweden Kuno",
2417210,
"gmq-eas",
"Latn, Runr",
strip_diacritics = {remove_diacritics = c.macron},
}
m["gmq-pro"] = {
"Norse Purba",
1671294,
"gmq",
"Runr",
translit = "Runr-translit",
}
m["gmq-scy"] = {
"Scanian",
768017,
"gmq-eas",
"Latn",
}
m["gmw-bgh"] = {
"Bergish",
329030,
"gmw-frk",
"Latn",
}
m["gmw-cfr"] = {
"Franconia Tengah",
572197,
"gmw-hgm",
"Latn",
ancestors = "gmh",
wikimedia_codes = "ksh",
}
m["gmw-ecg"] = {
"Jerman Tengah Timur",
499344, -- merangkumi Q699284, Q152965
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gmw-fin"] = {
"Fingallian",
3072588,
"gmw-ian",
"Latn",
}
m["gmw-gts"] = {
"Gottscheerish",
533109,
"gmw-hgm",
"Latn",
ancestors = "bar",
}
m["gmw-jdt"] = {
"Belanda Jersey",
1687911,
"gmw-frk",
"Latn",
ancestors = "nl",
}
m["gmw-msc"] = {
"Scots Pertengahan",
3327000,
"gmw-ang",
"Latn",
ancestors = "enm-esc",
}
m["gmw-pro"] = {
"Jermanik Barat Purba",
78079021,
"gmw",
"Latn, Runr",
-- type = "reconstructed",
-- sebahagian besarnya tetapi tidak sepenuhnya direkonstruksi (seperti Proto-Norse); lihat BP Apr '24, tetapkan kembali kepada direkonstruksi (?) jika 'anti-asterisk' ditambah
sort_key = "gmw-pro-sortkey",
}
m["gmw-rfr"] = {
"Franconia Rhine",
707007,
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gmw-stm"] = {
"Schwaben Szatmár",
2223059,
"gmw-hgm",
"Latn",
ancestors = "swg",
}
m["gmw-tsx"] = {
"Saxon Transylvania",
260942,
"gmw-hgm",
"Latn",
ancestors = "gmw-cfr",
}
m["gmw-vog"] = {
"Jerman Volga",
312574,
"gmw-hgm",
"Latn",
ancestors = "gmw-rfr",
}
m["gmw-zps"] = {
"Jerman Zipser",
205548,
"gmw-hgm",
"Latn",
ancestors = "gmh",
}
m["gn-cls"] = {
"Guarani Klasik",
17478065,
"gn",
"Latn",
}
m["grk-cal"] = {
"Yunani Calabria",
1146398,
"grk",
"Latn, Grek",
ancestors = "grk-ita",
translit = {
Grek = "el-translit",
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-ita"] = {
"Yunani Italiot",
19720507,
"grk",
"Latn, Grek",
ancestors = "gkm",
translit = {
Grek = "el-translit",
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-mar"] = {
"Yunani Mariupol",
4400023,
"grk",
"Cyrl, Latn, Grek",
ancestors = "gkm",
translit = {
Cyrl = "grk-mar-translit",
Grek = "grk-mar-translit",
},
override_translit = true,
strip_diacritics = {
Cyrl = {remove_diacritics = c.acute},
},
-- translit, display_text, strip_diacritics, sort_key Grek dalam [[Module:scripts/data]]
}
m["grk-pro"] = {
"Hellenik Purba",
1231805,
"grk",
"Latn, Polyt",
type = "reconstructed",
sort_key = {Latn = {
from = {"ʰ", "ʷ"},
to = {"h", "w"},
remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.caron .. c.CGJ
}},
display_text = {Latn = {
from = {"([dlLt])" .. c.caron},
to = {"%1" .. c.CGJ .. c.caron},
}},
strip_diacritics = {Latn = {
from = {"([dlLt])" .. c.caron},
to = {"%1" .. c.CGJ .. c.caron},
}},
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
-- NOTA: dahulunya tiada translit ditentukan untuk Polyt; mungkin ketinggalan secara tidak sengaja; jika tidak, tetapkan Polyt = false dalam
-- bahagian translit
}
m["hmn-pro"] = {
"Hmongik Purba",
116773210,
"hmn",
"Latn",
type = "reconstructed",
}
m["hmx-mie-pro"] = {
"Mienik Purba",
116773229,
"hmx-mie",
"Latn",
type = "reconstructed",
}
m["hmx-pro"] = {
"Hmong-Mien Purba",
7251846,
"hmx",
"Latn",
type = "reconstructed",
}
m["hyx-pro"] = {
"Armenia Purba",
3848498,
"hyx",
"Latn",
type = "reconstructed",
}
m["iir-nur-pro"] = {
"Nuristani Purba",
116773248,
"iir-nur",
"Latn",
type = "reconstructed",
}
m["iir-pro"] = {
"Indo-Iran Purba",
966439,
"iir",
"Latn",
type = "reconstructed",
}
m["ijo-pro"] = {
"Ijoid Purba",
116773766,
"ijo",
"Latn",
type = "reconstructed",
}
m["inc-apa"] = {
"Apabhramsa",
616419,
"inc-mid",
"Deva, Shrd, Sidd",
ancestors = "pra",
translit = {
Deva = "sa-translit",
-- translit Shrd dalam [[Module:scripts/data]]
-- translit Sidd dalam [[Module:scripts/data]]
},
}
m["inc-ash"] = {
"Prakrit Ashoka",
104854379,
"inc-mid",
"Brah, Khar",
ancestors = "sa",
translit = {
-- translit Brah dalam [[Module:scripts/data]]
Khar = "Khar-translit",
},
}
m["inc-dng-pro"] = {
"Dangari Purba",
nil,
"inc-dng",
"Latn",
type = "reconstructed",
}
m["inc-kam"] = {
"Prakrit Kamarupi",
6356097,
"inc-bas",
"Brah, Sidd",
-- translit Brah, Sidd dalam [[Module:scripts/data]]
}
m["inc-kho"] = {
"Kholosi",
24952008,
"inc-snd",
"Latn",
}
m["inc-khr"] = {
"Khortha",
13406670,
"inc-sad",
"Deva, Kthi",
translit = {
Deva = "bho-translit",
Kthi = "bho-Kthi-translit",
},
}
m["inc-krd-pro"] = {
"Kamta Purba",
128816843,
"inc-bas",
"Latn",
ancestors = "inc-kam",
type = "reconstructed",
}
m["inc-mas"] = {
"Assam Pertengahan",
128806836,
"inc-bas",
"as-Beng",
ancestors = "inc-oas",
translit = "inc-mas-translit",
}
m["inc-mbn"] = {
"Benggali Pertengahan",
113559927,
"inc-bas",
"Beng",
ancestors = "inc-obn",
translit = "inc-mbn-translit",
}
m["inc-mgu"] = {
"Gujarati Pertengahan",
24907429,
"inc-wes",
"Deva",
ancestors = "inc-ogu",
}
m["inc-mor"] = {
"Odia Pertengahan",
128810882,
"inc-eas",
"Orya",
ancestors = "inc-oor",
}
m["inc-oas"] = {
"Assam Awal",
85758237,
"inc-bas",
"as-Beng",
ancestors = "inc-kam",
translit = "inc-oas-translit",
}
m["inc-oaw"] = {
"Awadhi Kuno",
nil,
"inc-hie",
"Deva, Kthi, Aran",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Deva = "sa-translit",
Kthi = "sa-Kthi-translit",
Aran = "inc-ohi-translit",
},
}
m["inc-obn"] = {
"Benggali Kuno",
113559926,
"inc-bas",
"Beng",
}
m["inc-ogu"] = {
"Gujarati Kuno",
24907427,
"inc-wes",
"Deva",
translit = "sa-translit",
}
m["inc-ohi"] = {
"Hindi Kuno",
48767781,
"inc-hiw",
"Deva, Aran",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Deva = "sa-translit",
Aran = "inc-ohi-translit",
},
}
m["inc-oor"] = {
"Odia Kuno",
128807801,
"inc-eas",
"Orya",
}
m["inc-opa"] = {
"Punjabi Kuno",
115270971,
"inc-pan",
"Guru, Aran",
translit = {
Guru = "inc-opa-Guru-translit",
Aran = "pa-Aran-translit",
},
strip_diacritics = {remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun},
}
m["inc-pro"] = {
"Indo-Arya Purba",
23808344,
"inc",
"Latn",
type = "reconstructed",
}
m["inc-sar"] = {
"Sarazi",
85799728,
"him",
"Aran, Deva, Takr",
strip_diacritics = {
from = {"هٔ", "ۂ"}, -- aksara "ۂ" kod U+06C2 kepada "ه" dan "هٔ" (U+0647 + U+0654) kepada "ه"
to = {"ہ", "ہ"},
remove_diacritics = c.fathatan .. c.dammatan .. c.kasratan .. c.fatha .. c.damma .. c.kasra .. c.shadda .. c.sukun .. c.nunghunna .. c.superalef
},
translit = {
Aran = "ur-translit",
Deva = "hi-translit",
-- Takr = "Takr-translit",
},
}
m["ine-ana-pro"] = {
"Anatolia Purba",
7251833,
"ine-ana",
"Latn",
type = "reconstructed",
}
m["ine-bsl-pro"] = {
"Balto-Slavik Purba",
1703347,
"ine-bsl",
"Latn",
type = "reconstructed",
sort_key = {
from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", c.acute, c.macron, "ˀ"},
to = {"a", "e", "i", "o", "u"}
},
}
m["ine-kal"] = {
"Kalašma",
122770439,
"ine-ana",
"Xsux",
}
m["ine-pae"] = {
"Paeonia",
2705672,
"ine",
"Polyt",
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["ine-pro"] = {
"Indo-Eropah Purba",
37178,
"ine",
"Latn",
type = "reconstructed",
sort_key = {
from = {"[áā]", "[éēḗ]", "[íī]", "[óōṓ]", "[úū]", "ĺ", "ḿ", "ń", "ŕ", "ǵ", "ḱ", "ʰ", "ʷ", "₁", "₂", "₃", c.ringbelow, c.acute, c.macron},
to = {"a", "e", "i", "o", "u", "l", "m", "n", "r", "g'", "k'", "¯h", "¯w", "1", "2", "3"}
},
}
m["ine-toc-pro"] = {
"Tocharia Purba",
104841462,
"ine-toc",
"Latn",
type = "reconstructed",
}
m["xme-old"] = {
"Median Kuno",
36461,
"xme",
"Polyt, Latn",
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["xme-mid"] = {
"Median Pertengahan",
12836150,
"xme",
"Latn",
}
m["xme-ker"] = {
"Kerman",
129850,
"xme",
"Arab, Latn, Hebr",
ancestors = "xme-mid",
-- display_text, strip_diacritics, sort_key Hebr dalam [[Module:scripts/data]]
}
m["xme-taf"] = {
"Tafreshi",
nil,
"xme",
"Arab, Latn",
ancestors = "xme-mid",
}
m["xme-ttc-pro"] = {
"Tatik Purba",
122973870,
"xme-ttc",
"Latn",
ancestors = "xme-mid",
}
m["xme-kls"] = {
"Kalasuri",
nil,
"xme-ttc",
ancestors = "xme-ttc-nor",
}
m["xme-klt"] = {
"Kilit",
3612452,
"xme-ttc",
"Cyrl", -- dan Arab?
}
m["xme-ott"] = {
"Tati Kuno",
434697,
"xme-ttc",
"Arab, Latn",
}
m["ira-kms-pro"] = {
"Komisenia Purba",
116773777,
"ira-kms",
"Latn",
type = "reconstructed",
}
m["ira-mpr-pro"] = {
"Medo-Parthia Purba",
116773227,
"ira-mpr",
"Latn",
type = "reconstructed",
}
m["ira-pat-pro"] = {
"Pathan Purba",
116773255,
"ira-pat",
"Latn",
type = "reconstructed",
}
m["ira-pro"] = {
"Iran Purba",
4167865,
"ira",
"Latn",
type = "reconstructed",
}
m["ira-zgr-pro"] = {
"Zaza-Gorani Purba",
116775031,
"ira-zgr",
"Latn",
type = "reconstructed",
}
m["xsc-pro"] = {
"Scythia Purba",
116773273,
"xsc",
"Latn",
type = "reconstructed",
}
m["xsc-sar-pro"] = {
"Sarmatia Purba",
116773249,
"xsc-sar",
"Latn",
type = "reconstructed",
}
m["xsc-skw-pro"] = {
"Saka-Wakhi Purba",
116773267,
"xsc-skw",
"Latn",
type = "reconstructed",
}
m["xsc-sak-pro"] = {
"Saka Purba",
116773264,
"xsc-sak",
"Latn",
type = "reconstructed",
}
m["ira-sym-pro"] = {
"Shughni-Yazghulami-Munji Purba",
116773813,
"ira-sym",
"Latn",
type = "reconstructed",
}
m["ira-sgi-pro"] = {
"Sanglechi-Ishkashimi Purba",
116773808,
"ira-sgi",
"Latn",
type = "reconstructed",
}
m["ira-mny-pro"] = {
"Munji-Yidgha Purba",
116773792,
"ira-mny",
"Latn",
type = "reconstructed",
}
m["ira-shy-pro"] = {
"Shughni-Yazghulami Purba",
116773812,
"ira-shy",
"Latn",
type = "reconstructed",
}
m["ira-shr-pro"] = {
"Shughni-Roshani Purba",
116773811,
"ira-shr",
"Latn",
type = "reconstructed",
}
m["ira-sgc-pro"] = {
"Sogdia Purba",
116773276,
"ira-sgc",
"Latn",
type = "reconstructed",
}
m["ira-wnj"] = {
"Vanji",
3398419,
"ira-shy",
"Latn",
}
m["iro-ere"] = {
"Erie",
5388365,
"iro-nor",
"Latn",
}
m["iro-min"] = {
"Mingo",
128531,
"iro-nor",
"Latn",
ietf_subtag = "i-mingo", -- tag IETF yang diwarisi
}
m["iro-nor-pro"] = {
"Iroquois Utara Purba",
116773242,
"iro-nor",
"Latn",
type = "reconstructed",
}
m["iro-pro"] = {
"Iroquois Purba",
7251852,
"iro",
"Latn",
type = "reconstructed",
}
m["itc-pro"] = {
"Italik Purba",
17102720,
"itc",
"Latn",
type = "reconstructed",
}
m["itc-psa"] = {
"Pra-Samnit",
7239186,
"itc-sbl",
"Ital, Polyt, Latn",
-- translit Ital dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja)
-- translit, display_text, strip_diacritics, sort_key Polyt dalam [[Module:scripts/data]]
}
m["jpx-hcj"] = {
"Hachijō",
5637049,
"jpx",
"Jpan",
ancestors = "ojp-eas",
translit = s["jpx-translit"],
display_text = s["jpx-displaytext"],
strip_diacritics = s["jpx-stripdiacritics"],
sort_key = s["jpx-sortkey"],
}
m["jpx-pro"] = {
"Jepunik Purba",
3924309,
"jpx",
"Latn",
type = "reconstructed",
}
m["jpx-ryu-pro"] = {
"Ryukyu Purba",
56349069,
"jpx-ryu",
"Latn",
type = "reconstructed",
}
m["kar-pro"] = {
"Karen Purba",
85794783,
"kar",
"Latn",
type = "reconstructed",
}
m["kca-eas"] = {
"Khanty Timur",
30304622,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
-- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka)
sort_key = { Cyrl = { from = {""}, to = {""} } },
}
m["kca-nor"] = {
"Khanty Utara",
30304527,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
-- TODO sementara sehingga MediaWiki menyokong Unicode 16 (mungkin memerlukan kemas kini PHP dari pihak mereka)
sort_key = { Cyrl = { from = {""}, to = {""} } },
}
m["kca-pro"] = {
"Khanty Purba",
127505171,
"kca",
"Latn",
type = "reconstructed",
}
m["kca-sou"] = {
"Khanty Selatan",
30304618,
"kca",
"Cyrl",
translit = "kca-translit",
override_translit = true,
}
m["khi-kho-pro"] = {
"Khoe Purba",
116773218,
"khi-kho",
"Latn",
type = "reconstructed",
}
m["khi-kun"] = {
"ǃKung",
32904,
"khi-kxa",
"Latn",
}
m["ko-ear"] = {
"Korea Moden Awal",
756014,
"qfa-kor",
"Kore",
ancestors = "okm",
translit = "okm-translit",
-- strip_diacritics Kore dalam [[Module:scripts/data]]
}
m["kro-pro"] = {
"Kru Purba",
116773778,
"kro",
"Latn",
type = "reconstructed",
}
m["ku-pro"] = {
"Kurdi Purba",
116773221,
"ku",
"Latn",
type = "reconstructed",
}
m["map-ata-pro"] = {
"Atayalik Purba",
116773151,
"map-ata",
"Latn",
type = "reconstructed",
}
m["map-bms"] = {
"Banyumasan",
33219,
"map",
"Latn, Java",
}
m["map-pro"] = {
"Austronesia Purba",
49230,
"map",
"Latn",
type = "reconstructed",
}
m["mis-hkl"] = {
"Hokkien Peranakan Kelantan",
108794818,
"qfa-mix",
ancestors = "nan-hbl, sou, mfa",
}
m["mis-idn"] = {
"Idiom Neutral",
35847,
"art",
"Latn",
type = "appendix-constructed",
}
m["mis-isa"] = {
"Isauria",
16956868,
nil,
-- "Xsux, Hluw, Latn",
}
m["mis-jie"] = {
"Jie",
124424186,
nil,
"Hani",
sort_key = "Hani-sortkey",
}
m["mis-jzh"] = {
"Jizhao",
45242758,
"qfa-bej",
"Latn",
}
m["mis-kas"] = {
"Kassite",
35612,
nil,
"Xsux",
}
m["mis-mmd"] = {
"Mimi Decorse",
6862206,
nil,
"Latn",
}
m["mis-mmn"] = {
"Mimi Nachtigal",
6862207,
nil,
"Latn",
}
m["mis-phi"] = {
"Filistin",
2230924,
nil,
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]] (NOTA: tidak hadir sebelum ini, mungkin ketinggalan secara tidak sengaja)
}
m["mis-rou"] = {
"Rouran",
48816637,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-tdl"] = {
"Turdulia",
133176492,
}
m["mis-tdt"] = {
"Turdetania",
133176461,
}
m["mis-tnw"] = {
"Tangwang",
7683179,
"qfa-mix",
"Latn",
ancestors = "cmn, sce",
}
m["mis-tuh"] = {
"Tuyuhun",
48816625,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-tuo"] = {
"Tuoba",
48816629,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-wuh"] = {
"Wuhuan",
118976867,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-xbi"] = {
"Xianbei",
4448647,
"qfa-xgx",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mis-xnu"] = {
"Xiongnu",
10901674,
nil,
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mjg-mgl"] = {
"Mongghul",
53765528,
"mjg",
"Latn", -- juga Mong, Cyrl?
}
m["mjg-mgr"] = {
"Mangghuer",
56285392,
"mjg",
"Latn", -- juga Mong, Cyrl?
}
m["mkh-asl-pro"] = {
"Asli Purba",
55630680,
"mkh-asl",
"Latn",
type = "reconstructed",
}
m["mkh-ban-pro"] = {
"Bahnar Purba",
116773189,
"mkh-ban",
"Latn",
type = "reconstructed",
}
m["mkh-kat-pro"] = {
"Katuik Purba",
116773772,
"mkh-kat",
"Latn",
type = "reconstructed",
}
m["mkh-khm-pro"] = {
"Khmuik Purba",
116773774,
"mkh-khm",
"Latn",
type = "reconstructed",
}
m["mkh-kmr-pro"] = {
"Khmer Purba",
55630684,
"mkh-kmr",
"Latn",
type = "reconstructed",
}
m["mkh-mmn"] = {
"Mon Pertengahan",
121337926,
"mkh-mnc",
"Latn, Mymr", --dan juga Pallava
ancestors = "omx",
}
m["mkh-mnc-pro"] = {
"Monik Purba",
116773231,
"mkh-mnc",
"Latn",
type = "reconstructed",
}
m["mkh-mvi"] = {
"Vietnam Pertengahan",
9199,
"mkh-vie",
"Hani, Latn",
sort_key = {Hani = "Hani-sortkey"},
}
m["mkh-pal-pro"] = {
"Palaungik Purba",
104847372,
"mkh-pal",
"Latn",
type = "reconstructed",
}
m["mkh-pea-pro"] = {
"Pearik Purba",
116773804,
"mkh-pea",
"Latn",
type = "reconstructed",
}
m["mkh-pkn-pro"] = {
"Pakanik Purba",
116773803,
"mkh-pkn",
"Latn",
type = "reconstructed",
}
m["mkh-pro"] = { --Ini akan digabungkan ke dalam aav-pro 2015.
"Mon-Khmer Purba",
7251859,
"mkh",
"Latn",
type = "reconstructed",
}
m["mnw-tha"] = { -- Untuk dibuang.
"Mon Thailand",
nil,
"mkh-mnc",
"Mymr, Thai",
ancestors = "mkh-mmn",
sort_key = {
from = {"[%p]", "ျ", "ြ", "ွ", "ှ", "ၞ", "ၟ", "ၠ", "ၚ", "ဿ", "[็-๎]", "([เแโใไ])([ก-ฮ])ฺ?"},
to = {"", "္ယ", "္ရ", "္ဝ", "္ဟ", "္န", "္မ", "္လ", "င", "သ္သ", "", "%2%1"}
},
}
m["mkh-vie-pro"] = {
"Vietik Purba",
109432616,
"mkh-vie",
"Latn",
type = "reconstructed",
}
m["mns-cen"] = {
"Mansi Tengah",
128810384,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mns-nor"] = {
"Mansi Utara",
30304537,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mns-pro"] = {
"Mansi Purba",
128883093,
"mns",
"Latn",
type = "reconstructed",
}
m["mns-sou"] = {
"Mansi Selatan",
30304629,
"mns",
"Cyrl",
translit = "mns-translit",
override_translit = true,
}
m["mun-pro"] = {
"Munda Purba",
105102373,
"mun",
"Latn",
type = "reconstructed",
}
m["myn-chl"] = { -- peringkat selepas ''emy''
"Ch'olti'",
873995,
"myn",
"Latn",
}
m["myn-pro"] = {
"Maya Purba",
3321532,
"myn",
"Latn",
type = "reconstructed",
}
m["nai-ala"] = {
"Alazapa",
128810233,
nil,
"Latn",
}
m["nai-bay"] = {
"Bayogoula",
1563704,
nil,
"Latn",
}
m["nai-cal"] = {
"Calusa",
51782,
nil,
"Latn",
}
m["nai-chi"] = {
"Chiquimulilla",
25339627,
"nai-xin",
"Latn",
}
m["nai-chu-pro"] = {
"Chumash Purba",
116773736,
"nai-chu",
"Latn",
type = "reconstructed",
}
m["nai-cig"] = {
"Ciguayo",
20741700,
nil,
"Latn",
}
m["nai-ckn-pro"] = {
"Chinook Purba",
116773735,
"nai-ckn",
"Latn",
type = "reconstructed",
}
m["nai-guz"] = {
"Guazacapán",
19572028,
"nai-xin",
"Latn",
}
m["nai-hit"] = {
"Hitchiti",
1542882,
"nai-mus",
"Latn",
}
m["nai-ipa"] = {
"Ipai",
3027474,
"nai-yuc",
"Latn",
}
m["nai-jtp"] = {
"Jutiapa",
nil,
"nai-xin",
"Latn",
}
m["nai-jum"] = {
"Jumaytepeque",
25339626,
"nai-xin",
"Latn",
}
m["nai-kat"] = {
"Kathlamet",
6376639,
"nai-ckn",
"Latn",
}
m["nai-klp-pro"] = {
"Kalapuya Purba",
116773771,
"nai-klp",
"Latn",
type = "reconstructed",
}
m["nai-knm"] = {
"Konomihu",
3198734,
"nai-shs",
"Latn",
}
m["nai-kum"] = {
"Kumeyaay",
4910139,
"nai-yuc",
"Latn",
}
m["nai-mac"] = {
"Macoris",
21070851,
nil,
"Latn",
}
m["nai-mdu-pro"] = {
"Maidu Purba",
116773784,
"nai-mdu",
"Latn",
type = "reconstructed",
}
m["nai-miz-pro"] = {
"Mixe-Zoque Purba",
7251858,
"nai-miz",
"Latn",
type = "reconstructed",
}
m["nai-mus-pro"] = {
"Muskogi Purba",
116775368,
"nai-mus",
"Latn",
type = "reconstructed",
}
m["nai-nao"] = {
"Naolan",
6964594,
nil,
"Latn",
}
m["nai-nrs"] = {
"Shasta Sungai Baru",
7011254,
"nai-shs",
"Latn",
}
m["nai-okw"] = {
"Okwanuchu",
3350126,
"nai-shs",
"Latn",
}
m["nai-per"] = {
"Pericú",
3375369,
nil,
"Latn",
}
m["nai-pic"] = {
"Picuris",
7191257,
"nai-kta",
"Latn",
}
m["nai-plp-pro"] = {
"Penuti Penara Purba",
116773806,
"nai-plp",
"Latn",
type = "reconstructed",
}
m["nai-pom-pro"] = {
"Pomo Purba",
116773262,
"nai-pom",
"Latn",
type = "reconstructed",
}
m["nai-qng"] = {
"Quinigua",
36360,
nil,
"Latn",
}
m["nai-sca-pro"] = { -- PERHATIAN 'sio-pro' "Proto-Siouan" iaitu Proto-Sioux Barat
"Siouan-Catawba Purba",
116773275,
"nai-sca",
"Latn",
type = "reconstructed",
}
m["nai-sin"] = {
"Sinacantán",
24190249,
"nai-xin",
"Latn",
}
m["nai-sln"] = {
"Lenca Salvador",
3229434,
"nai-len",
"Latn",
}
m["nai-spt"] = {
"Sahaptin",
3833015,
"nai-shp",
"Latn",
}
m["nai-tap"] = {
"Tapachultec",
7684401,
"nai-miz",
"Latn",
}
m["nai-taw"] = {
"Tawasa",
7689233,
nil,
"Latn",
}
m["nai-teq"] = {
"Tequistlatec",
2964454,
"nai-tqn",
"Latn",
}
m["nai-tip"] = {
"Tipai",
3027471,
"nai-yuc",
"Latn",
}
m["nai-tot-pro"] = {
"Totozoquean Purba",
116773285,
"nai-tot",
"Latn",
type = "reconstructed",
}
m["nai-tsi-pro"] = {
"Tsimshianik Purba",
nil,
"nai-tsi",
"Latn",
type = "reconstructed",
}
m["nai-utn-pro"] = {
"Utik Purba",
116773290,
"nai-utn",
"Latn",
type = "reconstructed",
}
m["nai-wai"] = {
"Waikuri",
3118702,
nil,
"Latn",
}
m["nai-wji"] = {
"Jicaque Barat",
3178610,
"nai-jcq",
"Latn",
}
m["nai-yup"] = {
"Yupiltepeque",
25339628,
"nai-xin",
"Latn",
}
m["nan-dat"] = {
"Min Datian",
19855572,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-hbl"] = {
"Hokkien",
1624231,
"zhx-nan",
"Hants, Latn, Bopo, Kana",
wikimedia_codes = "zh-min-nan",
generate_forms = "zh-generateforms",
sort_key = {
Hani = "Hani-sortkey",
Kana = "Kana-sortkey"
},
}
m["nan-hlh"] = {
"Min Hailufeng",
120755728,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-lnx"] = {
"Min Longyan",
6674568,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-tws"] = {
"Teochew",
36759,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["nan-zhe"] = {
"Min Zhenan",
3846710,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["nan-zsh"] = {
"Min Sanxiang",
7420769,
"zhx-nan",
"Hants",
generate_forms = "zh-generateforms",
sort_key = "Hani-sortkey",
}
m["ngf-bin-pro"] = {
"Binandere Purba",
137881672,
"ngf-bin",
"Latn",
type = "reconstructed",
}
m["ngf-pro"] = {
"Trans-New Guinea Purba",
85794785,
"ngf",
"Latn",
type = "reconstructed",
}
m["nic-bco-pro"] = {
"Benue-Congo Purba",
116773194,
"nic-bco",
"Latn",
type = "reconstructed",
}
m["nic-bod-pro"] = {
"Bantoid Purba",
116773190,
"nic-bod",
"Latn",
type = "reconstructed",
}
m["nic-eov-pro"] = {
"Oti-Volta Timur Purba",
116773753,
"nic-eov",
"Latn",
type = "reconstructed",
}
m["nic-gns-pro"] = {
"Gurunsi Purba",
116773759,
"nic-gns",
"Latn",
type = "reconstructed",
}
m["nic-grf-pro"] = {
"Grassfields Purba",
116773755,
"nic-grf",
"Latn",
type = "reconstructed",
}
m["nic-gur-pro"] = {
"Gur Purba",
116773758,
"nic-gur",
"Latn",
type = "reconstructed",
}
m["nic-jkn-pro"] = {
"Jukunoid Purba",
116773769,
"nic-jkn",
"Latn",
type = "reconstructed",
}
m["nic-lcr-pro"] = {
"Cross River Hilir Purba",
116773782,
"nic-lcr",
"Latn",
type = "reconstructed",
}
m["nic-ogo-pro"] = {
"Ogoni Purba",
116773799,
"nic-ogo",
"Latn",
type = "reconstructed",
}
m["nic-ovo-pro"] = {
"Oti-Volta Purba",
116773802,
"nic-ovo",
"Latn",
type = "reconstructed",
}
m["nic-plt-pro"] = {
"Plateau Purba",
116773805,
"nic-plt",
"Latn",
type = "reconstructed",
}
m["nic-pro"] = {
"Niger-Congo Purba",
108000748,
"nic",
"Latn",
type = "reconstructed",
}
m["nic-ubg-pro"] = {
"Ubangi Purba",
116773818,
"nic-ubg",
"Latn",
type = "reconstructed",
}
m["nic-ucr-pro"] = {
"Cross River Hulu Purba",
116773819,
"nic-ucr",
"Latn",
type = "reconstructed",
}
m["nic-vco-pro"] = {
"Volta-Congo Purba",
116773293,
"nic-vco",
"Latn",
type = "reconstructed",
}
m["njo-jgl"] = {
"Ao Chungli",
55607615,
"njo",
"Latn",
}
m["njo-mng"] = {
"Ao Mongsen",
85383221,
"njo",
"Latn",
}
m["nub-har"] = {
"Haraza",
19572059,
"nub",
"Arab, Latn",
}
m["nub-pro"] = {
"Nubia Purba",
116773246,
"nub",
"Latn",
type = "reconstructed",
}
m["omq-cha-pro"] = {
"Chatino Purba",
116773202,
"omq-cha",
"Latn",
type = "reconstructed",
}
m["omq-maz-pro"] = {
"Mazatec Purba",
116773790,
"omq-maz",
"Latn",
type = "reconstructed",
}
m["omq-mix-pro"] = {
"Mixtecan Purba",
21573423,
"omq-mix",
"Latn",
type = "reconstructed",
}
m["omq-mxt-pro"] = {
"Mixtec Purba",
21573424,
"omq-mxt",
"Latn",
type = "reconstructed",
}
m["omq-otp-pro"] = {
"Oto-Pamean Purba",
116773251,
"omq-otp",
"Latn",
type = "reconstructed",
}
m["omq-pro"] = {
"Oto-Manguean Purba",
33669,
"omq",
"Latn",
type = "reconstructed",
}
m["omq-sjq"] = {
"Chatino San Juan Quiahije",
138330751,
"omq-cha",
"Latn",
}
m["omq-tel"] = {
"Mixtec Teposcolula",
nil,
"omq-mxt",
"Latn",
}
m["omq-teo"] = {
"Chatino Teojomulco",
25340451,
"omq-cha",
"Latn",
}
m["omq-tri-pro"] = {
"Triqui Purba",
116773817,
"omq-tri",
"Latn",
type = "reconstructed",
}
m["omq-zap-pro"] = {
"Zapotecan Purba",
116773297,
"omq-zap",
"Latn",
type = "reconstructed",
}
m["omq-zpc-pro"] = {
"Zapotec Purba",
116773296,
"omq-zpc",
"Latn",
type = "reconstructed",
}
m["omv-aro-pro"] = {
"Aroid Purba",
116773721,
"omv-aro",
"Latn",
type = "reconstructed",
}
m["omv-diz-pro"] = {
"Dizoid Purba",
116773750,
"omv-diz",
"Latn",
type = "reconstructed",
}
m["omv-pro"] = {
"Omotik Purba",
116773800,
"omv",
"Latn",
type = "reconstructed",
}
m["oto-otm-pro"] = {
"Otomi Purba",
5908710,
"oto-otm",
"Latn",
type = "reconstructed",
}
m["oto-pro"] = {
"Otomian Purba",
116773252,
"oto",
"Latn",
type = "reconstructed",
}
m["paa-kmn"] = {
"Kómnzo",
18344310,
"paa-wko",
"Latn",
}
m["paa-kwn"] = {
"Kuwani",
6449056,
"qfa-unc", -- kurang dibuktikan, mungkin sama dengan atau berkaitan dengan Kalabra
"Latn",
}
m["paa-lei"] = {
"Leitre",
85776228,
"paa-isk",
}
m["paa-nha-pro"] = {
"Halmahera Utara Purba",
116773241,
"paa-nha",
"Latn",
type = "reconstructed"
}
m["paa-nun"] = {
"Nungon",
128807788,
"ngf-ynu",
"Latn",
}
m["phi-din"] = {
"Agta Dinapigue",
16945774,
"phi",
"Latn",
}
m["phi-kal-pro"] = {
"Kalamian Purba",
116773213,
"phi-kal",
"Latn",
type = "reconstructed",
}
m["phi-nag"] = {
"Agta Nagtipunan",
16966111,
"phi",
"Latn",
}
m["phi-pro"] = {
"Filipina Purba",
18204898,
"phi",
"Latn",
type = "reconstructed",
}
m["poz-abi"] = {
"Abai",
19570729,
"poz-san",
"Latn",
}
m["poz-bal"] = {
"Baliledo",
4850912,
"poz",
"Latn",
}
m["poz-btk-pro"] = {
"Bungku-Tolaki Purba",
116773724,
"poz-btk",
"Latn",
type = "reconstructed",
}
m["poz-cet-pro"] = {
"Melayu-Polinesia Tengah-Timur Purba",
2269883,
"poz-cet",
"Latn",
type = "reconstructed",
}
m["poz-hce-pro"] = {
"Halmahera-Cenderawasih Purba",
116773209,
"poz-hce",
"Latn",
type = "reconstructed",
}
m["poz-lgx-pro"] = {
"Lampung Purba",
116773222,
"poz-lgx",
"Latn",
type = "reconstructed",
}
m["poz-mcm-pro"] = {
"Melayu-Chamik Purba",
116773225,
"poz-mcm",
"Latn",
type = "reconstructed",
}
m["poz-mic-pro"] = {
"Mikronesia Purba",
111939079,
"poz-mic",
"Latn",
type = "reconstructed",
}
m["poz-mly-pro"] = {
"Melayik Purba",
98057728,
"poz-mly",
"Latn",
type = "reconstructed",
}
m["poz-msa-pro"] = {
"Melayu-Sumbawa Purba",
116773226,
"poz-msa",
"Latn",
type = "reconstructed",
}
m["poz-nes"] = {
"Nese",
2157412,
"poz-vnc",
"Latn",
}
m["poz-oce-pro"] = {
"Oceania Purba",
141741,
"poz-oce",
"Latn",
type = "reconstructed",
}
m["poz-pcc-pro"] = {
"Pasifik Tengah Purba",
111962726,
"poz-pcc",
"Latn",
type = "reconstructed",
}
m["poz-pep-pro"] = {
"Polinesia Timur Purba",
113988745,
"poz-pep",
"Latn",
type = "reconstructed",
}
m["poz-pnp-pro"] = {
"Polinesia Teras Purba",
113988746,
"poz-pnp",
"Latn",
type = "reconstructed",
}
m["poz-pol-pro"] = {
"Polinesia Purba",
1658709,
"poz-pol",
"Latn",
type = "reconstructed",
}
m["poz-pro"] = {
"Melayu-Polinesia Purba",
3832960,
"poz",
"Latn",
type = "reconstructed",
}
m["poz-sml"] = {
"Melayu Sarawak",
4251702,
"poz-mly",
"Latn, Arab",
}
m["poz-ssw-pro"] = {
"Sulawesi Selatan Purba",
116773279,
"poz-ssw",
"Latn",
type = "reconstructed",
}
m["poz-swa-pro"] = {
"Sarawak Utara Purba",
116773243,
"poz-swa",
"Latn",
type = "reconstructed",
}
m["poz-ter"] = {
"Melayu Terengganu",
4207412,
"poz-mly",
"Latn, Arab",
}
m["pqe-pro"] = {
"Melayu-Polinesia Timur Purba",
2269883,
"pqe",
"Latn",
type = "reconstructed",
}
m["pra-niy"] = {
"Prakrit Niya",
11991601,
"inc-mid",
"Khar",
ancestors = "inc-ash",
translit = "Khar-translit",
}
m["qfa-adm-pro"] = {
"Andaman Besar Purba",
116773756,
"qfa-adm",
"Latn",
type = "reconstructed",
}
m["qfa-bet-pro"] = {
"Be-Tai Purba",
116773193,
"qfa-bet",
"Latn",
type = "reconstructed",
}
m["qfa-cka-pro"] = {
"Chukotko-Kamchatka Purba",
7251837,
"qfa-cka",
"Latn",
type = "reconstructed",
}
m["qfa-hur-pro"] = {
"Hurro-Urartia Purba",
116773211,
"qfa-hur",
"Latn",
type = "reconstructed",
}
m["qfa-kad-pro"] = {
"Kadu Purba",
116773770,
"qfa-kad",
"Latn",
type = "reconstructed",
}
m["qfa-kms-pro"] = {
"Kam-Sui Purba",
55630682,
"qfa-kms",
"Latn",
type = "reconstructed",
}
m["qfa-kor-pro"] = {
"Korea Purba",
467883,
"qfa-kor",
"Latn",
type = "reconstructed",
}
m["qfa-kra-pro"] = {
"Kra Purba",
7251854,
"qfa-kra",
"Latn",
type = "reconstructed",
}
m["qfa-lic-pro"] = {
"Hlai Purba",
7251845,
"qfa-lic",
"Latn",
type = "reconstructed",
}
m["qfa-onb-pro"] = {
"Be Purba",
116773192,
"qfa-onb",
"Latn",
type = "reconstructed",
}
m["qfa-ong-pro"] = {
"Onga Purba",
116773801,
"qfa-ong",
"Latn",
type = "reconstructed",
}
m["qfa-tak-pro"] = {
"Kra-Dai Purba",
104901616,
"qfa-tak",
"Latn",
type = "reconstructed",
}
m["qfa-yen-pro"] = {
"Yenisei Purba",
27639,
"qfa-yen",
"Latn",
type = "reconstructed",
}
m["qfa-yuk-pro"] = {
"Yukaghir Purba",
116773294,
"qfa-yuk",
"Latn",
type = "reconstructed",
}
m["qwe-kch"] = {
"Kichwa",
1740805,
"qwe",
"Latn",
ancestors = "qu",
}
m["qwe-pro"] = {
"Quechua Purba",
5575757,
"qwe",
"Latn",
type = "reconstructed",
}
m["roa-ang"] = {
"Angevin",
56782,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-bbn"] = {
"Bourbonnais-Berrichon",
2899128,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-brg"] = {
"Bourguignon",
508332,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-can"] = {
"Cantabrian",
917021,
"roa-asl",
"Latn",
}
m["roa-cha"] = {
"Champenois",
430018,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-fcm"] = {
"Franc-Comtois",
510561,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-gal"] = {
"Gallo",
37300,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-gib"] = {
"Gallo-Italik Basilicata",
3094838,
"roa-git",
ancestors = "pms-old",
"Latn",
}
m["roa-gis"] = {
"Gallo-Italik Sicily",
2629019,
"roa-git",
"Latn",
ancestors = "pms-old",
}
m["roa-leo"] = {
"Leon",
34108,
"roa-asl",
"Latn",
}
m["roa-lor"] = {
"Lorrain",
671198,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-oca"] = {
"Catalonia Kuno",
15478520,
"roa-ocr",
"Latn",
sort_key = {remove_diacritics = c.grave .. c.acute .. c.diaer .. c.cedilla .. "·"},
}
m["roa-ole"] = {
"Leon Kuno",
125977465,
"roa-asl",
"Latn",
}
m["roa-ona"] = {
"Navarro-Aragon Kuno",
2736184,
"roa-nar",
"Latn",
}
m["roa-opt"] = {
"Galicia-Portugis Kuno",
1072111,
"roa-gap",
"Latn",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.circ},
}
m["roa-orl"] = {
"Orléanais",
28497058,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-poi"] = {
"Poitevin-Saintongeais",
514123,
"roa-oil",
"Latn",
sort_key = s["roa-oil-sortkey"],
}
m["roa-tar"] = {
"Tarantino",
695526,
"roa-itr",
"Latn",
wikimedia_codes = "roa-tara",
}
m["sai-all"] = {
"Allentiac",
19570789,
"sai-hrp",
"Latn",
}
m["sai-and"] = {
"Andoquero",
16828359,
"sai-wit",
"Latn",
}
m["sai-ayo"] = {
"Ayomán",
16937754,
"sai-jir",
"Latn",
}
m["sai-bae"] = {
"Baenan",
3401998,
"qfa-unc", -- pupus, kurang dibuktikan; hanya dikenali melalui 9 perkataan
"Latn",
}
m["sai-bag"] = {
"Bagua",
5390321,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin bahasa Carib
"Latn",
}
m["sai-bet"] = {
"Betoi",
926551,
"qfa-iso",
"Latn",
}
m["sai-bor-pro"] = {
"Bora Purba",
nil,
"sai-bor",
"Latn",
}
m["sai-cac"] = {
"Cacán",
945482,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-caq"] = {
"Caranqui",
2937753,
"sai-bar",
"Latn",
}
m["sai-car-pro"] = {
"Carib Purba",
116773196,
"sai-car",
"Latn",
type = "reconstructed",
}
m["sai-cat"] = {
"Catacao",
5051136,
"sai-ctc",
"Latn",
}
m["sai-cer-pro"] = {
"Cerrado Purba",
116773200,
"sai-cer",
"Latn",
type = "reconstructed",
}
m["sai-chi"] = {
"Chirino",
5390321,
"qfa-unc", -- pupus, hanya empat perkataan diketahui; mungkin berkaitan dengan Candoshi-Shapra (cbu)
"Latn",
}
m["sai-chn"] = {
"Chaná",
5072718,
"sai-crn",
"Latn",
}
m["sai-chp"] = {
"Chapacura",
5072884,
"sai-cpc",
"Latn",
}
m["sai-chr"] = {
"Charrua",
5086680,
"sai-crn",
"Latn",
}
m["sai-chu"] = {
"Churuya",
5118339,
"sai-guh",
"Latn",
}
m["sai-cje-pro"] = {
"Jê Tengah Purba",
116773198,
"sai-cje",
"Latn",
type = "reconstructed",
}
m["sai-cmg"] = {
"Comechingon",
6644203,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-cno"] = {
"Chono",
5104704,
"qfa-unc", -- pupus, kurang dibuktikan; tiada konsensus mengenai pengelasan, mungkin palsu
"Latn",
}
m["sai-cnr"] = {
"Cañari",
5055572,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Chimuan atau Barbacoan
"Latn",
}
m["sai-coe"] = {
"Coeruna",
6425639,
"sai-wit",
"Latn",
}
m["sai-col"] = {
"Colán",
5141893,
"sai-ctc",
"Latn",
}
m["sai-cop"] = {
"Copallén",
5390321,
"qfa-unc", -- pupus, hanya empat perkataan dibuktikan; mungkin Cholonan
"Latn",
}
m["sai-crd"] = {
"Coroado Puri",
24191321,
"sai-mje",
"Latn",
}
m["sai-ctq"] = {
"Catuquinaru",
16858455,
"qfa-unc", -- pupus, kurang dibuktikan; kosa kata tidak menyerupai bahasa lain
"Latn",
}
m["sai-cul"] = {
"Culli",
2879660,
"qfa-unc", -- pupus, kurang dibuktikan; sering dianggap sebagai pencilan
"Latn",
}
m["sai-cva"] = {
"Cueva",
5192644,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Chocoan
"Latn",
}
m["sai-esm"] = {
"Esmeralda",
3058083,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin berkaitan dengan Yaruro
"Latn",
}
m["sai-ewa"] = {
"Ewarhuyana",
16898104,
nil,
"Latn",
}
m["sai-gam"] = {
"Gamela",
5403661,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin pencilan
"Latn",
}
m["sai-gay"] = {
"Gayón",
5528902,
"sai-jir",
"Latn",
}
m["sai-gmo"] = {
"Guamo",
5613495,
"qfa-unc", -- pupus; "Kaufman (1990) mendapati hubungan dengan bahasa-bahasa Chapacuran meyakinkan." [Wikipedia] Dianggap sebagai pencilan oleh Campbell (2024).
"Latn",
}
m["sai-gua"] = {
"Guachí",
5613172,
"sai-guc",
"Latn",
}
m["sai-gue"] = {
"Güenoa",
5626799,
"sai-crn",
"Latn",
}
m["sai-hau"] = {
"Haush",
3128376,
"sai-cho",
"Latn",
}
m["sai-jee-pro"] = {
"Jê Purba",
116773212,
"sai-jee",
"Latn",
type = "reconstructed",
}
m["sai-jko"] = {
"Jeikó",
6176527,
"sai-mje",
"Latn",
}
m["sai-jrj"] = {
"Jirajara",
6202966,
"sai-jir",
"Latn",
}
m["sai-kat"] = { -- kontras xoo, kzw, sai-xoc
"Katembri",
6375925,
"qfa-unc", -- pupus, kurang dibuktikan; "Kaufman (1990) telah menghubungkannya dengan bahasa Taruma yang hampir pupus, walaupun ini tidak diterima oleh sarjana lain." [Wikipedia]
"Latn",
}
m["sai-mal"] = {
"Malalí",
6741212,
"sai-mje", -- dianggap sebagai bahasa Maxakalían yang paling divergen (subbahagian kepada Macro-Jê), yang mana kami tiada entri
"Latn",
}
m["sai-mar"] = {
"Maratino",
6755055,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Uto-Aztecan
"Latn",
}
m["sai-mat"] = {
"Matanawi",
6786047,
"qfa-unc", -- pupus; sama ada pencilan atau berkait jauh dengan bahasa-bahasa Muran; Campbell (2024) menyenarainya sebagai pencilan, Glottolog memberikannya sebagai tidak terkelas
"Latn",
}
m["sai-mcn"] = {
"Mocana",
3402048,
"qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad)
"Latn",
}
m["sai-men"] = {
"Menien",
16890110,
"sai-mje",
"Latn",
}
m["sai-mil"] = {
"Millcayac",
19573012,
"sai-hrp",
"Latn",
}
m["sai-mlb"] = {
"Malibu",
134374036,
"qfa-unc", -- pupus, kurang dibuktikan; diberikan sebagai sebahagian daripada bahasa Malibu (pengumpulan geografi; bukan klad)
"Latn",
}
m["sai-msk"] = {
"Masakará",
6782426,
"sai-mje",
"Latn",
}
m["sai-muc"] = {
"Mucuchí",
6931290,
nil, -- lazimnya dianggap sebagai Timotean, yang mana kami tiada entri
"Latn",
}
m["sai-mue"] = {
"Muellama",
16886936,
"sai-bar",
"Latn",
}
m["sai-muz"] = {
"Muzo",
6644203,
"qfa-unc", -- bahasa pupus di Colombia, kurang dibuktikan; mungkin Pijao (Cariban)
"Latn",
}
m["sai-mys"] = {
"Maynas",
16919393,
"sai-cah", -- mengikut Campbell (2024); dahulu dianggap tidak terkelas
"Latn",
}
m["sai-nat"] = {
"Natú",
9006749,
"qfa-unc", -- pupus, kurang dibuktikan; "hanya Greenberg yang berani mengelaskannya".[Wikipedia, memetik Moseley, Christopher; Asher, R. E.; Tait, Mary (1994), Atlas of the world's languages]
"Latn",
}
m["sai-nje-pro"] = {
"Jê Utara Purba",
116773245,
"sai-nje",
"Latn",
type = "reconstructed",
}
m["sai-opo"] = {
"Opón",
7099152,
"sai-car",
"Latn",
}
m["sai-oto"] = {
"Otomaco",
16879234,
"sai-otm",
"Latn",
}
m["sai-pal"] = {
"Palta",
3042978,
"qfa-unc", -- pupus, tidak terkelas; mungkin Chicham
"Latn",
}
m["sai-pam"] = {
"Pamigua",
5908689,
"sai-tin",
"Latn",
}
m["sai-par"] = {
"Paratió",
16890038,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Xukuruan
"Latn",
}
m["sai-peb"] = {
"Peba",
3373890,
"sai-pey",
"Latn",
}
m["sai-pnz"] = {
"Panzaleo",
3123275,
"qfa-unc", -- pupus, tidak terkelas; mungkin Paezan
"Latn",
}
m["sai-prh"] = {
"Puruhá",
3410994,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin dalam keluarga dengan Cañari
"Latn",
}
m["sai-ptg"] = {
"Patagón",
128807870,
"sai-tar", -- pupus, hanya diketahui daripada 4 perkataan, yang mencadangkan susur galur Cariban (Campbell 2024)
"Latn",
}
m["sai-pur"] = {
"Purukotó",
7261622,
"sai-pem",
"Latn",
}
m["sai-pyg"] = {
"Payaguá",
7156643,
"sai-guc",
"Latn",
}
m["sai-pyk"] = {
"Pykobjê",
98113977,
"sai-nje",
"Latn",
}
m["sai-qmb"] = {
"Quimbaya",
7272043,
"qfa-unc", -- pupus, mungkin tidak wujud; sedikit perkataan yang diketahui
"Latn",
}
m["sai-qtm"] = {
"Quitemo",
7272651,
"sai-cpc",
"Latn",
}
m["sai-rab"] = {
"Rabona",
6644203,
"qfa-unc", -- pupus, kurang dibuktikan, kebanyakan nama tumbuhan; mungkin Candoshi-Shapra
"Latn",
}
m["sai-ram"] = {
"Ramanos",
16902824,
"qfa-unc", -- pupus, kurang dibuktikan, mungkin pencilan; mengikut Glottolog: "senarai perkataan yang kerdil ... tidak menunjukkan persamaan yang meyakinkan dengan bahasa sekeliling"
"Latn",
}
m["sai-sac"] = {
"Sácata",
5390321,
"qfa-unc", -- pupus, hanya 3 perkataan diketahui; mungkin Candoshí atau Arawak
"Latn",
}
m["sai-san"] = {
"Sanaviron",
16895999,
"qfa-unc", -- pupus, tidak terkelas; tiada konsensus mengenai pengelasan
"Latn",
}
m["sai-sap"] = {
"Sapará",
7420922,
"sai-car",
"Latn",
}
m["sai-sec"] = {
"Sechura",
7442912,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan
"Latn",
}
m["sai-sin"] = {
"Sinúfana",
7525275,
"qfa-unc", -- hampir pupus, kurang dibuktikan; mungkin Chocoan
"Latn",
}
m["sai-sje-pro"] = {
"Jê Selatan Purba",
116773814,
"sai-sje",
"Latn",
type = "reconstructed",
}
m["sai-tab"] = {
"Tabancale",
5390321,
"qfa-unc", -- pupus, hanya 5 perkataan diketahui; tiada kaitan yang jelas, mungkin pencilan
"Latn",
}
m["sai-tal"] = {
"Tallán",
16910468,
"qfa-unc", -- pupus, kurang dibuktikan; mungkin Catacaoan
"Latn",
}
m["sai-tap"] = {
"Tapayuna",
30719984,
"sai-nje",
"Latn",
}
m["sai-tar-pro"] = {
"Taranoan Purba",
116773816,
"sai-tar",
"Latn",
type = "reconstructed",
}
m["sai-teu"] = {
"Teushen",
3519243,
"qfa-unc", -- mungkin pupus menjelang 1950-an; mungkin Chonan
"Latn",
}
m["sai-tim"] = {
"Timote",
7806995,
nil, -- mungkin dalam keluarga Timote kecil
"Latn",
}
m["sai-tpr"] = {
"Taparita",
7684460,
"sai-otm",
"Latn",
}
m["sai-trr"] = {
"Tarairiú",
7685313,
"qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan
"Latn",
}
m["sai-wai"] = {
"Waitaká",
16918610,
"qfa-unc", -- pupus, mungkin Purian
"Latn",
}
m["sai-way"] = {
"Wayumara",
7960726,
"sai-car",
"Latn",
}
m["sai-wit-pro"] = {
"Witotoan Purba",
116773823,
"sai-wit",
"Latn",
type = "reconstructed",
}
m["sai-wnm"] = {
"Wanham",
16879440,
"sai-cpc",
"Latn",
}
m["sai-xoc"] = { -- kontras xoo, kzw, sai-kat
"Xocó",
12953620,
"qfa-unc", -- pupus dan kurang dibuktikan; tidak jelas sama ada satu atau tiga bahasa
"Latn",
}
m["sai-yao"] = {
"Yao (Amerika Selatan)",
16979655,
"sai-ven",
"Latn",
}
m["sai-yar"] = { -- bukan keluarga yang sama dengan 'suy'
"Yarumá",
3505859,
"sai-pek",
"Latn",
}
m["sai-yri"] = {
"Yuri",
2669157,
"sai-tyu",
"Latn",
}
m["sai-yup"] = {
"Yupua",
8061430,
"sai-tuc",
"Latn",
}
m["sai-yur"] = {
"Yurumanguí",
1281291,
"qfa-unc", -- pupus, terlalu kurang dibuktikan untuk dikelaskan
"Latn",
}
m["sal-pro"] = {
"Salish Purba",
116773269,
"sal",
"Latn",
type = "reconstructed",
}
m["sdv-daj-pro"] = {
"Daju Purba",
116773739,
"sdv-daj",
"Latn",
type = "reconstructed",
}
m["sdv-eje-pro"] = {
"Jebel Timur Purba",
116773751,
"sdv-eje",
"Latn",
type = "reconstructed",
}
m["sdv-nil-pro"] = {
"Nilotik Purba",
116773794,
"sdv-nil",
"Latn",
type = "reconstructed",
}
m["sdv-nyi-pro"] = {
"Nyima Purba",
116773796,
"sdv-nyi",
"Latn",
type = "reconstructed",
}
m["sdv-tmn-pro"] = {
"Taman Purba",
116773815,
"sdv-tmn",
"Latn",
type = "reconstructed",
}
m["sel-nor"] = {
"Selkup Utara",
30304565,
"sel",
"Cyrl",
translit = "sel-nor-translit",
}
m["sel-pro"] = {
"Selkup Purba",
128884235,
"sel",
"Latn",
type = "reconstructed",
}
m["sel-sou"] = {
"Selkup Selatan",
30304639,
"sel",
"Cyrl",
translit = "sel-sou-translit",
}
m["sem-amm"] = {
"Ammon",
279181,
"sem-can",
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]]
}
m["sem-amo"] = {
"Amor",
35941,
"sem-nwe",
"Xsux, Latn",
}
m["sem-cha"] = {
"Chaha",
35543,
"sem-eth",
"Ethi",
translit = "Ethi-translit",
}
m["sem-dad"] = {
"Dadan",
21838040,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-dum"] = {
"Dumait",
128810397,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-has"] = {
"Hasait",
3541433,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-his"] = {
"Hisma",
22948260,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-mhr"] = {
"Muher",
33743,
"sem-eth",
"Latn",
}
m["sem-pro"] = {
"Samiah Purba",
1658554,
"sem",
"Latn",
type = "reconstructed",
}
m["sem-saf"] = {
"Safait",
472586,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-sam"] = {
"Samal",
85847147,
"sem-nwe",
"Phnx",
-- translit Phnx dalam [[Module:scripts/data]]
}
m["sem-srb"] = {
"Arab Selatan Kuno",
35025,
"sem-osa",
"Sarb",
-- translit Sarb dalam [[Module:scripts/data]]
}
m["sem-tay"] = {
"Tayman",
24912301,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-tha"] = {
"Thamud",
843030,
"sem-cen",
"Narb",
-- translit Narb dalam [[Module:scripts/data]]
}
m["sem-wes-pro"] = {
"Samiah Barat Purba",
98021726,
"sem-wes",
"Latn",
type = "reconstructed",
}
m["sio-pro"] = { -- PERHATIAN ini bukan 'nai-sca-pro' "Proto-Siouan-Catawban" iaitu Proto-Sioux Barat
"Sioux Purba",
34181,
"sio",
"Latn",
type = "reconstructed",
}
m["sit-aao-pro"] = {
"Naga Tengah Purba",
nil,
"sit-aao",
"Latn",
type = "reconstructed",
}
m["sit-bai-pro"] = {
"Bai Purba",
nil,
"sit-bai",
"Latn",
type = "reconstructed",
}
m["sit-ban"] = {
"Bangru",
56071779,
"sit-hrs",
"Latn",
}
m["sit-bdi-pro"] = {
"Bodish Purba",
nil,
"sit-bdi",
"Latn",
type = "reconstructed",
}
m["sit-bok"] = {
"Bokar",
4938727,
"sit-tan",
"Latn, Tibt",
override_translit = true,
-- translit, display_text, strip_diacritics, sort_key Tibt dalam [[Module:scripts/data]]
}
m["sit-cai"] = {
"Caijia",
5017528,
"sit-cln",
"Latn"
}
m["sit-cha"] = {
"Chairel",
5068066,
"sit-luu",
"Latn",
}
m["sit-ers-pro"] = {
"Ersu Purba",
nil,
"sit-ers",
"Latn",
type = "reconstructed",
}
m["sit-hrs-pro"] = {
"Hrusish Purba",
116773762,
"sit-hrs",
"Latn",
type = "reconstructed",
}
m["sit-jap"] = {
"Japhug",
3162245,
"sit-egy",
"Latn",
}
m["sit-kha-pro"] = {
"Kham Purba",
116773773,
"sit-kha",
"Latn",
type = "reconstructed",
}
m["sit-khb-pro"] = {
"Kho-Bwa Purba",
nil,
"sit-khb",
"Latn",
type = "reconstructed",
}
m["sit-khp-pro"] = {
"Puroik Purba",
nil,
"sit-khb",
"Latn",
type = "reconstructed",
}
m["sit-khw-pro"] = {
"Kho-Bwa Barat Purba",
nil,
"sit-khw",
"Latn",
type = "reconstructed",
}
m["sit-kon-pro"] = {
"Naga Utara Purba",
nil,
"sit-kon",
"Latn",
type = "reconstructed",
}
m["sit-liz"] = {
"Lizu",
6660653,
"sit-ers",
"Latn", -- dan Ersu Shaba
}
m["sit-lnj"] = {
"Longjia",
17096251,
"sit-cln",
"Latn"
}
m["sit-lrn"] = {
"Luren",
16946370,
"sit-cln",
"Latn"
}
m["sit-luu-pro"] = {
"Luish Purba",
116773783,
"sit-luu",
"Latn",
type = "reconstructed",
}
m["sit-nas-pro"] = {
"Naish Purba",
nil,
"sit-nas",
"Latn",
type = "reconstructed",
}
m["sit-prn"] = {
"Puiron",
7259048,
"sit-zem",
}
m["sit-pro"] = {
"Sino-Tibet Purba",
24839178,
"sit",
"Latn",
type = "reconstructed",
}
m["sit-sit"] = {
"Situ",
19840830,
"sit-egy",
"Latn",
}
m["sit-tam-pro"] = {
"Tamang Purba",
117469295,
"sit-tam",
"Latn",
type = "reconstructed",
}
m["sit-tan-pro"] = {
"Tani Purba",
116773284,
"sit-tan",
"Latn", -- memerlukan pengesahan
type = "reconstructed",
}
m["sit-tgm"] = {
"Tangam",
17041370,
"sit-tan",
"Latn",
}
m["sit-tng-pro"] = {
"Tangkhul Purba",
nil,
"sit-tng",
"Latn",
type = "reconstructed",
}
m["sit-tos"] = {
"Tosu",
7827899,
"sit-ers",
"Latn", -- juga Ersu Shaba
}
m["sit-tsh"] = {
"Tshobdun",
19840950,
"sit-egy",
"Latn",
}
m["sit-zbu"] = {
"Zbu",
19841106,
"sit-egy",
"Latn",
}
m["sla-pro"] = {
"Slav Purba",
747537,
"sla",
"Latn",
type = "reconstructed",
strip_diacritics = {
remove_diacritics = c.grave .. c.acute .. c.tilde .. c.macron .. c.dgrave .. c.invbreve,
remove_exceptions = {'ś'},
},
sort_key = {
from = {"č", "ď", "ě", "ę", "ь", "ľ", "ň", "ǫ", "ř", "š", "ś", "ť", "ъ", "ž"},
to = {"c²", "d²", "e²", "e³", "i²", "l²", "nj", "o²", "r²", "s²", "s³", "t²", "u²", "z²"},
}
}
m["smi-pro"] = {
"Sami Purba",
7251862,
"smi",
"Latn",
type = "reconstructed",
sort_key = {
from = {"ā", "č", "δ", "[ëē]", "ŋ", "ń", "ō", "š", "θ", "%([^()]+%)"},
to = {"a", "c²", "d", "e", "n²", "n³", "o", "s²", "t²"}
},
}
m["son-pro"] = {
"Songhai Purba",
116773277,
"son",
"Latn",
type = "reconstructed",
}
m["sqj-pro"] = {
"Albania Purba",
18210846,
"sqj",
"Latn",
type = "reconstructed",
}
m["ssa-klk-pro"] = {
"Kuliak Purba",
116773779,
"ssa-klk",
"Latn",
type = "reconstructed",
}
m["ssa-kom-pro"] = {
"Koma Purba",
116773775,
"ssa-kom",
"Latn",
type = "reconstructed",
}
m["ssa-pro"] = {
"Nilo-Sahara Purba",
116773236,
"ssa",
"Latn",
type = "reconstructed",
}
m["syd-pro"] = {
"Samoyed Purba",
7251863,
"syd",
"Latn",
type = "reconstructed",
}
m["tai-pro"] = {
"Tai Purba",
6583709,
"tai",
"Latn",
type = "reconstructed",
}
m["tai-swe-pro"] = {
"Tai Barat Daya Purba",
116773280,
"tai-swe",
"Latn",
type = "reconstructed",
}
m["tbq-bdg-pro"] = {
"Bodo-Garo Purba",
116773195,
"tbq-bdg",
"Latn",
type = "reconstructed",
}
m["tbq-blg"] = {
"Bailang",
2879843,
"tbq-lob",
"Hani",
sort_key = "Hani-sortkey",
}
m["tbq-brm-pro"] = {
"Burma Purba",
nil,
"tbq-brm",
"Latn",
type = "reconstructed",
}
m["tbq-gkh"] = {
"Gokhy",
5578069,
"tbq-sil",
"Latn",
}
m["tbq-kuk-pro"] = {
"Kuki-Chin Purba",
116773220,
"tbq-kuk",
"Latn",
type = "reconstructed",
}
m["tbq-lal-pro"] = {
"Lalo Purba",
116773781,
"tbq-lal",
"Latn",
type = "reconstructed",
}
m["tbq-laz"] = {
"Laze",
17007626,
"sit-nas",
"Latn",
}
m["tbq-lob-pro"] = {
"Lolo-Burma Purba",
116773224,
"tbq-lob",
"Latn",
type = "reconstructed",
}
m["tbq-lol-pro"] = {
"Lolo Purba",
7251855,
"tbq-lol",
"Latn",
type = "reconstructed",
}
m["tbq-mil"] = {
"Milang",
6850761,
"sit-gsi",
"Deva, Latn",
}
m["tbq-mor"] = {
"Moran",
6909216,
"tbq-bdg",
"Latn",
}
m["tbq-ngo"] = {
"Ngochang",
56582,
"tbq-brm",
"Latn",
}
-- tbq-pro kini khusus etimologi
m["trk-dkh"] = {
"Dukhan",
12809273,
"trk-ssb",
"Latn, Cyrl, Mong",
-- translit, display_text dan strip_diacritics Mong dalam [[Module:scripts/data]]
}
-- Seperti yang diuraikan dalam ''Dīwān Lughāt al-Turk'' karya Mahmud al-Kashgari abad ke-11.
m["trk-eog"] = {
"Oghuz Kuno Awal",
nil,
"trk-ogz",
"Arab",
strip_diacritics = {Arab = "ar-stripdiacritics"},
}
m["trk-oat"] = {
"Turki Anatolia Kuno",
7083390,
"trk-ogz",
"Arab",
strip_diacritics = {Arab = "ar-stripdiacritics"},
ancestors = "trk-eog",
}
m["trk-pro"] = {
"Turkik Purba",
3657773,
"trk",
"Latn",
type = "reconstructed",
standard_chars = {
Latn = " ()-abdegiklmnoprstuxyzïöüāčēīĺŋōŕšūǖȫẹ" .. c.macron,
}
}
m["tup-gua-pro"] = {
"Tupi-Guarani Purba",
116773288,
"tup-gua",
"Latn",
type = "reconstructed",
}
m["tup-kab"] = {
"Kabishiana",
15302988,
"tup",
"Latn",
}
m["tup-kaw"] = {
"Kawahiva",
6346712,
"tup-gua",
"Latn",
}
m["tup-pro"] = {
"Tupi Purba",
10354700,
"tup",
"Latn",
type = "reconstructed",
}
m["tuw-alk"] = {
"Alchuka",
113553616,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-bal"] = {
"Bala",
86730632,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-kkl"] = {
"Kyakala",
118875708,
"tuw-jrc",
"Latn, Hans",
sort_key = {Hans = "Hani-sortkey"},
}
m["tuw-kli"] = {
"Kili",
6406892,
"tuw-ewe",
"Cyrl",
}
m["tuw-pro"] = {
"Tungus Purba",
85872335,
"tuw",
"Latn",
type = "reconstructed",
}
m["tuw-sol"] = {
"Solon",
30004,
"tuw-ewe",
}
m["urj-fin-pro"] = {
"Finnik Purba",
11883720,
"urj-fin",
"Latn",
type = "reconstructed",
}
m["urj-koo"] = {
"Komi Kuno",
86679962,
"kv",
"Perm, Cyrs",
translit = "urj-koo-translit",
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]; sebelum ini, strip_diacritics Cyrs tidak hadir
}
m["urj-kuk"] = {
"Kukkuzi",
107410460,
"urj-fin",
"Latn",
ancestors = "vot",
}
m["urj-kya"] = {
"Komi-Yazva",
2365210,
"kv",
"Cyrl",
translit = "kv-translit",
override_translit = true,
strip_diacritics = {remove_diacritics = c.acute},
}
m["urj-mdv-pro"] = {
"Mordvinik Purba",
116773232,
"urj-mdv",
"Latn",
type = "reconstructed",
}
m["urj-prm-pro"] = {
"Permik Purba",
116773257,
"urj-prm",
"Latn",
type = "reconstructed",
}
m["urj-pro"] = {
"Uralik Purba",
288765,
"urj",
"Latn",
type = "reconstructed",
}
m["urj-ugr-pro"] = {
"Ugrik Purba",
156631,
"urj-ugr",
"Latn",
type = "reconstructed",
}
m["xnd-pro"] = {
"Na-Dene Purba",
116773233,
"xnd",
"Latn",
type = "reconstructed",
}
m["xgn-pro"] = {
"Mongol Purba",
2493677,
"xgn",
"Latn",
type = "reconstructed",
sort_key = {
from = {"č", "i", "ï", "ǰ", "ŋ", "ö", "š", "ü"},
to = {"c", "i" .. p[1], "i", "j", "n" .. p[1], "o" .. p[1], "s" .. p[1], "u" .. p[1]},
},
}
m["yok-bvy"] = {
"Yokuts Buena Vista",
4985474,
"yok",
"Latn",
}
m["yok-dly"] = {
"Yokuts Delta",
70923266,
"yok",
"Latn",
}
m["yok-gsy"] = {
"Yokuts Gashowu",
3098708,
"yok",
"Latn",
}
m["yok-kry"] = {
"Yokuts Sungai Kings",
6413014,
"yok",
"Latn",
}
m["yok-nvy"] = {
"Yokuts Lembah Utara",
85789777,
"yok",
"Latn",
}
m["yok-ply"] = {
"Yokuts Palewyami",
2387391,
"yok",
"Latn",
}
m["yok-svy"] = {
"Yokuts Lembah Selatan",
12642473,
"yok",
"Latn",
}
m["yok-tky"] = {
"Yokuts Tule-Kaweah",
7851988,
"yok",
"Latn",
}
m["ypk-pro"] = {
"Yupik Purba",
116773295,
"ypk",
"Latn",
type = "reconstructed",
}
m["yrk-for"] = {
"Nenets Hutan",
1295107,
"yrk",
"Cyrl",
translit = "yrk-for-translit",
strip_diacritics = {remove_diacritics = c.grave .. c.acute .. c.macron .. c.breve .. c.dotabove},
}
m["yrk-tun"] = {
"Nenets Tundra",
36452,
"yrk",
"Cyrl",
strip_diacritics = {
from = {"ӑ", "а̄", "э̇", "ӣ", "ы̄", "ӯ", "ю̄", "я̆", "я̄"},
to = {"а", "а", "э", "и", "ы", "у", "ю", "я", "я"},
},
translit = "yrk-tun-translit",
}
m["zhx-min-pro"] = {
"Min Purba",
19646347,
"zhx-min",
"Latn",
type = "reconstructed",
}
m["zhx-sht"] = {
"Tuhua Shaozhou",
1920769,
"zhx",
"Nshu, Hants",
generate_forms = "zh-generateforms",
sort_key = {Hani = "Hani-sortkey"},
}
m["zhx-sic"] = {
"Sichuan",
2278732,
"zhx-man",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["zhx-tai"] = {
"Taishan",
2208940,
"zhx-yue",
"Hants",
generate_forms = "zh-generateforms",
translit = "zh-translit",
sort_key = "Hani-sortkey",
}
m["zle-ono"] = {
"Novgorod Kuno",
162013,
"zle",
"Cyrs, Glag",
translit = {Cyrs = "Cyrs-translit", Glag = "Glag-translit"},
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zle-ort"] = {
"Ruthenia Kuno",
13211,
"zle",
"Arab, Cyrs, Latn",
ancestors = "orv",
translit = {
Cyrs = "zle-ort-translit",
Arab = "zle-ort-Arab-translit",
},
strip_diacritics = {
Cyrs = {
remove_diacritics = m_langdata.chars_substitutions["Cyrs_remove_diacritics"],
remove_exceptions = {"Ї", "ї"},
},
Arab = "ar-stripdiacritics",
},
-- sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zls-chs"] = {
"Slav Gereja",
33251,
"zls",
"Cyrs, Glag, Latn",
ancestors = "cu",
translit = {
Cyrs = "Cyrs-translit",
Glag = "Glag-translit"
},
-- strip_diacritics, sort_key Cyrs dalam [[Module:scripts/data]]
}
m["zlw-ocs"] = {
"Czech Kuno",
593096,
"zlw",
"Latn",
}
m["zlw-opl"] = {
"Poland Kuno",
149838,
"zlw-lch",
"Latn",
strip_diacritics = {remove_diacritics = c.ringabove},
}
m["zlw-osk"] = {
"Slovak Kuno",
12776676,
"zlw",
"Latn",
}
m["zlw-slv"] = {
"Slovincia",
36822,
"zlw-pom",
"Latn",
strip_diacritics = {remove_diacritics = c.macron .. c.breve},
}
-- Kod tambahan untuk bahasa-bahasa yang digunakan di Malaysia, yang tidak wujud di Wikikamus Bahasa Inggeris
m["zlm-coa"] = {
"Melayu Terengganu Pesisir",
4207412,
"poz-mly",
"Latn, ms-Arab",
}
m["zlm-pah"] = {
"Melayu Pahang",
7310370,
"poz-mly",
"Latn",
}
return require("Module:languages").finalizeData(m, "language")
dluvfqb7mu3yt1ue6qh0s4xr04usufo
Modul:languages/data/3/k/extra
828
33767
375378
373746
2026-09-22T05:36:05Z
Hakimi97
2668
[[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]]
375378
Scribunto
text/plain
local m = {}
m["kaa"] = {
aliases = {"Qaraqalpaq"},
}
m["kab"] = {
aliases = {"Kabylian"},
}
m["kac"] = {
aliases = {"Kachin"},
}
m["kad"] = {
}
m["kae"] = {
}
m["kaf"] = {
otherNames = {"Kazhuo"},
}
m["kag"] = {
}
m["kah"] = {
otherNames = {"Kara"},
}
m["kai"] = {
}
m["kaj"] = {
}
m["kak"] = {
}
m["kam"] = {
otherNames = {"Kikamba", "Kamba (Kenya)"},
}
m["kao"] = {
otherNames = {"Khasonke", "Kasonke", "Khassonké"},
}
m["kap"] = {
otherNames = {"Bezheta", "Kapucha", "Bezhita"},
}
m["kaq"] = {
otherNames = {"Kapanawa"},
}
m["kaw"] = {
aliases = {"Kawi"},
}
m["kax"] = {
}
m["kay"] = {
}
m["kba"] = {
}
m["kbb"] = {
otherNames = {"Kachuyana", "Kaxuiana", "Kaxuiâna", "Kashuyana"},
}
m["kbc"] = {
otherNames = {"Caduveo", "Ediu-Adig", "Guaicurú", "Kadiweu", "Mbayá", "Mbayá-Guaycuru", "Waikurú"},
}
m["kbd"] = {
aliases = {"East Circassian"},
}
m["kbe"] = {
otherNames = {"Kaanytju", "Kandju", "Kaantyu", "Gandju", "Gandanju", "Kamdhue", "Kandyu", "Kanyu"},
}
m["kbh"] = {
}
m["kbi"] = {
}
m["kbj"] = {
otherNames = {"Kare", "Kare (Central African Republic)", "Bantoid Kare"},
}
m["kbk"] = {
otherNames = {"Koiari"},
}
m["kbm"] = {
}
m["kbn"] = {
otherNames = {"Kare (Central African Republic)", "Mbum Kare"},
}
m["kbo"] = {
}
m["kbp"] = {
otherNames = {"Kabiye", "Kabye"},
}
m["kbq"] = {
}
m["kbr"] = {
}
m["kbs"] = {
}
m["kbt"] = {
}
m["kbu"] = {
}
m["kbv"] = {
otherNames = {"Dera", "Dera (New Guinea)"},
}
m["kbw"] = {
}
m["kbx"] = {
}
m["kbz"] = {
}
m["kcb"] = {
}
m["kcc"] = {
}
m["kcd"] = {
}
m["kce"] = {
}
m["kcf"] = {
}
m["kcg"] = {
}
m["kch"] = {
}
m["kci"] = {
}
m["kcj"] = {
}
m["kck"] = {
}
m["kcl"] = {
otherNames = {"Kela", "Gela"},
}
m["kcm"] = {
}
m["kcn"] = {
otherNames = {"Ki-Nubi"},
}
m["kco"] = {
}
m["kcp"] = {
}
m["kcq"] = {
}
m["kcr"] = {
}
m["kcs"] = {
}
m["kct"] = {
}
m["kcu"] = {
otherNames = {"Kami"},
}
m["kcv"] = {
}
m["kcw"] = {
}
m["kcx"] = {
}
m["kcy"] = {
}
m["kcz"] = {
}
m["kda"] = {
otherNames = {"Gadang", "Gadhang", "Gadjang", "Kattang", "Kutthung"},
}
m["kdc"] = {
}
m["kdd"] = {
}
m["kde"] = {
}
m["kdf"] = {
}
m["kdg"] = {
}
m["kdh"] = {
}
m["kdi"] = {
otherNames = {"Kuman"},
}
m["kdj"] = {
}
m["kdk"] = {
}
m["kdl"] = {
}
m["kdm"] = {
}
m["kdn"] = {
}
m["kdp"] = {
}
m["kdq"] = {
}
m["kdr"] = {
}
m["kdt"] = {
}
m["kdu"] = {
otherNames = {"Kedaru", "Debri"}, -- Debri is subsumed for now as it lacks an ISO code, may need to be split
}
m["kdv"] = {
otherNames = {"Kadu"},
}
m["kdw"] = {
}
m["kdx"] = {
}
m["kdy"] = {
}
m["kdz"] = {
otherNames = {"Ndaktup", "Ncha", "Bitwi"},
}
m["kea"] = {
otherNames = {"Cape Verdean Creole", "Kriolu", "Creole", "Barlavento", "Sotavento"},
}
m["keb"] = {
}
m["kec"] = {
}
m["ked"] = {
}
m["kee"] = {
}
m["kef"] = {
}
m["keg"] = {
}
m["keh"] = {
}
m["kei"] = {
}
m["kej"] = {
}
m["kek"] = {
aliases = {"Qʼeqchi"}
}
m["kel"] = {
otherNames = {"Kela", "Yela"},
}
m["kem"] = {
}
m["ken"] = {
}
m["keo"] = {
}
m["kep"] = {
}
m["keq"] = {
}
m["ker"] = {
}
m["kes"] = {
}
m["ket"] = {
}
m["keu"] = {
}
m["kev"] = {
}
m["kew"] = {
otherNames = {"West Kewa", "East Kewa", "South Kewa", "Erave", "Pasuma"},
}
m["kex"] = {
}
m["key"] = {
}
m["kez"] = {
}
m["kfa"] = {
}
m["kfb"] = {
}
m["kfc"] = {
}
m["kfd"] = {
}
m["kfe"] = {
otherNames = {"Kota"},
}
m["kff"] = {
}
m["kfg"] = {
}
m["kfh"] = {
}
m["kfi"] = {
}
m["kfj"] = {
}
m["kfk"] = {
}
m["kfl"] = {
}
m["kfn"] = {
}
m["kfo"] = {
otherNames = {"Koro", "Koro Jula"}, -- the last name is misleading, as Jula is a diff. language
}
m["kfp"] = {
}
m["kfq"] = {
}
m["kfr"] = {
aliases = {"Kutchi", "Cutchi", "Kachchhi", "Kutchhi"},
}
m["kfs"] = {
}
m["kft"] = {
}
m["kfu"] = {
}
m["kfv"] = {
}
m["kfw"] = {
otherNames = {"Kharam"},
}
m["kfx"] = {
otherNames = {"Kullu"},
}
m["kfy"] = {
}
m["kfz"] = {
}
m["kga"] = {
}
m["kgb"] = {
}
m["kgd"] = {
}
m["kge"] = {
}
m["kgf"] = {
}
m["kgg"] = {
}
m["kgi"] = {
}
m["kgj"] = {
}
m["kgk"] = {
}
m["kgl"] = {
}
m["kgn"] = {
otherNames = {"Keringani"},
}
m["kgo"] = {
}
m["kgp"] = {
}
m["kgq"] = {
}
m["kgr"] = {
}
m["kgs"] = {
}
m["kgt"] = {
}
m["kgu"] = {
}
m["kgv"] = {
}
m["kgw"] = {
}
m["kgx"] = {
}
m["kgy"] = {
}
m["kha"] = {
}
m["khb"] = {
aliases = {"Lue", "Tai Lü", "Tai Lue", "Dai Lue"},
}
m["khc"] = {
}
m["khd"] = {
}
m["khe"] = {
}
m["khf"] = {
}
m["khh"] = {
}
m["khj"] = {
}
m["khl"] = {
}
m["kho"] = {
}
m["khp"] = {
}
m["khq"] = {
otherNames = {"Western Songhay", "Koyra Chiini Songhay"},
}
m["khr"] = {
}
m["khs"] = {
}
m["kht"] = {
aliases = {"Tai Khamti"},
}
m["khu"] = {
}
m["khv"] = {
otherNames = {"Khwarshi", "Xvarshi", "Inkhokvari"},
}
m["khw"] = {
}
m["khx"] = {
}
m["khy"] = {
otherNames = {"Kele", "Kele (Congo)", "Kele (Democratic Republic of the Congo)", "Lokele"},
}
m["khz"] = {
}
m["kia"] = {
}
m["kib"] = {
}
m["kic"] = {
}
m["kid"] = {
}
m["kie"] = {
}
m["kif"] = {
}
m["kig"] = {
}
m["kih"] = {
}
m["kii"] = {
otherNames = {"Kichai"},
}
m["kij"] = {
}
m["kil"] = {
}
m["kim"] = {
otherNames = {"Tofalar", "Karagas"},
}
m["kio"] = {
}
m["kip"] = {
}
m["kiq"] = {
}
m["kis"] = {
}
m["kit"] = {
}
m["kiv"] = {
}
m["kiw"] = {
}
m["kix"] = {
}
m["kiy"] = {
otherNames = {"Faia"},
}
m["kiz"] = {
}
m["kja"] = {
}
m["kjb"] = {
}
m["kjc"] = {
}
m["kjd"] = {
}
m["kje"] = {
}
m["kjg"] = {
}
m["kjh"] = {
}
m["kji"] = {
}
m["kjj"] = {
otherNames = {"Khinalig", "Xinalug", "Xinalugh", "Khinalugh"},
}
m["kjk"] = {
}
m["kjl"] = {
}
m["kjm"] = {
}
m["kjn"] = {
otherNames = {"Uw Oykangand", "Uw Olkola", "Olkol", "Olgolo", "Uw-Oykangand", "Uw-Olgol", "Koko Wanggara", "Ogh-Undjan", "Undjan", "Kawarrangg", "Athima", "Uw", "Kunjen-Undjan-Athima"},
}
m["kjo"] = {
}
m["kjp"] = {
aliases = {"Phlou", "Eastern Pwo Karen"},
}
m["kjq"] = {
}
m["kjr"] = {
}
m["kjs"] = {
}
m["kjt"] = {
aliases = {"Phrae Pwo Karen", "Northeastern Pwo", "Northeastern Pwo Karen"},
}
m["kju"] = {
}
m["kjx"] = {
otherNames = {"Keriaka"},
}
m["kjy"] = {
}
m["kjz"] = {
}
m["kka"] = {
}
m["kkb"] = {
}
m["kkc"] = {
}
m["kkd"] = {
}
m["kke"] = {
}
m["kkf"] = {
}
m["kkg"] = {
}
m["kkh"] = {
aliases = {"Tai Khün", "Dai Kun"},
}
m["kki"] = {
otherNames = {"Kaguru"},
}
m["kkj"] = {
}
m["kkk"] = {
}
m["kkl"] = {
}
m["kkm"] = {
}
m["kkn"] = {
}
m["kko"] = {
otherNames = {"Kithonirishe"},
}
m["kkp"] = {
otherNames = {"Kok-Kaper", "Gugubera", "Koko-Pera"},
}
m["kkq"] = {
}
m["kkr"] = {
otherNames = {"Kir"},
}
m["kks"] = {
otherNames = {"Giiwo"},
}
m["kkt"] = {
}
m["kku"] = {
}
m["kkv"] = {
}
m["kkw"] = {
}
m["kkx"] = {
}
m["kky"] = {
}
m["kkz"] = {
}
m["kla"] = {
otherNames = {"Klamath"},
}
m["klb"] = {
}
m["klc"] = {
}
m["kld"] = {
otherNames = {"Kamilaroi", "Kamilarai", "Kamalarai", "Gamilaroi"},
}
m["kle"] = {
}
m["klf"] = {
}
m["klg"] = {
}
m["klh"] = {
}
m["kli"] = {
}
m["klj"] = {
otherNames = {"Turkic Khalaj", "Arghu"},
}
m["klk"] = {
otherNames = {"Kono"},
}
m["kll"] = {
}
m["klm"] = {
otherNames = {"Migum"},
}
m["kln"] = {
}
m["klo"] = {
}
m["klp"] = {
}
m["klq"] = {
}
m["klr"] = {
}
m["kls"] = {
}
m["klt"] = {
}
m["klu"] = {
}
m["klv"] = {
}
m["klw"] = {
otherNames = {"Tado"},
}
m["klx"] = {
}
m["kly"] = {
}
m["klz"] = {
}
m["kma"] = {
}
m["kmb"] = {
otherNames = {"North Mbundu"},
}
m["kmc"] = {
aliases = {"Southern Gam", "Southern Dong"},
}
m["kmd"] = {
}
m["kme"] = {
}
m["kmf"] = {
otherNames = {"Kare", "Kare (Papua New Guinea)"},
}
m["kmg"] = {
}
m["kmh"] = {
}
m["kmi"] = {
}
m["kmj"] = {
otherNames = {"Kumarbhag", "Kumarbhag Pahariya", "Kumar Paharia", "Malto"},
}
m["kmk"] = {
}
m["kml"] = {
otherNames = {"Lower Tanudan Kalinga", "Upper Tanudan Kalinga"},
}
m["kmm"] = {
otherNames = {"Kom"},
}
m["kmn"] = {
}
m["kmo"] = {
}
m["kmp"] = {
}
m["kmq"] = {
}
m["kmr"] = {
aliases = {"Kurmanji"},
}
m["kms"] = {
}
m["kmt"] = {
}
m["kmu"] = {
}
m["kmv"] = {
otherNames = {"Karipúna French Creole", "Amapá French Creole"},
}
m["kmw"] = {
otherNames = {"Kikomo", "Komo (Democratic Republic of the Congo)", "Komo", "Kikumu"},
}
m["kmx"] = {
}
m["kmy"] = {
}
m["kmz"] = {
otherNames = {"Khorasani Turkic"},
}
m["kna"] = {
otherNames = {"Dera", "Dera (Nigeria)"},
}
m["knb"] = {
}
m["knd"] = {
}
m["kne"] = {
aliases = {"Kankana-ey"},
}
m["knf"] = {
}
m["kni"] = {
}
m["knj"] = {
otherNames = {"Acateco", "Western Kanjobal"},
}
m["knk"] = {
}
m["knl"] = {
}
m["knm"] = { -- two unrelated lects have this name; this is the Katukinian one
otherNames = {"Kanamarí", "Katukina-Kanamari", "Kanamare", "Katukína", "Katukina"},
}
m["kno"] = {
otherNames = {"Kono", "Konnoh"},
}
m["knp"] = {
}
m["knq"] = {
}
m["knr"] = {
}
m["kns"] = {
}
m["knt"] = {
otherNames = {"Panoan Katukína", "Katukína", "Catuquina", "Waninawa", "Waninnawa", "Kamanawa", "Kamannaua", "Katukina do Jurua", "Katukina of Olinda", "Katukina of Sete Estreles", "Kanamari"},
}
m["knu"] = { -- a dialect of 'kpe'
otherNames = {"Kono"},
}
m["knv"] = {
}
m["knx"] = {
otherNames = {"Salako", "Selako", "Ahe"},
}
m["kny"] = {
}
m["knz"] = {
}
m["koa"] = {
}
m["koc"] = {
}
m["kod"] = {
}
m["koe"] = {
}
m["kof"] = {
}
m["kog"] = {
otherNames = {"Kogi", "Cogi", "Kagaba", "Cagaba", "Cágaba"},
}
m["koh"] = {
}
m["koi"] = {
}
m["kok"] = {
}
m["kol"] = {
otherNames = {"Kol", "Kol (Papua New Guina)"},
}
m["koo"] = {
}
m["kop"] = {
aliases = {"Waupe", "Kwato"},
}
m["koq"] = {
aliases = {"iKota", "Ikota", "Kota"},
}
m["kos"] = {
}
m["kot"] = {
aliases = {"Logone"},
}
m["kou"] = {
}
m["kov"] = {
}
m["kow"] = {
}
m["koy"] = {
otherNames = {"Denaakk'e"},
}
m["koz"] = {
}
m["kpa"] = {
}
m["kpb"] = {
}
m["kpc"] = {
otherNames = {"Kurripako"},
}
m["kpd"] = {
}
m["kpe"] = {
}
m["kpf"] = {
}
m["kpg"] = {
}
m["kph"] = {
}
m["kpi"] = {
}
m["kpj"] = {
}
m["kpk"] = {
}
m["kpl"] = {
}
m["kpm"] = {
aliases = {"K'Ho"},
}
m["kpn"] = {
}
m["kpo"] = {
}
m["kpq"] = {
}
m["kpr"] = {
}
m["kps"] = {
}
m["kpt"] = {
}
m["kpu"] = {
}
m["kpv"] = {
otherNames = {"Komi"},
}
m["kpw"] = {
}
m["kpx"] = {
otherNames = {"Mountain Koiali"},
}
m["kpy"] = {
}
m["kpz"] = {
}
m["kqa"] = {
}
m["kqb"] = {
}
m["kqc"] = {
}
m["kqd"] = {
}
m["kqe"] = {
}
m["kqf"] = {
}
m["kqg"] = {
}
m["kqh"] = {
}
m["kqi"] = {
}
m["kqj"] = {
}
m["kqk"] = {
}
m["kql"] = {
}
m["kqm"] = {
}
m["kqn"] = {
otherNames = {"Chikaonde", "Kawonde"},
}
m["kqo"] = {
}
m["kqp"] = {
}
m["kqq"] = {
}
m["kqr"] = {
}
m["kqs"] = {
}
m["kqt"] = {
}
m["kqu"] = {
}
m["kqv"] = {
}
m["kqw"] = {
}
m["kqx"] = {
}
m["kqy"] = {
}
m["kqz"] = {
}
m["kra"] = {
}
m["krb"] = {
}
m["krc"] = {
}
m["krd"] = {
}
m["kre"] = {
}
m["krf"] = {
otherNames = {"Koro"},
}
m["krh"] = {
}
m["kri"] = {
otherNames = {"Sierra Leonean Creole"},
}
m["krj"] = {
}
m["krk"] = {
}
m["krl"] = {
varieties = {
{ "North Karelian", "Northern Karelian" },
{ "South Karelian", "Southern Karelian" },
{ "Tver Karelian" }
}
}
m["krm"] = {
}
m["krn"] = {
}
m["krp"] = {
}
m["krr"] = {
otherNames = {"Krung", "Kreung", "Krüng"},
}
m["krs"] = {
otherNames = {"Gbaya"},
}
m["kru"] = {
aliases = {"Kurux"},
}
m["krv"] = {
otherNames = {"Kravet"},
}
m["krw"] = {
}
m["krx"] = {
}
m["kry"] = {
otherNames = {"Kryc", "Kryz"},
varieties = {"Jek", "Dzhek", "Cek", "Khaput", "Yergyudzh", "Alyk"},
}
m["krz"] = {
}
m["ksa"] = {
}
m["ksb"] = {
otherNames = {"Shambaa"},
}
m["ksc"] = {
}
m["ksd"] = {
otherNames = {"Kuanua"},
}
m["kse"] = {
}
m["ksf"] = {
}
m["ksg"] = {
}
m["ksi"] = {
}
m["ksj"] = {
}
m["ksk"] = {
}
m["ksl"] = {
}
m["ksm"] = {
}
m["ksn"] = {
}
m["kso"] = {
}
m["ksp"] = {
}
m["ksq"] = {
}
m["ksr"] = {
}
m["kss"] = {
}
m["kst"] = {
}
m["ksu"] = {
}
m["ksv"] = {
}
m["ksw"] = {
aliases = {"S'gaw Kayin", "S'gaw", "Sgaw", "White Karen"},
}
m["ksx"] = {
}
m["ksy"] = {
}
m["ksz"] = {
}
m["kta"] = {
}
m["ktb"] = {
}
m["ktc"] = {
}
m["ktd"] = {
}
m["ktf"] = {
}
m["ktg"] = {
otherNames = {"Kalkutungu", "Galgadungu", "Kalkutung", "Kalkadoon", "Galgaduun"},
}
m["kth"] = {
}
m["kti"] = {
otherNames = {"Kati"},
}
m["ktj"] = {
}
m["ktk"] = {
}
m["ktl"] = {
}
m["ktm"] = {
}
m["ktn"] = {
otherNames = {"Caritiana"},
}
m["kto"] = {
}
m["ktp"] = {
otherNames = {"Khatu"},
}
m["ktq"] = {
}
m["kts"] = {
}
m["ktt"] = {
}
m["ktu"] = {
otherNames = {"Munukutuba", "Kikongo-Kituba", "Kikongo", "Kikongo ya leta", "Kibulamatadi", "Kikwango", "Ikeleve", "Kizabave"},
}
m["ktv"] = {
}
m["ktw"] = {
otherNames = {"Cahto"},
}
m["ktx"] = {
}
m["kty"] = {
otherNames = {"Kango (Bas-Uélé District)"}, -- distinct in name, but not necessarily in identity, from 'kzy'
}
m["ktz"] = {
otherNames = {"Zhuǀ'hoan", "ǂKxʼauǁʼein", "ǁAuǁei", "ǁAuǁen", "Auen", "Kaukau", "Koko", "Kung-Gobabis", "‡Kx'auǁ'ei", "ǂKx'auǁ'ein", "ǁX'auǁ'e", "Juǀ'hoansi", "Juǀʼhoan"},
}
m["kub"] = {
}
m["kuc"] = {
}
m["kud"] = {
otherNames = {"'Auhelawa"},
}
m["kue"] = {
otherNames = {"Simbu", "Chimbu"},
}
m["kuf"] = {
}
m["kug"] = {
}
m["kuh"] = {
}
m["kui"] = {
otherNames = {"Kuikúro-Kalapálo", "Kuikuro", "Apalakiri"},
}
m["kuj"] = {
}
m["kuk"] = {
}
m["kul"] = {
otherNames = {"Tof", "Korom Boye", "Akandi", "Akande", "Kande", "Richa"},
}
m["kum"] = {
}
m["kun"] = {
}
m["kuo"] = {
}
m["kup"] = {
}
m["kuq"] = {
}
m["kus"] = {
}
m["kut"] = {
}
m["kuu"] = {
}
m["kuv"] = {
}
m["kuw"] = {
}
m["kux"] = {
}
m["kuy"] = {
}
m["kuz"] = {
}
m["kva"] = {
}
m["kvb"] = {
}
m["kvc"] = {
}
m["kvd"] = {
otherNames = {"Kui"},
}
m["kve"] = {
}
m["kvf"] = {
}
m["kvg"] = {
}
m["kvh"] = {
}
m["kvi"] = {
}
m["kvj"] = {
}
m["kvk"] = {
}
m["kvl"] = {
}
m["kvm"] = {
}
m["kvn"] = {
}
m["kvo"] = {
}
m["kvp"] = {
}
m["kvq"] = {
}
m["kvr"] = {
}
m["kvt"] = {
}
m["kvu"] = {
}
m["kvv"] = {
}
m["kvw"] = {
}
m["kvx"] = {
}
m["kvy"] = {
}
m["kvz"] = {
}
m["kwa"] = {
}
m["kwb"] = {
otherNames = {"Kwa"},
}
m["kwc"] = {
}
m["kwd"] = {
}
m["kwe"] = {
}
m["kwf"] = {
}
m["kwg"] = {
}
m["kwh"] = {
}
m["kwi"] = {
otherNames = {"Awa", "Cuaiquer", "Awa Pit", "Awapit", "Kwaiker", "Coaiquer", "Quaiquer"},
}
m["kwj"] = {
}
m["kwk"] = {
aliases = {"Kwakʼwala"}
}
m["kwl"] = {
}
m["kwm"] = {
}
m["kwn"] = {
}
m["kwo"] = {
}
m["kwp"] = {
}
m["kwq"] = {
}
m["kwr"] = {
}
m["kws"] = {
}
m["kwt"] = {
}
m["kwu"] = {
}
m["kwv"] = {
otherNames = {"Sara Dunjo"},
}
m["kww"] = {
}
m["kwx"] = {
}
m["kwz"] = {
}
m["kxa"] = {
}
m["kxb"] = {
}
m["kxc"] = {
}
m["kxd"] = {
aliases = {"Brunei"},
}
m["kxe"] = {
}
m["kxf"] = {
}
m["kxh"] = {
}
m["kxi"] = {
otherNames = {"Nabay", "Nabaay"},
}
m["kxj"] = {
}
m["kxk"] = {
}
m["kxm"] = {
aliases = {"Thai Khmer", "Surin Khmer"},
}
m["kxn"] = {
otherNames = {"Tanjong", "Kanowit-Tanjong Melanau"},
}
m["kxo"] = {
}
m["kxp"] = {
}
m["kxq"] = {
}
m["kxr"] = {
otherNames = {"Koro (Papua New Guinea)", "Koro"},
}
m["kxs"] = {
}
m["kxt"] = {
}
m["kxu"] = {
otherNames = {"Kui", "Kuy"},
}
m["kxv"] = {
}
m["kxw"] = {
}
m["kxx"] = {
}
m["kxy"] = {
}
m["kxz"] = {
}
m["kya"] = {
}
m["kyb"] = {
}
m["kyc"] = {
}
m["kyd"] = {
}
m["kye"] = {
}
m["kyf"] = {
}
m["kyg"] = {
}
m["kyh"] = {
otherNames = {"Karuk"},
}
m["kyi"] = {
}
m["kyj"] = {
}
m["kyk"] = {
}
m["kyl"] = {
}
m["kym"] = {
}
m["kyn"] = {
}
m["kyo"] = {
}
m["kyp"] = {
}
m["kyq"] = {
}
m["kyr"] = {
otherNames = {"Caravare", "Curuaia", "Kuruaia"},
}
m["kys"] = {
}
m["kyt"] = {
}
m["kyu"] = {
}
m["kyv"] = {
}
m["kyw"] = {
otherNames = {"Kurmali"},
}
m["kyx"] = {
otherNames = {"Konua"},
}
m["kyy"] = {
}
m["kyz"] = {
}
m["kza"] = {
}
m["kzb"] = {
}
m["kzc"] = {
}
m["kzd"] = {
}
m["kzf"] = {
otherNames = {"Tado", "Inde", "Pekava", "West Kaili"},
}
m["kzg"] = {
}
m["kzh"] = {
otherNames = {"Kenuzi-Dongola", "Andaandi", "Kenzi", "Mattoki"},
}
m["kzi"] = {
}
m["kzj"] = {
}
m["kzk"] = {
otherNames = {"Dororo", "Guliguli"},
}
m["kzl"] = {
}
m["kzm"] = {
}
m["kzn"] = {
}
m["kzo"] = {
}
m["kzp"] = {
}
m["kzq"] = {
}
m["kzr"] = {
aliases = {"Mbum East", "Lakka"},
}
m["kzs"] = {
}
m["kzt"] = {
}
m["kzu"] = {
}
m["kzv"] = {
}
m["kzw"] = { -- contrast xoo, sai-kat, sai-xoc, the last of which the ISO conflated into this code
otherNames = {"Kipeá", "Quipea", "Kamurú", "Camuru", "Dzubukuá", "Dzubucua", "Karirí", "Sabujá", "Sapoyá", "Pedra Branca"},
}
m["kzx"] = {
}
m["kzy"] = {
otherNames = {"Kango", "Kango (Tshopo District)"}, -- distinct in name, but not necessarily in identity, from 'kty'
}
m["kzz"] = {
}
return m
3noq5crf33bzu48k0iw2jk0rbdgjeup
Modul:languages/data/exceptional/extra
828
33778
375380
374164
2026-09-22T05:36:21Z
Hakimi97
2668
[[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]]
375380
Scribunto
text/plain
local m = {}
m["aav-khs-pro"] = {
aliases = {"Proto-Khasic"},
}
m["aav-nic-pro"] = {
}
m["aav-pkl-pro"] = {
}
m["aav-pro"] = { -- mkh-pro will merge into this.
}
m["afa-pro"] = {
aliases = {"Proto-Afro-Asiatic", "Hamito-Semitic"},
}
m["alg-aga"] = {
aliases = {"Agwam", "Agaam"},
}
m["alg-pro"] = {
}
m["alv-ama"] = {
}
m["alv-bgu"] = {
otherNames = {"Gubëeher", "Nyun Gubëeher", "Nun Gubëeher"},
}
m["alv-bua-pro"] = {
}
m["alv-cng-pro"] = {
}
m["alv-edk-pro"] = {
}
m["alv-edo-pro"] = {
}
m["alv-fli-pro"] = {
}
m["alv-gbe-pro"] = {
}
m["alv-gng-pro"] = {
}
m["alv-gtm-pro"] = {
aliases = {"Proto-Ghana-Togo Mountain"},
}
m["alv-gwa"] = {
}
m["alv-hei-pro"] = {
}
m["alv-ido-pro"] = {
}
m["alv-igb-pro"] = {
}
m["alv-kwa-pro"] = {
}
m["alv-mum-pro"] = {
}
m["alv-nup-pro"] = {
}
m["alv-pro"] = {
}
m["alv-von-pro"] = {
}
m["alv-yor-pro"] = {
}
m["alv-yrd-pro"] = {
}
m["apa-pro"] = {
aliases = {"Proto-Apache", "Proto-Southern Athabaskan"},
}
m["aql-pro"] = {
}
m["art-adu"] = {
aliases = {"Westron"},
}
m["art-bel"] = {
}
m["art-blk"] = {
}
m["art-bsp"] = {
}
m["art-com"] = {
}
m["art-dtk"] = {
}
m["art-elo"] = {
}
m["art-gld"] = {
}
m["art-lap"] = {
}
m["art-man"] = {
}
m["art-mun"] = {
}
m["art-nav"] = {
}
m["art-vlh"] = {
}
m["ath-nic"] = {
}
m["ath-pro"] = {
}
m["auf-pro"] = {
aliases = {"Proto-Arawan", "Proto-Arauan"},
}
m["aus-alu"] = {
otherNames = {"Ogh-Alungul", "Alngula"},
}
m["aus-and"] = {
aliases = {"Adithinngithigh"},
}
m["aus-ang"] = {
otherNames = {"Ogh-Anggula", "Anggula", "Ogh-Anggul", "Anggul"},
}
m["aus-arn-pro"] = {
}
m["aus-bra"] = {
aliases = {"Barranbinja", "Baranbinya", "Burranbinya", "Burrumbiniya", "Burrunbinya", "Barrumbinya", "Barren-binya", "Parran-binye"},
}
m["aus-brm"] = {
}
m["aus-cww-pro"] = {
}
m["aus-dal-pro"] = {
}
m["aus-guw"] = {
otherNames = {"Gowar", "Goowar", "Gooar", "Guar", "Gowr-burra", "Ngugi", "Mugee", "Wogee", "Gnoogee", "Chunchiburri", "Booroo-geen-merrie"},
}
m["aus-lsw"] = {
aliases = {"Little Swanport Tasmanian"},
}
m["aus-mbi"] = {
otherNames = {"Mbeiwum"},
}
m["aus-ngk"] = {
otherNames = {"Ngkot", "Nggoth"},
}
m["aus-nyu-pro"] = {
}
m["aus-pam-pro"] = {
}
m["aus-tul"] = {
otherNames = {"Dappil", "Dapil", "Toolooa", "Dulua", "Narung", "Dandan"},
}
m["aus-uwi"] = {
otherNames = {"Uwinjmil"},
}
m["aus-wdj-pro"] = {
}
m["aus-won"] = {
}
m["aus-wul"] = {
otherNames = {"Manbara", "Wulgurugaba", "Wulgurukaba", "Nhawalgaba"},
}
m["aus-ynk"] = { -- contrast nny
}
m["awd-amc-pro"] = {
otherNames = {"Western Maipuran"},
}
m["awd-kmp-pro"] = {
otherNames = {"Campa", "Kampan", "Campan", "Pre-Andine Maipurean"},
}
m["awd-prw-pro"] = {
otherNames = {"Paresí-Waurá", "Parecí–Xingú", "Paresí–Xingu", "Central Arawak", "Central Maipurean"},
}
m["awd-ama"] = {
}
m["awd-ana"] = {
aliases = {"Anauya"},
}
m["awd-apo"] = {
otherNames = {"Lapachu"},
}
m["awd-cab"] = {
aliases = {"Cabere", "Cávere", "Cavere"},
}
m["awd-gnu"] = {
otherNames = {"Guinao", "Inao", "Guniare", "Quinhau", "Guiano"},
}
m["awd-kar"] = {
aliases = {"Kariaí", "Kariai", "Cariyai", "Carihiahy"},
}
m["awd-kaw"] = {
aliases = {"Cawishana", "Cayuishana", "Kaishana", "Cauixana"},
}
m["awd-kus"] = {
aliases = {"Kustenaú", "Custenau", "Kutenabu"},
}
m["awd-man"] = {
}
m["awd-mar"] = {
aliases = {"Marawán"},
}
m["awd-mpr"] = {
aliases = {"Maypure", "Mejepure"},
}
m["awd-mrt"] = {
aliases = {"Mariate"},
}
m["awd-nwk-pro"] = {
aliases = {"Proto-Newiki"},
}
m["awd-pai"] = {
aliases = {"Paiconeca", "Paikone", "Paicone"},
}
m["awd-pas"] = {
aliases = {"Passé", "Pazé"},
}
m["awd-pro"] = {
otherNames = {"Proto-Arawakan", "Proto-Maipurean", "Proto-Maipuran"},
}
m["awd-she"] = {
aliases = {"Shebaya", "Shebaye"},
}
m["awd-taa-pro"] = {
otherNames = {"Proto-Ta-Arawakan", "Proto-Caribbean Northern Arawak"},
}
m["awd-wai"] = {
otherNames = {"Wainuma", "Wai", "Waima", "Wainumi", "Wainambí", "Waiwana", "Waipi", "Yanuma"},
}
m["awd-war"] = {
}
m["awd-yum"] = {
aliases = {"Jumana"},
}
m["azc-caz"] = {
aliases = {"Caxcan", "Kaskán"},
}
m["azc-cup-pro"] = {
}
m["azc-ktn"] = {
aliases = {"Gitanemuk"},
}
m["azc-nah-pro"] = {
}
m["azc-nic"] = {
}
m["azc-num-pro"] = {
}
m["azc-pro"] = {
}
m["azc-tak-pro"] = {
}
m["azc-tat"] = {
}
m["ber-fog"] = {
otherNames = {"El-Fogaha", "El-Foqaha", "Foqaha", "Fuqaha"},
}
m["ber-pro"] = {
}
m["ber-zuw"] = {
}
m["bnt-bal"] = {
}
m["bnt-bon"] = {
}
m["bnt-boy"] = {
}
m["bnt-bwa"] = {
}
m["bnt-cmw"] = {
otherNames = {"Bravanese", "Mwiini", "Mwini", "Chimwini", "Chimini", "Brava"},
}
m["bnt-ind"] = {
otherNames = {"Kɔlɔmɔnyi", "Kɔlɛ", "Kasaï Oriental"},
}
m["bnt-lal"] = {
}
m["bnt-mpi"] = {
}
m["bnt-mpu"] = {
}
m["bnt-ngu-pro"] = {
}
m["bnt-phu"] = {
aliases = {"Siphuthi"},
}
m["bnt-pro"] = {
}
m["bnt-sab-pro"] = {
}
m["bnt-sbo"] = {
}
m["bnt-sts-pro"] = {
}
m["btk-pro"] = {
}
m["cau-abz-pro"] = {
otherNames = {"Proto-Abazgi", "Proto-Abkhaz-Tapanta"},
}
m["cau-and-pro"] = {
aliases = {"Proto-Andi", "Proto-Andic"},
}
m["cau-ava-pro"] = {
aliases = {"Proto-Avar-Andian", "Proto-Avar-Andi", "Proto-Avar-Andic"},
}
m["cau-cir-pro"] = {
otherNames = {"Proto-Adyghe-Kabardian", "Proto-Adyghe-Circassian"},
}
m["cau-drg-pro"] = {
otherNames = {"Proto-Dargin"},
}
m["cau-lzg-pro"] = {
aliases = {"Proto-Lezgi", "Proto-Lezgian", "Proto-Lezgic"},
}
m["cau-nec-pro"] = {
}
m["cau-nkh-pro"] = {
}
m["cau-nwc-pro"] = {
}
m["cau-tsz-pro"] = {
otherNames = {"Proto-Tsezic", "Proto-Didoic"},
}
m["cba-ata"] = {
otherNames = {"Atanque", "Cancuamo", "Kankuamo", "Kankwe", "Kankuí", "Atanke"},
}
m["cba-cat"] = {
otherNames = {"Catio Chibcha", "Old Catio"},
}
m["cba-dor"] = {
otherNames = {"Chumulu", "Changuena", "Changuina", "Chánguena", "Gualaca"},
}
m["cba-dui"] = {
}
m["cba-hue"] = {
otherNames = {"Güetar", "Guetar", "Brusela"},
}
m["cba-nut"] = {
otherNames = {"Nutabane"},
}
m["cba-pro"] = {
}
m["ccs-pro"] = {
}
m["ccs-gzn-pro"] = {
aliases = {"Proto-Karto-Zan"},
}
m["cdc-cbm-pro"] = {
otherNames = {"Proto-Central-Chadic", "Proto-Biu-Mandara"},
}
m["cdc-mas-pro"] = {
}
m["cdc-pro"] = {
}
m["cdd-pro"] = {
}
m["cel-bry-pro"] = {
aliases = {"Proto-Brittonic", "Common Brythonic", "Common Brittonic"},
}
m["cel-gal"] = {
}
m["cel-gau"] = {
}
m["cel-pro"] = {
}
m["chi-pro"] = {
}
m["chm-pro"] = {
}
m["cmc-pro"] = {
}
m["crp-bip"] = {
}
m["crp-cpr"] = {
}
m["crp-gep"] = {
aliases = {"Greenlandic Pidgin", "Greenlandic Eskimo Pidgin"},
}
m["crp-kia"] = {
}
m["crp-mar"] = {
otherNames = {"Jamaican Maroon Spirit Possession Language"},
}
m["crp-mpp"] = {
aliases = {"Macao Pidgin Portuguese"},
}
m["crp-rsn"] = {
}
m["crp-slb"] = {
otherNames = {"Solombala-English", "Solombala English-Russian Pidgin"},
}
m["crp-spp"] = {
}
m["crp-tpr"] = {
}
m["csu-bba-pro"] = {
}
m["csu-maa-pro"] = {
}
m["csu-pro"] = {
}
m["csu-sar-pro"] = {
}
m["cus-ash"] = {
otherNames = {"Ashraf", "Af-Ashraaf"},
varieties = { {"Marka, Lower Shabelle"}, "Shingani"},
}
m["cus-hec-pro"] = {
}
m["cus-som-pro"] = {
otherNames = {"Proto-Sam", "Proto-Macro-Somali"},
}
m["cus-sou-pro"] = {
otherNames = {"Proto-Rift"},
}
m["cus-pro"] = {
}
m["dmn-dam"] = {
}
m["dra-bry"] = {
aliases = {"Byari"},
}
m["dra-cen-pro"] = {
}
m["dra-mkn"] = {
aliases = {"Nadugannada"},
}
m["dra-nor-pro"] = {
}
m["dra-okn"] = {
aliases = {"Halegannada"},
}
m["dra-ote"] = {
}
m["dra-pro"] = {
}
m["dra-sdo-pro"] = {
aliases = {"Proto-South Dravidian"},
}
m["dra-sdt-pro"] = {
aliases = {"Proto-South-Central Dravidian"},
}
m["dra-sou-pro"] = {
aliases = {"Proto-Southern Dravidian"},
}
m["egx-dem"] = {
aliases = {"Demotic Egyptian", "Enchorial"},
}
m["dmn-pro"] = {
}
m["dmn-mdw-pro"] = {
}
m["dru-pro"] = {
}
m["ero-gsz"] = {
}
m["ero-nya"] = {
}
m["ero-tau"] = {
}
m["esx-esk-pro"] = {
}
m["esx-ink"] = {
}
m["esx-inq"] = {
}
m["esx-inu-pro"] = {
}
m["esx-pro"] = {
}
m["esx-tut"] = {
}
m["euq-pro"] = {
aliases = {"Proto-Vasconic"},
}
m["gba-pro"] = {
}
m["gem-pro"] = {
aliases = {"Common Germanic"},
}
m["gme-bur"] = {
aliases = {"Burgundish", "Burgundic"},
}
m["gme-cgo"] = {
}
m["gmq-gut"] = {
}
m["gmq-jmk"] = {
aliases = {"Jamtlandic"},
}
m["gmq-mno"] = {
}
m["gmq-oda"] = {
}
m["gmq-ogt"] = {
aliases = {"Old Gotlandic"},
}
m["gmq-osw"] = {
}
m["gmq-pro"] = {
aliases = {"Proto-Scandinavian", "Primitive Norse", "Proto-Nordic",
"Ancient Nordic", "Ancient Scandinavian", "Old Nordic", "Old Scandinavian",
"Proto-North Germanic", "North Proto-Germanic", "Common Scandinavian"},
}
m["gmq-scy"] = {
}
m["gmw-bgh"] = {
}
m["gmw-cfr"] = {
varieties = {"Mittelfränkisch", "Ripuarian", "Moselle Franconian", "Colognian", "Kölsch"},
}
m["gmw-ecg"] = {
varieties = {"Thuringian", "Thüringisch", "Upper Saxon", "Upper Saxon German", "Obersächsisch", "Lusatian", "Erzgebirgisch", "Silesian", "Silesian German", "High Prussian"},
}
m["gmw-fin"] = {
aliases = {"Fingal"},
}
m["gmw-gts"] = {
aliases = {"Gottscheerisch"},
}
m["gmw-jdt"] = {
}
m["gmw-msc"] = {
}
m["gmw-pro"] = {
}
m["gmw-rfr"] = {
aliases = {"Rheinfränkisch", "Rhenish Franconian"},
varieties = {"Hessian", "Lorraine Franconian", "Lorrainian", "Lothringisch", "Palatine German", "Pfälzisch", "Pälzisch", "Palatinate German"},
}
m["gmw-stm"] = {
aliases = {"Satu Mare Swabian", "Sathmarschwäbisch", "Sathmarisch"},
}
m["gmw-tsx"] = {
aliases = {"Siebenbürger Saxon"},
}
m["gmw-vog"] = {
}
m["gmw-zps"] = {
aliases = {"Zipser", "Zipserisch", "Outzäpsersch"},
}
m["gn-cls"] = {
}
m["grk-cal"] = {
aliases = {"Italian Greek", "Bova"},
}
m["grk-ita"] = {
aliases = {"Griko", "Grico", "Grecanic"},
}
m["grk-mar"] = {
aliases = {"Mariupolitan Greek", "Rumeíka", "Rumeika"},
}
m["grk-pro"] = {
aliases = {"Proto-Greek"},
}
m["hmn-pro"] = {
}
m["hmx-mie-pro"] = {
}
m["hmx-pro"] = {
}
m["hyx-pro"] = {
}
m["iir-nur-pro"] = {
}
m["iir-pro"] = {
}
m["ijo-pro"] = {
aliases = {"Proto-Ijaw"},
}
m["inc-apa"] = {
aliases = {"Apabhraṃśa"},
}
m["inc-ash"] = {
aliases = {"Asokan Prakrit", "Aśokan Prakrit"},
}
m["inc-dng-pro"] = {
}
m["inc-kam"] = {
}
m["inc-kho"] = {
}
m["inc-khr"] = {
}
m["inc-krd-pro"] = {
}
m["inc-mas"] = {
}
m["inc-mbn"] = {
}
m["inc-mgu"] = {
}
m["inc-mor"] = {
aliases = {"Middle Oriya"},
}
m["inc-oas"] = {
}
m["inc-oaw"] = {
aliases = {"Early Awadhi"},
}
m["inc-obn"] = {
}
m["inc-ogu"] = {
otherNames = {"Old Western Rajasthani"},
}
m["inc-ohi"] = {
aliases = {"Dehlavi"},
}
m["inc-oor"] = {
aliases = {"Old Oriya"},
}
m["inc-opa"] = {
}
m["inc-pro"] = {
}
m["inc-sar"] = {
}
m["ine-ana-pro"] = {
}
m["ine-bsl-pro"] = {
}
m["ine-kal"] = {
aliases = {"Kalašmaic", "Kalasmaic"},
}
m["ine-pae"] = {
}
m["ine-pro"] = {
}
m["ine-toc-pro"] = {
}
m["itc-psa"] = {
}
m["mis-idn"] = {
}
m["mis-tdl"] = {
}
m["mis-tdt"] = {
}
m["mis-xnu"] = {
}
m["ngf-bin-pro"] = {
}
m["njo-jgl"] = {
}
m["njo-mng"] = {
}
m["paa-kmn"] = {
}
m["paa-lei"] = {
}
m["poz-nes"] = {
}
m["poz-pcc-pro"] = {
}
m["roa-can"] = {
}
m["roa-ona"] = {
}
m["sai-gua"] = {
}
m["sai-peb"] = {
}
m["sem-sam"] = {
}
m["sit-aao-pro"] = {
}
m["sit-ban"] = {
}
m["sit-bdi-pro"] = {
}
m["sit-ers-pro"] = {
}
m["sit-khb-pro"] = {
}
m["sit-khp-pro"] = {
}
m["sit-khw-pro"] = {
}
m["sit-kon-pro"] = {
}
m["sit-nas-pro"] = {
}
m["sit-tng-pro"] = {
}
m["tbq-brm-pro"] = {
}
m["tup-kaw"] = {
}
m["xme-old"] = {
}
m["xme-mid"] = {
aliases = {"Atropatenian"},
}
m["xme-ker"] = {
otherNames = {"Kermanian", "Central Iranian Dialects", "Central Plateau Dialects", "Central Iranian", "South Median", "Gazi", "Soi", "Sohi", "Abuzeydabadi", "Abyanehi", "Farizandi", "Jowshaqani", "Nashalji", "Qohrudi", "Yarandi", "Tari", "Sedehi", "Ardestani", "Zefrehi", "Isfahani", "Kafroni", "Varzenehi", "Khuri", "Nayini", "Anaraki", "Zoroastrian Dari", "Behdināni", "Behdinani", "Gabri", "Gavrŭni", "Gavruni", "Gabrōni", "Gabroni", "Kermani", "Yazdi", "Bidhandi", "Bijagani", "Chimehi", "Hanjani", "Komjani", "Naraqi", "Qalhari", "Varani", "Zori"},
}
m["xme-taf"] = {
}
m["xme-ttc-pro"] = {
}
m["xme-kls"] = {
aliases = {"Kalāsuri", "Kalasur", "Kalāsur"},
}
m["xme-klt"] = {
}
m["xme-ott"] = {
otherNames = {"Old Tatic", "Old Azeri", "Azari", "Azeri", "Āḏarī", "Adari", "Adhari"},
}
m["ira-kms-pro"] = {
}
m["ira-mpr-pro"] = {
}
m["ira-pat-pro"] = {
}
m["ira-pro"] = {
}
m["ira-zgr-pro"] = {
}
m["xsc-pro"] = {
}
m["xsc-sar-pro"] = {
}
m["xsc-skw-pro"] = {
}
m["xsc-sak-pro"] = {
aliases = {"Proto-Sakan"},
}
m["ira-sym-pro"] = {
}
m["ira-sgi-pro"] = {
}
m["ira-mny-pro"] = {
}
m["ira-shy-pro"] = {
}
m["ira-shr-pro"] = {
}
m["ira-sgc-pro"] = {
aliases = {"Proto-Sogdian"},
}
m["ira-wnj"] = {
aliases = {"Old Vanji", "Vanchi", "Vanži", "Wanji"},
}
m["iro-ere"] = {
}
m["iro-min"] = {
}
m["iro-nor-pro"] = {
}
m["iro-pro"] = {
}
m["itc-pro"] = {
}
m["jpx-hcj"] = {
aliases = {"Hachijo"},
}
m["jpx-pro"] = {
}
m["jpx-ryu-pro"] = {
}
m["kar-pro"] = {
}
m["kca-eas"] = {
}
m["kca-nor"] = {
}
m["kca-pro"] = {
}
m["kca-sou"] = {
}
m["khi-kho-pro"] = {
}
m["khi-kun"] = {
otherNames = {"ǃOǃKung", "ǃ'OǃKung", "Kung", "Ekoka ǃKung", "Ekoka Kung", "Sekele"},
}
m["ko-ear"] = {
}
m["kro-pro"] = {
}
m["ku-pro"] = {
}
m["map-ata-pro"] = {
}
m["map-bms"] = {
}
m["map-pro"] = {
}
m["mis-hkl"] = {
aliases = {"Kelantan Peranakan Chinese", "Kelantan Peranakan Hokkien", "Hokkien Kelantan", "Kelantan Local Hokkien"}
}
m["mis-isa"] = {
}
m["mis-jie"] = {
aliases = {"Chieh", "Kjet"},
}
m["mis-jzh"] = {
aliases = {"Haihua"},
}
m["mis-kas"] = {
aliases = {"Cassite", "Kassitic", "Kaššite"},
}
m["mis-mmd"] = {
otherNames = {"Mimi of Gaudefroy-Demombynes", "Mimi-D"},
}
m["mis-mmn"] = {
otherNames = {"Mimi-N"},
}
m["mis-phi"] = {
aliases = {"Philistian", "Philistinian"},
}
m["mis-rou"] = {
aliases = {"Ruanruan", "Ruan-ruan", "Juan-juan"},
}
m["mis-tnw"] = {
aliases = {"Tangwanghua"},
}
m["mis-tuh"] = {
aliases = {"'Azha"},
}
m["mis-tuo"] = {
aliases = {"Tabghach", "Taghbach"},
}
m["mis-wuh"] = {
aliases = {"Wuwan", "Awar"},
}
m["mis-xbi"] = {
aliases = {"Serbi", "Shirwi"},
}
m["mjg-mgl"] = {
aliases = {"Huzhu", "Huzhu Monguor"},
}
m["mjg-mgr"] = {
aliases = {"Minhe", "Minhe Monguor"},
}
m["mkh-asl-pro"] = {
}
m["mkh-ban-pro"] = {
}
m["mkh-kat-pro"] = {
}
m["mkh-khm-pro"] = {
}
m["mkh-kmr-pro"] = {
}
m["mkh-mmn"] = {
}
m["mkh-mnc-pro"] = {
}
m["mkh-mvi"] = {
}
m["mkh-pal-pro"] = {
}
m["mkh-pea-pro"] = {
}
m["mkh-pkn-pro"] = {
}
m["mkh-pro"] = { --This will be merged into 2015 aav-pro.
}
m["mnw-tha"] = {
aliases = {"Raman", "Thai Raman", "Siamese Mon"},
}
m["mkh-vie-pro"] = {
}
m["mns-cen"] = {
}
m["mns-nor"] = {
}
m["mns-pro"] = {
}
m["mns-sou"] = {
}
m["mun-pro"] = {
aliases = {"Proto-Mundan"},
}
m["myn-chl"] = { -- the stage after ''emy''
otherNames = {"Cholti", "Colonial Ch'olti'", "Colonial Cholti"},
}
m["myn-pro"] = {
aliases = {"Proto-Maya"},
}
m["nai-ala"] = {
otherNames = {"Alasapa", "Pinto"},
}
m["nai-bay"] = {
otherNames = {"Bayougoula", "Bayou Goula", "Ischenoca"}, -- tribe merged with "Mougulasha", "Mongoulacha", "Mugulasha", "Mougulasha", "Muglahsa", "Muglasha", "Muguasha", "Imongolosha", "Houma", "Acolapissa"
}
m["nai-cal"] = {
}
m["nai-chi"] = {
}
m["nai-chu-pro"] = {
aliases = {"Proto-Chumashan"},
}
m["nai-cig"] = {
}
m["nai-ckn-pro"] = {
aliases = {"Proto-Chinook"},
}
m["nai-guz"] = {
aliases = {"Guazacapan"},
}
m["nai-hit"] = {
otherNames = {"Atcik-hata", "At-pasha-shliha"},
}
m["nai-ipa"] = {
otherNames = {"'Iipay 'aa", "Northern Diegueño", "Diegueño"},
}
m["nai-jtp"] = {
otherNames = {"Xutiapa", "Jalapa", "Xalapa"},
}
m["nai-jum"] = {
aliases = {"Jumaitepeque", "Jumaytepec"},
}
m["nai-kat"] = {
otherNames = {"Kathlamet Chinook"},
}
m["nai-klp-pro"] = {
}
m["nai-knm"] = {
}
m["nai-kum"] = {
otherNames = {"Kumiai", "Central Diegueño", "Diegueño"},
}
m["nai-mac"] = {
aliases = {"Macorís", "Macorix", "Mazorij", "Mazorig", "Mazoriges"},
}
m["nai-mdu-pro"] = {
aliases = {"Proto-Maiduan"},
}
m["nai-miz-pro"] = {
aliases = {"Proto-Mixe-Zoquean"},
}
m["nai-mus-pro"] = {
aliases = {"Proto-Muskhogean", "Proto-Muskogee"},
}
m["nai-nao"] = {
}
m["nai-nrs"] = {
}
m["nai-okw"] = {
}
m["nai-per"] = {
}
m["nai-pic"] = {
}
m["nai-plp-pro"] = {
}
m["nai-pom-pro"] = {
aliases = {"Proto-Pomoan"},
}
m["nai-qng"] = {
}
m["nai-sca-pro"] = { -- NB 'sio-pro' "Proto-Siouan" which is Proto-Western Siouan
}
m["nai-sin"] = {
aliases = {"Sinacantan", "Zinacantán", "Zinacantan"},
}
m["nai-sln"] = {
}
m["nai-spt"] = {
aliases = {"Shahaptin"},
}
m["nai-tap"] = {
otherNames = {"Tapachulteca", "Tapachulteco", "Tapachula"},
}
m["nai-taw"] = {
}
m["nai-teq"] = {
otherNames = {"Tequistlateco", "Tequistlateca", "Chontal", "Chontol of Oaxaca", "Oaxaca Chontal", "Oaxacan Chontal"},
}
m["nai-tip"] = {
otherNames = {"Tipay", "Tiipai", "Tiipay", "Jamul Tiipay", "Southern Digueño", "Diegueño"},
}
m["nai-tot-pro"] = {
}
m["nai-tsi-pro"] = {
}
m["nai-utn-pro"] = {
otherNames = {"Proto-Miwok-Costanoan"},
}
m["nai-wai"] = {
aliases = {"Guaycura", "Waicura"},
}
m["nai-wji"] = {
otherNames = {"Jicaque of El Palmar", "Sula"},
}
m["nai-yup"] = {
aliases = {"Jupiltepeque", "Yupiltepec", "Jupiltepec", "Xupiltepec"},
}
m["nan-dat"] = {
aliases = {"Datian"},
}
m["nan-hbl"] = {
aliases = {"Hokkienese", "Quanzhang", "Fukien", "Banlam", "Banlamese", "Ban-lam"},
}
m["nan-hlh"] = {
aliases = {"Hailufeng", "Hoklo Min", "Hai Lok Hong"},
}
m["nan-lnx"] = {
aliases = {"Longyan", "Liongna"},
}
m["nan-tws"] = {
aliases = {"Teochew Min", "Chiuchow", "Teo-Swa", "Teo-Swa Min", "Tio-Sua"},
}
m["nan-zhe"] = {
aliases = {"Zhenan"},
}
m["nan-zsh"] = {
aliases = {"Sanxiang", "Samheung", "Sahiu"},
}
m["ngf-pro"] = {
}
m["nic-bco-pro"] = {
}
m["nic-bod-pro"] = {
}
m["nic-eov-pro"] = {
}
m["nic-gns-pro"] = {
}
m["nic-grf-pro"] = {
}
m["nic-gur-pro"] = {
}
m["nic-jkn-pro"] = {
}
m["nic-lcr-pro"] = {
}
m["nic-ogo-pro"] = {
}
m["nic-ovo-pro"] = {
}
m["nic-plt-pro"] = {
}
m["nic-pro"] = {
}
m["nic-ubg-pro"] = {
}
m["nic-ucr-pro"] = {
}
m["nic-vco-pro"] = {
}
m["nub-har"] = {
aliases = {"Ḥarāza"},
}
m["nub-pro"] = {
}
m["omq-cha-pro"] = {
}
m["omq-maz-pro"] = {
aliases = {"Proto-Mazatecan"},
}
m["omq-mix-pro"] = {
}
m["omq-mxt-pro"] = {
}
m["omq-otp-pro"] = {
}
m["omq-pro"] = {
aliases = {"Proto-Otomanguean", "Proto-Oto-Mangue"},
}
m["omq-sjq"] = {
aliases = {"Chatino Sign Language", "San Juan Quiahije Chatino Sign Language"},
}
m["omq-tel"] = {
}
m["omq-teo"] = {
}
m["omq-tri-pro"] = {
}
m["omq-zap-pro"] = {
}
m["omq-zpc-pro"] = {
}
m["omv-aro-pro"] = {
}
m["omv-diz-pro"] = {
aliases = {"Proto-Maji"},
}
m["omv-pro"] = {
}
m["oto-otm-pro"] = {
}
m["oto-pro"] = {
}
m["paa-kwn"] = {
}
m["paa-nha-pro"] = {
}
m["paa-nun"] = {
}
m["phi-din"] = {
}
m["phi-kal-pro"] = {
aliases = {"Proto-Calamian"},
}
m["phi-nag"] = {
}
m["phi-pro"] = {
}
m["poz-abi"] = {
otherNames = {"Sembuak", "Tubu"},
}
m["poz-bal"] = {
}
m["poz-btk-pro"] = {
}
m["poz-cet-pro"] = {
}
m["poz-hce-pro"] = {
otherNames = {"Proto-South Halmahera - West New Guinea"},
}
m["poz-lgx-pro"] = {
}
m["poz-mcm-pro"] = {
}
m["poz-mic-pro"] = {
}
m["poz-mly-pro"] = {
}
m["poz-msa-pro"] = {
}
m["poz-oce-pro"] = {
}
m["poz-pep-pro"] = {
aliases = {"Proto-Eastern-Polynesian", "Proto-East Polynesian", "Proto-East-Polynesian"},
}
m["poz-pnp-pro"] = {
}
m["poz-pol-pro"] = {
}
m["poz-pro"] = {
otherNames = {"Proto-Western Malayo-Polynesian"}, -- Western is subsumed into general Proto-MP
}
m["poz-sml"] = {
aliases = {"Sarawak"},
}
m["poz-ssw-pro"] = {
}
m["poz-swa-pro"] = {
}
m["poz-ter"] = {
aliases = {"Terengganu"},
}
m["pqe-pro"] = {
}
m["pra-niy"] = {
}
m["qfa-adm-pro"] = {
}
m["qfa-bet-pro"] = {
aliases = {"Proto-Tai-Be"},
}
m["qfa-cka-pro"] = {
}
m["qfa-hur-pro"] = {
}
m["qfa-kad-pro"] = {
}
m["qfa-kms-pro"] = {
}
m["qfa-kor-pro"] = {
}
m["qfa-kra-pro"] = {
}
m["qfa-lic-pro"] = {
}
m["qfa-onb-pro"] = {
aliases = {"Proto-Ong-Be", "Proto-Bê"},
}
m["qfa-ong-pro"] = {
}
m["qfa-tak-pro"] = {
aliases = {"Proto-Tai-Kadai"},
}
m["qfa-yen-pro"] = {
}
m["qfa-yuk-pro"] = {
}
m["qwe-kch"] = {
otherNames = {"Kichwa shimi", "Runashimi", "Runa", "Quichua", "Quecha", "Inga", "Chimborazo", "Imbabura Highland Kichwa", "Cañar Highland Quecha", "Quechua"},
}
m["qwe-pro"] = {
}
m["roa-ang"] = {
otherNames = {"Craonnais", "Baugeois", "Saumurois"},
}
m["roa-bbn"] = {
otherNames = {"Bourbonnais", "Berrichon", "Moulins", "Allier", "Nivernais", "Haut-Berrichon", "Bas-Berrichon"},
}
m["roa-brg"] = {
otherNames = {"Burgundian", "Bregognon", "Dijonnais", "Morvandiau", "Morvandeau", "Morvan", "Bourguignon-Morvandiau", "Mâconnais", "Brionnais", "Brionnais-Charolais", "Auxerrois", "Beaunois", "Langrois", "Valsaônois", "Verduno-Chalonnais", "Sédelocien"},
}
m["roa-cha"] = {
otherNames = {"Bassignot", "Langrois", "Sennonais", "Vallage", "Troyen", "Briard", "Der", "Perthois", "Rémois", "Argonnais", "Porcien", "Ardennais", "Sugny"},
}
m["roa-fcm"] = {
otherNames = {"Frainc-Comtou", "Comtois", "Jurassien", "Ajoulot", "Vâdais", "Taignon", "Bisontin", "Bousbot"},
}
m["roa-gal"] = {
}
m["roa-gib"] = {
}
m["roa-gis"] = {
}
m["roa-leo"] = {
}
m["roa-lor"] = {
otherNames = {"Gaumais", "Vosgien", "Welche", "Argonnais", "Longovicien", "Messin", "Nancéien", "Spinalien", "Déodatien"},
}
m["roa-oca"] = {
}
m["roa-ole"] = {
}
m["roa-opt"] = {
aliases = {"Galician-Portuguese", "Galician Portuguese", "Medieval Galician", "Medieval Portuguese", "Old Galician", "Old Portuguese"},
}
m["roa-orl"] = {
otherNames = {"Beauceron", "Solognot", "Gâtinais", "Blaisois", "Vendômois"},
}
m["roa-poi"] = {
otherNames = {"Poitevin", "Saintongeais", "Maraîchin"},
}
m["roa-tar"] = {
}
m["sai-all"] = {
otherNames = {"Alyentiyak", "Huarpe", "Warpe"},
}
m["sai-and"] = { -- not to be confused with 'cbc' or 'ano'
otherNames = {"Miranya", "Miranha", "Miranha Carapana-Tapuya", "Miraña-Carapana-Tapuyo", "Andokero", "Miranya-Karapana-Tapuyo", "Miraña", "Carapana"},
}
m["sai-ayo"] = {
aliases = {"Ayoman", "Ayamán", "Ayaman"},
}
m["sai-bae"] = {
aliases = {"Baenã", "Baenán", "Baena"},
}
m["sai-bag"] = {
otherNames = {"Patagón de Bagua"},
}
m["sai-bet"] = {
otherNames = {"Betoy", "Betoya", "Betoye", "Betoi-Jirara", "Jirara"},
}
m["sai-bor-pro"] = {
otherNames = {"Proto-Bora-Muinane", "Proto-Bora-Muiname"},
}
m["sai-cac"] = {
otherNames = {"Kakán", "Diaguita", "Cacan", "Kakan", "Calchaquí", "Chaka", "Kaka", "Kaká", "Caca", "Caca-Diaguita", "Catamarcano", "Capayán", "Capayana", "Yacampis"},
}
m["sai-caq"] = {
otherNames = {"Cara", "Kara"},
}
m["sai-car-pro"] = {
}
m["sai-cat"] = {
}
m["sai-cer-pro"] = {
otherNames = {"Proto-Amazonian Jê"},
}
m["sai-chi"] = {
}
m["sai-chn"] = {
aliases = {"Chana"},
}
m["sai-chp"] = {
aliases = {"Txapacura", "Xapacura", "Guapore", "Šapakura", "Txapakura", "Txapakúra", "Xapakúra"},
}
m["sai-chr"] = {
aliases = {"Charrúa", "Charruá"},
}
m["sai-chu"] = {
aliases = {"Churoya"},
}
m["sai-cje-pro"] = {
otherNames = {"Proto-Akuwẽ"},
}
m["sai-cmg"] = {
aliases = {"Comechingón", "Comechingona", "Comechingone"},
}
m["sai-cno"] = {
otherNames = {"Chonos", "Caucau"},
}
m["sai-cnr"] = {
aliases = {"Cañar"},
}
m["sai-coe"] = {
aliases = {"Koeruna"},
}
m["sai-col"] = {
aliases = {"Colan"},
}
m["sai-cop"] = {
}
m["sai-crd"] = {
otherNames = {"Coroado"},
}
m["sai-ctq"] = {
aliases = {"Catuquinarú", "Katukinaru"},
}
m["sai-cul"] = {
otherNames = {"Culle", "Kulyi", "Ilinga", "Linga"},
}
m["sai-cva"] = {
}
m["sai-esm"] = {
otherNames = {"Esmeraldeño", "Atacame", "Takame"},
}
m["sai-ewa"] = {
}
m["sai-gam"] = {
aliases = {"Gamella", "Acobu", "Curinsi", "Barbados"},
}
m["sai-gay"] = {
aliases = {"Gayon"},
}
m["sai-gmo"] = {
otherNames = {"Wamo", "Santa Rosa", "San Jose", "Barinas", "Guamotey", "Guama"},
}
m["sai-gue"] = {
aliases = {"Guenoa"},
}
m["sai-hau"] = {
otherNames = {"Manek'enk"},
}
m["sai-jee-pro"] = {
otherNames = {"Proto-Gê", "Proto-Jean", "Proto-Gean", "Proto-Jê-Kaingang", "Proto-Ye"},
}
m["sai-jko"] = {
aliases = {"Geicó", "Jeicó", "Jaikó", "Geikó", "Yeikó", "Jeiko", "Geico", "Jeico", "Jaiko", "Geiko", "Yeiko", "Eyco"},
}
m["sai-jrj"] = {
}
m["sai-kat"] = { -- contrast xoo, kzw, sai-xoc
otherNames = {"Catrimbi", "Catembri", "Kariri de Mirandela", "Mirandela", "Kariri", "Kiriri"},
}
m["sai-mal"] = {
aliases = {"Malali"},
}
m["sai-mar"] = {
}
m["sai-mat"] = {
otherNames = {"Matanauí", "Matanaui", "Matanawü", "Mitandua", "Moutoniway"},
}
m["sai-mcn"] = {
aliases = {"Mokana"},
}
m["sai-men"] = {
aliases = {"Menién"},
}
m["sai-mil"] = {
otherNames = {"Milykayak", "Huarpe", "Warpe"},
}
m["sai-mlb"] = {
aliases = {"Malibú", "Malebú"},
}
m["sai-msk"] = {
aliases = {"Masakara", "Masacará", "Masacara"},
}
m["sai-muc"] = {
otherNames = {"Mucuchi", "Mokochi", "Mocochí", "Mirripú", "Maripú", "Mucuchí-Maripú"},
}
m["sai-mue"] = {
aliases = {"Muellamués"},
}
m["sai-muz"] = {
}
m["sai-mys"] = {
otherNames = {"Mayna", "Maina", "Rimachu"},
}
m["sai-nat"] = {
otherNames = {"Natu", "Peagaxinan"},
}
m["sai-nje-pro"] = {
otherNames = {"Proto-Core Jê"},
}
m["sai-opo"] = {
otherNames = {"Opon", "Opón-Karare", "Opón-Carare", "Carare", "Carare-Opón"},
}
m["sai-oto"] = {
aliases = {"Otomako", "Otomacan", "Otomac", "Otomak"},
}
m["sai-pal"] = {
}
m["sai-pam"] = {
aliases = {"Pamiwa"},
}
m["sai-par"] = {
aliases = {"Paratio", "Prarto"},
}
m["sai-pnz"] = {
aliases = {"Pansaleo"},
}
m["sai-prh"] = {
}
m["sai-ptg"] = {
otherNames = {"Patagón de Perico"},
}
m["sai-pur"] = {
aliases = {"Purukoto", "Purucotó", "Purucoto"},
}
m["sai-pyg"] = {
aliases = {"Payawá", "Payagua"},
}
m["sai-pyk"] = {
aliases = {"Gavião-Pykobjê", "Pykobjê-Gavião", "Gavião", "Pyhcopji", "Gavião-Pyhcopji"},
}
m["sai-qmb"] = {
otherNames = {"Kimbaya", "Quindío", "Quindio", "Quindo"},
}
m["sai-qtm"] = {
aliases = {"Quitemoca"},
}
m["sai-rab"] = {
}
m["sai-ram"] = {
}
m["sai-sac"] = {
otherNames = {"Sacata", "Zácata", "Chillao"},
}
m["sai-san"] = {
aliases = {"Sanavirón", "Sanabirón", "Sanabiron", "Sanavirona", "Zanavirona"},
}
m["sai-sap"] = {
aliases = {"Zapará", "Zapara"},
}
m["sai-sec"] = {
otherNames = {"Sek", "Sec"},
}
m["sai-sin"] = {
otherNames = {"Cenúfana", "Zenúfana", "Cinifaná", "Sinufana", "Sinú", "Cenú", "Zenú", "Finzenú", "Fincenú", "Pancenú", "Sutagao"},
}
m["sai-sje-pro"] = {
}
m["sai-tab"] = {
otherNames = {"Aconipa"},
}
m["sai-tal"] = {
otherNames = {"Atalán", "Tallan", "Tallanca", "Atalan", "Sek"},
}
m["sai-tap"] = {
otherNames = {"Tapayúna", "Kajkwakhrattxi"},
}
m["sai-tar-pro"] = {
}
m["sai-teu"] = {
aliases = {"Tehues", "Teuéx"},
}
m["sai-tim"] = {
otherNames = {"Cuica", "Timote-Cuica"},
}
m["sai-tpr"] = {
aliases = {"Taparito"},
}
m["sai-trr"] = {
otherNames = {"Caratiú"},
}
m["sai-wai"] = {
aliases = {"Waitaka", "Waitacá", "Waitaca", "Goytacá", "Goitacá", "Guaitacá", "Guiatacá", "Guiatacás", "Goiatacá", "Goiatacás", "Guaiatacá", "Goytacaz", "Goitacaz", "Goyataca", "Aitacaz", "Uetacaz", "Uetacá", "Outacá", "Ouetacá", "Eutacá", "Itacaz", "Vaitacá"},
}
m["sai-way"] = {
aliases = {"Wajumará", "Wajumara", "Wayumará", "Azumara", "Guimara"},
}
m["sai-wit-pro"] = {
otherNames = {"Proto-Huitotoan", "Proto-Uitotoan"},
}
m["sai-wnm"] = {
otherNames = {"Wañam", "Wanyam", "Huanyam", "Uanham", "Abitana"},
}
m["sai-xoc"] = { -- contrast xoo, kzw, sai-kat
otherNames = {"Xoco", "Chocó", "Shokó", "Shoko", "Shocó", "Shoco", "Choco", "Chocaz", "Kariri-Xocó", "Kariri-Xoco", "Kariri-Shoko", "Cariri-Chocó", "Xukuru-Kariri", "Xucuru-Kariri", "Xucuru-Cariri", "Xukurú-Kirirí"},
}
m["sai-yao"] = {
aliases = {"Yao", "Jaoi", "Yaoi", "Yaio", "Anacaioury"},
}
m["sai-yar"] = { -- not the same family as 'suy'
aliases = {"Yaruma"},
}
m["sai-yri"] = {
aliases = {"Jurí"},
}
m["sai-yup"] = {
otherNames = {"Yupuá", "Yupúa", "Jupua", "Jupuá", "Jupúa", "Hiupiá", "Yupuá-Duriña", "Duriña"},
}
m["sai-yur"] = {
aliases = {"Yurumangui", "Yurimangí", "Yurimangi", "Yurimanguí", "Yurimangui"},
}
m["sal-pro"] = {
aliases = {"Proto-Salishan"},
}
m["sdv-daj-pro"] = {
}
m["sdv-eje-pro"] = {
}
m["sdv-nil-pro"] = {
}
m["sdv-nyi-pro"] = {
}
m["sdv-tmn-pro"] = {
}
m["sel-nor"] = {
aliases = {"Taz Selkup"},
}
m["sel-pro"] = {
}
m["sel-sou"] = {
}
m["sem-amm"] = {
}
m["sem-amo"] = {
aliases = {"Amoritic"},
}
m["sem-cha"] = {
aliases = {"Cheha", "Čäha", "Čäxa"},
}
m["sem-dad"] = {
otherNames = {"Dadanite", "Lihyanite", "Lihyanitic"},
}
m["sem-dum"] = {
}
m["sem-has"] = {
}
m["sem-his"] = {
otherNames = {"Thamudic E"},
}
m["sem-mhr"] = {
otherNames = {"Muher Gurage", "Muxar", "Muxər", "Muhər", "Muḫər"},
}
m["sem-pro"] = {
}
m["sem-saf"] = {
}
m["sem-srb"] = {
}
m["sem-tay"] = {
otherNames = {"Taymanite", "Thamudic A"},
}
m["sem-tha"] = {
}
m["sem-wes-pro"] = {
}
m["sio-pro"] = { -- NB this is not Proto-Siouan-Catawban 'nai-sca-pro'
}
m["sit-bok"] = {
otherNames = {"Ramo", "Pailibo"},
}
m["sit-bai-pro"] = {
}
m["sit-cai"] = {
}
m["sit-cha"] = {
}
m["sit-hrs-pro"] = {
}
m["sit-jap"] = {
otherNames = {"Chabao", "Kuru"},
}
m["sit-kha-pro"] = {
}
m["sit-liz"] = {
}
m["sit-lnj"] = {
}
m["sit-lrn"] = {
}
m["sit-luu-pro"] = {
}
m["sit-prn"] = {
}
m["sit-pro"] = {
}
m["sit-sit"] = {
otherNames = {"Eastern rGyalrong", "rGyalrong", "Rgyalrong", "rGyalrongic", "Gyalrong", "Gyarong", "rGyarong", "Gyarung", "Jiarong", "Jiarongyu", "Jyarong", "Jyarung", "Yelong", "Kuru"},
}
m["sit-tam-pro"] = {
aliases = {"Proto-Tamang"},
}
m["sit-tan-pro"] = {
}
m["sit-tgm"] = {
}
m["sit-tos"] = {
}
m["sit-tsh"] = {
otherNames = {"Caodeng", "Sidaba", "rGyalrong", "Rgyalrong", "Jiarong", "Gyarung", "Kuru"},
}
m["sit-zbu"] = {
otherNames = {"Ribu", "Rdzong'bur", "Rdzongmbur", "Showu", "rGyalrong", "Rgyalrong", "Jiarong", "Gyarung", "Kuru"},
}
m["sla-pro"] = {
aliases = {"Common Slavic"},
}
m["smi-pro"] = {
aliases = {"Proto-Sami"},
}
m["son-pro"] = {
aliases = {"Proto-Songhai"},
}
m["sqj-pro"] = {
}
m["ssa-klk-pro"] = {
aliases = {"Proto-Rub"},
}
m["ssa-kom-pro"] = {
}
m["ssa-pro"] = {
}
m["syd-pro"] = {
}
m["tai-pro"] = {
}
m["tai-swe-pro"] = {
}
m["tbq-bdg-pro"] = {
}
m["tbq-blg"] = {
aliases = {"Pai-lang", "Pailang"},
}
m["tbq-gkh"] = {
aliases = {"Gɔkhý", "Gɔkhy", "Gouke"},
}
m["tbq-kuk-pro"] = {
otherNames = {"Proto-Kukish"},
}
m["tbq-lal-pro"] = {
}
m["tbq-laz"] = {
otherNames = {"Lare", "Shuitianhua"},
}
m["tbq-lob-pro"] = {
}
m["tbq-lol-pro"] = {
otherNames = {"Proto-Yi", "Proto-Ngwi", "Proto-Nisoic"},
}
m["tbq-mil"] = {
}
m["tbq-mor"] = {
aliases = {"Morān"},
}
m["tbq-ngo"] = {
otherNames = {"Ngachang", "Achang"},
}
-- tbq-pro is now etymology-only
m["trk-dkh"] = {
aliases = {"Dukha"},
}
m["trk-eog"] = {
}
m["trk-oat"] = {
}
m["trk-pro"] = {
}
m["tup-gua-pro"] = {
}
m["tup-kab"] = {
aliases = {"Kabixiana", "Cabixiana", "Cabishiana", "Kapishana", "Capishana", "Kapišana", "Cabichiana", "Capichana", "Capixana"},
}
m["tuw-alk"] = {
aliases = {"Alechuka"},
}
m["tuw-bal"] = {
}
m["tuw-kkl"] = {
aliases = {"Chinese Kyakala"},
}
m["tuw-kli"] = {
aliases = {"Kilen", "Kirin", "Kila", "Hezhe", "Qile'en"},
}
m["tup-pro"] = {
}
m["tuw-pro"] = {
}
m["tuw-sol"] = {
}
m["urj-fin-pro"] = {
}
m["urj-koo"] = {
aliases = {"Old Permian"},
}
m["urj-kuk"] = {
aliases = {"Kukkuzi Votic", "Kukkuzi Ingrian", "Kukkusi"},
}
m["urj-kya"] = {
}
m["urj-mdv-pro"] = {
}
m["urj-prm-pro"] = {
}
m["urj-pro"] = {
otherNames = {"Proto-Finno-Ugric", "Proto-Finno-Permic"}, -- PFU and PFP are subsumed into PU per [[Wiktionary:Beer parlour/2015/January#Merging Finno-Volgaic, Finno-Samic, Finno-Permic and Finno-Ugric into Uralic]]
}
m["urj-ugr-pro"] = {
}
m["xgn-pro"] = {
}
m["xnd-pro"] = {
otherNames = {"Proto-Na-Dené", "Proto-Athabaskan-Eyak-Tlingit"},
}
m["yok-bvy"] = {
otherNames = {"Tulamni-Hometwoli", "Tulamni", "Tulamne", "Tuolumne", "Tawitchi", "Hometwoli", "Taneshach"},
}
m["yok-dly"] = {
otherNames = {"Far Northern Valley Yokuts", "Yachikumne", "Yachikumni", "Chulamni", "Lower San Joaquin", "Lakisamni", "Tawalimni"},
}
m["yok-gsy"] = {
}
m["yok-kry"] = {
otherNames = {"Choinimni", "Choynimni", "Ayticha", "Kocheyali", "Ayitcha", "Michahay", "Chukaymina", "Chukaimina"},
}
m["yok-nvy"] = {
otherNames = {"Chukchansi", "Kechayi", "Dumna", "Chawchila", "Noptinte", "Nopṭinṭe", "Nopthrinthre", "Nopchinchi", "Takin"},
}
m["yok-ply"] = {
otherNames = {"Paleuyami", "Altinin", "Poso Creek", "Poso Creek Yokuts"},
}
m["yok-svy"] = {
otherNames = {"Yawelmani", "Tachi", "Koyeti", "Nutunutu", "Chunut", "Wo'lasi", "Choynok", "Choinok", "Wechihit"},
}
m["yok-tky"] = {
otherNames = {"Wikchamni", "Wukchamni", "Wukchumni", "Yawdanchi"},
}
m["ypk-pro"] = {
}
m["yrk-for"] = {
}
m["yrk-tun"] = {
}
m["zhx-min-pro"] = {
}
m["zhx-sht"] = {
otherNames = {"Xiangnan Tuhua", "Yuebei Tuhua", "Shipo", "Shina"},
}
m["zhx-sic"] = {
otherNames = {"Sichuanese Mandarin"},
}
m["zhx-tai"] = {
aliases = {"Toishanese"},
}
m["zle-ono"] = {
}
m["zle-ort"] = {
}
m["zlm-coa"] = {
}
m["zlm-pah"] = {
}
m["zls-chs"] = {
}
m["zlw-ocs"] = {
}
m["zlw-opl"] = {
}
m["zlw-osk"] = {
}
m["zlw-slv"] = {
}
return m
k4oi3h042u52duml7dvr2ovzu55yu2m
Modul:languages/data/3/t/extra
828
33780
375379
373750
2026-09-22T05:36:14Z
Hakimi97
2668
[[MediaWiki:UpdateLanguageNameAndCode.js|kemas kini menggunakan gajet bahasa]]
375379
Scribunto
text/plain
local m = {}
m["taa"] = {
otherNames = {"Tanana", "Middle Tanana"},
}
m["tab"] = {
aliases = {"Tabassaran"},
}
m["tac"] = {
}
m["tad"] = {
}
m["tae"] = {
}
m["taf"] = {
}
m["tag"] = {
}
m["taj"] = {
}
m["tak"] = {
}
m["tal"] = {
}
m["tan"] = {
}
m["tao"] = {
otherNames = {"Tao"},
}
m["tap"] = {
}
m["tar"] = {
}
m["tas"] = {
otherNames = {"Tay Boi Pidgin French", "Vietnamese Pidgin French"},
}
m["tau"] = {
otherNames = {"Tabesna", "Nabesna"},
}
m["tav"] = {
}
m["taw"] = {
}
m["tax"] = {
}
m["tay"] = {
}
m["taz"] = {
}
m["tba"] = {
}
m["tbc"] = {
}
m["tbd"] = {
}
m["tbe"] = {
}
m["tbf"] = {
}
m["tbg"] = {
}
m["tbh"] = {
}
m["tbi"] = {
otherNames = {"Ingessana", "Gaahmg"},
}
m["tbj"] = {
}
m["tbk"] = {
}
m["tbl"] = {
aliases = {"Tagabili"},
}
m["tbm"] = {
}
m["tbn"] = {
}
m["tbo"] = {
}
m["tbp"] = {
otherNames = {"Diebroud", "Dabra"},
}
m["tbr"] = {
}
m["tbs"] = {
}
m["tbt"] = {
otherNames = {"Tembo"},
}
m["tbu"] = {
otherNames = {"Tubare"},
}
m["tbv"] = {
}
m["tbw"] = {
}
m["tbx"] = {
otherNames = {"Middle Watut"},
}
m["tby"] = {
}
m["tbz"] = {
}
m["tca"] = {
otherNames = {"Tikuna"},
}
m["tcb"] = {
}
m["tcc"] = {
}
m["tcd"] = {
}
m["tce"] = {
}
m["tcf"] = {
}
m["tcg"] = {
}
m["tch"] = {
}
m["tci"] = {
}
m["tck"] = {
}
m["tcl"] = {
otherNames = {"Taman", "Taman (Burma)"},
}
m["tcm"] = {
}
m["tco"] = {
}
m["tcp"] = {
otherNames = {"Tawr"},
}
m["tcq"] = {
}
m["tcs"] = {
otherNames = {"Big Thap", "Blaikman", "Brokan", "Broken", "Broken English", "Cape York Creole", "Lockhart Creole", "Papuan Pidgin English", "Torres Strait Brokan", "Torres Strait Broken", "Torres Strait Pidgin", "Yumplatok"},
}
m["tct"] = {
}
m["tcu"] = {
}
m["tcw"] = {
}
m["tcx"] = {
}
m["tcy"] = {
}
m["tcz"] = {
otherNames = {"Thado"},
}
m["tda"] = {
}
m["tdb"] = {
}
m["tdc"] = {
}
m["tdd"] = {
aliases = {"Tai Nuea", "Dehong Dai", "Tai Dehong", "Tai Le", "Chinese Shan", "Chinese Tai"},
}
m["tde"] = {
}
m["tdf"] = {
otherNames = {"Taliang", "Tariang", "Kasseng"},
}
m["tdg"] = {
}
m["tdh"] = {
}
m["tdi"] = {
}
m["tdj"] = {
}
m["tdk"] = {
}
m["tdl"] = {
otherNames = {"Tapshin"},
}
m["tdm"] = {
otherNames = {"Taruamá"},
}
m["tdn"] = {
}
m["tdo"] = {
}
m["tdq"] = {
}
m["tdr"] = {
}
m["tds"] = {
otherNames = {"Taori"},
}
m["tdt"] = {
otherNames = {"Tetum Dili", "Tetun Prasa", "Tétum Praça", "Tetun-Dili", "Tetun-Prasa"},
}
m["tdv"] = {
}
m["tdy"] = {
}
m["tea"] = {
}
m["teb"] = {
}
m["tec"] = {
}
m["ted"] = {
}
m["tee"] = {
}
m["tef"] = {
}
m["teg"] = {
}
m["teh"] = {
otherNames = {"Patagón", "Chon", "Chon Patagón", "Chon Patagon", "Aoniken", "Aonikenk", "Inaquen", "Aonek'o 'ajen"},
}
m["tei"] = {
}
m["tek"] = {
}
m["tem"] = {
otherNames = {"Timne", "Themne", "KaThemne"},
}
m["ten"] = {
otherNames = {"Tama"},
}
m["teo"] = {
}
m["tep"] = {
}
m["teq"] = {
}
m["ter"] = {
}
m["tes"] = {
}
m["tet"] = {
otherNames = {"Tetun"},
}
m["teu"] = {
}
m["tev"] = {
}
m["tew"] = {
otherNames = {"Tano", "Santa Clara Tewa", "San Ildefonso Tewa", "Tesuque Tewa", "Nambe Tewa", "Ohkay Owingeh", "Pojoaque"},
}
m["tex"] = {
}
m["tey"] = {
}
m["tez"] = {
otherNames = {"Tin Sert"},
}
m["tfi"] = {
}
m["tfn"] = {
otherNames = {"Tanaina"},
}
m["tfo"] = {
}
m["tfr"] = {
}
m["tft"] = {
}
m["tga"] = {
}
m["tgb"] = {
}
m["tgc"] = {
}
m["tgd"] = {
}
m["tge"] = {
}
m["tgf"] = {
otherNames = {"Chalikha", "Chalipkha", "Tshali", "Tshalingpa"},
}
m["tgh"] = {
}
m["tgi"] = {
}
m["tgn"] = {
}
m["tgo"] = {
}
m["tgp"] = {
}
m["tgq"] = {
}
m["tgr"] = {
}
m["tgs"] = {
}
m["tgt"] = {
}
m["tgu"] = {
}
m["tgv"] = {
}
m["tgw"] = {
}
m["tgx"] = {
}
m["tgy"] = {
}
m["thc"] = {
}
m["thd"] = {
otherNames = {"Thaayorre", "Thayore"},
}
m["the"] = {
}
m["thf"] = {
}
m["thh"] = {
}
m["thi"] = {
}
m["thk"] = {
}
m["thl"] = {
}
m["thm"] = {
aliases = {"Aheu", "So (Thavung)"},
}
m["thn"] = {
}
m["thp"] = {
}
m["thq"] = {
}
m["thr"] = {
}
m["ths"] = {
}
m["tht"] = {
}
m["thu"] = {
}
m["thy"] = {
}
m["tic"] = {
}
m["tif"] = {
}
m["tig"] = {
}
m["tih"] = {
}
m["tii"] = {
}
m["tij"] = {
}
m["tik"] = {
}
m["til"] = {
}
m["tim"] = {
}
m["tin"] = {
}
m["tio"] = {
}
m["tip"] = {
}
m["tiq"] = {
}
m["tis"] = {
}
m["tit"] = {
}
m["tiu"] = {
}
m["tiv"] = {
otherNames = {"Tivi"},
}
m["tiw"] = {
}
m["tix"] = {
otherNames = {"Isleta", "Isleta Tiwa", "Isleta Pueblo", "Sandia", "Sandia Tiwa", "Sandia Pueblo"},
}
m["tiy"] = {
aliases = {"Teduray"},
}
m["tiz"] = {
}
m["tja"] = {
}
m["tjg"] = {
}
m["tji"] = {
}
m["tjl"] = {
aliases = {"Red Tai (Myanmar)", "Red Shan", "Shan Bamar", "Shan Kalee", "Shan Ni", "Tai Laeng", "Tai Lai", "Tai Leng", "Tai Nai", "Tai Naing"},
}
m["tjm"] = {
}
m["tjn"] = {
}
m["tjs"] = {
}
m["tju"] = {
}
m["tjw"] = {
otherNames = {"Djabwurrung", "Djab Wurrung", "Tjapwurrung"},
}
m["tka"] = {
}
m["tkb"] = {
}
m["tkd"] = {
}
m["tke"] = {
}
m["tkf"] = {
}
m["tkl"] = {
}
m["tkm"] = {
}
m["tkn"] = {
aliases = {"Tokunoshima", "Toku-no-Shima"},
}
m["tkp"] = {
}
m["tkq"] = {
}
m["tkr"] = {
otherNames = {"Caxur", "Tsaxur"},
}
m["tks"] = {
otherNames = {"Takestani"},
}
m["tkt"] = {
}
m["tku"] = {
}
m["tkv"] = {
otherNames = {"Pano"},
}
m["tkw"] = {
}
m["tkx"] = {
}
m["tkz"] = {
}
m["tla"] = {
}
m["tlb"] = {
}
m["tlc"] = {
otherNames = {"Yecuatla Totonac"},
}
m["tld"] = {
}
m["tlf"] = {
}
m["tlg"] = {
}
m["tlh"] = {
}
m["tli"] = {
}
m["tlj"] = {
}
m["tlk"] = {
}
m["tll"] = {
varieties = {"Indanga"},
}
m["tlm"] = {
}
m["tln"] = {
}
m["tlo"] = {
}
m["tlp"] = {
}
m["tlq"] = {
}
m["tlr"] = {
}
m["tls"] = {
}
m["tlt"] = {
otherNames = {"Sou Nama"},
}
m["tlu"] = {
}
m["tlv"] = {
otherNames = {"Soboyo"},
}
m["tlx"] = {
}
m["tly"] = {
otherNames = {"Talyshi", "Talishi", "Taleshi", "Tolashi", "Asalemi", "Anbarani"},
}
m["tma"] = {
otherNames = {"Tama"},
}
m["tmb"] = {
}
m["tmc"] = {
}
m["tmd"] = {
}
m["tme"] = {
}
m["tmf"] = {
}
m["tmg"] = {
}
m["tmh"] = {
otherNames = {"Tamashek", "Tamahaq", "Tamajaq", "Tamasheq"},
}
m["tmi"] = {
}
m["tmj"] = {
}
m["tml"] = {
}
m["tmm"] = {
}
m["tmn"] = {
otherNames = {"Taman"},
}
m["tmo"] = {
}
m["tmq"] = {
}
m["tms"] = {
}
m["tmt"] = {
}
m["tmu"] = {
otherNames = {"Turu"},
}
m["tmv"] = {
otherNames = {"Tembo"},
}
m["tmw"] = {
}
m["tmy"] = {
}
m["tmz"] = {
}
m["tna"] = {
}
m["tnb"] = {
}
m["tnc"] = {
}
m["tnd"] = {
}
m["tne"] = {
}
m["tng"] = {
}
m["tnh"] = {
}
m["tni"] = {
}
m["tnk"] = {
}
m["tnl"] = {
}
m["tnm"] = {
}
m["tnn"] = {
}
m["tno"] = {
}
m["tnp"] = {
}
m["tnq"] = {
aliases = {"Taino"},
}
m["tnr"] = {
}
m["tns"] = {
}
m["tnt"] = {
}
m["tnu"] = {
}
m["tnv"] = {
aliases = {"Tangchangya", "Tonchongya", "Tongchongya"},
}
m["tnw"] = {
}
m["tnx"] = {
}
m["tny"] = {
}
m["tnz"] = {
otherNames = {"Tonga"},
}
m["tob"] = {
otherNames = {"Chaco Sur", "Namqom", "Qom", "Toba Qom"},
}
m["toc"] = {
}
m["tod"] = {
}
m["tof"] = {
}
m["tog"] = {
otherNames = {"Kitonga", "Chitonga", "Siska", "Sisya", "Tonga", "Western Nyasa"},
}
m["toh"] = {
otherNames = {"Gitonga", "Tonga"},
}
m["toi"] = {
otherNames = {"Tonga", "Chitonga", "Plateau Tonga", "Zambezi"},
}
m["toj"] = {
}
m["tok"] = {
}
m["tol"] = {
otherNames = {"Smith River", "Smith River Tolowa"},
}
m["tom"] = {
}
m["too"] = {
}
m["top"] = {
}
m["toq"] = {
}
m["tor"] = {
}
m["tos"] = {
}
m["tou"] = {
}
m["tov"] = {
}
m["tow"] = {
otherNames = {"Towa"},
}
m["tox"] = {
}
m["toy"] = {
}
m["toz"] = {
}
m["tpa"] = {
}
m["tpc"] = {
}
m["tpe"] = {
}
m["tpf"] = {
}
m["tpg"] = {
}
m["tpi"] = {
otherNames = {"Melanesian Pidgin English", "Neo-Melanesian", "New Guinea Pidgin"},
}
m["tpj"] = {
}
m["tpk"] = {
otherNames = {"Coastal Tupi", "Tupiniquim"},
}
m["tpl"] = {
}
m["tpm"] = {
}
m["tpn"] = {
}
m["tpo"] = {
}
m["tpp"] = {
}
m["tpq"] = {
}
m["tpr"] = {
}
m["tpt"] = {
}
m["tpu"] = {
}
m["tpv"] = {
}
m["tpw"] = {
aliases = {"Classical Tupi"},
}
m["tpx"] = {
}
m["tpy"] = {
}
m["tpz"] = {
}
m["tqb"] = {
}
m["tql"] = {
}
m["tqm"] = {
}
m["tqn"] = {
}
m["tqo"] = {
}
m["tqp"] = {
}
m["tqq"] = {
}
m["tqr"] = {
}
m["tqt"] = {
}
m["tqu"] = {
}
m["tqw"] = {
}
m["tra"] = {
}
m["trb"] = {
}
m["trc"] = {
}
m["trd"] = {
}
m["tre"] = {
}
m["trf"] = {
}
m["trg"] = {
}
m["trh"] = {
}
m["tri"] = {
otherNames = {"Trio", "Tiriyó", "Tarano"},
}
m["trj"] = {
}
m["trl"] = {
}
m["trm"] = {
}
m["trn"] = {
otherNames = {"Trinitario Moxos", "Moxo", "Moxos", "Mojo", "Moxa"},
}
m["tro"] = {
otherNames = {"Tarao Naga", "Taraotrong", "Tarau"},
}
m["trp"] = {
}
m["trq"] = {
}
m["trr"] = {
}
m["trs"] = {
}
m["trt"] = {
}
m["tru"] = {
}
m["trv"] = {
otherNames = {"Seediq"},
}
m["trw"] = {
}
m["trx"] = {
otherNames = {"Tringus", "Tringgus-Sembaan Bidayuh"},
}
m["try"] = {
aliases = {"Tai Turung"},
}
m["trz"] = {
}
m["tsa"] = {
}
m["tsb"] = {
}
m["tsc"] = {
}
m["tsd"] = {
}
m["tse"] = {
}
m["tsg"] = {
aliases = {"Sūg"},
}
m["tsh"] = {
}
m["tsi"] = {
}
m["tsj"] = {
otherNames = {"Sharchop"},
}
m["tsl"] = {
}
m["tsm"] = {
}
m["tsp"] = {
}
m["tsq"] = {
}
m["tsr"] = {
}
m["tss"] = {
}
m["tsu"] = {
}
m["tsv"] = {
}
m["tsw"] = {
}
m["tsx"] = {
}
m["tsy"] = {
}
m["tta"] = {
}
m["ttb"] = {
}
m["ttc"] = {
}
m["ttd"] = {
}
m["tte"] = {
otherNames = {"Tubetube"},
}
m["ttf"] = {
}
m["ttg"] = {
}
m["tth"] = {
}
m["tti"] = {
}
m["ttj"] = {
aliases = {"Rutooro"},
}
m["ttk"] = {
aliases = {"Totoró"},
}
m["ttl"] = {
}
m["ttm"] = {
}
m["ttn"] = {
}
m["tto"] = {
}
m["ttp"] = {
}
m["ttr"] = {
}
m["tts"] = {
aliases = {"Isanese", "Isaan", "Issan", "Northeastern Thai"},
}
m["ttt"] = {
otherNames = {"Caucasian Tat", "Muslim Tat", "Armeno-Tat"},
}
m["ttu"] = {
}
m["ttv"] = {
}
m["ttw"] = {
otherNames = {"Tutoh"},
}
m["tty"] = {
}
m["ttz"] = {
}
m["tua"] = {
}
m["tub"] = {
}
m["tuc"] = {
}
m["tud"] = {
}
m["tue"] = {
}
m["tuf"] = {
}
m["tug"] = {
}
m["tuh"] = {
}
m["tui"] = {
}
m["tuj"] = {
}
m["tul"] = {
}
m["tum"] = {
}
m["tun"] = {
}
m["tuo"] = {
}
m["tuq"] = {
otherNames = {"Teda"},
}
m["tus"] = {
}
m["tuu"] = {
}
m["tuv"] = {
}
m["tux"] = {
}
m["tuy"] = {
}
m["tuz"] = {
}
m["tva"] = {
}
m["tvd"] = {
}
m["tve"] = {
}
m["tvk"] = {
}
m["tvl"] = {
}
m["tvm"] = {
}
m["tvn"] = {
}
m["tvo"] = {
}
m["tvs"] = {
}
m["tvt"] = {
}
m["tvu"] = {
otherNames = {"Tunen-Aling'a"},
}
m["tvw"] = {
}
m["tvx"] = {
}
m["tvy"] = {
otherNames = {"Bidau Creole Portuguese"},
}
m["twa"] = {
}
m["twb"] = {
}
m["twc"] = {
}
m["twe"] = {
otherNames = {"Tewa"},
}
m["twf"] = {
aliases = {"Northern Tiwa"},
}
m["twg"] = {
}
m["twh"] = {
aliases = {"Tai Khao", "White Tai"},
}
m["twm"] = {
}
m["twn"] = {
}
m["two"] = {
}
m["twp"] = {
}
m["twq"] = {
}
m["twr"] = {
}
m["twt"] = {
}
m["twu"] = {
}
m["tww"] = {
}
m["twy"] = {
otherNames = {"Taboyan"},
}
m["txa"] = {
}
m["txb"] = {
otherNames = {"West Tocharian", "Kuchean"},
}
m["txc"] = {
}
m["txe"] = {
}
m["txg"] = {
}
m["txj"] = {
}
m["txh"] = {
}
m["txi"] = {
}
m["txm"] = {
}
m["txn"] = {
}
m["txo"] = {
}
m["txq"] = {
}
m["txr"] = {
}
m["txs"] = {
}
m["txt"] = {
}
m["txu"] = {
}
m["txx"] = {
}
m["tya"] = {
}
m["tye"] = {
}
m["tyh"] = {
}
m["tyi"] = {
}
m["tyj"] = {
aliases = {"Tai Yo", "Tai Mène", "Tai Maen"},
}
m["tyl"] = {
}
m["tyn"] = {
}
m["typ"] = {
otherNames = {"Gugu Thaypan", "Thaypan", "Kuku Thaypan", "Agu Alaya", "Awu Alaya", "Alaya", "Gugu-Rarmul", "Koko-Rarmul", "Rarmul"},
}
m["tyr"] = {
aliases = {"Red Tai (Vietnam)"},
}
m["tys"] = {
aliases = {"Sa Pa", "Tày Sa Pa", "Tai Sapa"},
}
m["tyt"] = {
}
m["tyu"] = {
}
m["tyv"] = {
aliases = {"Tyvan"},
}
m["tyx"] = {
}
m["tyz"] = {
aliases = {"Tay", "Tho", "Bao Yen", "Cao Bang"}, -- Both Bao Lac and Trung Khanh are located in Cao Bang.
varieties = {"Central Tày", "Eastern Tày", "Northern Tày", "Southern Tày", "Tày Bao Lac", "Tày Trung Khanh"},
}
m["tza"] = {
}
m["tzh"] = {
}
m["tzj"] = {
aliases = {"Tzutujil"},
}
m["tzl"] = {
}
m["tzm"] = {
}
m["tzn"] = {
}
m["tzo"] = {
}
m["tzx"] = {
otherNames = {"Karawari"},
}
return m
t8jijj9iofqeq1rud9560uhc9d65jn2
Modul:gender and number/data
828
33823
375354
231400
2026-09-22T03:13:31Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/89447487|89447487]])
375354
Scribunto
text/plain
local data = {}
local insert = table.insert
-- A list of all possible "parts" that a specification can be made out of. For each part, we list the class it's in
-- (gender, animacy, etc.), the associated category (if any) and the display form. In a given gender/number spec, only
-- one part of each class is allowed. `display` is how the code is diplayed to the user and should normally be wrapped
-- in <abbr title="tooltip">...</abbr> with an explanatory tooltip. If not, it will automatically be wrapped in this
-- fashion. If `req` is true, a category "Requests for TYPE in LANG entries" will be generated, except for the code "?",
-- which is special-cased; TYPE is "gender" unless the POS is "verb", in which case it is "aspect".
data.codes = {
["?"] = {type = "other", req = true, display = '<abbr title="gender incomplete">?</abbr>'},
-- FIXME: The following should be either eliminated in favor of g! or converted to a general "gender/number unattested".
["?!"] = {type = "other", display = "genus tidak disahkan"},
-- Genders
["m"] = {type = "gender", cat = "POS maskulin", display = '<abbr title="masculine gender">m</abbr>'},
["f"] = {type = "gender", cat = "POS feminin", display = '<abbr title="feminine gender">f</abbr>'},
["n"] = {type = "gender", cat = "POS neuter", display = '<abbr title="neuter gender">n</abbr>'},
["c"] = {type = "gender", cat = "POS genus umum", display = '<abbr title="common gender">c</abbr>'},
["gneut"] = {type = "gender", cat = "POS bebas genus", display = "gender-neutral"},
["g!"] = {type = "gender", display = "genus tidak disahkan"},
["g?"] = {type = "gender", req = true, display = "genus tidak khusus"},
-- Animacy
-- Animate = either animal or personal (for Russian, etc.)
["an"] = {type = "animacy", cat = "POS bernyawa", display = '<abbr title="animate">anim</abbr>'},
["in"] = {type = "animacy", cat = "POS tidak bernyawa", display = '<abbr title="inanimate">inan</abbr>'},
-- Animal (for Ukrainian, Belarusian, Polish, etc.)
["anml"] = {type = "animacy", cat = "POS haiwan", display = "animal"},
-- Personal (for Ukrainian, Belarusian, Polish, etc.)
["pr"] = {type = "animacy", cat = "POS peribadi", display = '<abbr title="personal">pers</abbr>'},
["np"] = {type = "animacy", cat = "POS bukan peribadi", display = '<abbr title="nonpersonal">npers</abbr>'},
["an!"] = {type = "animacy", display = "kebernyawaan tidak disahkan"},
["an?"] = {type = "animacy", req = true, display = "kebernyawaan tidak khusus"},
-- Definiteness
["def"] = {type = "definiteness", cat = "POS muktamad", display = '<abbr title="definite">def</abbr>'},
["indef"] = {type = "definiteness", cat = "POS tak muktamad", display = '<abbr title="indefinite">indef</abbr>'},
-- Virility (for Polish)
["vr"] = {type = "virility", cat = "POS jantan", display = '<abbr title="virile (= masculine personal)">vir</abbr>'},
["nv"] = {type = "virility", cat = "POS bukan jantan", display = '<abbr title="nonvirile (= other than masculine personal)">nvir</abbr>'},
-- Numbers
["s"] = {type = "number", display = '<abbr title="singular number">sg</abbr>'},
["d"] = {type = "number", cat = "dualia tantum", display = '<abbr title="dual number">du</abbr>'},
["p"] = {type = "number", cat = "pluralia tantum", display = '<abbr title="plural number">pl</abbr>'},
["num!"] = {type = "number", display = "bilangan tidak disahkan"},
["num?"] = {type = "number", req = true, display = "bilangan tidak khusus"},
-- Verb qualifiers
["impf"] = {type = "aspect", cat = "POS tak sempurna", display = '<abbr title="imperfective aspect">impf</abbr>'},
["pf"] = {type = "aspect", cat = "POS sempurna", display = '<abbr title="perfective aspect">pf</abbr>'},
["asp!"] = {type = "aspect", display = "aspek tidak disahkan"},
["asp?"] = {type = "aspect", req = true, display = "aspek tidak khusus"},
}
-- Combined codes that are equivalent to giving multiple specs. `mf` is the same as specifying two separate specs,
-- one with `m` in it and the other with `f`. `mfbysense` is similar but is used for nouns that can be either masculine
-- or feminine according as to whether they refer to masculine or feminine beings.
local combinations = {
["biasp"] = {codes = {"impf", "pf"}},
["anin"] = {codes = {"an", "in"}}, -- "bianimate" doesn't exist as a linguistic term
}
for _, comb in ipairs{"mf", "mn", "fm", "fn", "cn", "nm", "nf", "nc", "mfn", "mnf", "fmn", "fnm", "nmf", "nfm"} do
local codes = {}
for ch in comb:gmatch(".") do
insert(codes, ch)
end
combinations[comb] = {codes = codes}
combinations[comb .. "equiv"] = {codes = codes, display = '<abbr title="different genders do not affect the meaning">same meaning</abbr>'}
if comb == "mf" or comb == "fm" then
combinations[comb .. "bysense"] = {codes = codes, cat = "masculine and feminine POS by sense",
display = '<abbr title="according to the gender of the referent">by sense</abbr>'}
end
end
data.combinations = combinations
-- Categories when multiple gender/number codes of a given type occur in different specs (two or more of the same type
-- cannot occur in a single spec).
data.multicode_cats = {
["gender"] = "POS dengan berbilang genus",
["animacy"] = "POS dengan berbilang kebernyawaan",
["aspect"] = "POS dwiaspek",
}
return data
1znket0vk37umsxisddunv6ol7cgsjt
Modul:module categorization
828
35046
375363
255151
2026-09-22T04:27:35Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92367848|92367848]])
375363
Scribunto
text/plain
local export = {}
local put_module = "Modul:parse utilities"
local rsplit = mw.text.split
local rfind = mw.ustring.find
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local insert = table.insert
local keyword_to_module_type = {
common = "Language-specific utility",
utilities = "Language-specific utility",
headword = "Headword-line",
translit = "transliterasi",
infl = "fleksi",
inflection = "fleksi",
decl = "fleksi",
declension = "fleksi",
adecl = "fleksi",
conj = "fleksi",
conjugation = "fleksi",
noun = "fleksi",
nouns = "fleksi",
pronoun = "fleksi",
pronouns = "fleksi",
verb = "fleksi",
verbs = "fleksi",
adjective = "fleksi",
adjectives = "fleksi",
adj = "fleksi",
nominal = "fleksi",
nominals = "fleksi",
pron = "sebutan",
pronun = "sebutan",
pronunc = "sebutan",
pronunciation = "sebutan",
IPA = "sebutan",
stripdiacritics = "penyiatan diakritik",
sortkey = "penjana kunci isih",
}
-- If a module type is here, we will generate a lang-specific module-type category such as
-- [[:Category:Pali inflection modules]].
local module_type_generates_lang_specific_cat = {
["fleksi"] = true,
["data"] = true,
["kes ujian"] = true,
}
-- If a module type is here, we will fetch extra languages to generate categories for. The value is a module that
-- returns a function that fetches all the languages that use a given module for
-- transliteration/diacritic-stripping/sortkey generation.
local languages_from_module_type = {
["transliterasi"] = "Modul:languages/byTranslitModule",
["kes ujian transliterasi"] = "Modul:languages/byTranslitModule",
["penyiatan diakritik"] = "Modul:languages/byStripDiacriticsModule",
["penjana kunci isih"] = "Modul:languages/bySortkeyModule",
}
local function get_transliteration_family_cats(objs)
local families_to_check = {"aav", "afa", "dra", "inc", "ira", "map", "sit", "sla", "trk", "urj"}
local result = {}
for _, famcode in ipairs(families_to_check) do
for _, obj in ipairs(objs) do
if obj:hasType("language") and obj:inFamily(famcode) then
local fam = require("Modul:families").getByCode(famcode)
insert(result, "Modul transliterasi bahasa-bahasa " .. fam:getCanonicalName())
break
end
end
end
return result
end
-- If a module type is here, we will fetch extra categories given the generated languages/scripts/families.
local extra_cats_from_module_type = {
["transliterasi"] = get_transliteration_family_cats,
}
local module_type_patterns = {
{"/data%f[-/%z]", "data"},
{"/testcases%f[-/%z]", function(typ)
if typ == "sebutan" then
return "kes ujian sebutan"
elseif typ == "transliterasi" then
return "kes ujian transliterasi"
else
return "kes ujian"
end
end},
}
-- Split an argument on comma, but not comma followed by whitespace.
local function split_on_comma(val)
if val:find(",%s") then
return require(put_module).split_on_comma(val)
else
return rsplit(val, ",")
end
end
local function get_lang_or_script(code)
return code == "-" and code or
require("Modul:languages").getByCode(code, nil, "allow etym") or
require("Modul:languages").getByCode(code .. "-pro", nil, "allow etym") or
require("Modul:scripts").getByCode(code)
end
local function obj_code(obj)
if obj == "-" then
return obj
end
return obj:getCode()
end
local function infer_lang_or_script_code(name)
local hyphen_parts = rsplit(name, "%-")
for i = #hyphen_parts - 1, 1, -1 do
local code = table.concat(hyphen_parts, "-", 1, i)
local obj = get_lang_or_script(code)
if obj then
local rest = table.concat(hyphen_parts, "-", i + 1)
return obj, rest
end
end
return nil, nil
end
local function infer_lang_and_script_codes(name)
local objs = {}
while true do
local obj, rest = infer_lang_or_script_code(name)
if not obj then
return objs, name
end
if #objs > 0 and obj:getCode() == "to" then
-- skip 'to' in e.g. [[Modul:ks-Arab-to-Deva-translit]]; it's not Tongan
else
insert(objs, obj)
end
name = rest
end
end
--[==[
Main entry point called from another module.
]==]
function export.categorize_module(data)
local pagename, return_raw, noerror = data.pagename, data.return_raw, data.noerror
local langlist, module_type, return_cats = data.langlist, data.module_type, data.return_cats
local title
if pagename then
title = mw.title.new(pagename, 'Modul')
else
title = mw.title.getCurrentTitle()
-- Fuckme, sometimes this function is called with a faked frame and a title with the namespace already chopped out,
-- so this test cannot be done in that case.
if title.nsText ~= "Modul" then
error(("This template should only be used in the Module namespace, not on page '%s'."):format(title.fullText))
end
pagename = title.fullText
end
local subpage = title.subpageText
local null_return_value = return_raw and {} or ""
-- To ensure no categories are added on documentation pages.
if subpage == "doc" then
return null_return_value
end
local categories = {}
local function insert_cat(cat, sortkey)
for _, existing_cat in ipairs(categories) do
if existing_cat.name == cat then
return
end
end
insert(categories, {name = cat, sort = sortkey})
end
local root_pagename
if subpage ~= pagename then
root_pagename = title.rootText
else
root_pagename = pagename
end
root_pagename = root_pagename:gsub("^Modul:", "")
-- Take the module type(s) from type= if given, or infer from the pagename.
local module_types
if module_type then
module_types = {}
local module_type_specs = split_on_comma(module_type)
for _, spec in ipairs(module_type_specs) do
local modtype, sortkey = spec:match("^(.-):(.*)$")
modtype = modtype or spec
sortkey = sortkey and sortkey:gsub("_", " ") or nil
insert(module_types, {type = modtype, sort = sortkey})
end
else
local module_type_keyword = root_pagename:match("[-%a]+[- ]([^/]+)%f[/%z]")
if not module_type_keyword then
if noerror then
return null_return_value
else
error(("Could not extract module type from root pagename '%s'"):format(root_pagename))
end
end
module_type = keyword_to_module_type[module_type_keyword]
if not module_type then
if noerror then
return null_return_value
else
error(("Did not recognize inferred module-type keyword '%s' from root pagename '%s'"):format(
module_type_keyword, root_pagename))
end
end
module_types = {{type = module_type}}
end
-- Look for additional module type(s) inferred by pattern.
for _, pattern_spec in ipairs(module_type_patterns) do
local pattern, inferred_type = unpack(pattern_spec)
if rfind(pagename, pattern) then
local function insert_module_type(typ)
require("Modul:table").insertIfNot(module_types, typ, {key = function(obj) return obj.type end})
end
if type(inferred_type) == "string" then
insert_module_type({type = inferred_type})
else
local addl_types = {}
for _, typ in ipairs(module_types) do
insert(addl_types, {type = inferred_type(typ.type), sort = typ.sort})
end
for _, typ in ipairs(addl_types) do
insert_module_type(typ)
end
end
end
end
-- If 1= specified, take the languages/scripts directly from there. Otherwise, (a) try to extract one or more
-- languages/scripts from the pagename (e.g. [[Modul:uk-be-headword]] -> Ukrainian and Belarusian (languages);
-- [[Modul:bho-Kthi-translit]] -> Bhojpuri (language) and Kaithi (script); [[Modul:Deva-Kthi-translit]] ->
-- Devanagari and Kaithi (scripts)); and (b) if the specified or inferred module type(s) contain a type listed in
-- languages_from_module_type[], use the function referenced there to extract additional languages (i.e. all the
-- languages that use the module we are processing).
local inferred_objs
if langlist then
inferred_objs = {}
for _, code in ipairs(rsplit(langlist, ",")) do
-- We need to have an indicator of families because we allow bare family codes to stand for proto-languages.
if code:find("^fam:") then
code = code:gsub("^fam:", "")
local family = require("Modul:families").getByCode(code) or
error(("Unrecognized family code '%s' in [[Modul:module categorization]]"):format(code))
local descendants = family:getDescendantCodes()
for _, desc in ipairs(descendants) do
local obj = get_lang_or_script(desc)
if obj then
-- make sure we skip families without proto-languages
insert(inferred_objs, obj)
end
end
else
local obj = get_lang_or_script(code)
if not obj then
error(("Unrecognized language or script code '%s'"):format(code))
end
insert(inferred_objs, obj)
end
end
else
inferred_objs = infer_lang_and_script_codes(root_pagename)
for _, modtype in ipairs(module_types) do
local languages_extractor = languages_from_module_type[modtype.type]
if languages_extractor then
local langs = require(languages_extractor)(root_pagename)
if langs then
for _, obj in ipairs(langs) do
require("Modul:table").insertIfNot(inferred_objs, obj, {key = obj_code})
end
end
end
end
if #inferred_objs == 0 then
if noerror then
return null_return_value
else
error(("Could not infer any languages or scripts from root pagename '%s'"):format(root_pagename))
end
end
end
if pagename:find("^Modul:Pengguna:") then
insert_cat("Modul kotak pasir pengguna")
elseif pagename:find("/sandbox") then
insert_cat("Modul kotak pasir")
else
for _, modtype in ipairs(module_types) do
for _, obj in ipairs(inferred_objs) do
local function insert_overall_module_type_cat(sortkey)
if modtype.type ~= "-" then
insert_cat("Modul " .. modtype.type, modtype.sort or sortkey)
end
end
if obj == "-" then
insert_overall_module_type_cat()
else
if obj:hasType("script") and modtype.type ~= "-" then
insert_cat("Modul " .. modtype.type .. " mengikut tulisan", obj:getCanonicalName())
end
local function construct_lang_or_sc_cat(obj, suffix)
local prefix
if obj:hasType("language") then
prefix = obj:getFullName()
else
prefix = obj:getCategoryName()
end
return suffix .. " bahasa " .. prefix
end
insert_cat(construct_lang_or_sc_cat(obj, "Modul"), modtype.type)
insert_overall_module_type_cat(obj:getCanonicalName())
if module_type_generates_lang_specific_cat[modtype.type] then
insert_cat(construct_lang_or_sc_cat(obj, "Modul bahasa " .. mw.getContentLanguage():lcfirst(modtype.type)))
end
end
end
if extra_cats_from_module_type[modtype.type] then
local extra_cats = extra_cats_from_module_type[modtype.type](inferred_objs)
for _, cat in ipairs(extra_cats) do
insert_cat(cat)
end
end
end
end
for i, catspec in ipairs(categories) do
if catspec.sort then
categories[i] = ("%s|%s"):format(catspec.name, catspec.sort)
else
categories[i] = catspec.name
end
end
if return_cats then
return table.concat(categories, ",")
elseif return_raw then
return categories
else
for i, cat in ipairs(categories) do
categories[i] = "[[Kategori:" .. cat .. "]]"
end
return table.concat(categories)
end
end
--[==[
Main entry point called from a template.
]==]
function export.categorize(frame)
local params = {
[1] = true, -- comma-separated list of languages; by default, inferred from module name
type = true,
[2] = {alias_of = "type"},
pagename = true, -- for testing
return_cats = {type = "boolean"}, -- for testing
}
local parent_args = frame:getParent().args
local args = require("Modul:parameters").process(parent_args, params)
return export.categorize_module {
pagename = args.pagename,
langlist = args[1],
module_type = args.type,
return_cats = args.return_cats,
}
end
--[==[Table used in the documentation to {{tl|module cat}}.]==]
function export.keyword_to_module_type_table()
local parts = {}
local function ins(text)
insert(parts, text)
end
ins('{|class="wikitable"')
ins("! Kata kunci !! Jenis modul disimpulkan")
local keywords = {}
for k, v in pairs(keyword_to_module_type) do
insert(keywords, k)
end
table.sort(keywords)
for _, keyword in ipairs(keywords) do
ins("|-")
ins(("| <code>%s</code> || <code>%s</code>"):format(keyword, keyword_to_module_type[keyword]))
end
ins("|}")
return table.concat(parts, "\n")
end
return export
nsy77mhqd1amfc11l0mjbxjw2vhai2t
Modul:audio
828
48711
375391
226683
2026-09-22T07:36:06Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/88355452|88355452]])
375391
Scribunto
text/plain
local export = {}
local headword_data_module = "Module:headword/data"
local IPA_module = "Module:IPA"
local labels_module = "Module:labels"
local links_module = "Module:links"
local parameters_module = "Module:parameters"
local qualifier_module = "Module:qualifier"
local references_module = "Module:references"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local template_styles_module = "Module:TemplateStyles"
local utilities_module = "Module:utilities"
local audio_styles_css = "audio/styles.css"
local function track(page)
require("Module:debug/track")("audio/" .. page)
return true
end
local function wrap_qualifier_css(text, suffix)
return require(qualifier_module).wrap_qualifier_css(text, suffix)
end
--[==[
Display a box that can be used to play an audio file. `data` is a table containing the following fields:
* `lang` ('''required'''): language object for the audio files;
* `file` ('''required'''): file containing the audio;
* `caption`: Caption to display before the audio box; normally {"Audio"}, and does not usually need to be changed;
* `nocaption`: If specified, don't display the caption;
* `q`: {nil} or a list of left regular qualifier strings, formatted using {format_qualifier()} in [[Module:qualifier]]
and displayed before the audio box and after the caption (and any accent qualifiers);
* `qq`: {nil} or a list of right regular qualifier strings, displayed directly after the audio box (and after any
accent qualifiers);
* `a`: {nil} or a list of left accent qualifier strings, formatted using {format_qualifiers()} in
[[Module:accent qualifier]] and displayed before the audio box and after the caption;
* `aa`: {nil} or a list of right accent qualifier strings, displayed directly after the homophone in question;
* `refs`: {nil} or a list of references or reference specs to add directly after the audio box; the value of a list item
is either a string containing the reference text (typically a call to a citation template such as {{tl|cite-book}}, or
a template wrapping such a call), or an object with fields `text` (the reference text), `name` (the name of the
reference, as in {{cd|<nowiki><ref name="foo">...</ref></nowiki>}} or {{cd|<nowiki><ref name="foo" /></nowiki>}})
and/or `group` (the group of the reference, as in {{cd|<nowiki><ref name="foo" group="bar">...</ref></nowiki>}} or
{{cd|<nowiki><ref name="foo" group="bar"/></nowiki>}}); this uses a parser function to format the reference
appropriately and insert a footnote number that hyperlinks to the actual reference, located in the
{{cd|<nowiki><references /></nowiki>}} section;
* `text`: Text of the audio snippet; if specified, should be an object of the form passed to {full_link()} in
[[Module:links]], including a `lang` field containing the language of the text (usually the same as `data.lang`);
displayed before the audio box, after any regular and accent qualifiers;
* `IPA`: IPA of the audio snippet, or a list of IPA specs; if specified, should be surrounded by slashes or brackets,
and will be processed using {format_IPA_multiple()} in [[Module:IPA]] and displayed before the audio box, after any
regular and accent qualifiers and after the text of the audio snippet, if given;
* `nocat`: If true, suppress categorization;
* `sort`: Sort key for categorization.
]==]
function export.format_audio(data)
local langname = data.lang:getFullName()
local cats = { "Perkataan bahasa " .. langname .. " dengan sebutan audio" }
local function format_a(a)
if a and a[1] then
return require(labels_module).show_labels {
lang = data.lang,
labels = a,
mode = "accent",
nocat = true,
open = false,
close = false,
no_track_already_seen = true,
}
end
return nil
end
local function format_q(q)
if q and q[1] then
return require(qualifier_module).format_qualifier(q, false, false)
end
return nil
end
local function make_td_if(text)
if text == "" then
return text
end
return "<td>" .. text .. "</td>"
end
-- Generate the full text preceding the audio box.
local pretext_parts = {}
local function ins(text)
table.insert(pretext_parts, text)
end
local formatted_accent_labels, formatted_qualifiers, formatted_text, formatted_ipa
formatted_accent_labels = format_a(data.a)
formatted_qualifiers = format_q(data.q)
if data.text then
formatted_text = require(links_module).full_link(data.text, "term", true)
end
if data.IPA then
local ipa_cats
local ipa = data.IPA
if type(ipa) == "string" then
ipa = {ipa}
end
local ipa_items = {}
for _, ipa_item in ipairs(ipa) do
table.insert(ipa_items, {pron = ipa_item})
end
formatted_ipa, ipa_cats = require(IPA_module).format_IPA_multiple(data.lang, ipa_items, nil, "no count", "raw")
if ipa_cats[1] then
require(table_module).extend(cats, ipa_cats)
end
end
local has_qual = formatted_accent_labels or formatted_qualifiers
if not data.nocaption then
-- Track uses of caption (3=). Over time as we eliminate most of them, we can use this to find and
-- eliminate the remainder.
if data.caption then
track("caption")
end
ins(data.caption or "Audio")
if has_qual then
ins(" " .. wrap_qualifier_css("(", "brac"))
end
end
if formatted_accent_labels then
ins(formatted_accent_labels)
if formatted_qualifiers then
ins(wrap_qualifier_css(",", "comma") .. " ")
end
end
if formatted_qualifiers then
ins(formatted_qualifiers)
end
if has_qual then
if not data.nocaption then
ins(wrap_qualifier_css(")", "brac"))
end
end
if (formatted_text or formatted_ipa) and (has_qual or not data.nocaption) then
ins(wrap_qualifier_css(";", "semicolon") .. " ")
end
if formatted_text then
ins(formatted_text)
if formatted_ipa then
ins(" ")
end
end
ins(formatted_ipa)
if not data.nocaption then
ins(wrap_qualifier_css(":", "colon"))
end
local pretext = make_td_if(table.concat(pretext_parts))
-- Generate the full text following the audio box.
local posttext_parts = {}
local function ins(text)
table.insert(posttext_parts, text)
end
local formatted_post_accent_labels = format_a(data.aa)
local formatted_post_qualifiers = format_q(data.qq)
local formatted_references = data.refs and require(references_module).format_references(data.refs) or nil
if formatted_references then
ins(formatted_references)
end
if formatted_post_accent_labels or formatted_post_qualifiers then
if formatted_references then
ins(" ")
end
ins(wrap_qualifier_css("(", "brac"))
if formatted_post_accent_labels then
ins(formatted_post_accent_labels)
if formatted_post_qualifiers then
ins(wrap_qualifier_css(",", "comma") .. " ")
end
end
if formatted_post_qualifiers then
ins(formatted_post_qualifiers)
end
ins(wrap_qualifier_css(")", "brac"))
end
if data.bad then
table.insert(cats, langname .. " terms with nonstandard or incorrect audio pronunciations")
ins(" " .. require(qualifier_module).wrap_css("Note: this pronunciation may be nonstandard or incorrect: " .. data.bad, "bad-audio-note"))
end
local posttext = make_td_if(table.concat(posttext_parts))
local template = [=[
<tr>%s<td class="audiofile">[[File:%s|noicon|175px]]</td><td class="audiometa" style="font-size: 80%%;">([[:File:%s|file]])</td>%s</tr>]=]
local text = template:format(pretext, data.file, data.file, posttext)
text = '<table class="audiotable" style="vertical-align: middle; display: inline-block; list-style: none; line-height: 1em; border-collapse: collapse; margin: 0;">' .. text .. "</table>"
local stylesheet = require(template_styles_module)(audio_styles_css)
local categories =
data.nocat and "" or
cats[1] and require(utilities_module).format_categories(cats, data.lang, data.sort) or ""
return stylesheet .. text .. categories
end
--[==[
FIXME: Old entry point for formatting multiple audios in a single table. Not used anywhere and needs rewriting to the
standard of format_audio().
Meant to be called from a module. `data` is a table containing the following fields:
<pre>
{
lang = LANGUAGE_OBJECT,
audios = {{file = "FILENAME", qualifiers = nil or {"QUALIFIER", "QUALIFIER", ...}}, ...},
caption = nil or "CAPTION"
}
</pre>
Here:
* `lang` is a language object.
* `audios` is the list of audio files to display. FILENAME is the name of the audio file without a namespace.
QUALIFIER is a qualifier string to display after the specific audio file in question, formatted using
{format_qualifier()} in [[Module:qualifier]].
* `caption`, if specified, adds a caption before the audio file.
]==]
function export.format_multiple_audios(data)
local audiocats = { "Perkataan bahasa " .. data.lang:getFullName() .. " dengan sebutan audio" }
local rows = { }
local caption = data.caption
for _, audio in ipairs(data.audios) do
local qualifiers = audio.qualifiers
local function repl(key)
if key == "file" then
return audio.file
elseif key == "caption" then
if not caption then return "" end
return "<td rowspan=" .. #data.audios .. ">" .. caption .. ":</td>"
elseif key == "qualifiers" then
if not qualifiers or not qualifiers[1] then return "" end
return "<td>" .. require(qualifier_module).format_qualifier(qualifiers) .. "</td>"
end
end
local template = [=[
<tr>{{{caption}}}
<td class="audiofile">[[File:{{{file}}}|noicon|175px]]</td>
<td class="audiometa" style="font-size: 80%;">([[:File:{{{file}}}|file]])</td>
{{{qualifiers}}}</tr>]=]
local text = (mw.ustring.gsub(template, "{{{([a-z0-9_:]+)}}}", repl))
table.insert(rows, text)
caption = nil
end
local function repl(key)
if key == "rows" then
return table.concat(rows, "\n")
end
end
local template = [=[
<table class="audiotable" style="vertical-align: middle; display: inline-block; list-style: none; line-height: 1em; border-collapse: collapse;">
{{{rows}}}
</table>
]=]
local stylesheet = require(template_styles_module)(audio_styles_css)
local text = mw.ustring.gsub(template, "{{{([a-z0-9_:]+)}}}", repl)
local categories =
data.nocat and "" or
#audiocats > 0 and require(utilities_module).format_categories(audiocats, data.lang, data.sort) or ""
-- remove newlines due to HTML generator bug in MediaWiki(?) - newlines in tables cause list items to not end correctly
text = mw.ustring.gsub(text, "\n", "")
return stylesheet .. text .. categories
end
--[==[
Construct the `text` object passed into {format_audio()}, from raw-ish arguments (essentially, the output of {process()}
in [[Module:parameters]]). On entry, `args` contains the following fields:
* `lang` ('''required'''): Language object.
* `text`: Text. If this isn't defined and neither are any of `gloss`, `tr`, `ts`, `pos`, `lit` or `genders`, the
function returns {nil}.
* `gloss`: Gloss of text.
* `tr`: Manual transliteration of text.
* `ts`: Transcription of text.
* `pos`: Part of speech of text.
* `lit`: Literal meaning of text.
* `genders`: List of gender/number spec(s) of text.
* `sc`: Optional script object of text (rarely needs to be set).
* `pagename`: Pagename; used in place of `text` when `text` is unset but other text-related parameters are set.
If not specified, taken from the actual pagename.
]==]
function export.construct_audio_textobj(args)
local textobj
if args.text or args.gloss or args.tr or args.ts or args.pos or args.lit or args.genders and args.genders[1] then
local text = args.text or args.pagename or mw.loadData("Module:headword/data").pagename
textobj = {
lang = args.lang,
alt = wrap_qualifier_css("“", "quote") .. text .. wrap_qualifier_css("”", "quote"),
gloss = args.gloss,
tr = args.tr,
ts = args.ts,
pos = args.pos,
lit = args.lit,
genders = args.genders,
sc = args.sc,
}
end
return textobj
end
--[==[
Entry point for {{tl|audio}} template.
]==]
function export.show(frame)
local parent_args = frame:getParent().args
local compat = parent_args.lang
local offset = compat and 0 or 1
local params = {
[compat and "lang" or 1] = {required = true, type = "language", default = "en"},
[1 + offset] = {required = true, default = "Example.ogg"},
[2 + offset] = {},
["q"] = {type = "qualifier"},
["qq"] = {type = "qualifier"},
["a"] = {type = "labels"},
["aa"] = {type = "labels"},
["ref"] = {type = "references"},
["IPA"] = {sublist = true},
["text"] = {},
["t"] = {},
["gloss"] = {alias_of = "t"},
["tr"] = {},
["ts"] = {},
["pos"] = {},
["lit"] = {},
["g"] = {sublist = true},
["sc"] = {type = "script"},
["bad"] = {},
["nocat"] = {type = "boolean"},
["sort"] = {},
["pagename"] = {},
}
local args = require(parameters_module).process(parent_args, params)
local lang = args[compat and "lang" or 1]
-- Needed in construct_audio_textobj().
args.lang = lang
local textobj = export.construct_audio_textobj(args)
local caption = args[2 + offset]
local nocaption
if caption == "-" then
caption = nil
nocaption = true
end
if caption then
-- Remove final colon if given, to avoid two colons.
caption = caption:gsub(":$", "")
end
local data = {
lang = lang,
file = args[1 + offset],
caption = caption,
nocaption = nocaption,
q = args.q,
qq = args.qq,
a = args.a,
aa = args.aa,
refs = args.ref,
text = textobj,
IPA = args.IPA,
bad = args.bad,
nocat = args.nocat,
sort = args.sort,
}
return export.format_audio(data)
end
return export
anshp52gfmccr0u5s0khcpl3vgtaamp
Modul:headword utilities
828
54854
375352
223188
2026-09-22T03:13:11Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[:en:Special:Diff/92717011|92717011]])
375352
Scribunto
text/plain
local export = {}
local require_when_needed = require("Module:utilities/require when needed")
local affix_module = "Module:affix"
local debug_track_module = "Module:debug/track"
local decorations_module = "Module:decorations"
local en_utilities_module = "Module:en-utilities"
local fun_is_callable_module = "Module:fun/isCallable"
local headword_module = "Module:headword"
local headword_data_module = "Module:headword/data"
local languages_module = "Module:languages"
local links_module = "Module:links"
local parameters_module = "Module:parameters"
local parse_interface_module = "Module:parse interface"
local parse_utilities_module = "Module:parse utilities"
local string_pattern_escape_module = "Module:string/patternEscape"
local string_replacement_escape_module = "Module:string/replacementEscape"
local string_utilities_module = "Module:string utilities"
local table_module = "Module:table"
local yesno_module = "Module:yesno"
local dump = mw.dumpObject
local unpack = unpack or table.unpack -- Lua 5.2 compatibility
local insert = table.insert
local concat = table.concat
local remove = table.remove
local sort = table.sort
local deep_equals = require_when_needed(table_module, "deepEquals")
local extend = require_when_needed(table_module, "extend")
local insert_if_not = require_when_needed(table_module, "insertIfNot")
local list_to_set = require_when_needed(table_module, "listToSet")
local serial_comma_join = require_when_needed(table_module, "serialCommaJoin")
local shallow_copy = require_when_needed(table_module, "shallowCopy")
local split = require_when_needed(string_utilities_module, "split")
local ugsub = require_when_needed(string_utilities_module, "gsub")
local umatch = require_when_needed(string_utilities_module, "match")
local pattern_escape = require_when_needed(string_pattern_escape_module)
local replacement_escape = require_when_needed(string_replacement_escape_module)
local escape_wikicode = require_when_needed(parse_utilities_module, "escape_wikicode")
local parse_inline_modifiers = require_when_needed(parse_utilities_module, "parse_inline_modifiers")
local term_contains_top_level_html = require_when_needed(parse_utilities_module, "term_contains_top_level_html")
local get_lang_by_code = require_when_needed(languages_module, "getByCode")
local is_callable = require_when_needed(fun_is_callable_module)
local format_decorations = require_when_needed(decorations_module, "format_decorations")
local function split_on_comma(val)
if val:find(",") then
return require(parse_interface_module).split_on_comma(val)
else
return {val}
end
end
local function ine(val)
if val == "" then return nil else return val end
end
--[=[
Add decorations to a term. `termobj` is the object describing the term, which should optionally contain:
* left qualifiers in `q`, an array of strings;
* right qualifiers in `qq`, an array of strings;
* left labels in `l`, an array of strings;
* right labels in `ll`, an array of strings;
* references in `refs`, an array either of strings (formatted reference text) or objects containing fields `text`
(formatted reference text) and optionally `name` and/or `group`;
`text` is the text of the term itself, and `lang` is the language object.
]=]
local function add_decorations(text, termobj, lang)
local function field_non_empty(field)
local list = termobj[field]
if not list then
return nil
end
if type(list) ~= "table" then
error(("Internal error: Wrong type for `termobj.%s`=%s, should be \"table\""):format(
field, mw.dumpObject(list)))
end
return list[1]
end
if field_non_empty("q") or field_non_empty("qq") or field_non_empty("l") or field_non_empty("ll") or
field_non_empty("refs") then
text = format_decorations {
lang = lang,
text = text,
q = termobj.q,
qq = termobj.qq,
l = termobj.l,
ll = termobj.ll,
refs = termobj.refs,
}
end
return text
end
local param_mods = {
id = {}, -- disabled when `is_head = true`
q = {type = "qualifier"},
qq = {type = "qualifier"},
l = {type = "labels"},
ll = {type = "labels"},
-- [[Module:headword]] expects part references in `.refs`.
ref = {item_dest = "refs", type = "references", store = "insert-flattened"},
}
local optional_param_mods = {
g = {item_dest = "genders", type = "genders"},
alt = {},
lang = {type = "language"},
sc = {type = "script"},
t = {item_dest = "gloss"},
gloss = {},
pos = {},
lit = {},
tr = {},
ts = {},
face = {},
nolinkinfl = {type = "boolean"},
}
local optional_headword_param_mods = {
sc = {type = "script"},
tr = {},
ts = {},
}
--[==[
Parse a single inflection or headword form or list of such forms. In either case, inline modifiers may be attached.
`data` is an object with the following fields:
* `val`: The raw value to parse. Required.
* `paramname`: The name of the parameter from which the value was taken; used in error messages. Required.
* `is_head`: We are parsing a headword parameter (a value which goes into the `heads` field of `data`). This changes
the allowed modifiers, disabling the `id` modifier and only allowing a subset of optional modifiers.
* `frob`: An optional function of one value to apply to the form after inline modifiers have been removed (i.e. to
apply to the `.term` field of the returned object).
* `include_mods`: List of extra inline modifiers to include, besides the default ones (see below). Each list item is
either a string specifying a recognized extra inline modifier (see `optional_param_mods` in the code), or a two-item
list of modifier name and modifier spec, where the spec should follow the syntax for modifier specs in
`parse_inline_modifiers` in [[Module:parse utilities]].
* `exclude_mods`: List of default inline modifiers to not include.
* `splitchar`: If specified, the value in `val` can be a list of forms to parse, separated by the value of `splitchar`
(which is a Lua pattern, as in `parse_inline_modifiers` in [[Module:parse utilities]]). Most commonly, `splitchar` is
a single comma and the values are comma-separated (in this case, splitting will not happen if a space follows the
comma).
* `parse_lang_prefix`: If specified, allow a language prefix to precede a form, and if found, store into the `.lang`
field of the returned object.
* `preserve_splitchar`, `delimiter_key`, `escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in
`parse_inline_modifiers` in [[Module:parse utilities]].
Returns an object suitable for storing as one element of one of the lists in `headdata.inflections`, where `headdata`
is the structure passed to [[Module:headword]]. If `splitchar` is specified, howeve, the return value is a list of such
objects.
The following default inline modifiers are currently recognized:
* `q`: Left qualifier.
* `qq`: Right qualifier.
* `l`: Comma-separated list of left labels. No space should follow the comma.
* `ll`: Comma-separated list of right labels. No space should follow the comma.
* `ref`: Reference or references. See {{tl|IPA}} for the syntax.
* `id`: Sense ID, in case there are multiple senses. See {{tl|l}}.
The following are the recognized additional inline modifiers:
* `g`: Comma-separated list of genders.
* `alt`: Display text.
* `lang`: Language code of language of the form, if different from the language of the headword.
* `sc`: Script code of script of the form. Almost never needed.
* `t`: Gloss for the form.
* `gloss`: Gloss for the form (alias for `t`).
* `pos`: Part of speech of the form.
* `lit`: Literal meaning of the form.
* `tr`: Manual transliteration of the form.
* `ts`: Transcription of the form, for languages where the transliteration differs markedly from the pronunciation.
* `face`: Face to display the form in, e.g. {"hypothetical"} for a hypothetical form (unlinkable and displayed in italics).
* `nolinkinfl`: Make the form unlinkable.
]==]
function export.parse_term_with_modifiers(data)
local paramname, val, frob = data.paramname, data.val, data.frob
local function generate_obj(term, parse_err)
if frob then
term = frob(term, parse_err)
end
if data.parse_lang_prefix and term:find(":") then
return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix {
term = term,
paramname = paramname,
parse_lang_prefix = true,
parse_err = parse_err,
}
else
return {term = term}
end
end
-- Check for inline modifier, e.g. מרים<tr:Miryem>. But exclude top-level HTML entry with <span ...>,
-- <sup> or similar in it.
if (val:find("<", nil, true) or data.splitchar) and not term_contains_top_level_html(val) and
-- don't parse inline modifiers if is_head and the value begins with a ~ (link modifier syntax)
(not data.is_head or not val:find("^~")) then
local param_mods = param_mods
if data.is_head then
param_mods = shallow_copy(param_mods)
param_mods.id = nil
end
if data.include_mods or data.exclude_mods then
if not data.is_head then
-- already copied when data.is_head
param_mods = shallow_copy(param_mods)
end
if data.include_mods then
local optional_mods = data.is_head and optional_headword_param_mods or optional_param_mods
for _, mod in ipairs(data.include_mods) do
if type(mod) == "table" then
if #mod ~= 2 then
error(("Internal error: Modifier spec %s in `include_mods` should be of length 2"):format(
dump(mod)))
end
local modkey, modvalue = unpack(mod)
param_mods[modkey] = modvalue
elseif not optional_mods[mod] then
error(("Internal error: Unrecognized modifier spec %s in `include_mods`"):format(
dump(mod)))
else
param_mods[mod] = optional_mods[mod]
end
end
end
if data.exclude_mods then
for _, mod in ipairs(data.exclude_mods) do
if not param_mods[mod] then
error(("Internal error: Modifier spec %s in `exclude_mods` not found among existing modifiers"
):format(dump(mod)))
else
param_mods[mod] = nil
end
end
end
end
return parse_inline_modifiers(val, {
paramname = paramname,
param_mods = param_mods,
generate_obj = generate_obj,
splitchar = data.splitchar,
preserve_splitchar = data.preserve_splitchar,
delimiter_key = data.delimiter_key,
escape_fun = data.escape_fun,
unescape_fun = data.unescape_fun,
pre_normalize_modifiers = data.pre_normalize_modifiers,
})
else
local retval = generate_obj(val)
if data.splitchar then
retval = {retval}
end
return retval
end
end
--[==[
Parse a list of inflection forms that may have inline modifiers attached. `data` is an object with the following fields:
* `forms`: The list of raw values to parse. Required.
* `paramname`: The name of the first parameter from which the value was taken; used in error messages. If this is a
two-element list, the first element is the first parameter and the second element is the prefix of the remaining
parameters. Parameter names that are numbers are handled correctly, as are those with \1 in it marking where the
parameter index goes. Required.
* `qualifiers`: If specified, a possibly gappy list of left qualifiers to add to the parsed terms (for compatibility
purposes).
* `splitchar`: As in `parse_term_with_modifiers()`. The resulting per-term lists will be flattened.
* `frob`, `include_mods`, `exclude_mods`, `is_head`, `preserve_splitchar`, `parse_lang_prefix`, `delimiter_key`,
`escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_with_modifiers()`.
Returns a list of objects, suitable for storing as one of the lists in `headdata.inflections` (once a label is added),
where `headdata` is the structure passed to [[Module:headword]].
]==]
function export.parse_term_list_with_modifiers(data)
local paramname, forms = data.paramname, data.forms
local qualifiers = data.qualifiers
local first, restpref
if type(paramname) == "table" then
first = paramname[1]
restpref = paramname[2]
else
first = paramname
restpref = paramname
end
local terms = {}
data = shallow_copy(data)
for i, val in ipairs(forms) do
data.paramname = i == 1 and first or type(restpref) == "number" and restpref + i - 1 or
restpref:find("\1", nil, true) and restpref:gsub("\1", tostring(i)) or restpref .. i
data.val = val
local parsed = export.parse_term_with_modifiers(data)
if qualifiers and qualifiers[i] then
if data.splitchar then
for _, term in ipairs(parsed) do
term.q = {qualifiers[i]}
end
else
parsed.q = {qualifiers[i]}
end
end
if data.splitchar then
extend(terms, parsed)
else
terms[i] = parsed
end
end
return terms
end
--[==[
Construct a link to [[Appendix:Glossary]] for `entry`. If `text` is specified, it is the display text; otherwise,
`entry` is used.
]==]
function export.glossary_link(entry, text)
text = text or entry
return "[[Lampiran:Glosari#" .. entry .. "|" .. text .. "]]"
end
function export.replace_glossary_links_in_label(label)
if label:find("<<", nil, true) then
label = label:gsub("<<(.-)|(.-)>>", export.glossary_link):gsub("<<(.-)>>", export.glossary_link)
end
return label
end
--[==[
Insert a fixed inflection (a label not associated with any inflection values) into an `inflections` field. The
`inflections` field will be initialized if needed. `data` is an object with the following fields:
* `headdata`: The headword structure passed to [[Module:headword]]. Required.
* `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs
to be set for nested inflections, which are specified for an inflection object rather than the headword data
structure as a whole.
* `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the
glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required.
* `originating_term`: The term object from which this label is derived. If specified, decorations will be taken from
this object.
]==]
function export.insert_fixed_inflection(data)
local headdata, origterm, label = data.headdata, data.originating_term, data.label
local inflobj = data.inflobj or headdata
inflobj.inflections = inflobj.inflections or {}
if not origterm then
insert(inflobj.inflections, {
label = export.replace_glossary_links_in_label(label)
})
else
if origterm.id then
error(("It doesn't make sense to pass in an ID '%s' for label '%s' in conjunction with a term value '%s'"
):format(origterm.id, label, origterm.term))
end
origterm = shallow_copy(origterm)
-- Preserve decorations
origterm.term = nil
origterm.label = export.replace_glossary_links_in_label(label)
insert(inflobj.inflections, origterm)
end
end
--[==[
Insert previously-parsed terms into an `inflections` field. The `inflections` field will be initialized if needed.
`data` is an object with the following fields:
* `headdata`: The headword structure passed to [[Module:headword]]. Required.
* `inflobj`: The object whose `inflections` field the terms are inserted into. Defaults to `headdata`. Only needs
to be set for nested inflections, which are specified for an inflection object rather than the headword data
structure as a whole.
* `terms`: The list of parsed terms. If {nil} or omitted, nothing happens unless `request` is set.
* `label`: The label that the inflections are given; any parts of the label surrounded in `<<...>>` are linked to the
glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.) Required.
* `no_label`: If the term is {"-"} and there are no other terms, insert a fixed label with this value. Defaults to
{"no "} plus the label.
* `usually_no_label`: If the term is {"-"} and there are other terms, insert a fixed label with this value. Defaults to
{"usually no "} plus the label.
* `cats`: List of categories to insert when terms are given that are not {"-"}. Each category is a string naming a full
category to insert (including the appropriate language name prefixed).
* `no_cats`: List of categories to insert when a term is given as {"-"}.
* `usually_no_cats`: List of categories to insert when a term is given as {"-"} and additional terms are specified as
well (representing, e.g. for the inflection {"plural"}, a term which usually has no plural but does under some
circumstances). If omitted, both the categories in `cats` and `no_cats` are inserted.
* `accel`: If specified, a full accelerator object to add to the inflections.
* `request`: If specified and no terms are given, insert a label with a request for inflections to be given.
* `enable_auto_translit`: If specified and terms are given, display automatic transliteration of the terms.
The return value indicates whether the inflection exists and how many terms are in it. It is an object with the
following fields:
* `exists`: {"yes"} if one or more terms were specified; {"no"} if the value was given as {"-"}; {"usually no"} if
the first value was given as {"-"} but additional terms were supplied; otherwise {nil}, indicating that the status
is unspecified.
* `numterms`: Number of terms in the inflection. Will be 0 unless `exists` has the value {"yes"} or {"usually no"}.
* `request`: True if no terms were specified but a term request was inserted into the inflection (because
`data.request` was specified). Otherwise {nil}.
]==]
function export.insert_inflection(data)
local headdata, terms, label = data.headdata, data.terms, data.label
local inflobj = data.inflobj or headdata
local retval = {}
local accel = data.accel
if data.accel_form then
if accel then
error("Internal error: can't specify both data.accel and data.accel_form")
end
if headdata.heads then
local lemmas = {}
local lemma_translits = {}
for i, headobj in ipairs(headdata.heads) do
lemmas[i] = headobj.term
if lemmas[i] == "+" then
error("Internal error: If you use data.accel_form, you should have resolved all occurrences of + in heads appropriately")
end
lemma_translits[i] = headobj.tr
end
accel = {
lemma = lemmas,
lemma_translit = lemma_translits,
form = data.accel_form,
}
else
accel = {
form = data.accel_form,
}
end
end
local function insert_cats(cats)
for _, cat in ipairs(cats) do
insert(headdata.categories, cat)
end
end
if terms and terms[1] then
terms = shallow_copy(terms)
if terms[1].term == "-" then
if terms[2] then
export.insert_fixed_inflection {
headdata = headdata,
inflobj = inflobj,
originating_term = terms[1],
label = data.usually_no_label or "biasanya tiada " .. label,
}
remove(terms, 1)
retval.numterms = #terms
retval.exists = "usually no"
if data.usually_no_cats then
insert_cats(data.usually_no_cats)
else
if data.no_cats then
insert_cats(data.no_cats)
end
if data.cats then
insert_cats(data.cats)
end
end
else
export.insert_fixed_inflection {
headdata = headdata,
inflobj = inflobj,
originating_term = terms[1],
label = data.no_label or "tiada " .. label,
}
retval.numterms = 0
retval.exists = "no"
if data.no_cats then
insert_cats(data.no_cats)
end
return retval
end
else
retval.numterms = #terms
retval.exists = "yes"
if data.cats then
insert_cats(data.cats)
end
end
if data.check_missing then
error("Internal error: check_missing support removed; use checkredlinks=true in [[Module:headword]]")
end
terms.label = export.replace_glossary_links_in_label(label)
if accel then
terms.accel = accel
end
terms.enable_auto_translit = data.enable_auto_translit
inflobj.inflections = inflobj.inflections or {}
insert(inflobj.inflections, terms)
elseif data.request then
inflobj.inflections = inflobj.inflections or {}
insert(inflobj.inflections, {
label = export.replace_glossary_links_in_label(label),
request = true,
})
retval.numterms = 0
-- retval.exists = nil
retval.request = true
else
retval.numterms = 0
-- retval.exists = nil
end
return retval
end
--[==[
Parse raw arguments from `forms` for inline modifiers, and insert the resulting terms (which should not require
significant additional processing) into `headdata.inflections`. `data` is an object with the following fields:
* `forms`: The list of raw values to parse. If {nil} or omitted, nothing happens.
* `headdata`: The headword structure passed to [[Module:headword]]. Required.
* `paramname`: As in `parse_term_list_with_modifiers()`. Required.
* `label`: As in `insert_inflection()`. Required.
* `qualifiers`, `frob`, `include_mods`, `exclude_mods`, `is_head`, `splitchar`, `preserve_splitchar`, `delimiter_key`,
`escape_fun`, `unescape_fun`, `pre_normalize_modifiers`: As in `parse_term_list_with_modifiers()`.
* `accel`: As in `insert_inflection()`.
Return value is as in `insert_inflection()`.
]==]
function export.parse_and_insert_inflection(data)
local forms = data.forms
if forms and forms[1] then
data = shallow_copy(data)
data.forms = forms
data.terms = export.parse_term_list_with_modifiers(data)
return export.insert_inflection(data)
end
return {
numterms = 0
}
end
--[==[
Canonicalize a single term or term-like object or a list of either into a list of term-like objects. `abterms` is the
term or list to canonicalize, and `field` is the name of the field holding the term (defaulting to {"term"}). This
does the minimal work necessary, meaning that the return value may partly or completely share memory with the value
passed in. As a special case, if `abterms` is {nil}, {nil} is returned. If `origin_val` is specified, add a field
`origin` containing the value of `origin_val` to each resulting term-like object (in this case, the object will be
copied a necessary to avoid side-effecting the passed-in objects).
]==]
function export.canonicalize_termobj_list(abterms, field, origin_val)
if abterms == nil then
return nil
end
field = field or "term"
if type(abterms) == "string" then
return {{[field] = abterms, origin = origin_val}}
elseif not abterms[1] then
if origin_val ~= nil then
abterms = shallow_copy(abterms)
abterms.origin = origin_val
end
return {abterms}
else
-- Check if already in full list term and return directly if so (unless `origin_val` is given, in which case we
-- need to shallow-copy both the list and each term in it).
local must_convert = false
for _, term in ipairs(abterms) do
if type(term) == "string" then
must_convert = true
break
end
end
if not must_convert then
if origin_val ~= nil then
abterms = shallow_copy(abterms)
for i, abterm in ipairs(abterms) do
abterms[i] = shallow_copy(abterm)
abterms[i].origin = origin_val
end
end
return abterms
end
end
local retval = {}
for _, term in ipairs(abterms) do
if type(term) == "string" then
insert(retval, {[field] = term, origin = origin_val})
else
if origin_val ~= nil then
term = shallow_copy(term)
term.origin = origin_val
end
insert(retval, term)
end
end
return retval
end
--[==[
Combine two sets of decorations. If either is {nil}, just return the other, and if both are {nil}, return {nil}.
]==]
function export.combine_decorations(decs1, decs2)
if not decs1 and not decs2 then
return nil
end
if not decs1 then
return decs2
end
if not decs2 then
return decs1
end
local combined = shallow_copy(decs1)
for _, dec in ipairs(decs2) do
insert_if_not(combined, dec)
end
return combined
end
function export.combine_qualifiers_or_labels(...)
-- FIXME: Added 2026-09-17. Remove after a month.
error("Use combine_decorations instead")
end
--[==[
Combine the decorations (qualifiers, labels, references) and ID's of two term objects. `destobj` is the "destination
term object" into which the combined properties are written, and `srcobj` is the "source object" into which the
properties are merged. `destobj` is side-effected (but the lists inside of `destobj` are not); if this is undesirable,
make sure to shallow-copy `destobj` first. If both objects have values for a given decoration, the values of `destobj`
come first. If both objects have a value for `id`, the values must match or an error is thrown; otherwise, the resulting
value of `id` comes from whichever one is defined.
'''NOTE:''' This may not be the correct behavior when deduplicating a list of term objects. See
`insert_termobj_combining_duplicates` for a different approach.
]==]
function export.combine_termobj_decorations(destobj, srcobj)
destobj.q = export.combine_decorations(destobj.q, srcobj.q)
destobj.qq = export.combine_decorations(destobj.qq, srcobj.qq)
destobj.l = export.combine_decorations(destobj.l, srcobj.l)
destobj.ll = export.combine_decorations(destobj.ll, srcobj.ll)
destobj.refs = export.combine_decorations(destobj.refs, srcobj.refs)
if destobj.id and srcobj.id and destobj.id ~= srcobj.id then
-- FIXME: We probably want to pass in an error function
error(("Can't specify two different ID's %s and %s when combining objects"):format(srcobj.id, destobj.id))
end
destobj.id = destobj.id or srcobj.id
return destobj
end
function export.combine_termobj_qualifiers_labels(...)
-- FIXME: Added 2026-09-17. Remove after a month.
error("Use combine_termobj_decorations instead")
end
function export.termobj_has_decorations(obj)
return obj.q and obj.q[1] or obj.qq and obj.qq[1] or obj.l and obj.l[1] or obj.ll and obj.ll[1] or
obj.refs and obj.refs[1]
end
function export.termobj_has_qualifiers_or_labels(...)
-- FIXME: Added 2026-09-17. Remove after a month.
error("Use termobj_has_decorations instead")
end
local function one_decoration_equal(prop1, prop2)
local prop1_is_nil = not prop1 or not prop1[1]
local prop2_is_nil = not prop2 or not prop2[1]
if prop1_is_nil and prop2_is_nil then
return true
end
if prop1_is_nil or prop2_is_nil then
return false
end
return deep_equals(prop1, prop2)
end
function export.termobj_decorations_equal(obj1, obj2)
return one_decoration_equal(obj1.q, obj2.q) and
one_decoration_equal(obj1.qq, obj2.qq) and
one_decoration_equal(obj1.l, obj2.l) and
one_decoration_equal(obj1.ll, obj2.ll) and
one_decoration_equal(obj1.refs, obj2.refs) and
obj1.id == obj2.id
end
function export.termobj_ancillary_properties_equal(...)
-- FIXME: Added 2026-09-17. Remove after a month.
error("Use termobj_decorations_equal instead")
end
function export.convert_termobj_to_formobj(termobj)
local formobj = {
form = termobj.term,
translit = termobj.tr,
}
local footnotes
local function mods_to_footnote(mod_prefix, mod_vals)
if mod_vals and mod_vals[1] then
footnotes = footnotes or {}
for _, val in ipairs(mod_vals) do
insert(footnotes, "[" .. mod_prefix .. ":" .. val .. "]")
end
end
end
mods_to_footnote("q", termobj.q)
mods_to_footnote("qq", termobj.qq)
mods_to_footnote("l", termobj.l)
mods_to_footnote("ll", termobj.ll)
mods_to_footnote("ref", termobj.refs)
mods_to_footnote("id", termobj.id and {termobj.id} or nil)
formobj.footnotes = footnotes
return formobj
end
local recognized_multi_mods = {
q = "q",
qq = "qq",
l = "l",
ll = "ll",
ref = "refs",
}
local recognized_single_mods = {
id = "id",
}
function export.add_footnote_to_termobj(termobj, footnote)
local stripped_footnote = footnote:match("^%[(.*)%]$")
if not stripped_footnote then
error("Internal error: Footnote should be surrounded by brackets at this stage: " .. footnote)
end
local prefix, rest = stripped_footnote:match("^([a-z]+):(.+)$")
local field, is_multi
if prefix then
if recognized_multi_mods[prefix] then
field = recognized_multi_mods[prefix]
is_multi = true
elseif recognized_single_mods[prefix] then
field = recognized_single_mods[prefix]
is_multi = false
end
end
if not field then
rest = stripped_footnote
field = "l"
is_multi = true
end
if is_multi then
if not termobj[field] then
termobj[field] = {}
end
insert(termobj[field], rest)
else
if termobj[field] and termobj[field] ~= rest then
error(("Can't set two values for '%s': '%s' and '%s'"):format(field, termobj[field], rest))
end
termobj[field] = rest
end
end
function export.convert_formobj_to_termobj(formobj)
local termobj = {
term = formobj.form,
tr = formobj.translit,
}
if formobj.footnotes then
for _, footnote in ipairs(formobj.footnotes) do
export.add_footnote_to_termobj(termobj, footnote)
end
end
return termobj
end
local function extract_termobj_field_modifiers(fieldval)
return fieldval:match("^([*+]?)(.*)$")
end
function export.remove_termobj_field_modifiers(termobj)
local function remove_field_modifiers(field)
if termobj[field] and termobj[field][1] then
local any_field_modifiers = false
for _, val in ipairs(termobj[field]) do
local field_mods, _ = extract_termobj_field_modifiers(val)
if field_mods ~= "" then
any_field_modifiers = true
break
end
end
local new_field = {}
if any_field_modifiers then
for _, val in ipairs(termobj[field]) do
local _, field_without_mods = extract_termobj_field_modifiers(val)
insert_if_not(new_field, field_without_mods)
end
termobj[field] = new_field
end
end
end
remove_field_modifiers("q")
remove_field_modifiers("qq")
remove_field_modifiers("l")
remove_field_modifiers("ll")
remove_field_modifiers("refs")
end
function export.insert_termobj_combining_duplicates(destobjs, termobj)
for _, destobj in ipairs(destobjs) do
if destobj.term == termobj.term and destobj.tr == termobj.tr then
-- Form already present; maybe combine footnotes.
local function combine_field_values(field)
if termobj[field] and termobj[field][1] then
-- Check to see if there are existing values with *; if so, remove them.
if destobj[field] and destobj[field][1] then
local any_values_with_asterisk = false
for _, val in ipairs(destobj[field]) do
local field_mods, _ = extract_termobj_field_modifiers(val)
if field_mods:find("%*") then
any_values_with_asterisk = true
break
end
end
if any_values_with_asterisk then
local filtered_values = {}
for _, val in ipairs(destobj[field]) do
local field_mods, _ = extract_termobj_field_modifiers(val)
if not field_mods:find("%*") then
insert(filtered_values, val)
end
end
if filtered_values[1] then
destobj[field] = filtered_values
else
destobj[field] = nil
end
end
end
local any_footnotes_with_plus = false
for _, val in ipairs(termobj[field]) do
local field_mods, _ = extract_termobj_field_modifiers(val)
if field_mods:find("%+") then
any_footnotes_with_plus = true
break
end
end
if any_footnotes_with_plus then
if not destobj[field] then
destobj[field] = {}
else
destobj[field] = shallow_copy(destobj[field])
end
for _, val in ipairs(termobj[field]) do
local already_seen = false
local field_mods, field_without_mods = extract_termobj_field_modifiers(val)
if field_mods:find("%+") then
for _, existing_val in ipairs(destobj[field]) do
local _, existing_field_without_mods =
extract_termobj_field_modifiers(existing_val)
if existing_field_without_mods == field_without_mods then
already_seen = true
break
end
end
if not already_seen then
insert(destobj[field], val)
end
end
end
end
end
end
combine_field_values("q")
combine_field_values("qq")
combine_field_values("l")
combine_field_values("ll")
combine_field_values("refs")
if destobj.id and termobj.id and destobj.id ~= termobj.id then
-- FIXME: We probably want to pass in an error function
error(("Can't specify two different ID's %s and %s when combining objects"):format(termobj.id, destobj.id))
end
destobj.id = destobj.id or termobj.id
return
end
end
insert(destobjs, termobj)
end
export.allowed_special_indicators = {
["first"] = true,
["first-second"] = true,
["first-last"] = true,
["second"] = true,
["last"] = true,
["each"] = true,
["+"] = true, -- requests the default behavior with preposition handling
}
--[==[
Check for special indicators (values such as {"+first"} or {"+first-last"} that are used in a `pl`, `f`, etc. argument
and indicate how to inflect a multiword term). If `form` is such an indicator, the return value is `form` minus
the initial `+` sign; otherwise, if form begins with a `+` sign, an error is thrown; otherwise the return value is nil.
]==]
function export.get_special_indicator(form, noerror)
if form:find("^%+") then
form = form:gsub("^%+", "")
if not export.allowed_special_indicators[form] then
if noerror then
return nil
end
local indicators = {}
for indic, _ in pairs(export.allowed_special_indicators) do
insert(indicators, "+" .. indic)
end
sort(indicators)
error("Special inflection indicator beginning with '+' can only be " ..
mw.text.listToText(indicators) .. ": +" .. form)
end
return form
end
return nil
end
local function add_endings(bases, endings)
local retval = {}
if type(bases) ~= "table" then
bases = {bases}
end
if type(endings) ~= "table" then
endings = {endings}
end
for _, base in ipairs(bases) do
for _, ending in ipairs(endings) do
insert(retval, base .. ending)
end
end
return retval
end
--[==[
Inflect a possibly multiword or hyphenated term `form` using the function `inflect`, which is a function of one argument
that is called on a single word to inflect and should return either the inflected word or a list of inflected words.
`special` indicates how to inflect the multiword term and should be e.g. {"first"} to inflect only the first word,
{"first-last"} to inflect the first and last words, {"each"} to inflect each word, etc. See `allowed_special_indicators`
above for the possibilities. If `special` is `+`, or is omitted and the term is multiword (i.e. containing a space
character), and `prepositions` is supplied, the function checks for multiword or hyphenated terms containing the
prepositions in `prepositions`, e.g. Italian [[senso di marcia]] or [[medaglia d'oro]] or Portuguese
[[tartaruga-do-mar]]. If such a term is found, only the first word is inflected. Otherwise, the default is
{"first-last"}. `prepositions` is a list of Lua patterns matching prepositions. The patterns will automatically have the
separator character (space or hyphen) added to the left side but not the right side, so they should contain a space
character (which will automatically be converted to the appropriate separator) on the right side unless the preposition
is joined on the right side with an apostrophe. Examples of preposition patterns for Italian are {"di "}, {"sull'"} and
{"d?all[oae] "} (which matches {"dallo "}, {"dalle "}, {"alla "}, etc.).
The return value is always either a list of inflected multiword or hyphenated terms, or nil if `special` is omitted
and `form` is not multiword. (If `special` is specified and `form` is not multiword or hyphenated, an error results.)
]==]
function export.handle_multiword(form, special, inflect, prepositions, sep)
sep = sep or form:find(" ") and " " or "%-"
local raw_sep = sep == " " and " " or "-"
-- Used to add regex version of separator in the replacement portion of ugsub() or :gsub()
local sep_replacement = sep == " " and " " or "%%-"
-- Given a Lua pattern, replace space with the appropriate separator.
local function hack_re(re)
if sep == " " then
return re
end
return (re:gsub(" ", sep_replacement))
end
if special == "first" then
local first, rest = form:match(hack_re("^(.-)( .*)$"))
if not first then
error("Special indicator 'first' can only be used with a multiword term: " .. form)
end
return add_endings(inflect(first), rest)
elseif special == "second" then
local first, second, rest = form:match(hack_re("^([^ ]+ )([^ ]+)( .*)$"))
if not first then
error("Special indicator 'second' can only be used with a term with three or more words: " .. form)
end
return add_endings(add_endings({first}, inflect(second)), rest)
elseif special == "first-second" then
local first, space, second, rest = form:match(hack_re("^([^ ]+)( )([^ ]+)( .*)$"))
if not first then
error("Special indicator 'first-second' can only be used with a term with three or more words: " .. form)
end
return add_endings(add_endings(add_endings(inflect(first), space), inflect(second)), rest)
elseif special == "each" then
local terms = split(form, sep)
if #terms < 2 then
error("Special indicator 'each' can only be used with a multiword term: " .. form)
end
for i, term in ipairs(terms) do
terms[i] = inflect(term)
if i > 1 then
terms[i] = add_endings(raw_sep, terms[i])
end
end
local result = ""
for _, term in ipairs(terms) do
result = add_endings(result, term)
end
return result
elseif special == "first-last" then
local first, middle, last = form:match(hack_re("^(.-)( .* )(.-)$"))
if not first then
first, middle, last = form:match(hack_re("^(.-)( )(.*)$"))
end
if not first then
error("Special indicator 'first-last' can only be used with a multiword term: " .. form)
end
return add_endings(add_endings(inflect(first), middle), inflect(last))
elseif special == "last" then
local rest, last = form:match(hack_re("^(.* )(.-)$"))
if not rest then
error("Special indicator 'last' can only be used with a multiword term: " .. form)
end
return add_endings(rest, inflect(last))
elseif special and special ~= "+" then
error("Unrecognized special=" .. special)
end
-- Only do default behavior if special indicator '+' explicitly given or separator is space; otherwise we will
-- break existing behavior with hyphenated words.
if (special == "+" or sep == " ") and form:find(sep) then
if prepositions then
-- check for prepositions in the middle of the word; do it this way so we can handle
-- more than one word before the preposition (and usually inflect each word)
for _, prep in ipairs(prepositions) do
local first, space_prep_rest = umatch(form, hack_re("^(.-)( " .. prep .. ".*)$"))
if first then
return add_endings(inflect(first), space_prep_rest)
end
end
end
-- multiword or hyphenated expressions default to first-last; we need to pass in the separator to avoid
-- problems with multiword terms containing hyphens in the individual words
return export.handle_multiword(form, "first-last", inflect, prepositions, sep)
end
return nil
end
local function link_hyphen_split_component(word, data)
if data.link_hyphen_split_component then
return data.link_hyphen_split_component(word)
else
return "[[" .. word .. "]]"
end
end
-- Default function to split a word on apostrophes. Don't split apostrophes at the beginning or end of a word (e.g.
-- [['ndrangheta]] or [[po']]). Handle multiple apostrophes correctly, e.g. [[l'altr'ieri]] -> [[l']][altr']][[ieri]].
function export.default_split_apostrophe(word, data)
local apostrophe_parts = split(word, "'", true, true)
local linked_apostrophe_parts = {}
local apostrophes_at_beginning = ""
local i = 1
-- Apostrophes at beginning get attached to the first word after (which will always exist but may
-- be blank if the word consists only of apostrophes).
while i < #apostrophe_parts do -- <, not <=, in case the word consists only of apostrophes
local apostrophe_part = apostrophe_parts[i]
i = i + 1
if apostrophe_part == "" then
apostrophes_at_beginning = apostrophes_at_beginning .. "'"
else
break
end
end
apostrophe_parts[i] = apostrophes_at_beginning .. apostrophe_parts[i]
-- Now, do the remaining parts. A blank part indicates more than one apostrophe in a row; we join
-- all of them to the preceding word.
while i <= #apostrophe_parts do
local apostrophe_part = apostrophe_parts[i]
if apostrophe_part == "" then
linked_apostrophe_parts[#linked_apostrophe_parts] =
linked_apostrophe_parts[#linked_apostrophe_parts] .. "'"
elseif i == #apostrophe_parts then
insert(linked_apostrophe_parts, apostrophe_part)
else
insert(linked_apostrophe_parts, apostrophe_part .. "'")
end
i = i + 1
end
for j, tolink in ipairs(linked_apostrophe_parts) do
linked_apostrophe_parts[j] = link_hyphen_split_component(tolink, data)
end
return concat(linked_apostrophe_parts)
end
--[=[
Auto-add links to a word that should not have spaces but may have hyphens and/or apostrophes. We split off final
punctuation, then split on hyphens if `data.split_hyphen` is given, and also split on apostrophes if
`data.split_apostrophe` is given. We only split on hyphens if they are in the middle of the word, not at the beginning
or end (hyphens at the beginning or end indicate suffixes or prefixes, respectively). `include_hyphen_prefixes`, if
given, is a set of prefixes (not including the final hyphen) where we should include the final hyphen in the prefix.
Hence, e.g. if "anti" is in the set, a Portuguese word like [[anti-herói]] "anti-hero" will be split [[anti-]][[herói]]
(whereas a word like [[código-fonte]] "source code" will be split as [[código]]-[[fonte]]).
If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and
the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the
default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in
Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French
[[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified
but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word
to split, and the return value should be the split and linked word.
]=]
local function add_single_word_links(space_word, data, term_has_spaces)
local space_word_no_punct, punct
local punct_pattern = data.punctuation
if punct_pattern and is_callable(punct_pattern) then
space_word_no_punct, punct = punct_pattern(space_word)
else
if punct_pattern == nil then
punct_pattern = "[,;:?!]"
end
space_word_no_punct, punct = umatch(space_word, "^(.*)(" .. punct_pattern .. ")$")
end
space_word_no_punct = space_word_no_punct or space_word
punct = punct or ""
local words
if space_word_no_punct:sub(1, 1) == "-" or space_word_no_punct:sub(-1) == "-" then
-- don't split prefixes and suffixes
words = {space_word_no_punct}
else
local splitter
if term_has_spaces then
splitter = data.split_hyphen_when_space
else
splitter = data.split_hyphen_when_no_space
end
if is_callable(splitter) then
words = splitter(space_word_no_punct)
if type(words) == "string" then
return words .. punct
end
end
end
if not words then
local split_hyphen
if term_has_spaces then
split_hyphen = data.split_hyphen_when_space
else
split_hyphen = data.split_hyphen_when_no_space
if split_hyphen == nil then -- default to true; use `false` to avoid this
split_hyphen = true
end
end
if split_hyphen then
words = split(space_word_no_punct, "-", true, true)
else
words = {space_word_no_punct}
end
end
local linked_words = {}
for j, word in ipairs(words) do
if j < #words and data.include_hyphen_prefixes and data.include_hyphen_prefixes[word] then
word = "[[" .. word .. "-]]"
elseif j > 1 and data.include_hyphen_suffixes and data.include_hyphen_suffixes[word] then
word = "[[-" .. word .. "]]"
else
-- Don't split on apostrophes if the word is in `no_split_apostrophe_words`.
if (not data.no_split_apostrophe_words or not data.no_split_apostrophe_words[word]) and
data.split_apostrophe and word:find("'", nil, true) then
if data.split_apostrophe == true then
word = export.default_split_apostrophe(word, data)
else -- custom apostrophe splitter/linker
word = data.split_apostrophe(word)
end
elseif word ~= "" then -- avoid -[[]]- (e.g. f--k)
word = link_hyphen_split_component(word, data)
end
if j < #words then
word = word .. "-"
end
end
insert(linked_words, word)
end
return concat(linked_words) .. punct
end
--[=[
Auto-add links to a multiword term. `data` contains fields customizing how to do this. By default we proceed as follows:
(1) If the term already has embedded links in it, they are left unchanged.
(2) Otherwise, if there are spaces present, we split on spaces and link each word separately.
(3) If a given space-separated component ends in punctuation (defaulting to [,;:?!]), it is separated off, the remainder
of the algorithm run, and the punctuation pasted back on.
(4) If there are hyphens in a given space-separated component, we may link each hyphenated term separately depending
on the settings in `data`. Normally the hyphens are not included in the linked terms, but this can be overridden
for specific prefixes and/or suffixes. By default, if there are spaces in the multiword term, we do not link
hyphenated components (because of cases like "boire du petit-lait" where "petit-lait" should be linked as a whole),
but do so otherwise (e.g. for "avant-avant-hier"); this can overridden for cases like "croyez-le ou non".
Cases where only some of the hyphens should be split can always be handled by explicitly specifying the head (e.g.
"Nord-Pas-de-Calais" given as head=[[Nord]]-[[Pas-de-Calais]]).
(5) If there are apostrophes in a given component, we may link each apostrophe-separated term separately depending
on the settings in `data`, including the apostrophe in the link to its left (so we split "de l'eau" as
"[[de]] [[l']][[eau]]").
The settings in `data` are as follows:
`split_hyphen_when_no_space`: Whether to split on hyphens when the term has no spaces. Defaults to true if set to `nil`.
This can be a function of one argument, to implement a custom splitting algorithm for hyphen-separated terms. If
this returns [FIXME: FINISH ME ...]
If `data.split_apostrophe` is specified, we split on apostrophes unless `data.no_split_apostrophe_words` is given and
the word is in the specified set, such as French [[c'est]] and [[quelqu'un]]. If `data.split_apostrophe` is true, the
default algorithm applies, which splits on all apostrophes except those at the beginning and end of a word (as in
Italian [['ndrangheta]] or [[po']]), and includes the apostrophe in the link to its left (so we auto-split French
[[l'eau]] as [[l']][[eau]] and [[l'altr'ieri]] as [[l']][altr']][[ieri]]). If `data.split_apostrophe` is specified
but not `true`, it should be a function of one argument that does custom apostrophe-splitting. The argument is the word
to split, and the return value should be the split and linked word.
We don't always split on hyphens because of cases like "boire du petit-lait" where "petit-lait" should be linked as a
whole, but provide the option to do it for cases like "croyez-le ou non". If there's no space, however, then it makes
sense to split on hyphens by `no_split_apostrophe_words` and `include_hyphen_prefixes` allow for special-case handling
of particular words and are as described in the comment above add_single_word_links().
]=]
function export.add_links_to_multiword_term(term, data)
if term:match("[%[%]]") then
return term
end
local words = split(term, " ", true, true)
local term_has_spaces = #words > 1
local linked_words = {}
for _, word in ipairs(words) do
insert(linked_words, add_single_word_links(word, data, term_has_spaces))
end
local retval = concat(linked_words, " ")
-- If we ended up with a single link consisting of the entire term,
-- remove the link.
return retval:match("^%[%[([^%[%]]*)%]%]$") or retval
end
local function canonicalize_begin_end_spec(spec)
local from, to = spec:match("^(.-):(.*)$")
if not from then
from = spec
to = ""
end
return from, to
end
--[==[
Given a `linked_term` that is the output of add_links_to_multiword_term(), apply modifications as given in
`modifier_spec` to change the link destination of subterms (normally single-word non-lemma forms; sometimes
collections of adjacent words). This is usually used to link non-lemma forms to their corresponding lemma, but can
also be used to replace a span of adjacent separately-linked words to a single multiword lemma. The format of
`modifier_spec` is one or more semicolon-separated subterm specs, where each such spec is of the form
SUBTERM:DEST, where SUBTERM is one or more words in the `linked_term` but without brackets in them, and DEST is the
corresponding link destination to link the subterm to. Any occurrence of ~ in DEST is replaced with SUBTERM.
Alternatively, a single modifier spec can be of the form BEGIN[FROM:TO], which is equivalent to writing
BEGINFROM:BEGINTO (see example below).
For example, given the source phrase [[il bue che dice cornuto all'asino]] "the pot calling the kettle black"
(literally "the ox that calls the donkey horned/cuckolded"), the result of calling add_links_to_multiword_term()
is [[il]] [[bue]] [[che]] [[dice]] [[cornuto]] [[all']][[asino]]. With a modifier_spec of 'dice:dire', the result
is [[il]] [[bue]] [[che]] [[dire|dice]] [[cornuto]] [[all']][[asino]]. Here, based on the modifier spec, the
non-lemma form [[dice]] is replaced with the two-part link [[dire|dice]].
Another example: given the source phrase [[chi semina vento raccoglie tempesta]] "sow the wind, reap the whirlwind"
(literally (he) who sows wind gathers [the] tempest"). The result of calling add_links_to_multiword_term() is
[[chi]] [[semina]] [[vento]] [[raccoglie]] [[tempesta]], and with a modifier_spec of 'semina:~re; raccoglie:~re',
the result is [[chi]] [[seminare|semina]] [[vento]] [[raccogliere|raccoglie]] [[tempesta]]. Here we use the ~
notation to stand for the non-lemma form in the destination link.
A more complex example is [[se non hai altri moccoli puoi andare a letto al buio]], which becomes
[[se]] [[non]] [[hai]] [[altri]] [[moccoli]] [[puoi]] [[andare]] [[a]] [[letto]] [[al]] [[buio]] after calling
add_links_to_multiword_term(). With the following modifier_spec:
'hai:avere; altr[i:o]; moccol[i:o]; puoi: potere; andare a letto:~; al buio:~', the result of applying the spec is
[[se]] [[non]] [[avere|hai]] [[altro|altri]] [[moccolo|moccoli]] [[potere|puoi]] [[andare a letto]] [[al buio]].
Here, we rely on the alternative notation mentioned above for e.g. 'altr[i:o]', which is equivalent to 'altri:altro',
and link multiword subterms using e.g. 'andare a letto:~'. (The code knows how to handle multiword subexpressions
properly, and if the link text and destination are the same, only a single-part link is formed.)
]==]
function export.apply_link_modifiers(linked_term, modifier_spec, lang)
local split_modspecs = split(modifier_spec, "%s*;%s*")
for j, modspec in ipairs(split_modspecs) do
local id
if modspec:find("<") then
local rest
rest, id = modspec:match("^(.*)<id:(.-)>$")
if rest then
modspec = rest
end
end
local subterm, dest, otherlang
local begin_spec, rest, end_spec = modspec:match("^%[(.-)%]([^:]*)%[(.-)%]$")
if begin_spec then
local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec)
local end_from, end_to = canonicalize_begin_end_spec(end_spec)
subterm = begin_from .. rest .. end_from
dest = begin_to .. rest .. end_to
end
if not subterm then
rest, end_spec = modspec:match("^([^:]*)%[(.-)%]$")
if rest then
local end_from, end_to = canonicalize_begin_end_spec(end_spec)
subterm = rest .. end_from
dest = rest .. end_to
end
end
if not subterm then
begin_spec, rest = modspec:match("^%[(.-)%]([^:]*)$")
if begin_spec then
local begin_from, begin_to = canonicalize_begin_end_spec(begin_spec)
subterm = begin_from .. rest
dest = begin_to .. rest
end
end
if not subterm then
subterm, dest = modspec:match("^(.-)%s*:%s*(.*)$")
if subterm and subterm ~= "^" and subterm ~= "$" then
local langdest
-- Parse off an initial language code (e.g. 'en:Higgs', 'la:minūtia' or 'grc:σκατός'). Also handle
-- Wikipedia prefixes ('w:Abatemarco' or 'w:it:Colle Val d'Elsa').
otherlang, langdest = dest:match("^([A-Za-z0-9._-]+):([^ ].*)$")
if otherlang == "w" then
local foreign_wikipedia, foreign_term = langdest:match("^([A-Za-z0-9._-]+):([^ ].*)$")
if foreign_wikipedia then
otherlang = otherlang .. ":" .. foreign_wikipedia
langdest = foreign_term
end
dest = ("%s:%s"):format(otherlang, langdest)
otherlang = nil
elseif otherlang then
otherlang = get_lang_by_code(otherlang, true, "allow etym")
dest = langdest
end
end
end
if not subterm then
if modspec == "?" or modspec == "!" then
subterm = "$"
dest = modspec
elseif modspec == "..." or modspec == "...?" then
subterm = "$"
dest = " " .. modspec
elseif modspec:find("^[A-Z]$") then
-- X, Y, etc. by themselves are unlinked, to help with snowclones
subterm = modspec
dest = "_"
else
subterm = modspec
dest = "~"
end
end
if subterm == "^" then
linked_term = dest:gsub("_", " ") .. linked_term
elseif subterm == "$" then
linked_term = linked_term .. dest:gsub("_", " ")
else
if subterm:find("[", nil, true) then
error(("Subterm '%s' in modifier spec '%s' cannot have brackets in it"):format(
escape_wikicode(subterm), escape_wikicode(modspec)))
end
local escaped_subterm = pattern_escape(subterm)
local subterm_re = "%[%[" .. escaped_subterm:gsub("(%%?[ ',%-])", "%%]*%1%%[*") .. "%]%]"
local expanded_dest
if dest:find("~", nil, true) then
expanded_dest = dest:gsub("~", replacement_escape(subterm))
else
expanded_dest = dest
end
if otherlang then
expanded_dest = expanded_dest .. "#" .. otherlang:getCanonicalName()
end
local subterm_replacement
if expanded_dest == "_" then
subterm_replacement = subterm
if id then
error("Can't supply <id:...> with an unlinked subterm")
end
if otherlang then
error("Can't supply prefixed language with an unlinked subterm")
end
elseif id or otherlang then
if id and expanded_dest:find("[", nil, true) then
error("Can't supply <id:...> with destination with embedded brackets")
end
subterm_replacement = require(links_module).language_link {
lang = otherlang or lang,
term = expanded_dest,
alt = subterm,
id = id,
}
elseif expanded_dest:find("[", nil, true) then
-- Use the destination directly if it has brackets in it (e.g. to put brackets around parts of a word).
subterm_replacement = expanded_dest
elseif expanded_dest == subterm then
subterm_replacement = "[[" .. subterm .. "]]"
else
subterm_replacement = "[[" .. expanded_dest .. "|" .. subterm .. "]]"
end
local escaped_subterm_replacement = replacement_escape(subterm_replacement)
local replaced_linked_term = ugsub(linked_term, subterm_re, escaped_subterm_replacement)
if replaced_linked_term == linked_term then
mw.log(("Attempted to replace %s with %s in %s"):format(subterm_re, escaped_subterm_replacement, linked_term))
error(("Subterm '%s' could not be located in %slinked expression %s, or replacement same as subterm"):format(
subterm, j > 1 and "intermediate " or "", escape_wikicode(linked_term)))
else
linked_term = replaced_linked_term
end
end
end
return linked_term
end
local inflection_to_cats = {
plural = {
filter_plpos = function(plpos)
-- plurals also occur with determiners, adjectives etc. and we don't want to generate categories like
-- 'countable determiners', 'countable adjectives', etc. Note that the passed-in `plpos` has `proper nouns`
-- converted to `nouns`.
return plpos == "Kata nama"
end,
cats = {"{plpos} terhitung"},
no_cats = {"{plpos} tak terhitung"},
},
comparative = {
cats = {"{plpos} bandingan"},
no_cats = {"{plpos} bukan bandingan"},
},
["female equivalent"] = {
cats = {"{plpos} dengan padanan genus lain"},
},
["male equivalent"] = {
cats = {"{plpos} dengan padanan genus lain"},
},
}
--[=[
Validate the items in `items` against the list or set of valid items in `valid_items`. If `field` is given, fetch the
item to check from that-named field of each object in `items`; otherwise use the items in `items` directly. If an error
occurs, `item_type` specifies the type of item to mention in the error message, which will also list the allowed items
(either taken directly from `valid_items` if a list, or from the sorted keys if a set).
]=]
local function validate_items(data)
local items, field, valid_items, item_type =
data.items, data.field, data.valid_items, data.item_type
local valid_set
if valid_items[1] then
valid_set = list_to_set(valid_items)
else
valid_set = valid_items
end
for _, item in ipairs(items) do
if field then
item = item[field]
end
if not valid_set[item] then
local valid_list
if valid_items[1] then
valid_list = valid_items
else
valid_list = {}
for valid_item, _ in pairs(valid_items) do
insert(valid_list, valid_item)
end
table.sort(valid_list)
end
error(("Invalid %s: %s; expected one of %s"):format(item_type, item, mw.text.listToText(valid_list)))
end
end
end
local Headdata = {}
function Headdata:get_canonicalized_plpos()
return (self.pos_category:gsub("Kata nama khas", "Kata nama"))
end
--[==[
Canonicalize a category. The category string will have the full language name (i.e. the name of the L2 language under
which an entry is inserted, which may a parent language if the language in question is an etymology-only language)
prepended to it, and any occurrences of `{plpos}` in the string replaced with the actual plural part of speech (with
some canonicalization; specifically, `proper nouns` is converted to `nouns` when replacing `{plpos}`). To specify a
full category and not have the language name prepended to it, precede it with {"Category:"}, which will be removed.
]==]
function Headdata:canonicalize_category(category)
if category:find("{plpos}") then
local plpos = self:get_canonicalized_plpos()
category = category:gsub("{plpos}", plpos)
end
if category:find("^Kategori:") then
return (category:gsub("^Kategori:", ""))
else
return category .. " bahasa " .. self.langfullname
end
end
--[==[
Canonicalize a list of categories according to the process described in `Headdata:canonicalize_category`. This simply
loops over each category in `categories` and calls `Headdata:canonicalize_category` on each one.
]==]
function Headdata:canonicalize_categories(categories)
if not categories then
return categories
end
local canon_cats = {}
for _, cat in ipairs(categories) do
insert(canon_cats, self:canonicalize_category(cat))
end
return canon_cats
end
--[==[
Insert a category into the `categories` list in the headword `data` structure. `category` is normally a string naming
the category, which will have the full language name prepended to it and any occurrences of `{plpos}` in the string
replaced with the actual plural part of speech (with some canonicalization; specifically, `proper nouns` is converted to
`nouns` when replacing `{plpos}`). To specify a full category and not have the language name prepended to it, precede it
with {"Category:"}.
]==]
function Headdata:insert_category(category)
insert(self.categories, self:canonicalize_category(category))
end
--[==[
Validate the genders in `genders` (a list of gender spec objects, as produced by {type = "genders"} in
[[Module:parameters]] and accepted by [[Module:gender and number]]), checking that all specified genders are in the list
given in `valid_genders`. Optional `props` controls how the validation happens. In particular, unless `props.no_augment`
is given, then for any gender beginning with `m`, if a corresponding gender beginning with `f` occurs, analogous genders
beginning with `mf`, `mfbysense` and `mfequiv` are also allowed. For example, if `m-d` (masculine dual) and `f-d`
(feminine dual) both occur, genders `mf-d`, `mfbysense-d` and `mfequiv-d` are also allowed. If a disallowed gender is
given, an error occurs, giving the disallowed gender along with the list of all allowed genders.
]==]
function Headdata:validate_genders(genders, valid_genders, props)
if not genders then
return
end
props = props or {}
local gender_type, no_augment = props.gender_type, props.no_augment
gender_type = gender_type or "headword"
local valid_gender_set = list_to_set(valid_genders)
local augmented_gender_set
if no_augment then
augmented_gender_set = valid_gender_set
else
augmented_gender_set = {}
for g, _ in pairs(valid_gender_set) do
augmented_gender_set[g] = true
if g:find("^m") and not g:find("^mf") and valid_gender_set[g:gsub("^m", "f")] then
augmented_gender_set[g:gsub("^m", "mf")] = true
augmented_gender_set[g:gsub("^m", "mfbysense")] = true
augmented_gender_set[g:gsub("^m", "mfequiv")] = true
end
end
end
validate_items {
items = genders,
field = "spec",
valid_items = augmented_gender_set,
item_type = ("%s gender"):format(gender_type),
}
end
--[==[
Parse an inflection specified in `field`, the name of a parameter holding an inflection. If the parameter is numeric,
the field should be given as a number (as with the `params` structure passed to [[Module:parameters]]), not a string
containing the representation of a number. The field can specify multiple comma-separated terms, and each term can have
associated inline modifiers that will be parsed (unless there is top-level HTML in the parameter, i.e. HTML not
contained inside an inline modifier, e.g. as may be generated by using {{tl|l}} or similar template inside a parameter).
This is a wrapper around the top-level `parse_term_with_modifiers()` function. `props` is an optional structure
containing additional properties, including all additional properties documented for the top-level
`parse_term_with_modifiers()` function.
If the parameter in `field` is unspecified, the return value of this function will be an empty list, not {nil}, so it
is always safe to iterate over the return value.
By default, the allowed modifiers are the same as for `parse_term_with_modifiers()`, except that (normally) the
`<tr:...>` modifier will be allowed if `include_tr` was specified in the original call to `process_headword()`; likewise
for the `<ts:...>` modifier if `include_ts` was specified and the `<sc:...>` modifier if `include_sc` was specified. If
If you pass in your own `include_mods` list of additional allowed modifiers, it will (normally) automatically be
augmented with {"tr"}, {"ts"} and/or {"sc"} if `include_tr`, `include_ts` and/or `include_sc` was specified when calling
`process_headword()`. To disable automatic augmentation of these modifiers (whether or not you specify an `include_mods`
property), specify {no_augment_include_mods = true} in `props`.
]==]
function Headdata:parse_inflection(field, props)
local val = self.process_props.args[field]
if not val then
return {}
end
props = props and shallow_copy(props) or {}
local include_mods = props.include_mods
local data = self.process_props.data
if not props.no_augment_include_mods and (data.include_tr or data.include_ts or data.include_sc) then
include_mods = include_mods and shallow_copy(include_mods) or {}
if data.include_tr then
insert_if_not(include_mods, "tr")
end
if data.include_ts then
insert_if_not(include_mods, "ts")
end
if data.include_sc then
insert_if_not(include_mods, "sc")
end
end
props.val = val
props.paramname = field
props.splitchar = props.splitchar or ","
props.include_mods = include_mods
return export.parse_term_with_modifiers(props) or {}
end
--[==[
Insert previously-parsed terms into the `inflections` of the headword `data` structure. This is a wrapper around
the top-level `insert_inflection()` function. `terms` is the list of parsed terms. (If {nil}, nothing happens unless
`request` is set in `props`.) `label` is the the label that the inflections are given; any parts of the label surrounded
in `<<...>>` are linked to the glossary. (If the contents of `<<...>>` contain a `|` in them, they are a two-part link.)
`props` is an optional structure containing additional properties, including all additional properties documented for
the top-level `insert_inflection()` function.
Unless `no_auto_cats` is given in `props`, certain labels automatically trigger the insertion of additional
categories in specific circumstances. This is controlled by the `inflection_to_cats` structure in
[[Module:headword utilities]]. For example, if the part of speech is {"nouns"} or {"proper nouns"} and the label (after
removing any links and `<<...>>` glossary specs) is {"plural"}, an additional category
<code><var>lang</var> countable nouns</code> will be added if a plural value is given (i.e. the value is not {"-"}). If
the value is {"-"} (which indicates that there is no plural and triggers the insertion of the fixed inflection label
{"no plural"}), <code><var>lang</var> uncountable nouns</code> will be inserted instead, and if both {"-"} and a value
are given (which triggers the insertion of the {"usually no plural"} fixed inflection label), both categories are added.
Similar categories are inserted when a comparative is given (with a label {"comparative"}), and if the label is
{"female equivalent"} or {"male equivalent"} and the value is not {"-"}, a category such as
<code><var>lang</var> nouns with other-gender equivalents</code> is inserted.
]==]
function Headdata:insert_inflection(terms, label, props)
props = props and shallow_copy(props) or {}
if not props.no_auto_cats then
local bare_label = label
if bare_label:find("[[", nil, true) then
bare_label = require(links_module).remove_links(bare_label)
end
if bare_label:find("<<", nil, true) then
bare_label = bare_label:gsub("<<.-|(.-)>>", "%1"):gsub("<<(.-)>>", "%1")
end
local cats = inflection_to_cats[bare_label]
if cats then
if not cats.filter_plpos or cats.filter_plpos(self:get_canonicalized_plpos()) then
if props.cats == nil then
props.cats = self:canonicalize_categories(cats.cats)
end
if props.usually_no_cats == nil then
props.usually_no_cats = self:canonicalize_categories(cats.usually_no_cats)
end
if props.no_cats == nil then
props.no_cats = self:canonicalize_categories(cats.no_cats)
end
end
end
end
props.headdata = self
props.terms = terms
props.label = label
return export.insert_inflection(props)
end
--[==[
Insert a "fixed" inflection (a label without associated values) into the `inflections` table of the headword `data`
structure, labeled according to `label` (which can have glossary links in it specified using `<<...>>`, exactly as for
`:insert_inflection()`). An example label (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"},
specifying that the noun belongs to the hidden-''g'' declension. This is a direct wrapper around the top-level function
`insert_fixed_inflection()`; see that function for more details on optional `props`.
]==]
function Headdata:insert_fixed_inflection(label, props)
props = props and shallow_copy(props) or {}
props.headdata = self
props.label = label
export.insert_fixed_inflection(props)
end
--[==[
Parse the inflection(s) specified in `field` and insert them into the `inflections` table of the headword `data`
structure, labeled according to `label`. This is equivalent to calling {terms = data:parse_inflection(field, props)}
followed by {return data:insert_inflection(terms, label, props)} and behaves the same as the combination of those two
functions. See their documentation for more details.
]==]
function Headdata:parse_and_insert_inflection(field, label, props)
local terms = self:parse_inflection(field, props)
return self:insert_inflection(terms, label, props)
end
--[==[
Generate an inflection that may be specified explicitly or defaulted (which involves looping over the specified or
defaulted heads and determining the script of each one, since the formation of the default depends on the script).
`data` is the data object passed into the POS handler. `terms` is the list of terms to process. Those where the term
itself is not `+` will be returned unchanged, while those where the term is `+` will be handled by generating the
appropriate inflections from the headwords using `make_inflection` (which is passed three arguments, `head`, `tr` and
`sccode`, i.e. the script code of `head`) and should return two values, term and translit, either of which can be
nil. A nil head will be ignored, and otherwise the decorations specified on the `+` term will be combined with the
decorations specified on the head. The return value is a list of inflections where no requests for the default
inflection remain.
]==]
function Headdata:resolve_special(terms, handle_special, props)
props = props or {}
local infls = {}
local is_special = props.is_special or function(_data, infl) return infl.term == "+" end
for _, termobj in ipairs(terms) do
if not is_special(self, termobj) then
insert(infls, termobj)
else
for _, headobj in ipairs(self.heads) do
local head = headobj.term or self.pagename
local head_no_links
if props.with_links then
head = head:find("%[") and head or require(headword_module).add_multiword_links(head, not headobj.term)
head_no_links = require(links_module).remove_links(head)
else
head = require(links_module).remove_links(head)
head_no_links = head
end
local newterms = handle_special {
head = head,
tr = headobj.tr,
infl = termobj,
sc = self.lang:findBestScript(head_no_links),
}
if newterms then
newterms = export.canonicalize_termobj_list(newterms, "term", "resolve_special")
for _, newterm in ipairs(newterms) do
if not props.no_combine_handle_special_retval_with_origin then
export.combine_termobj_decorations(newterm, termobj)
end
if not props.no_combine_handle_special_retval_with_head then
export.combine_termobj_decorations(newterm, headobj)
end
insert(infls, newterm)
end
end
end
end
end
return infls
end
--[==[
Add the current page to a tracking page named `Wiktionary:Tracking/``lang``-headword/``page```, where ``lang`` is the
language code of the current language. For example, if the current language is `mak` and `page` is {"redundant-lon"},
the current page will get added to the tracking page `Wiktionary:Tracking/mak-headword/redundant-lon`. All pages added
to that tracking page can be seen by going to [[Special:WhatLinksHere/Wiktionary:Tracking/mak-headword/redundant-lon]].
This is typically used to track issues occurring in user-specified parameters that do not rise to the level of errors
(e.g. redundant parameters, deprecated usages or other dispreferred values).
]==]
function Headdata:track(page)
return require(debug_track_module)(self.langcode .. "-headword/" .. page)
end
local boolean_param = {type = "boolean"}
--[==[
Process an arbitrary headword in an arbitrary language, handling generic and language-specific arguments and calling
`full_headword()` in [[Module:headword]]. This is intended for use in implementing headword modules (e.g.
[[Module:uz-headword]] for Uzbek, [[Module:mn-headword]] for Uzbek, [[Module:gsw-headword]] for Alemannic German, etc.)
and provides a general implementation of such modules. On input, `data` is an object with the following fields:
* `lang`: The language object of the language being handled. '''Required.''' Use the special value {false} to indicate
that the language is specified by the user in {{para|1}}.
* `frame`: The frame object passed into the `show()` function of your module, which implements headword-handling for
all parts of speech in the module, including a generic POS-handling template (e.g. {{tl|uz-head}} or {{tl|mn-head}}),
which allows arbitrary parts of speech to be handled. '''Required.'''
* `pos_functions`: A table listing, for each part of speech requiring special handling, the extra parameters (if any)
that the part of speech accepts, along with how to handle them. See the examples below. '''Required.'''
* `validate_lang`: If {lang = true} is specified, this is a function of one argument (a language object, based on the
language specified in {{para|1}}) that should throw an error if the language object is disallowed. If omitted, all
languages are allowed.
* `numbered_head`: If true, explicit headwords are specified in a numbered param instead of in {{para|head}}. The param
used is usually {{para|1}}, but is {{para|2}} for generic POS templates such as {{tl|mn-head}} or for POS-specific
templates when {lang = true} is specified (e.g. {{tl|arb-noun}}), and is {{para|3}} for generic POS templates when
{lang = true} is spcified (e.g. {{tl|arb-head}}).
* `include_tr`: If true, allow explicit transliterations to be specified. The transliteration(s) for the headword(s)
themselves is/are specified in {{para|tr}} or through the {{cd|<tr:...>}} inline modifier on headwords, and
transliterations of inflections are specified through the {{cd|<tr:...>}} inline modifier. This should generally be
given when a headword for the language may be in a script other than Latin.
* `include_ts`: If true, allow explicit transcriptions to be specified. The transcription(s) for the headword(s)
themselves is/are specified in {{para|ts}} or through the {{cd|<ts:...>}} inline modifier on headwords, and
transcriptions of inflections are specified through the {{cd|<ts:...>}} inline modifier. This should generally only
be given for certain languages where the spelling is radically different from the pronunciation (e.g. in cuneiform
languages such as Hittite and Akkadian, and potentially in Tibetan), and represents a pronunciation-based rendering
(usually not direct IPA).
* `include_sc`: If true, allow an explicit script code to be specified. The overall script code for the headword(s)
themselves can be specified using {{para|sc}}, and per-headword or per-inflection script codes are specified using the
{{cd|<sc:...>}} inline modifier. This should generally be given when a language supports multiple scripts.
* `infls`: An inflections structure specifying extra generic parameters that apply to all parts of speech and how to
handle them. The format is the same as for the `infls` structure in `pos_functions`.
* `augment_params`: A callback function to add extra generic parameters, or modify existing generic parameters in the
`params` structure; but in general, extra generic parameters should be added through the `infls` structure instead.
This callback should not be used to add part-of-speech-specific parameters; those are handled through the appropriate
setting in `pos_functions`. It is passed two single arguments, the `headdata` object and the `params` table to be
augmented. See below for extra fields stored in the `process_props` structure of the `headdata` object. This function
is called after initializing the `params` table and processing the overall `infls` structure, but just before adding
part-of-speech-specific parameters (from `pos_functions`) to `params`. Thus, it can override any generic parameters
but may itself be overridden by a part-of-speech-specific parameter.
* `augment_headdata`: A callback function to modify the `headdata` object passed to `full_headword()` in
[[Module:headword]]. This should not be used to for part-of-speech-specific parameter handling; this is handled
through the appropriate setting in `pos_functions`. It is passed a single argument, the `headdata` object, as for
`augment_params`; but the `process_props` structure and other fields will be more filled out, as this callback is
called later. This can be used, for example, to override the value of a generic setting in `headdata` (e.g.
[[Module:uz-headword]] uses this to mark non-Latin terms as variant forms by setting `headdata.var`, being careful not
to override a value already set by the user). This function is called after initializing the `headdata` table with all
information taken from generic parameters and processing generic parameters specified in the overall `infls`
structure, and just before processing the appropriate part-of-speech-specific `infls` structure in `pos_functions`
(which in turn is followed by any handler function in `pos_functions`). Thus, it can override any value set during
generic parameter processing but may itself be overridden by a part-of-speech-specific parameter or handler.
* `force_cat`: If true, add the headword to the appropriate categories even on non-mainspace pages. This can be used for
testing category handling in sample template calls on userspace test pages or template documentation pages. It should
not be set in production code.
* `enable_auto_translit`: If true, turn on automatic transliteration of inflections at a global level (i.e. applying to
all inflections). This has no effect on headwords, which are automatically transliterated by default if in a non-Latin
script and automatic transliteration is available for the language. You can also set this value for particular
inflections in the `insert_inflection()` function.
The `headdata` headword data structure has an extra field in it called `process_props` that is specific to the
`process_headword()` function, containing various extra properies. As the operation of `process_headword()` proceeds,
this object gets filled out with more fields. For example, once parameter parsing happens, the resulting values are
available in the `args` field of `process_props`. The following fields are found in `process_props` (note that `poscat`,
the canonicalized plural part of speech of the headword being processed, is *not* present here; it's directly on
`headdata`):
* `namespace`: The name of the current namespace; an empty string for the mainspace. This references the namespace of
the actual page and isn't affected by the {{para|pagename}} parameter.
* `indexing_poscat`: The canonicalized part of speech of the headword used to index into `pos_functions`. This is the
same as `poscat` for specific part-of-speech templates such as {{tl|uz-noun}}, but has the value {"head"} for generic
part-of-speech templates such as {{tl|uz-head}}. (Note that `poscat` is directly available on `headdata`.)
* `generic_pos_template`: True if a generic POS templates like {{tl|uz-head}} or {{tl|mn-head}} was used. (This is
signaled by omitting the invocation parameter {{para|1}} to `process_headword`.)
* `lang_in_1`: True if the language code is to be fetched from {{para|1}}.
* `pos_param`: The parameter holding the part of speech, if a generic POS tempalte like {{tl|mn-head}} is being
processed (i.e. `generic_pos_template` is set). In such a case, it will have the value of {1} or {2}, depending on
whether the language code is being fetched from {{para|1}} (see `lang_in_1`). Otherwise it will be {nil}.
* `head_param`: The parameter holding the explicit headword. If `numbered_head` was specified (as for Mongolian headword
templates), this has the value {1}, {2} or {3} depending on whether the language code is being fetched from {{para|1}}
(see `lang_in_1`) and whether a generic POS template like {{tl|mn-head}} is being processed (see
`generic_pos_template`). Otherwise, it has the value {"head"}. Also see the `lang` and `numbered_head` properties in
the `data` structure sent to `process_headword()`.
* `is_suffix`: True if the current term is a suffix. This is set when processing the `suffix`, `nosuffix` and `clitic`
parameters; it is always {false} beforehand (i.e. during `augment_params` and processing of the general `infls`
structure).
* `insert_specs`: This is a table mapping parameter names to the return value of `Headdata:insert_inflection()`, filled
out as parameter values are processed. This lets a given parameter processing function in `infls` gain access to the
result of calling `insert_inflection()` on previous parameters (which indicates the number of items inserted as well
as whether `-` was specified).
The `pos_functions` table contains an entry for each part of speech needing special handling, where the key is the
canonical plural part of speech (e.g. {"adverbs"} or {"proper nouns"}). The value associated with each key is a table
normally containing a field `infls`, listing the extra part-of-speech-specific inflection and other parameters along
with how to handle them. The specs in `infls` are used in three ways:
# to augment the `params` object passed to the `process()` function in [[Module:parameters]], specifying how to parse
the appropriate inflectional parameters;
# to specify how to process any inflectional parameters given and insert them into the `headdata` object passed to
`full_headword()` in [[Module:headword]];
# to generate appropriate documentation for the parameters and other changes made by the headword template (e.g.
inserting categories).
Alternatively, you can separately control the augmentation of the `params` object and the procesing of the resulting
arguments. This is done by specifying two fields in place of `infls`, named `params` and `func`. `params` is a table
containing extra parameters to add to the overall `params` object passed to the `process()` function in
[[Module:parameters]]. `func` is a function of two arguments, normally called `data` (the headword data structure
`headdata`) and `args` (the processed arguments table). However, this alternative method is not normally recommended
because it leads to duplication between the `params and `func` fields and the documentation, which must be manually
specified.
A simple example, as used to handle pronouns for Turoyo, is
{
local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"}
pos_functions["pronouns"] = {
infls = {
{2, type = "genders", validate = valid_genders},
{"f", label = "feminine"},
{"pl", label = "plural"},
},
}
}
The equivalent using `params` and `func` is
{
local valid_genders = {"m", "f", "m-p", "f-p", "p", "?"}
pos_functions["pronouns"] = {
params = {
[2] = {type = "genders"},
f = true,
pl = true,
},
func = function(data, args)
data:validate_genders(args[2], valid_genders)
data.genders = args[2]
data:parse_and_insert_inflection("f", "feminine")
data:parse_and_insert_inflection("pl", "plural")
end
}
}
Note how the version with separate `params and `func` is longer and splits information on the parameters between the
two fields. The `params` structure sets extra user-specifiable parameters {{para|2}} for genders (since the headword is
in {{para|1}}) as well a {{para|f}} and {{para|pl}}, and the `func` handler processes those parameters. Note how this is
done by calling methods on the headword `data` structure. Each such parameter can have multiple comma-separated values,
and each value can have inline modifiers attached to it to specify further properties of the value. The `infls` version
ends up making the same method calls, but does it for you instead of you having to do it yourself.
These methods are implemented through a metatable set on the headword `data` structure, which is removed before calling
`full_headword()` in [[Module:headword]]. The methods access extra information related to headword processing (such as
the `args` table) that is stored in the `process_props` field of the headword `data` strucuture. This field is also
removed prior to calling `full_headword()`.
The methods available on the headword `data` structure are as follows. Each one also has its own documentation.
* {parse_inflection(field, props)}: Parse value(s) specified in `field` (a user-specified parameter in the `args` table)
and return a list of term objects. Optional `props` specifies additional properties controlling the parsing.
* {insert_inflection(terms, label, props)}: Insert the terms in `terms` (a list of term objects as returned by
`parse_inflection()`) into the `inflections` list in the headword `data` structure, giving the inflection the label as
specified in `label`. Optional `props` specifies additional properties controlling the parsing.
* {parse_and_insert_inflection(field, label, props)}: A combination of `parse_inflection()` and `insert_inflection()`,
if no further processing of the parsed values needs to be done before insertion.
* {insert_fixed_inflection(label, props)}: Insert a "fixed" inflection (a label without associated values) into the
`inflections` table. An example (from {{tl|mn-noun}} in [[Module:mn-headword]]) is {"hidden-g declension"}, specifying
that the noun belongs to the hidden-''g'' declension.
* {resolve_special(terms, handle_special, props)}: Resolve "special" indicators as specified by the user in an inflection
parameter. A typical example is {"+"}, requesting a default value. `terms` is the list of parsed term objects and
`handle_special` is a handler function to process special indicators and convert them to their actual values.
* {validate_genders(genders, valid_genders, props)}: Validate that the user-specified genders in `genders` all belong to
the list given in `valid_genders`, throwing an error if not.
* {insert_category(category)}: Insert a category into the `categories` list in the headword `data` structure. `category`
is normally a string naming the category, which will have the language prepended to it and any occurrences of
`{plpos}` in the string replaced with the actual plural part of speech.
]==]
function export.process_headword(data)
local lang, frame, pos_functions, validate_lang, numbered_head, include_tr, include_ts, include_sc, force_cat,
enable_auto_translit, infls, augment_params, augment_headdata =
data.lang, data.frame, data.pos_functions, data.validate_lang, data.numbered_head, data.include_tr,
data.include_ts, data.include_sc, data.force_cat, data.enable_auto_translit, data.infls, data.augment_params,
data.augment_headdata
local iparams = {
[1] = true,
def = true,
}
local iargs = require(parameters_module).process(frame.args, iparams)
local parargs = frame:getParent().args
local langcode
if not lang then
error("Internal error: `data.lang` must be specified; either a language object or `true` for a user-specified language")
end
local lang_in_1
if lang == true then
lang_in_1 = true
langcode = ine(parargs[1])
if langcode then
langcode = mw.text.trim(langcode)
lang = require(languages_module).getByCode(langcode, 1, true)
if validate_lang then
validate_lang(lang)
end
else
error("Language code (see [[WT:Language codes]]) must be specified in 1=")
end
else
langcode = lang:getCode()
if validate_lang then
error("Internal error: `data.validate_lang` must not be specified if a language code is given in `data.lang`")
end
end
local poscat = iargs[1]
local generic_pos_template = not poscat
local pos_param
if generic_pos_template then
pos_param = lang_in_1 and 2 or 1
poscat = ine(parargs[pos_param]) or
mw.title.getCurrentTitle().fullText == ("Templat:%s-head"):format(langcode) and "interjection" or
error(("Part of speech must be specified in %s="):format(pos_param))
poscat = require(headword_module).canonicalize_pos(poscat)
end
local head_param = numbered_head and (generic_pos_template and lang_in_1 and 3 or
(generic_pos_template or lang_in_1) and 2 or 1) or "head"
local indexing_poscat = generic_pos_template and "head" or poscat
local namespace = mw.loadData(headword_data_module).page.namespace
-- Partly initialize headdata now for use in generic infls callbacks. Will be further initialized later after
-- processing parameters.
local headdata = {
lang = lang,
langcode = langcode,
langfullcode = lang:getFullCode(),
langname = lang:getCanonicalName(),
langfullname = lang:getFullName(),
process_props = {
namespace = namespace,
data = data,
indexing_poscat = indexing_poscat,
generic_pos_template = generic_pos_template,
lang_in_1 = lang_in_1,
pos_param = pos_param,
head_param = head_param,
is_suffix = false,
insert_specs = {},
},
pos_category = poscat,
orig_poscat = poscat, -- preserve user-specified poscat in case pos_category is changed to 'suffixes'
categories = {},
inflections = {enable_auto_translit = enable_auto_translit},
force_cat_output = force_cat,
no_redundant_head_cat = true,
}
setmetatable(headdata, {__index = Headdata})
local params = {
[head_param] = {template_default = iargs.def},
head2 = {replaced_by = false, instead = ("use comma-separated |%s="):format(head_param)},
id = true,
sort = true,
cat = true,
nolink = boolean_param,
nolinkhead = {type = "boolean", alias_of = "nolink"},
suffix = boolean_param,
nosuffix = boolean_param,
clitic = true,
addlpos = true,
var = {type = "boolean", allow = {"both"}},
json = boolean_param,
pagename = true, -- for testing
}
if include_sc then
params.sc = {type = "script"}
end
if include_tr then
params.tr = true
params.tr2 = {replaced_by = false, instead = "use comma-separated |tr= or <tr:...> inline modifier on head"}
end
if include_ts then
params.ts = true
params.ts2 = {replaced_by = false, instead = "use comma-separated |ts= or <ts:...> inline modifier on head"}
end
if lang_in_1 then
params[1] = {required = true} -- required but ignored as already processed above
end
if generic_pos_template then
params[pos_param] = {required = true} -- required but ignored as already processed above
end
local function resolve_prop(prop, ...)
if type(prop) == "function" then
prop = prop(headdata, ...)
end
return prop
end
local function augment_params_from_infls(infls)
infls = resolve_prop(infls)
for _, infl in ipairs(infls) do
local function interr(txt)
error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl)))
end
local param = infl[1]
if param then
param = resolve_prop(param)
if type(param) ~= "string" and type(param) ~= "number" then
interr(("Parameter name %s must be a string or number"):format(dump(param)))
end
-- We handle defaults as well as validation ourselves.
local typ = resolve_prop(infl.type) or "string"
if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then
-- FIXME: Handle more types.
interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format(
dump(typ)))
end
params[param] = {type = typ, required = resolve_prop(infl.required), template_default = resolve_prop(infl.template_default)}
if typ ~= "boolean" and type(param) == "string" then
params[param .. "2"] = {replaced_by = false, instead = ("use comma-separated |%s="):format(param)}
end
end
end
end
if infls then
augment_params_from_infls(infls)
end
if augment_params then
augment_params(headdata, params)
end
if pos_functions[indexing_poscat] then
local pos_infls = pos_functions[indexing_poscat].infls
if pos_infls then
augment_params_from_infls(pos_infls)
end
local pos_params = pos_functions[indexing_poscat].params
if pos_params then
for key, val in pairs(pos_params) do
params[key] = val
end
end
end
local args = require("Module:parameters").process(parargs, params)
local pagename = args.pagename or mw.loadData(headword_data_module).pagename
local sc = args.sc or lang:findBestScript(pagename)
headdata.pagename = pagename
headdata.process_props.args = args
headdata.sc = sc
headdata.id = args.id
headdata.sort = args.sort
-- No redundant script cat unless the user explicitly gave sc=
headdata.no_script_code_cat = not args.sc
headdata.var = args.var
local extra_term_mods = {}
if include_tr then
insert(extra_term_mods, "tr")
end
if include_ts then
insert(extra_term_mods, "ts")
end
if include_sc then
insert(extra_term_mods, "sc")
end
if not extra_term_mods[1] then
extra_term_mods = nil
end
local trs = args.tr and split_on_comma(args.tr) or {}
local num_trs = #trs
local tss = args.ts and split_on_comma(args.ts) or {}
local num_tss = #tss
local heads = args[head_param] and export.parse_term_with_modifiers {
val = args[head_param],
paramname = head_param,
splitchar = ",",
is_head = true,
include_mods = extra_term_mods,
} or {}
local num_heads = #heads
if num_heads > 0 and num_trs > 0 and num_heads ~= num_trs then
error(("%s head%s specified explicitly but %s translit%s; they must match; use '+' to stand for the default head (the pagename) or default automatic translit and '-' to stand for no translit"):format(
num_heads, num_heads > 1 and "s" or "", num_trs, num_trs > 1 and "s" or ""))
end
if num_heads > 0 and num_tss > 0 and num_heads ~= num_tss then
error(("%s head%s specified explicitly but %s transcription%s; they must match; use '+' to stand for the default head (the pagename) and '-' to stand for no transcription"):format(
num_heads, num_heads > 1 and "s" or "", num_tss, num_tss > 1 and "s" or ""))
end
if num_trs > 0 and num_tss > 0 and num_trs ~= num_tss then
error(("%s translit%s specified explicitly but %s transcription%s; they must match; use '+' to stand for default automatic translit and '-' to stand for no translit or transcription"):format(
num_trs, num_trs > 1 and "s" or "", num_tss, num_tss > 1 and "s" or ""))
end
-- Be careful here not to overwrite user_specified_heads if it's empty so we can later check user_specified_heads
-- to see if the user provided any heads.
local max_tr_ts = math.max(num_trs, num_tss)
if num_heads == 0 and max_tr_ts > 0 then
heads = {}
for i = 1, max_tr_ts do
heads[i] = {term = "+"}
end
end
if not heads[1] then
heads = {{term = "+"}}
end
for i, headobj in ipairs(heads) do
if headobj.tr and trs[i] then
if headobj.tr ~= trs[i] then
error(("Saw two different translits '%s' and '%s' for head #%s"):format(
headobj.tr, trs[i], i))
end
else
headobj.tr = headobj.tr or trs[i]
end
if headobj.tr == "+" then
headobj.tr = nil
end
if headobj.ts and tss[i] then
if headobj.ts ~= tss[i] then
error(("Saw two different transcriptions '%s' and '%s' for head #%s"):format(
headobj.ts, tss[i], i))
end
else
headobj.ts = headobj.ts or tss[i]
end
if headobj.ts == "-" then
headobj.ts = nil
end
if headobj.term == "+" then
headobj.term = args.nolink and pagename or nil
if headobj.term and namespace == "Reconstruction" then
headobj.term = "*" .. headobj.term
end
end
end
headdata.heads = heads
local function pagename_is_suffix()
if sc:getCode() == "Latn" then
-- shortcut Latin terms to avoid unnecessarily loading [[Module:affix]]
return pagename:find("^%-") and not pagename:find("%-$")
else
local affix_type, _, _, _ = require(affix_module).parse_term_for_affixes(pagename, lang, sc)
return affix_type == "suffix"
end
end
local clitic_label
if args.clitic then
clitic_label = require(yesno_module)(args.clitic, args.clitic)
end
if clitic_label == true then
clitic_label = "klitik"
end
if clitic_label then
headdata:insert_category("Klitik")
headdata:insert_fixed_inflection(clitic_label)
elseif args.suffix or (
not args.nosuffix and pagename_is_suffix() and poscat ~= "Akhiran" and poscat ~= "Bentuk akhiran"
) then
headdata.process_props.is_suffix = true
local function handle_suffix_pos(pos, is_first)
local form_type = pos:match("^Bentuk (.*)$")
local actual_poscat
if form_type then
headdata:insert_category(("Bentuk akhiran %s"):format(form_type))
headdata:insert_fixed_inflection("Bentuk akhiran " .. form_type)
else
local singular_pos = require(en_utilities_module).singularize(pos)
headdata:insert_category(("Akhiran membentuk %s"):format(singular_pos))
headdata:insert_fixed_inflection("Akhiran membentuk " .. singular_pos)
end
local postype = require(headword_module).pos_lemma_or_nonlemma(pos)
if not postype then
error(("Unrecognized canonicalized part of speech '%s' in addlpos=, cannot determine whether lemma or non-lemma form"):format(
pos
))
end
actual_poscat = postype == "Lema" and "Akhiran" or "Bentuk akhiran"
if is_first then
headdata.pos_category = actual_poscat
elseif headdata.pos_category ~= actual_poscat then
error(("Cannot mix suffixes and suffix forms using addlpos=; '%s' is a %s while overall POS '%s' is a %s; use separate POS headers for the two"):
format(pos, actual_poscat, poscat, headdata.pos_category))
end
end
handle_suffix_pos(poscat, true)
if args.addlpos then
for _, addlpos in ipairs(split(args.addlpos, "%s*,%s*")) do
addlpos = require(headword_module).canonicalize_pos(addlpos)
handle_suffix_pos(addlpos, false)
end
end
end
if args.cat then
for _, cat in ipairs(split_on_comma(args.cat)) do
headdata:insert_category(cat)
end
end
local function augment_headdata_from_infls(infls)
infls = resolve_prop(infls)
for _, infl in ipairs(infls) do
local function interr(txt)
error(("Internal error: %s (coming from infls spec %s)"):format(txt, dump(infl)))
end
local function process_labelobjs(labelobjs, originating_term, handle_labelobj)
if labelobjs == nil then
return
end
if type(labelobjs) ~= "string" and type(labelobjs) ~= "table" then
interr(("Wrong type '%s' for label object(s) %s, expected string or table"):format(
type(labelobjs), dump(labelobjs)
))
end
if type(labelobjs) == "string" or type(labelobjs) == "table" and not labelobjs[1] then
labelobjs = {labelobjs}
end
for _, labelobj in ipairs(labelobjs) do
local label, termobj
if type(labelobj) == "string" then
label = labelobj
termobj = originating_term
elseif type(labelobj) ~= "table" then
interr(("Wrong type '%s' for label object %s, expected string or table"):format(
type(labelobj), dump(labelobj)
))
label = labelobj.term
if type(label) ~= "string" then
interr(("Wrong type '%s' for label %s from label object %s, expected string"):format(
type(label), dump(label), dump(labelobj)
))
end
termobj = labelobj
end
handle_labelobj(label, termobj)
end
end
local param = infl[1]
if param then
local vals
-- Fetch the param and make sure it's a string or number.
param = resolve_prop(param)
if type(param) ~= "string" and type(param) ~= "number" then
interr(("Parameter name %s must be a string or number"):format(dump(param)))
end
-- Fetch the type and validate.
local typ = resolve_prop(infl.type)
if typ == nil then
typ = "string"
end
if typ ~= "genders" and typ ~= "boolean" and typ ~= "string" then
-- FIXME: Handle more types.
interr(('Unrecognized type %s; can only currently handle "genders", "boolean" and "string" (the default)'):format(
dump(typ)))
end
-- Fetch the value(s).
if typ == "genders" or typ == "boolean" then
vals = args[param]
elseif typ == "string" then
local parse_inflection_props = resolve_prop(infl.parse_inflection_props)
local include_mods = resolve_prop(infl.include_mods)
local no_augment_include_mods = resolve_prop(infl.no_augment_include_mods)
if include_mods ~= nil or no_augment_include_mods ~= nil then
if parse_inflection_props == nil then
parse_inflection_props = {}
else
parse_inflection_props = shallow_copy(parse_inflection_props)
end
if include_mods ~= nil then
parse_inflection_props.include_mods = include_mods
end
if no_augment_include_mods ~= nil then
parse_inflection_props.no_augment_include_mods = no_augment_include_mods
end
end
vals = headdata:parse_inflection(param, parse_inflection_props)
-- Convert an empty list to nil for consistent checking below.
if not vals[1] then
vals = nil
end
else
interr(("Unrecognized type '%s"):format(typ))
end
-- If value(s) nil, fetch the default.
if vals == nil and infl.default ~= nil then
local default = resolve_prop(infl.default)
if typ == "genders" then
vals = export.canonicalize_termobj_list(default, "spec", "default")
elseif typ == "boolean" then
vals = default
elseif typ == "string" then
vals = export.canonicalize_termobj_list(default, "term", "default")
else
interr(("Unrecognized type '%s"):format(typ))
end
end
-- Resolve "special" values (special signals a string values, such as requesting the default with "+").
if vals ~= nil and infl.resolve_special then
if typ ~= "string" then
interr(("Cannot specify resolve_special= for type %s"):format(dump(typ)))
end
local resolve_special_props = resolve_prop(infl.resolve_special_props, vals)
if infl.is_special ~= nil then
if resolve_special_props == nil then
resolve_special_props = {}
else
resolve_special_props = shallow_copy(resolve_special_props)
end
resolve_special_props.is_special = infl.is_special
end
vals = headdata:resolve_special(vals, infl.resolve_special, resolve_special_props)
end
-- Validate the value(s).
if vals ~= nil and infl.validate ~= nil then
if typ == "boolean" then
interr('Cannot specify validate= when type is "boolean"')
elseif type(infl.validate) == "function" then
infl.validate(headdata, vals)
elseif typ == "genders" then
headdata:validate_genders(vals, infl.validate)
elseif typ == "string" then
validate_items {
items = vals,
field = "term",
valid_items = infl.validate,
item_type = ("values in |%s="):format(param),
}
else
interr(("Unrecognized type '%s"):format(typ))
end
end
-- Run the process_after_parse handler, if it exists.
if vals ~= nil and infl.process_after_parse ~= nil then
local intentionally_nil
vals, intentionally_nil = infl.process_after_parse(headdata, vals)
if vals == nil and not intentionally_nil then
interr("If you return nil from process_after_parse, you must return a second non-nil return " ..
"value to indicate that the nil return value was intentional")
end
end
-- "Implement" the values, if non-falsy (i.e. we don't want to fire on boolean false or empty list).
-- If a fixed label is specified, insert it. Then, depending on the type, attach the values to a label
-- as an inflection, set the `genders` field, or do nothing if boolean (throwing an error if there was
-- no fixed label).
if vals == true or type(vals) == "table" and vals[1] then
if infl.fixed_label and infl.all_fixed_label then
interr("Cannot specify both fixed_label= and all_fixed_label=; specify one or the other")
end
local function check_fixed_label_references_val(label)
if type(label) == "table" and label[1] then
for _, lab in ipairs(label) do
if check_fixed_label_references_val(lab) then
return true
end
end
return false
end
if type(label) == "table" then
if not label.term then
interr(("Fixed label structure %s does not have a value for `.term`"):format(dump(label)))
end
label = label.term
end
if type(label) ~= "string" then
interr(("Wrong type for fixed label %s, should be string"):format(type(label)))
end
return not not label:find("{val}")
end
local fixed_label = infl.fixed_label
local all_fixed_label = infl.all_fixed_label
-- If the value being processed is boolean, there's only one value so treat a fixed_label as an
-- all_fixed_label and output only once; likewise if the caller specified a fixed_label without
-- {val} in it.
if fixed_label and (typ == "boolean" or type(fixed_label) ~= "function" and not
check_fixed_label_references_val(fixed_label)) then
all_fixed_label = fixed_label
fixed_label = nil
end
local inserted_fixed_label
if fixed_label then
if typ == "boolean" then
interr("Boolean fixed_label values should have been converted to all_fixed_label")
end
for _, valobj in ipairs(vals) do
local labelobjs = resolve_prop(fixed_label, valobj)
process_labelobjs(labelobjs, valobj, function(label, termobj)
if label:find("{val}") then
if typ ~= "string" then
interr(('Cannot specify {val} in fixed_label %s when type is "%s"'):format(dump(label), typ))
end
label = label:gsub("{val}", replacement_escape(valobj.term))
end
headdata:insert_fixed_inflection(label, {
originating_term = termobj
})
inserted_fixed_label = true
end)
end
elseif all_fixed_label then
local labelobjs = resolve_prop(all_fixed_label, vals)
process_labelobjs(labelobjs, nil, function(label, termobj)
if label:find("{vals}") then
if typ ~= "string" then
interr(('Cannot specify {vals} in all_fixed_label %s when type is "%s"'):format(dump(label), typ))
end
local formatted_labels = {}
for _, valobj in ipairs(vals) do
insert(formatted_labels, add_decorations(valobj.term, valobj, lang))
end
label = label:gsub("{vals}", replacement_escape(serial_comma_join(formatted_labels)))
end
headdata:insert_fixed_inflection(label, termobj)
inserted_fixed_label = true
end)
end
local inserted_vals
if infl.label ~= nil then
if typ ~= "string" then
interr(("label=%s can only be specified for type 'string', not '%s'"):format(
dump(infl.label), typ
))
end
local label = resolve_prop(infl.label, vals)
if label ~= nil then
local insert_inflection_props = resolve_prop(infl.insert_inflection_props, vals)
local no_auto_cats = resolve_prop(infl.no_auto_cats, vals)
if no_auto_cats ~= nil then
if insert_inflection_props == nil then
insert_inflection_props = {}
else
insert_inflection_props = shallow_copy(insert_inflection_props)
end
insert_inflection_props.no_auto_cats = infl.no_auto_cats
end
local insert_spec = headdata:insert_inflection(vals, label, insert_inflection_props)
headdata.process_props.insert_specs[param] = insert_spec
inserted_vals = true
end
end
if typ == "genders" then
headdata.genders = vals
end
local inserted_cat
if infl.cat then
local allcats = {}
for _, valobj in ipairs(vals) do
local cats = resolve_prop(infl.cat, valobj)
if type(cats) == "string" then
cats = {cats}
end
if cats ~= nil then
for _, cat in ipairs(cats) do
if cat:find("{val}") then
cat = cat:gsub("{val}", replacement_escape(valobj.term))
end
insert_if_not(allcats, cat)
end
end
end
for _, cat in ipairs(allcats) do
headdata:insert_category(cat)
inserted_cat = true
end
end
if typ == "boolean" then
if not inserted_fixed_label and not inserted_cat then
interr(("User set boolean setting for %s= but no fixed label added and no category " ..
"inserted; if you took action in process_after_parse(), make sure to return " ..
"`nil, true`"):format(param))
end
elseif typ == "string" then
if not inserted_vals and not inserted_fixed_label then
interr(("User set value(s) %s for %s= but no inflection inserted and no fixed label " ..
"added; if you took action in process_after_parse(), make sure to return " ..
"`nil, true`"):format(dump(vals), param))
end
end
end
else -- no param specified
if infl.label or infl.all_fixed_label then
interr("Cannot have label= or all_fixed_label= without specifying a param")
end
if infl.fixed_label then
local labelobjs = resolve_prop(infl.fixed_label)
process_labelobjs(labelobjs, nil, function(label, termobj)
if label:find("{val}") then
interr("Cannot specify {val} in a fixed_label= value without specifying a param")
end
headdata:insert_fixed_inflection(label, {
originating_term = termobj
})
end)
end
if infl.cat then
local cats = resolve_prop(infl.cat)
if type(cats) == "string" then
cats = {cats}
end
if cats ~= nil then
for _, cat in ipairs(cats) do
if cat:find("{val}") then
interr("Cannot specify {val} in a cat= value without specifying a param")
end
headdata:insert_category(cat)
end
end
end
end
end
end
if infls then
augment_headdata_from_infls(infls)
end
if augment_headdata then
augment_headdata(headdata, args)
end
if pos_functions[indexing_poscat] then
local pos_infls = pos_functions[indexing_poscat].infls
if pos_infls then
augment_headdata_from_infls(pos_infls)
end
local func = pos_functions[indexing_poscat].func
if func then
func(headdata, args)
end
end
setmetatable(headdata, nil)
if args.json then
return require("Module:JSON").toJSON(headdata)
end
headdata.process_props = nil
return require(headword_module).full_headword(headdata)
end
return export
nrwzxukyx15w9bs1qysq8d6jicifovd
Modul:parameter utilities
828
55604
375382
280880
2026-09-22T07:04:25Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92708725|92708725]])
375382
Scribunto
text/plain
local export = {}
local debug_track_module = "Module:debug/track"
local functions_module = "Module:fun"
local parameters_module = "Module:parameters"
local parse_interface_module = "Module:parse interface"
local parse_utilities_module = "Module:parse utilities"
local table_module = "Module:table"
local dump = mw.dumpObject
local error = error
local insert = table.insert
local ipairs = ipairs
local next = next
local pairs = pairs
local require = require
local tonumber = tonumber
local type = type
--[==[
Loaders for functions in other modules, which overwrite themselves with the target function when called. This ensures
modules are only loaded when needed, retains the speed/convenience of locally-declared pre-loaded functions, and has no
overhead after the first call, since the target functions are called directly in any subsequent calls.
]==]
local function debug_track(...)
debug_track = require(debug_track_module)
return debug_track(...)
end
local function is_callable(...)
is_callable = require(functions_module).is_callable
return is_callable(...)
end
local function list_to_set(...)
list_to_set = require(table_module).listToSet
return list_to_set(...)
end
local function parse_inline_modifiers(...)
parse_inline_modifiers = require(parse_interface_module).parse_inline_modifiers
return parse_inline_modifiers(...)
end
local function process_params(...)
process_params = require(parameters_module).process
return process_params(...)
end
local function shallow_copy(...)
shallow_copy = require(table_module).shallowCopy
return shallow_copy(...)
end
local function table_len(...)
table_len = require(table_module).length
return table_len(...)
end
----------------- end loaders ----------------
local function track(page, track_module)
return debug_track((track_module or "parameter utilities") .. "/" .. page)
end
-- Throw an error prefixed with the words "Internal error" (and suffixed with a dumped version of `spec`, if provided).
-- This is for logic errors in the code itself rather than template user errors.
local function internal_error(msg, spec)
if spec then
msg = ("%s: %s"):format(msg, dump(spec))
end
error(("Internal error: %s"):format(msg))
end
-- Table listing the default recognized special separator arguments and how they display.
export.default_special_separators = {
[";"] = "; ",
["_"] = " ",
["~"] = " ~ ",
["→"] = " → ",
}
-- Table listing how subitem delimiters display. Unlike for `default_special_separators`, the presence of an item in
-- this table does not mean that the delimiter is recognized; only those specified by `data.splitchar` are recognized.
export.default_subitem_separator_map = {
[";"] = "; ",
[","] = ", ",
["/"] = "/",
["_"] = " ",
["~"] = " ~ ",
["→"] = " → ",
}
--[==[ intro:
The purpose of this module is to facilitate implementation of templates that can have arguments specified either through
inline modifiers or separate parameters. There are two types of templates supported: those that take a list of items
with associated properties, which can be specified either through indexed separate parameters (e.g. {{para|t2}},
{{para|pos3}}) or inline modifiers (`<t:...>`, `<pos:...>`, etc.); and those that take a single term, whose properties
can be specified through non-indexed separate parameters (e.g. {{para|t}} or {{para|pos}}) or inline modifiers. Both
types of templates can optionally have subitems in the term parameter(s), where the subitems are typically (but not
necessarily) separated with commas and each subitem can have its own inline modifiers.
Some examples of templates that take a list of items are {{tl|alter}}/{{tl|alt}}; {{tl|synonyms}}/{{tl|syn}},
{{tl|antonyms}}/{{tl|ant}}, and other "nyms" templates; {{tl|col}}, {{tl|col2}}, {{tl|col3}}, {{tl|col4}} and other
column templates; {{tl|descendant}}/{{tl|desc}}; {{tl|affix}}/{{tl|af}}, {{tl|prefix}}/{{tl|pre}} and related *fix
templates; {{tl|affixusex}}/{{tl|afex}} and related templates; {{tl|IPA}}; {{tl|homophones}}; {{tl|rhymes}}; and several
others.
Examples of templates that take a single item are form-of templates ({{tl|inflection of}}/{{tl|infl of}},
{{tl|form of}}, and specific templates such as {{tl|alt form}}/{{tl|alternative form of}},
{{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others); for etymology templates
({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates like {{tl|ellipsis}},
{{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and other templates that take an argument structure
similar to {{tl|l}} or {{tl|m}}.
This module can be thought of as a combination of [[Module:parameters]] (which parses template parameters, and in
particular handles the separate parameter versions of the properties) and `parse_inline_modifiers()` in
[[Module:parse utilities]] (which parses inline modifiers).
The two main entry points are `parse_list_with_inline_modifiers_and_separate_params()` (for templates that take a list
of items) and `parse_term_with_inline_modifiers_and_separate_params()` (for templates that take a single item). However,
there are other functions provided, e.g. to initialize the `param_mods` structure that is passed to the two entry
points.
The typical workflow for using `parse_list_with_inline_modifiers_and_separate_params()` looks as follows (a slightly
simplified version of the code in [[Module:nyms]]):
{
local export = {}
local parameter_utilities_module = "Module:parameter utilities"
...
-- Entry point to be invoked from a template.
function export.show(frame)
local parent_args = frame:getParent().args
-- Parameters that don't have corresponding inline modifiers. Note in particular that the parameter corresponding to
-- the items themselves must be specified this way, and must specify either `allow_holes = true` (if the user can
-- omit terms, typically by specifying the term using |altN= or <alt:...> so that they remain unlinked) or
-- `disallow_holes = true` (if omitting terms is not allowed). (If neither `allow_holes` nor `disallow_holes` is
-- specified, an error is thrown in parse_list_with_inline_modifiers_and_separate_params().)
local params = {
[1] = {required = true, type = "language", default = "und"},
[2] = {list = true, allow_holes = true, required = true, default = "term"},
}
local m_param_utils = require(parameter_utilities_module)
-- This constructs the `param_mods` structure by adding well-known groups of parameters (such as all the parameters
-- associated with based on full_link() in [[Module:links]], with default properties that can be overridden. This is
-- easier and less error-prone than manually specifying the `param_mods` structure (see below for how this would
-- look). Here, we specify the group "link" (consisting of all the link parameters for use with full_link()), group
-- "ref" (which adds the "ref" parameter for specifying references), group "l" (which adds the "l" and "ll"
-- parameters for specifying labels) and group "q" (which adds the "q" and "qq" parameters for specifying regular
-- qualifiers). By default, labels and qualifiers have `separate_no_index` set so that e.g. |q1= is distinct from
-- |q=, the former specifying the left qualifier for the first item and the latter specifying the overall left
-- qualifier. For compatibility, we override the `separate_no_index` setting for the group "q", which causes |q= and
-- |q1= to be the same, and likewise for |qq= and |qq1=. Finally, also for compatibility, we add an "lb" parameter
-- that is an alias of "ll" (in all respects; |lb= is the same as |ll=, |lb1= is the same as |ll1=, <lb:...> is the
-- same as <ll:...>, etc.).
local param_mods = m_param_utils.construct_param_mods {
{group = {"link", "ref", "l"}},
{group = "q", separate_no_index = false},
{param = "lb", alias_of = "ll"},
}
-- This processes the raw arguments in `parent_args`, parses inline modifiers and creates corresponding objects
-- containing the property values specified either through inline modifiers or separate parameters.
local items, args = m_param_utils.parse_list_with_inline_modifiers_and_separate_params {
params = params,
param_mods = param_mods,
raw_args = parent_args,
termarg = 2,
parse_lang_prefix = true,
track_module = "nyms",
lang = 1,
sc = "sc.default",
}
local lang = args[1]
-- Now do the actual implementation of the template. Generally this should be split into a separate function, often
-- in a separate module (if the implementation goes in [[Module:foo]], the template interface code goes in
-- [[Module:foo/templates]]).
...
}
The `param_mods` structure controls the properties that can be specified by the user for a given item, and is
conceptually very similar to the `param_mods` structure used by `parse_inline_modifiers()`. The key is the name of the
parameter (e.g. {"t"}, {"pos"}) and the value is a table with optional elements as follows:
* `item_dest`, `store`: Same as the corresponding fields in the `param_mods` structure passed to
`parse_inline_modifiers()`.
* `type`, `set`, `sublist`, `convert` and associated fields such as `family` and `method`: These control parsing and
conversion of the raw values specified by the user and have the same meaning as in [[Module:parameters]] and also in
`parse_inline_modifiers()` (which delegates the actual conversion to [[Module:parameters]]). These fields — and for
that matter, all fields other than `item_dest`, `store` and `overall` — are forwarded to the `process()` function in
[[Module:parameters]].
* `alias_of`: This parameter is an alias of some other parameter. This spec is recognized only by `process()` in
[[Module:parameters]], and not by `parse_inline_modifiers()`; to set up an alias in `parse_inline_modifiers()`, you
need to make sure (using `item_dest`) that both the alias and aliasee modifiers store their values in the same
location, and you need to copy the remaining properties from the aliasee's spec to the aliasing modifier's spec. All
of this happens automatically if you generate the `param_mods` structure using `construct_param_mods()`.
* `require_index`: This means that the non-indexed parameter version of the property is not recognized. E.g. in the
case of the {"sc"} property, use of the {{para|sc}} parameter would result in an error, while {{para|sc1}} is
recognized and specifies the {"sc"} property for the first item. The default, if neither `require_index` nor
`separate_no_index` is given, is for {{para|sc}} and {{para|sc1}} to mean the same thing (both would specify the
{"sc"} property of the first item). Note that `require_index` and `separate_no_index` are mutually exclusive, and if
either one is specified during processing by `construct_param_mods()`, the other one is automaticallly turned off.
* `separate_no_index`: This means that e.g. the {{para|sc}} parameter is distinct from the {{para|sc1}} parameter
(and thus from the `<sc:...>` inline modifier on the first item). This is typically used to distinguish an overall
version of a property from the corresponding item-specific property on the first item. (In this case, for example,
{{para|sc}} overrides the script code for all items, while {{para|sc1}} overrides the script code only for the
first item.) If not given, and if `require_index` is not given, {{para|sc}} and {{para|sc1}} would have the same
meaning and refer to the item-specific property on the first item. When this is given, the overall value can be
accessed using the `.default` field of the property value in `args`, e.g. in this case `args.sc.default`. Note that
(as mentioned above) `require_index` and `separate_no_index` are mutually exclusive, and if either one is specified
during processing by `construct_param_mods()`, the other one is automaticallly turned off.
* `list`, `allow_holes`, `disallow_holes`: These should '''not''' be given. `list` and `allow_holes` are automatically
set for all parameter specs added to the `params` structure used by `process()` in [[Module:parameters]], and
`disallow_holes` clashes with `allow_holes`.
For the above workflow example, the call to `construct_param_mods()` generates the following `param_mods` structure:
{
local param_mods = {
-- the parameters generated by group "link"
alt = {},
t = {
-- [[Module:links]] expects the gloss in "gloss".
item_dest = "gloss",
},
gloss = {
alias_of = "t",
},
tr = {},
ts = {},
g = {
-- [[Module:links]] expects the genders in "genders".
item_dest = "genders",
type = "genders",
},
pos = {},
ng = {},
lit = {},
id = {},
sc = {
separate_no_index = true,
type = "script",
},
-- the parameters generated by group "ref"
ref = {
item_dest = "refs",
type = "references",
},
-- the parameters generated by group "l"
l = {
type = "labels",
separate_no_index = true,
},
ll = {
type = "labels",
separate_no_index = true,
},
-- the parameters generated by group "q"; note that `separate_no_index = true` would be set, but is overridden
-- (specifying `separate_no_index = false` in the `param_mods` structure is equivalent to not specifying it at all)
q = {
type = "qualifier",
separate_no_index = false,
},
qq = {
type = "qualifier",
separate_no_index = false,
},
infl = {
type = "form of tags",
separate_no_index = true,
},
-- the parameter generated by the individual "lb" parameter spec; note that only `alias_of` was explicitly given,
-- while `item_dest` is automatically set so that inline modifier <lb:...> stores into the same place as <ll:...>,
-- and the other specs are copied from the `ll` spec so `lb` works like `ll` in all regards
lb = {
alias_of = "ll",
item_dest = "ll",
type = "labels",
separate_no_index = true,
},
}
}
]==]
local qualifier_spec = {
type = "qualifier",
separate_no_index = true,
}
local label_spec = {
type = "labels",
separate_no_index = true,
}
local form_of_spec = {
type = "form of tags",
separate_no_index = true,
}
local recognized_param_mod_groups = {
link = {
alt = {},
t = {
-- [[Module:links]] expects the gloss in "gloss".
item_dest = "gloss",
},
gloss = {
alias_of = "t",
},
tr = {},
ts = {},
g = {
-- [[Module:links]] expects the genders in "genders".
item_dest = "genders",
type = "genders",
},
pos = {},
ng = {},
lit = {},
id = {},
sc = {
separate_no_index = true,
type = "script",
},
},
lang = {
lang = {
require_index = true,
type = "language",
},
},
q = {
q = qualifier_spec,
qq = qualifier_spec,
},
a = {
a = label_spec,
aa = label_spec,
},
l = {
l = label_spec,
ll = label_spec,
},
infl = {
infl = form_of_spec,
},
ref = {
ref = {
item_dest = "refs",
type = "references",
},
},
}
local function merge_param_mod_settings(orig, additions)
local merged = shallow_copy(orig)
for k, v in pairs(additions) do
merged[k] = v
if k == "require_index" then
merged.separate_no_index = nil
elseif k == "separate_no_index" then
merged.require_index = nil
end
end
merged.default = nil
merged.group = nil
merged.param = nil
merged.exclude = nil
merged.include = nil
return merged
end
local function verify_type(spec, param, typ1, typ2)
if not spec[param] then
return
end
local val = spec[param]
if type(val) ~= typ1 and (not typ2 or type(val) ~= typ2) then
internal_error(("Parameter `%s` must be a %s%s but saw a %s"):format(param, typ1, typ2 and " or " .. typ2 or "",
type(val)), spec)
end
end
local function verify_well_constructed_spec(spec)
local num_control = (spec.default and 1 or 0) + (spec.group and 1 or 0) + (spec.param and 1 or 0)
if num_control == 0 then
internal_error(
"Spec passed to construct_param_mods() must have either the `default`, `group` or `param` keys set", spec)
end
if num_control > 1 then
internal_error(
"Exactly one of `default`, `group` or `param` must be set in construct_param_mods() spec", spec)
end
if spec.list or spec.allow_holes then
-- FIXME: We need to support list = "foo" for list parameters that are stored in e.g. 2=, foo2=, foo3=, etc.
internal_error("`list` and `allow_holes` may not be set; they are automatically set when constructing the " ..
"corresponding spec in the `params` object passed to [[Module:parameters]]", spec)
end
if spec.disallow_holes then
internal_error("`disallow_holes` may not be set; it conflicts with `allow_holes`, which is automatically " ..
"set when constructing the corresponding spec in the `params` object passed to [[Module:parameters]]", spec)
end
if spec.include and spec.exclude then
internal_error("Saw both `include` and `exclude` in the same spec", spec)
end
if (spec.include or spec.exclude) and not spec.group then
internal_error(
"`include` and `exclude` can only be specified along with `group`, not with `default` or `param`", spec)
end
verify_type(spec, "group", "string", "table")
verify_type(spec, "param", "string", "table")
verify_type(spec, "include", "table")
verify_type(spec, "exclude", "table")
end
--[==[
Construct the `param_mods` structure used in parsing arguments and inline modifiers from a list of specifications.
A sample invocation (a slightly simplified version of the actual invocation associated with {{tl|affix}} and related
templates) looks like this:
{
local param_mods = require("Module:parameter utilities").construct_param_mods {
-- We want to require an index for all params (or use separate_no_index, which also requires an index for the
-- param corresponding to the first item).
{default = true, require_index = true},
{group = {"link", "ref", "lang", "q", "l"}},
-- Override these two to have separate_no_index.
{param = {"lit", "pos"}, separate_no_index = true},
}
}
Each specification either sets the default value for further parameter specs or adds one or more parameters. Parameters
can be added directly using `param`, or groups of predefined parameters can be added using `group`. Specifications are
one of three types:
# Those that set the default properties for future-added parameters. These contain {default = true} as one of the
properties of the spec. Specs are processed in order and you can change the defaults mid-way through.
# Those that add the parameters associated with one or more pre-defined groups. These contain {group = "group"} or
{group = {"group1", "group2", ...}}. The pre-defined parameter groups and their associated properties are listed
below. The pre-defined properties of parameters in a group override properties associated with a {default = true}
spec, and are in turn overridden by any properties given directly in the spec itself. Note as well that setting the
`separate_no_index` property will automatically cause the `require_index` property to be unset and vice-versa, as the
two are mutually exclusive. (This happens in the example above, where the {separate_no_index = true} setting
associated with the params {"lit"} and {"pos"} cancels out the {require_index = true} default setting, as well as less
obviously with the pre-defined {"sc"} property of the {"link"} group, the {"q"} and {"qq"} properties of the {"q"}
group, and the {"l"} and {"ll"} properties of the {"l"} group, all of which have an associated pre-defined property
{separate_no_index = true}, which overrides and cancels out the {require_index = true} default setting. Finally, when
adding the parameters of a group, you can request the only a subset of the parameters be added using either the
`include` or `exclude` properties, each of whose values is a list of parameters that specify (respectively) the
parameters to include (all other parameters of the group are excluded) or to exclude (all other parameters of the
group are included). This is used, for example, in [[Module:romance etymology]] and [[Module:it-etymology]], which
specify {group = "link", exclude = {"tr", "ts", "sc"}} to exclude link parameters that aren't relevant to Latin-script
languages such as the Romance languages, and conversely in [[Module:IPA/templates]], which specifies
{group = "link", include = {"t", "gloss", "pos"}} to include only the specified parameters for use with {{tl|IPA}}.
# Those that add individual parameters. These contain {param = "param"} or {param = {"param1", "param2", ...}}, the
latter syntax used to control a set of parameters together. The resulting spec is formed by initializing the
parameter's settings with any previously-specified default properties (using a spec containing {default = true}) if
the parameter hasn't already been initialized, and then overriding the resulting settings with any settings given
directly in the specification. In the above example, the {"lit"} and {"pos"} parameters were previously initialized
through the {"link"} group (specified in the second of the three specifications) but ended up with
{require_index = true} due to the {default = true} spec (the first of the three specifications). We override these
two parameters to have {separate_no_index = true} (which, as mentioned above, cancels out {require_index = true}).
This is done so that {{tl|affix}} and related templates have {{para|pos}} and {{para|lit}} parameters distinct from
{{para|pos1}} and {{para|lit1}}, which are used to specify an overall part of speech (which applies to all parts of
the affix, as opposed to applying to just one element of the expression) or a literal definition for the entire
expression (instead of just for one element of the expression).
The built-in parameter groups are as follows:
{|class="wikitable"
! Group !! Group meaning !! Parameter !! Parameter meaning !! Default properties
|-
| rowspan=11| `link`
| rowspan=11| link parameters; same as those available on {{tl|l}}, {{tl|m}} and other linking templates
| `alt` || display text, overriding the term's display form || —
|-
| `t` || gloss (translation) of a non-English term || {item_dest = "gloss"}
|-
| `gloss` || gloss (translation); same as `t` || {alias_of = "t"}
|-
| `tr` || transliteration of a non-Latin-script term; only needed if the automatic transliteration is incorrect or unavailable (e.g. in Hebrew, which doesn't have automatic transliteration) || —
|-
| `ts` || transcription of a non-Latin-script term, if the transliteration is markedly different from the actual pronunciation; should not be used for IPA pronunciations || —
|-
| `g` || comma-separated list of genders; whitespace may surround the comma and will be ignored || {item_dest = "genders", type = "genders"}
|-
| `pos` || part of speech for the term || —
|-
| `ng` || arbitrary non-gloss descriptive text for the term || —
|-
| `lit` || literal meaning (translation) of the term || —
|-
| `id` || a sense ID for the term, which links to anchors on the page set by the {{tl|senseid}} template || —
|-
| `sc` || the script code (see [[Wiktionary:Scripts]]) for the script that the term is written in; rarely necessary, as the script is autodetected (in most cases, correctly) || {separate_no_index = true, type = "script"}
|-
| rowspan=2| `q`
| rowspan=2| left and right normal qualifiers (as displayed using {{tl|q}})
| `q` || left normal qualifier || {separate_no_index = true, type = "qualifier"}
|-
| `qq` || right normal qualifier || {separate_no_index = true, type = "qualifier"}
|-
| rowspan=2| `a`
| rowspan=2| left and right accent qualifiers (as displayed using {{tl|a}})
| `a` || comma-separated list of left accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"}
|-
| `aa` || comma-separated list of right accent qualifiers; whitespace must not surround the comma || {separate_no_index = true, type = "labels"}
|-
| rowspan=2| `l`
| rowspan=2| left and right labels (as displayed using {{tl|lb}}, but without categorizing)
| `l` || comma-separated list of left labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"}
|-
| `ll` || comma-separated list of right labels; whitespace must not surround the comma || {separate_no_index = true, type = "labels"}
|-
| `ref`
| reference(s) (in the format accepted by [[Module:references]]; see also the documentation for the {{para|ref}} parameter to {{tl|IPA}})
| `ref` || one or more references, in the format accepted by [[Module:references]] || {item_dest = "refs", type = "references"}
|-
| `lang`
| language for an individual term (provided for compatibility; it is preferred to specify languages for individual terms using language prefixes instead)
| `lang` || language code (see [[Wiktionary:Languages]]) for the term || {require_index = true, type = "language"}
|}
]==]
function export.construct_param_mods(specs)
local param_mods = {}
local default_specs = {}
for _, spec in ipairs(specs) do
verify_well_constructed_spec(spec)
if spec.default then
-- This will have an extra `default` field in it, but it will be erased by merge_param_mod_settings()
default_specs = spec
else
if spec.group then
local groups = spec.group
if type(groups) ~= "table" then
groups = {groups}
end
local include_set
if spec.include then
include_set = list_to_set(spec.include)
end
local exclude_set
if spec.exclude then
exclude_set = list_to_set(spec.exclude)
end
for _, group in ipairs(groups) do
local group_specs = recognized_param_mod_groups[group]
if not group_specs then
internal_error(("Unrecognized built-in param mod group '%s'"):format(group), spec)
end
for group_param, group_param_settings in pairs(group_specs) do
local include_param
if include_set then
include_param = include_set[group_param]
elseif exclude_set then
include_param = not exclude_set[group_param]
else
include_param = true
end
if include_param then
local merged_settings = merge_param_mod_settings(merge_param_mod_settings(
param_mods[group_param] or default_specs, group_param_settings), spec)
param_mods[group_param] = merged_settings
end
end
end
end
if spec.param then
local params = spec.param
if type(params) ~= "table" then
params = {params}
end
for _, param in ipairs(params) do
local settings = merge_param_mod_settings(param_mods[param] or default_specs, spec)
-- If this parameter is an alias of another parameter, we need to copy the specs from the other
-- parameter, since parse_inline_modifiers() doesn't know about `alias_of` and having the specs
-- duplicated won't cause problems for [[Module:parameters]]. We also need to set `item_dest` to
-- point to the `item_dest` of the aliasee (defaulting to the aliasee's value itself), so that
-- both modifiers write to the same location. Note that this works correctly in the common case of
-- <t:...> with `item_dest = "gloss"` and <gloss:...> with `alias_of = "t"`, because both will end
-- up with `item_dest = "gloss"`.
local aliasee = settings.alias_of
if aliasee then
local aliasee_settings = param_mods[aliasee]
if not aliasee_settings then
internal_error(("Undefined aliasee '%s'"):format(aliasee), spec)
end
for k, v in pairs(aliasee_settings) do
if settings[k] == nil then
settings[k] = v
end
end
if settings.item_dest == nil then
settings.item_dest = aliasee
end
end
param_mods[param] = settings
end
end
end
end
return param_mods
end
-- Return true if `k` is a "built-in" (specially recognized) key in a `param_mod` specification. All other keys
-- are forwarded to the structure passed to [[Module:parameters]].
local function param_mod_spec_key_is_builtin(k)
return k == "item_dest" or k == "overall" or k == "store"
end
--[==[
Convert the properties in `param_mods` into the appropriate structures for use by `process()` in [[Module:parameters]]
and store them in `params`. If `overall_only` is given, only store the properties in `param_mods` that correspond to
overall (non-item-specific) parameters. Currently this only happens when `separate_no_index` is specified.
]==]
function export.augment_params_with_modifiers(params, param_mods, overall_only)
if overall_only then
for param_mod, param_mod_spec in pairs(param_mods) do
if overall_only == "always" or param_mod_spec.separate_no_index then
local param_spec = {}
for k, v in pairs(param_mod_spec) do
if k ~= "separate_no_index" and k ~= "require_index" and not param_mod_spec_key_is_builtin(k) then
param_spec[k] = v
end
end
params[param_mod] = param_spec
end
end
else
local list_with_holes
-- Add parameters for each term modifier.
for param_mod, param_mod_spec in pairs(param_mods) do
local param_spec
for k, v in pairs(param_mod_spec) do
if not param_mod_spec_key_is_builtin(k) then
if param_spec == nil then
param_spec = {list = true}
end
param_spec[k] = v
end
end
if param_spec == nil then
if list_with_holes == nil then
list_with_holes = {list = true, allow_holes = true}
end
param_spec = list_with_holes
elseif param_spec.alias_of == nil then
param_spec.allow_holes = true
end
params[param_mod] = param_spec
end
end
end
--[==[
Return true if `k`, a key in an item, refers to a property of the item (is not one of the specially stored values).
Note that `lang` and `sc` are considered properties of the item, although `lang` is set when there's a language
prefix and both `lang` and `sc` may be set from default values specified in the `data` structure passed into
`parse_list_with_inline_modifiers_and_separate_params()` and `parse_term_with_inline_modifiers_and_separate_params()`.
If you don't want these treated as property keys, you need to check for them yourself.
]==]
function export.item_key_is_property(k)
return k ~= "term" and k ~= "termlang" and k ~= "termlangs" and k ~= "itemno" and k ~= "orig_index" and
k ~= "separator"
end
-- Fetch the argument in `args` corresponding to `index_or_value`, which may be a string of the form "foo.default"
-- (requesting the value of `args["foo"].default`); a string or number (requesting the value at that key); a function of
-- one argument (`args`), which returns the argument value; or the value itself. Return the resulting value and the
-- parameter in `args` that the value came from, or nil if unknown (i.e. a function or direct value was specified).
local function fetch_argument(args, index_or_value)
if not index_or_value then
return index_or_value, nil
end
local index_or_value_type = type(index_or_value)
if index_or_value_type == "string" then
if index_or_value:sub(-8) == ".default" then
local index_without_default = index_or_value:sub(1, -9)
local arg_obj = fetch_argument(args, index_without_default)
if type(arg_obj) ~= "table" then
internal_error(("Requested that the '.default' key of argument `%s` be fetched, but argument value is undefined or not a table"):
format(index_without_default), arg_obj)
end
return arg_obj.default, index_without_default
end
if index_or_value:match("^%d+$") then
index_or_value = tonumber(index_or_value)
end
return args[index_or_value], index_or_value
elseif index_or_value_type == "number" then
return args[index_or_value], index_or_value
elseif is_callable(index_or_value) then
return index_or_value(args), nil
end
return index_or_value, nil
end
function export.generate_obj_maybe_parsing_lang_prefix(data)
return require(parse_utilities_module).generate_obj_maybe_parsing_lang_prefix(data)
end
-- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and
-- parse_term_with_inline_modifiers_and_separate_params(), validating certain argument-related fields that are shared
-- among the two functions.
local function validate_argument_related_fields(data)
if not data.termarg then
internal_error("`data.termarg` must be given, indicating which argument contains the terms to be parsed", data)
end
if not data.param_mods then
internal_error("`data.param_mods` must be given, indicating the allowed inline modifiers and separate " ..
"parameters to copy", data)
end
local subitem_param_handling = data.subitem_param_handling or "only"
if subitem_param_handling ~= "only" and subitem_param_handling ~= "first" and subitem_param_handling ~= "last" then
internal_error("Unrecognized value for `data.subitem_param_handling`, should be 'first', 'last' or 'only'",
subitem_param_handling)
end
if data.raw_args then
if data.processed_args then
internal_error("Only one of `data.raw_args` and `data.processed_args` can be specified", data)
end
if not data.params then
internal_error("When `data.raw_args` is specified, so must `data.params`, so that the raw arguments " ..
"can be parsed", data)
end
if data.params[data.termarg] == nil then
internal_error("There must be a spec in `data.params` corresponding to `data.termarg`", data)
end
else
if not data.processed_args then
internal_error("Either `data.raw_args` or `data.processed_args` must be specified", data)
end
if data.params then
internal_error("When `data.processed_args` is specified, `data.params` should not be specified", data)
end
end
end
local function argval_missing(val)
return val == nil or type(val) == "table" and next(val) == nil
end
-- Subfunction of parse_list_with_inline_modifiers_and_separate_params() and
-- parse_term_with_inline_modifiers_and_separate_params(). After parsing inline modifiers, copy the separate parameters
-- to the generated object (or to the appropriate subobject if there are multiple). `data` contains the following
-- fields:
--
-- `args`: The separate-parameter argument structure.
-- `param_mods`: The structure describing the inline modifiers.
-- `itemno`: The logical item number of the term being processed, or nil if there's only a single term.
-- `termobj`: The object to store the inline modifiers into. If there are subitems, they are in the `terms` field;
-- otherwise the properties are stored directly into `termobj`.
-- `has_subitems`: True if there are subitems.
-- `subitem_separator_map`: If `has_subitems` and this is specified, controls the assignment of the `separator` field
-- in subitems. If not specified or a delimiter is not in the map, it is copied unchanged.
-- `lang`: Language object to store into all items.
-- `sc`: Script object to store into all items, or nil.
-- `subitem_param_handling`: "only", "first" or "last", indicating what to do if there are multiple subitems.
-- `allow_conflicting_inline_mods_and_separate_params`: If true, specifying a value for both an inline modifier and
-- corresponding separate parameter is allowed, and the inline modifier takes precedence. Otherwise, an error
-- occurs.
-- `postprocess_termobj`: Optional function called on all items at the end, to do any postprocessing. Called with one
-- argument, the object to postprocess.
-- `no_show_decorations`: If true, don't automatically set {show_decorations = true} on the object or subobject if there
-- are decorations (i.e. qualifiers, labels or references) specified for the object.
local function copy_separate_params_to_termobj_and_postprocess(data)
local args, param_mods, itemno, termobj = data.args, data.param_mods, data.itemno, data.termobj
local function set_lang_and_sc(termobj)
-- Set these after parsing inline modifiers, not in generate_obj(), otherwise we'll get an error in
-- parse_inline_modifiers() if we try to use <lang:...> or <sc:...> as inline modifiers.
termobj.lang = termobj.lang or data.lang
termobj.sc = termobj.sc or data.sc
end
local function set_show_decorations(termobj)
-- Need to set after parsing inline modifiers.
if not data.no_show_decorations and (termobj.q or termobj.qq or termobj.a or termobj.aa or termobj.l or termobj.ll or
termobj.refs) then
termobj.show_decorations = true
end
end
local function fetch_separate_param(args, paramkey, itemno)
local argval = args[paramkey]
-- Careful with argument values that may be `false`.
if argval and itemno then
argval = argval[itemno]
end
return argval
end
-- Copy separate parameters to a given object.
local function copy_separate_params_to_termobj(fetch_destobj)
for param_mod, param_mod_spec in pairs(param_mods) do
local dest = param_mod_spec.item_dest or param_mod
-- Don't do anything with the `sc` param, which will get overwritten below; we don't
-- want it to cause an error if there are multiple subitems.
if dest ~= "sc" then
local argval = fetch_separate_param(args, param_mod, itemno)
if not argval_missing(argval) then
local destobj = fetch_destobj(param_mod, param_mod_spec, dest)
-- Don't overwrite a value already set by an inline modifier.
if argval_missing(destobj[dest]) then
destobj[dest] = argval
elseif not data.allow_conflicting_inline_mods_and_separate_params then
error(("Can't specify a value for separate parameter %s%s= because there is " ..
"already an inline modifier <%s:...> specifying a value for the term"):format(
param_mod, itemno or "", param_mod))
end
end
end
end
end
if data.has_subitems then
-- If there are any separate indexed parameters, we need to copy them to the first, last or only
-- subitem, depending on the value of `data.subitem_param_handling` (which defaults to 'only',
-- meaning it's an error if there are multiple subitems). Do this before calling
-- postprocess_termobj() because the latter sets .lang and .sc and we want the user to be able to
-- set separate langN= and scN= parameters.
-- If there was no term, `termobj.terms` will not exist; make it exist to make the callers' lives easier.
if not termobj.terms then
termobj.terms = {}
end
-- Compute whether any of the separate indexed params exist for this index.
local any_param_at_index
for param_mod in pairs(param_mods) do
local argval = fetch_separate_param(args, param_mod, itemno)
if not argval_missing(argval) then
any_param_at_index = true
break
end
end
-- If there was no term, but there's a separate parameter, we need to create an empty subitem.
if any_param_at_index and not termobj.terms[1] then
termobj.terms[1] = {}
end
local function fetch_destobj(param_mod, param_mod_spec, dest)
if param_mod_spec.overall then
return termobj
end
if data.subitem_param_handling == "only" and termobj.terms[2] then
error(("Can't specify a value for separate parameter %s%s= because there are " ..
"multiple subitems (%s) in the term; use an inline modifier"):format(
param_mod, itemno or "", #termobj.terms))
end
local termind
-- q/a/l need to go at the beginning and qq/aa/ll/refs at the end, regardless; otherwise, respect
-- `data.subitem_param_handling`.
if dest == "q" or dest == "a" or dest == "l" then
termind = 1
elseif dest == "qq" or dest == "aa" or dest == "ll" or dest == "refs" then
termind = #termobj.terms
elseif data.subitem_param_handling == "only" or data.subitem_param_handling == "first" then
termind = 1
else
termind = #termobj.terms
end
return termobj.terms[termind]
end
copy_separate_params_to_termobj(fetch_destobj)
for i, subitem in ipairs(termobj.terms) do
set_lang_and_sc(subitem)
set_show_decorations(subitem)
if subitem.delimiter then
subitem.separator = i == 1 and "" or
data.subitem_separator_map and data.subitem_separator_map[subitem.delimiter] or
subitem.delimiter
end
if data.postprocess_termobj then
data.postprocess_termobj(subitem, data)
end
end
else
-- Copy all the parsed term-specific parameters into `termobj`.
copy_separate_params_to_termobj(function(param_mod, dest) return termobj end)
set_lang_and_sc(termobj)
set_show_decorations(termobj)
if data.postprocess_termobj then
data.postprocess_termobj(termobj, data)
end
end
end
local function postprocess_termobj(item, data)
if not (data.disallow_custom_separators or data.use_semicolon) then
if data.has_subitems and item.separator and item.separator:find(",", nil, true) then
data.use_semicolon = true
else
-- If the displayed term (from .term/etc. or .alt) has an embedded comma, use a semicolon to
-- join the terms.
local term_text = item[data.term_dest] or item.alt
if term_text and term_text:find(",", nil, true) then
data.use_semicolon = true
end
end
end
end
--[==[
Parse a list of terms, each of which may have properties specified using inline modifiers or separate parameters. This
function is intended for parsing the arguments of templates like {{tl|syn}}, {{tl|ant}} and related ''*nym'' templates;
alternative-form templates {{tl|alt}}/{{tl|alter}}; affix templates like {{tl|af}}/{{tl|affix}},
{{tl|com}}/{{tl|compound}}, etc.; affix usex templates like {{tl|afex}}/{{tl|affixusex}}; name templates like
{{tl|name translit}}; column templates like {{tl|col}}; pronunciation templates like {{tl|rhyme}}/{{tl|rhymes}} and
{{tl|hmp}}/{{tl|homophones}}; etc. In these templates there are one or more terms specified using numeric parameters, and
associated separate parameters specifying per-term properties such as {{para|t1}}, {{para|t2}}, {{para|t3}}, ... for the
gloss of the first, second, third, ... term respectively. All such properties can also be specified through inline
modifiers attached directly to each term (`<t:...>`, `<pos:...>`, etc.). Normally it is an error if both an inline
modifier and separate parameter for the same value are given, but this can be overridden (in which case inline modifiers
take precedence over separate parameters when both occur).
For an example of a typical workflow involving this function, see the comment at the top of this file.
Some notable properties of this function:
# Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this
function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent
args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]],
containing only the term list itself along with any other parameters that are '''not''' term properties (such as
a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is
''augmented'' with list parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the
latter case where raw argument processing is done by the caller, they must build the partial `params` structure;
augment it themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the
processed arguments. In both cases, the return value of this function contains three values: a list of objects, one
per term, specifying the term and all properties; the processed arguments structure, so that the non-term-property
arguments can be processed as appropriate; and an object containing miscellaneous global computed properties
(currently only `use_semicolon`; see below).
# Optionally, each term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the
possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality
is used, for example, by {{tl|col}} and variants, which allow each row to have comma-separated or tilde-separated
subitems. When this feature is invoked, the format of the per-term object changes; instead of directly being an object
describing the term and its properties, it is an object with a `terms` field containing a list of per-subitem objects
along with other top-level fields describing per-term properties. By default, if there are separate parameters
specified along with multiple subitems, an error occurs, but this is controllable; currently, you can request that the
parameters be assigned to the first or last subitem.
# By default, special ''separator'' arguments may be present, mixed in among regular term arguments. Examples of such
separator arguments are (by default; this can be overridden) a bare semicolon, specifying that the terms on either
side should be separated by a semicolon instead of a comma (indicating a higher-level grouping); a bare tilde,
replacing the comma separator with a tilde (indicating that the terms on either side are alternants); and a bare
underscore, replacing the comma separator with a space. Separator arguments are ignored when numbering the separate
parameters. You disable the separator argument handling entirely if it doesn't make sense to have this (e.g. in
{{tl|af}}/{{tl|affix}}, where the separator is always a {{cd|+}} sign).
`data` is an object containing several possible fields.
1. Fields that are required or recommended (usually related to argument processing):
* `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from
{frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw
arguments.
* `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both)
of `raw_args` and `processed_args` must be set.
* `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the
introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified
manually.
* `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters,
'''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically
"augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing
the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the
"augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by
[[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be
correctly handled.)
* `termarg` ('''required'''): The argument containing the first item with attached inline modifiers to be parsed.
Usually a numeric value such as {1} or {2}.
* `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used
internally to track pages containing template invocations with certain properties. Example properties tracked are
missing items with corresponding properties as well as missing items without corresponding properties (which are
skipped entirely). To find out the exact properties tracked and the name of the tracking pages, read the code.
* `lang` ('''recommended'''): The language object for the language of the items, or the name of the argument to fetch
the object from. It is not strictly necessary to specify this, as this function only initializes items based on inline
modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it is used
for certain purposes:
*# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language
prefix).
*# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same
language code may appear in several items, such as with {{tl|col}} and related templates).
The value of `lang` can be any of the following:
* If a string of the form "foo.default", it is assumed to be requesting the value of `args["foo"].default`.
* Otherwise, if a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the
string is in the form of a number (e.g. "3"), it is normalized to a number prior to fetching (this also happens with
a spec like "2.default").
* Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed
to the function as its only argument.
* Otherwise, it is used directly.
* `sc` ('''recommended'''): The script object for the items, or the name of the argument to fetch the object from. The
possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not strictly
necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of returned
items if not otherwise set (e.g. by the {{para|sc<var>N</var>}} parameter or `<sc:...>` inline modifier). The most
common value is {"sc.default"}.
2. Other argument-related fields:
* `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure
returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed
arguments. It should make modifications in-place.
* `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language
prefixes are stripped off. Defaults to {"term"}.
* `pre_normalize_modifiers`: As in `parse_inline_modifiers()`.
* `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for
a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence.
3. Fields related to language prefixes:
* `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such
as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to
such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in
[[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language
prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into
the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the
`termlangs` field, and the first specified object is stored in the `lang` field).
* `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can
be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above.
* `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not
trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is
given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space
follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an
error and will tend to remain unfixed.
4. Fields related to custom/special separators:
* `disallow_custom_separators`: If specified, disallow specifying custom separators (semicolon, underscore, tilde; see
the internal `default_special_separators` table, or the `special_separators` field) as an item value to override the
default separator. By default, the previous separator of each item is considered to be an empty string (for the first
item) and otherwise the value of the field `default_separator` (normally a comma + space), unless either the preceding
item is one of the values listed in `special_separators`, such as a bare semicolon (which causes the following item's
previous separator to be a semicolon + space) or an item has an embedded comma in it (which causes ''all'' items other
than the first to have their previous separator be a semicolon + space). The previous separator of each item is set on
the item's `separator` property. Bare semicolons and other separator arguments do not count when indexing items using
separate parameters.
For example, the following is correct:
** {{tl|template|lang|item 1|q1=qualifier 1|;|item 2|q2=qualifier 2}}
If `disallow_custom_separators` is specified, however, the `separator` property is not set and separator arguments are
not recognized.
* `default_separator`: Override the default separator (normally {", "}).
* `special_separators`: Table giving the special/custom separators that can be given, and how they should display. If
not specified, the default in `default_special_separators` is used. This is a table mapping separator values (such as
{"~"}) to the corresponding display string (such as {" ~ "}).
5. Fields related to multiple subitems in a given term:
* `splitchar`: A Lua pattern. If specified, each user-specified argument can consist of multiple delimiter-separated
subitems, each of which may be followed by inline modifiers. In this case, each element in the returned list of items
is no longer an object describing an item, but instead an object with a `terms` field, whose value is a list
describing the subitems (whose format is the same as the normal format of an item in the top-level list when
`splitchar` is not specified). Each subitem object will have a `delimiter` field holding the actual delimiter
occurring before the subitem, which is useful in the case where `splitchar` matches multiple possible characters. In
this case, it is possible to specify that a given modifier can only occur after the last subitem and effectively
modifies the whole collection of subitems by setting {overall = true} on the modifier. In this case, the modifier's
value will be stored in the top-level object (the object with the `terms` field specifying the subitems). Note that
splitting on delimiters will not happen in certain protected sequences (by default comma+whitespace; see below). In
addition, the algorithm to split on delimiters is sensitive to inline modifier syntax and will not be confused by
delimiters inside of inline modifiers or inside of square brackets, which do not trigger splitting (whether or not
contained within protected sequences).
* `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in
[[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified
(see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that
comma+whitespace sequences won't be split.
* `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems.
The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an
error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters
in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored
into all subitems.
* `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator`
field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an
item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar`
are recognized. Delimiters not in this map display as-is.
6. Other fields:
* `dont_skip_items`: Normally, items that are completely unspecified (have no term and no properties) are skipped and
not inserted into the returned list of items. (Such items cannot occur if {disallow_holes = true} is set on the term
specification in the `params` structure passed to `process()` in [[Module:parameters]]. It is generally recommended
to do so unless a specific meaning is associated the term value being missing.) If `dont_skip_items` is set, however,
items are never skipped, and completely unspecified items will be returned along with others. (They will not have
the term or any properties set, but will have the normal non-property fields set; see below.)
* `stop_when`: If specified, a function to determine when to prematurely stop processing items. It is passed a single
argument, an object containing the following fields:
** `term`: The raw term, prior to parsing off language prefixes and inline modifiers (since the processing of
`stop_when` happens before parsing the term).
** `any_param_at_index`: True if any separate property parameters exist for this item.
** `orig_index`: Same as `orig_index` below.
** `itemno`: Same as `itemno` below.
** `stored_itemno`: The index where this item will be stored into the returned items table. This may differ from
`itemno` due to skipped items (it will never be different if `dont_skip_items` is set).
The function should return true to stop processing items and return the ones processed so far (not including the item
currently being processed). This is used, for example, in [[Module:alternative forms]], where an unspecified item
signal the end of items and the start of labels.
* `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have
decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set,
causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem.
If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations.
Three values are returned: the list of items; the processed `args` structure; and an object of miscellaneous computed
global values (currently only `use_semicolon`, indicating that commas were found in individual arguments and so the
default separator should be a semicolon). In each returned item, there will be one field set for each specified property
(either through inline modifiers or separate parameters). If subitems are not allowed, each item directly has fields set
on it for the specified properties. If subitems ''are'' allowed, each item contains a `terms` field, which is a list of
subitem objects, each of which has fields set on it for the specified properties of that subitem. In addition, the
following fields may be set on each item or subitem:
* `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given.
* `orig_index`: The original index into the item in the items table returned by `process()` in [[Module:parameters]].
This may differ from `itemno` if there are raw semiclons and `disallow_custom_separators` is not given.
* `itemno`: The logical index of the item. The index of separate parameters corresponds to this index. This may be
different from `orig_index` in the presence of raw semicolons; see above.
* `termlang`: If there is a language prefix, the corresponding language object is stored here (only if
`parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set).
* `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are
set, the list of corresponding language objects is stored here.
* `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if
multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified;
(c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value.
* `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b)
`sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value.
* `delimiter`: If subitems are allowed, this is set on subitems and specifies the delimiter used prior to the given
subitem (e.g. {","}).
* `separator`: The separator to display before the item. Always set on subitems, and set on top-level items if
`disallow_custom_separators` is not given. Controlled by `special_separators` (for top-level items) and
`subitem_separator_map` (for subitems).
* `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified,
{show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is
called. Use {no_show_decorations = true} to prevent this.
]==]
function export.parse_list_with_inline_modifiers_and_separate_params(data)
validate_argument_related_fields(data)
local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods
if raw_args then
local params = data.params
local termarg_spec = params[termarg]
if termarg_spec == true or not termarg_spec.list then
internal_error("Term spec in `data.params` must have `list` set", termarg_spec)
end
if termarg_spec == true or not (termarg_spec.allow_holes or termarg_spec.disallow_holes) then
internal_error("Term spec in `data.params` must have either `allow_holes` or `disallow_holes` set",
termarg_spec)
end
export.augment_params_with_modifiers(params, param_mods)
args = process_params(raw_args, params)
else
args = data.processed_args
end
local process_args_before_parsing = data.process_args_before_parsing
if process_args_before_parsing then
process_args_before_parsing(args)
end
-- Find the maximum index among any of the list parameters.
local term_args = args[termarg]
-- As a special case, the term args might not have a `maxindex` field because they might have
-- been declared with `disallow_holes = true`, so fall back to the actual length of the list
-- using the table_len function, since # can be unpredictable with arbitrary tables.
local maxmaxindex = term_args.maxindex or table_len(term_args)
for _, v in pairs(args) do
if type(v) == "table" and v.maxindex and v.maxindex > maxmaxindex then
maxmaxindex = v.maxindex
end
end
local special_separators = data.special_separators or export.default_special_separators
local items, lang_cache, use_semicolon = {}, data.lang_cache or {}
local lang = fetch_argument(args, data.lang)
if lang then
lang_cache[lang:getCode()] = lang
end
local sc = fetch_argument(args, data.sc)
local term_dest = data.term_dest or "term"
-- FIXME: this is vulnerable to abusive inputs like 1000000=.
local itemno = 0
for i = 1, maxmaxindex do
local term = term_args[i]
if data.disallow_custom_separators or not special_separators[term] then
itemno = itemno + 1
-- Compute whether any of the separate indexed params exist for this index.
local any_param_at_index
for param_mod in pairs(param_mods) do
local argval = args[param_mod]
-- Careful with argument values that may be `false`.
if argval then
argval = argval[itemno]
end
if not argval_missing(argval) then
any_param_at_index = true
break
end
end
if data.stop_when and data.stop_when{
term = term,
-- FIXME, we should just pass in `any_param_at_index` directly.
any_param_at_index = term ~= nil or any_param_at_index,
orig_index = i,
itemno = itemno,
stored_itemno = #items + 1,
} then
break
end
-- If any of the params used for formatting this term is present, create a term and add it to the list.
if not data.dont_skip_items and term == nil and not any_param_at_index then
track("skipped-term", data.track_module)
else
if not term then
track("missing-term", data.track_module)
end
local termobj = {
itemno = itemno,
orig_index = i,
}
if not data.disallow_custom_separators then
termobj.separator = i == 1 and "" or special_separators[term_args[i - 1]]
end
-- Add 1 because first term index starts at 2.
local paramname = termarg + i - 1
if term then
local function generate_obj(term, parse_err)
return export.generate_obj_maybe_parsing_lang_prefix {
term = term,
termobj = data.splitchar and {} or termobj,
term_dest = term_dest,
paramname = paramname,
parse_lang_prefix = data.parse_lang_prefix,
parse_err = parse_err,
allow_bad_lang_prefix = data.allow_bad_lang_prefix,
allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes,
lang_cache = lang_cache,
}
end
parse_inline_modifiers(term, {
paramname = paramname,
param_mods = param_mods,
generate_obj = generate_obj,
splitchar = data.splitchar,
preserve_splitchar = true,
escape_fun = data.escape_fun,
unescape_fun = data.unescape_fun,
outer_container = data.splitchar and termobj or nil,
pre_normalize_modifiers = data.pre_normalize_modifiers,
})
end
-- FIXME: Make into an error, then remove after a month.
if data.no_show_qualifiers then
track("no_show_qualifiers")
end
local term_data = {
args = args,
param_mods = param_mods,
itemno = itemno,
termobj = termobj,
term_dest = term_dest,
has_subitems = not not data.splitchar,
lang = lang,
-- As a special case, if the caller defined a scN= separate param, set it on all subitems if there
-- are multiple, falling back to the overall sc= param.
sc = args.sc and args.sc[itemno] or sc,
subitem_param_handling = data.subitem_param_handling,
subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map,
allow_conflicting_inline_mods_and_separate_params =
data.allow_conflicting_inline_mods_and_separate_params,
postprocess_termobj = postprocess_termobj,
disallow_custom_separators = data.disallow_custom_separators,
use_semicolon = use_semicolon,
no_show_decorations = data.no_show_decorations or data.no_show_qualifiers,
}
copy_separate_params_to_termobj_and_postprocess(term_data)
use_semicolon = term_data.use_semicolon
insert(items, termobj)
end
end
end
if not data.disallow_custom_separators then
-- Set the default separator of all those items for which a separator wasn't explicitly given to the default
-- separator, defaulting to comma + space; but if any items have embedded commas, set the separator to
-- semicolon + space.
for _, item in ipairs(items) do
if not item.separator then
item.separator = use_semicolon and "; " or data.default_separator or ", "
end
end
end
return items, args, {use_semicolon = use_semicolon}
end
--[==[
Parse a single term that may have properties specified through inline modifiers or separate parameters. This differs
from `parse_list_with_inline_modifiers_and_separate_params()` in that the latter is for parsing a list of terms, each of
which may have properties specified through inline modifiers or separate parameters. Both functions optionally support
having multiple subitems in a single term. This function is used e.g. for form-of templates
({{tl|inflection of}}/{{tl|infl of}}, {{tl|form of}}, and specific templates such as
{{tl|alt form}}/{{tl|alternative form of}}, {{tl|abbr of}}/{{tl|abbreviation of}}, {{tl|clipping of}}, and many others);
for etymology templates ({{tl|bor}}/{{tl|borrowed}}, {{tl|der}}/{{tl|derived}}, etc. as well as `misc_variant` templates
like {{tl|ellipsis}}, {{tl|abbrev}}, {{tl|clipping}}, {{tl|reduplication}} and the like); and for other templates with
an argument structure similar to {{tl|l}} or {{tl|m}}. In these templates there is a term specified using a numeric
parameter and associated separate parameters specifying term properties such as {{para|t}} for the gloss or {{para|tr}}
for manual transliteration. All such properties can also be specified through inline modifiers attached directly to each
term (`<t:...>`, `<tr:...>`, etc.). Normally it is an error if both an inline modifier and separate parameter for the
same value are given, but this can be overridden (in which case inline modifiers take precedence over separate
parameters when both occur).
Some notable properties of this function:
# Processing of the raw frame parent args using `process()` in [[Module:parameters]] can occur either inside of this
function (the usual workflow) or outside of this function (for more complex cases). In the former case the raw parent
args are passed in along with a partially built `params` structure of the sort required by [[Module:parameters]],
containing only the term list itself along with any other parameters that are '''not''' term properties (such as
a language code in {{para|1}} and boolean flags like {{para|nocat}}, {{para|nocap}}, etc.). This structure is
''augmented'' with parameters, one for each per-term property, and [[Module:parameters]] is invoked. In the latter
case where raw argument processing is done by the caller, they must build the partial `params` structure; augment it
themselves using `augment_params_with_modifiers()`; call [[Module:parameters]] themselves; and pass in the processed
arguments. In both cases, the return value of this function contains two values, an object specifying the term and all
properties; and the processed arguments structure, so that the non-term-property arguments can be processed as
appropriate.
# Optionally, the term can consist of a number of ''subitems'' separated by delimiters (usually a comma, but the
possible delimiter or delimiters are controllable). Each subitem can have its own inline modifiers. This functionality
is used, for example, by form-of templates. When this feature is invoked, the format of the term object changes;
instead of directly being an object describing the term and its properties, it is an object with a `terms` field
containing a list of per-subitem objects along with other top-level fields describing per-term properties. By default,
if there are separate parameters specified along with multiple subitems, an error occurs, but this is controllable;
currently, you can request that the parameters be assigned to the first or last subitem.
`data` is an object containing several possible fields.
1. Fields that are required or recommended (usually related to argument processing):
* `raw_args` ('''required''' unless `processed_args` is specified): The raw arguments, normally fetched from
{frame:getParent().args}. They are parsed using `process()` in [[Module:parameters]]. Most callers pass in raw
arguments.
* `processed_args`: The object of parsed arguments returned by `process()` in [[Module:parameters]]. One (but not both)
of `raw_args` and `processed_args` must be set.
* `param_mods` ('''required'''): A structure describing the possible inline modifiers and their properties. See the
introductory comment above. Most often, this is generated using `construct_param_mods()` rather than specified
manually.
* `params` ('''required''' unless `processed_args` is specified): A structure describing the possible parameters,
'''other than''' the ones that are separate-parameter equivalents of inline modifiers. This is automatically
"augmented" with the separate-parameter equivalents of the inline modifiers described in `param_mods` prior to parsing
the raw arguments with [[Module:parameters]]. '''WARNING:''' This structure is destructively modified, both by the
"augmentation" process of adding separate-parameter equivalents of inline modifiers, and by the processing done by
[[Module:parameters]] itself. (Nonetheless, substructures can safely be shared in this structure, and will be
correctly handled.)
* `termarg` ('''required'''): The argument containing the item with attached inline modifiers to be parsed. Usually a
numeric value such as {1} or {2}.
* `track_module` ('''recommended'''): The name of the calling module, for use in adding tracking pages that are used
internally to track pages containing template invocations with certain properties.
* `lang` ('''recommended'''): The language object for the language of the item or subitems, or the name of the argument
to fetch the object from. It is not strictly necessary to specify this, as this function only initializes items based
on inline modifiers and separate arguments and doesn't actually format the resulting items. However, if specified, it
is used for certain purposes:
*# It specifies the default for the `lang` property of returned objects if not otherwise set (e.g. by a language
prefix).
*# It is used to initialize an internal cache for speeding up language-code parsing (primarily useful if the same
language code may appear in several subitems).
The value of `lang` can be any of the following:
* If a string or number, it is assumed to be requesting the value of `args` at that key. Note that if the string is in
the form of a number (e.g. "3"), it is normalized to a number prior to fetching.
* Otherwise, if a function, it is assumed to be a function to return the argument value given `args`, which is passed
to the function as its only argument.
* Otherwise, it is used directly.
* `sc` ('''recommended'''): The script object for the item or subitems, or the name of the argument to fetch the object
from. The possible values and their handling are the same as with `lang`. In general, as with `lang`, it is not
strictly necessary to specify this. However, if specified, it is used to supply the default for the `sc` property of
returned items if not otherwise set (e.g. by the {{para|sc}} parameter or `<sc:...>` inline modifier). The most common
value is {"sc"}.
* `make_separate_g_into_list`: Set this to {true} if separate gender parameters exist are are specified using
{{para|g}}, {{para|g2}}, etc. instead of using a single comma-separated {{para|g}} field.
2. Other argument-related fields:
* `adjust_params_before_arg_processing`: An optional function to further adjust the `params` structure prior to
calling `process()` in [[Module:parameters]]. This should be used when there are mismatches between the format of a
given property as an inline modifier and the corresponding property as a separate parameter (as with the {{para|g}}
parameter and {{cd|<g:...>}} modifier, but this particular case is handled by the `make_separate_g_into_list` field).
* `process_args_before_parsing`: An optional function to apply further processing to the processed `args` structure
returned by [[Module:parameters]], before parsing inline modifiers. This is passed one argument, the processed
arguments. It should make modifications in-place.
* `term_dest`: The field to store the value of the item itself into, after inline modifiers and (if allowed) language
prefixes are stripped off. Defaults to {"term"}.
* `pre_normalize_modifiers`: As in `parse_inline_modifiers()`.
* `allow_conflicting_inline_mods_and_separate_params`: If specified, don't throw an error if a value is specified for
a given property using both an inline modifier and separate param; in this case, the inline modifier takes precedence.
* `no_show_decorations`: If set, don't automatically set {show_decorations = true} on items or subitems that have
decorations (i.e. qualifiers, labels or references) attached to them. Normally, {show_decorations = true} is set,
causing `full_link()` in [[Module:links]] to appropriately display the decorations when showing the item or subitem.
If you handle this display yourself, set {no_show_decorations = true} to prevent double display of decorations.
3. Fields related to language prefixes:
* `parse_lang_prefix`: If true, allow and parse off a language code prefix attached to items followed by a colon, such
as {la:minūtia} or {grc:[[σκῶρ|σκατός]]}. Etymology-only languages are allowed. Inline modifiers can be attached to
such items. The exact syntax allowed is as specified in the `parse_term_with_lang()` function in
[[Module:parse utilities]]. If `allow_multiple_lang_prefixes` is given, a {{cd|+}}-sign-separated list of language
prefixes can be attached to an item. The resulting language object is stored into the `termlang` field, and also into
the `lang` field (or in the case of `allow_multiple_lang_prefixes`, the list of language objects is stored into the
`termlangs` field, and the first specified object is stored in the `lang` field).
* `allow_multiple_lang_prefixes`: If given in conjunction with `parse_lang_prefix`, multiple language code prefixes can
be given, separated by a {{cd|+}} sign. See `parse_lang_prefix` above.
* `allow_bad_lang_prefix`: If given in conjunction with `parse_lang_prefix`, unrecognized language prefixes do not
trigger an error, but are simply ignored (and not stripped off the item). Note that, regardless of whether this is
given, prefixes before a colon do not trigger an error if they do not have the form of a language prefix or if a space
follows the colon. It is not recommended that this be given because typos in language prefixes will not trigger an
error and will tend to remain unfixed.
4. Fields related to multiple subitems in the term:
* `splitchar`: A Lua pattern. If specified, the user-specified argument can consist of multiple delimiter-separated
subitems, each of which may be followed by inline modifiers. In this case, the first returned value is no longer an
object describing the item, but instead an object with a `terms` field, whose value is a list describing the subitems
(whose format is the same as the normal format of the item when `splitchar` is not specified). Each subitem object
will have a `delimiter` field holding the actual delimiter occurring before the subitem, which is useful in the case
where `splitchar` matches multiple possible characters. In this case, it is possible to specify that a given modifier
can only occur after the last subitem and effectively modifies the whole collection of subitems by setting
`overall = true` on the modifier. In this case, the modifier's value will be stored in the top-level object (the
object with the `terms` field specifying the subitems). Note that splitting on delimiters will not happen in certain
protected sequences (by default comma+whitespace; see below). In addition, the algorithm to split on delimiters is
sensitive to inline modifier syntax and will not be confused by delimiters inside of inline modifiers or inside of
square brackets, which do not trigger splitting (whether or not contained within protected sequences).
* `escape_fun` and `unescape_fun`: As in `split_escaping()` and `split_alternating_runs_escaping()` in
[[Module:parse utilities]]. They control the protected sequences that won't be split when `splitchar` is specified
(see previous item). By default, `escape_comma_whitespace` and `unescape_comma_whitespace` are used, so that
comma+whitespace sequences won't be split.
* `subitem_param_handling`: How to handle separate parameters that are specified in the presence of multiple subitems.
The possible values are {"only"} (only allow separate parameters if there aren't any subitems, otherwise throw an
error), {"first"} (store the separate parameters in the first subitem) and {"last"} (store the separate parameters
in the last subitem). The default is {"only"}. As a special case, an {{para|scN}} separate parameter will be stored
into all subitems.
* `subitem_separator_map`: Table mapping user-specified delimiters to displayed separators, stored in the `separator`
field of the subitem. If not specified, it defaults to `default_subitem_separator_map`. Note that the presence of an
item in this table does not mean that it can be used as a delimiter; only the delimiters specified using `splitchar`
are recognized. Delimiters not in this map display as-is.
Two values are returned, an object describing the item (or subitems) and the processed `args` structure. In the returned
item, there will be one field set for each specified property (either through inline modifiers or separate parameters).
If subitems are not allowed, the item directly has fields set on it for the specified properties. If subitems ''are''
allowed, the item contains a `terms` field, which is a list of subitem objects, each of which has fields set on it for
the specified properties of that subitem. In addition, the following fields may be set on the item or each subitem:
* `term`: The term portion of the item (minus inline modifiers and language prefixes). {nil} if no term was given.
* `termlang`: If there is a language prefix, the corresponding language object is stored here (only if
`parse_lang_prefix` is set and `allow_multiple_lang_prefixes` is not set).
* `termlangs`: If there is are language prefixes and both `parse_lang_prefix` and `allow_multiple_lang_prefixes` are
set, the list of corresponding language objects is stored here.
* `lang`: The language object of the item. This is set when either (a) there is a language prefix parsed off (if
multiple prefixes are allowed, this corresponds to the first one); (b) the `lang` property is allowed and specified;
(c) neither (a) nor (b) apply and the `lang` field of the overall `data` object is set, providing a default value.
* `sc`: The script object of the item. This is set when either (a) the `sc` property is allowed and specified; (b)
`sc` isn't otherwise set and the `sc` field of the overall `data` object is set, providing a default value.
* `delimiter`: If subitems are allowed, this specifies the delimiter used prior to the given subitem (e.g. {","}).
* `separator`: If subitems are allowed, this specifies the displayed form of the delimiter to be shown before a given
subitem. The mapping from user-specified delimiters to displayed separators is handled by `subitem_separator_map`;
see above. The first subitem always has a blank string in the `separator` field.
* `show_decorations`: If the item or subitem has any decorations (i.e. qualifiers, labels or references) specified,
{show_decorations = true} is normally set on the item, so that these decorations are displayed when `full_link()` is
called. Use {no_show_decorations = true} to prevent this.
]==]
function export.parse_term_with_inline_modifiers_and_separate_params(data)
validate_argument_related_fields(data)
local raw_args, termarg, param_mods, args = data.raw_args, data.termarg, data.param_mods
if raw_args then
local params = data.params
local termarg_spec = params[termarg]
if type(termarg_spec) == "table" and termarg_spec.list then
internal_error("Term spec in `data.params` must not have `list` set", termarg_spec)
end
export.augment_params_with_modifiers(params, param_mods, "always")
if data.make_separate_g_into_list then
-- HACK: g= is a list for compatibility, but sublist as an inline parameter.
params.g = {list = true, item_dest = "genders", type = "genders", flatten = true}
end
local adjust_params_before_arg_processing = data.adjust_params_before_arg_processing
if adjust_params_before_arg_processing then
adjust_params_before_arg_processing(params)
end
args = process_params(raw_args, params)
else
args = data.processed_args
end
local process_args_before_parsing = data.process_args_before_parsing
if process_args_before_parsing then
process_args_before_parsing(args)
end
local term, lang_cache = args[termarg], data.lang_cache
local lang = fetch_argument(args, data.lang)
if lang and lang_cache then
lang_cache[lang:getCode()] = lang
end
local sc = fetch_argument(args, data.sc)
local term_dest = data.term_dest or "term"
if not term then
track("missing-term", data.track_module)
end
local termobj, splitchar = {}, data.splitchar
if term then
local function generate_obj(term, parse_err)
return export.generate_obj_maybe_parsing_lang_prefix {
term = term,
termobj = splitchar and {} or termobj,
term_dest = term_dest,
paramname = termarg,
parse_lang_prefix = data.parse_lang_prefix,
parse_err = parse_err,
allow_bad_lang_prefix = data.allow_bad_lang_prefix,
allow_multiple_lang_prefixes = data.allow_multiple_lang_prefixes,
lang_cache = lang_cache,
}
end
parse_inline_modifiers(term, {
paramname = termarg,
param_mods = param_mods,
generate_obj = generate_obj,
splitchar = splitchar,
preserve_splitchar = true,
escape_fun = data.escape_fun,
unescape_fun = data.unescape_fun,
outer_container = splitchar and termobj or nil,
pre_normalize_modifiers = data.pre_normalize_modifiers,
})
end
-- FIXME: Make into an error, then remove after a month.
if data.no_show_qualifiers then
track("no_show_qualifiers")
end
copy_separate_params_to_termobj_and_postprocess {
args = args,
param_mods = param_mods,
termobj = termobj,
has_subitems = not not splitchar,
lang = lang,
sc = sc,
subitem_param_handling = data.subitem_param_handling,
subitem_separator_map = data.subitem_separator_map or export.default_subitem_separator_map,
allow_conflicting_inline_mods_and_separate_params = data.allow_conflicting_inline_mods_and_separate_params,
no_show_decorations = data.no_show_decorations or data.no_show_qualifiers,
}
if splitchar and termobj.terms[2] then
track("parse-term-multiple-subitems", data.track_module)
track("parse-term-multiple-subitems")
end
return termobj, args
end
return export
1kp3k2x2ejn7z1nhgqsp5b8fxr2qs7w
Modul:etymon
828
57903
375383
373456
2026-09-22T07:06:30Z
Hakimi97
2668
Mengemas kini mengikut padanan Wikikamus bahasa Inggeris (semakan [[en:Special:Diff/92736184|92736184]])
375383
Scribunto
text/plain
--[=[
This module implements the {{etymon}} template for structured etymology data on Wiktionary.
It enables the creation of etymology trees and text by parsing etymon chains,
scraping linked pages for their own {{etymon}} data, and recursively building a tree
of derivational relationships.
Authors:
- Original implementation: [[User:Ioaxxere]]
- Full refactor (September 2025): [[User:Fenakhay]] ([[Special:Diff/86717746]])
Modules:
- [[Module:etymon]]: main module handling parsing, validation, tree building, and page scraping
- [[Module:etymon/data]]: keyword definitions, configuration, and status constants
- [[Module:etymon/tree]]: etymology tree rendering
- [[Module:etymon/text]]: etymology text generation
- [[Module:etymon/categories]]: category generation logic
- [[Module:etymon/tracking]]: tracking
]=]
local export = {}
local __state = {
cached_etymon_args = {},
cached_etymon_pages = {},
cached_descendants_checks = {},
senseid_parent_etymon = {},
available_etymon_ids = {},
single_etymons = {},
entry_title = nil,
entry_lang_code = nil,
current_page_has_inline_etymology = false,
current_page_has_redundant_etymology = false,
used_idless_etymon = false,
toplevel_has_inline_etymology = false,
toplevel_redundant_etymology = false,
toplevel_idless_etymon = false,
has_mismatched_id = false,
linked_page_multiple_etymons_idless = false,
linked_page_partial_etymology_sections = false,
partial_etymology_targets = {},
skip_partial_etymology_category = false,
max_depth_reached = 0,
total_nodes = 0,
language_count = {},
toplevel_keyword_stats = {},
id_stats = nil,
warnings = {},
}
local function reset_invocation_state()
__state.current_page_has_inline_etymology = false
__state.current_page_has_redundant_etymology = false
__state.used_idless_etymon = false
__state.toplevel_has_inline_etymology = false
__state.toplevel_redundant_etymology = false
__state.toplevel_idless_etymon = false
__state.has_mismatched_id = false
__state.linked_page_multiple_etymons_idless = false
__state.linked_page_partial_etymology_sections = false
__state.max_depth_reached = 0
__state.total_nodes = 0
__state.language_count = {}
__state.toplevel_keyword_stats = {}
__state.warnings = {}
end
local M = require("Module:module loader").init({
require = {
data = "Module:etymon/data",
tree = "Module:etymon/tree",
text = "Module:etymon/text",
categories = "Module:etymon/categories",
tracking = "Module:etymon/tracking",
descendants = "Module:etymon/descendants",
anchors = "Module:anchors",
etydate = "Module:etydate",
etymology = "Module:etymology",
families = "Module:families",
languages = "Module:languages",
languages_errorgetby = "Module:languages/errorGetBy",
links = "Module:links",
pages = "Module:pages",
parameters = "Module:parameters",
string_utilities = "Module:string utilities",
template_parser = "Module:template parser",
utilities = "Module:utilities",
debug = "Module:debug",
en_utilities = "Module:en-utilities",
parameter_utilities = "Module:parameter utilities",
parse_utilities = "Module:parse utilities",
template_styles = "Module:TemplateStyles",
script_utilities = "Module:script utilities",
JSON = "Module:JSON",
yesno = "Module:yesno",
},
loadData = {
headword_data = "Module:headword/data",
parameters_data = "Module:parameters/data",
text_allowed = "Module:etymon/data/text_allowed",
},
})
local Util = {}
function Util.format_error(message, preview_only)
if preview_only and not M.pages.is_preview() then
return nil
end
return '<span class="error">' .. message .. '</span>'
end
function Util.add_warning(message, preview_only)
local formatted = Util.format_error(message, preview_only)
if formatted then
table.insert(__state.warnings, formatted)
end
end
function Util.is_text_param_allowed_for_lang(lang)
if not lang or type(lang) ~= "table" then
return false
end
local types = lang.getTypes and lang:getTypes()
if types and types.family then
local code = lang.getCode and lang:getCode()
return code and M.text_allowed.families[code] == true
end
local full_code = lang.getFullCode and lang:getFullCode()
if full_code and M.text_allowed.langs[full_code] then
return true
end
if lang.inFamily then
for family_code in pairs(M.text_allowed.families) do
if lang:inFamily(family_code) then
return true
end
end
end
return false
end
function Util.get_lang(code, no_error)
if no_error then
return M.languages.getByCode(code, nil, true)
end
return M.languages.getByCode(code, nil, true) or M.languages_errorgetby.code(code, true, true)
end
-- Match a term language against a text=:lang stop target (supports etymology-only codes).
function Util.lang_matches_stop_code(term_lang, stop_code)
if not term_lang or not stop_code or stop_code == "" then
return false
end
local stop_lang = Util.get_lang(stop_code, true)
if not stop_lang then
return false
end
if term_lang:getCode() == stop_lang:getCode() then
return true
end
if stop_lang:getFullCode() == stop_lang:getCode() then
return term_lang:getFullCode() == stop_lang:getCode()
end
return false
end
function Util.get_family(code)
return M.families.getByCode(code)
end
function Util.get_lang_exception(lang)
-- Families have no language-specific exceptions
if lang.getTypes and lang:getTypes().family then
return nil
end
local code = lang:getCode()
local lang_exceptions = M.data.config.lang_exceptions
if lang_exceptions[code] then
return lang_exceptions[code]
end
for norm_code, exc in pairs(lang_exceptions) do
if exc.normalize_to and code == exc.normalize_to then
return exc
end
if exc.normalize_from_families then
local should_normalize = false
for _, family in ipairs(exc.normalize_from_families) do
if lang:inFamily(family) then
should_normalize = true
break
end
end
if should_normalize and exc.normalize_exclude_families then
for _, family in ipairs(exc.normalize_exclude_families) do
if lang:inFamily(family) then
should_normalize = false
break
end
end
end
if should_normalize then
local ret = {}
for k, v in pairs(exc) do
ret[k] = v
end
ret.suppress_tr = nil
return ret
end
end
end
return nil
end
function Util.get_norm_lang(lang)
local exc = Util.get_lang_exception(lang)
if exc and exc.normalize_to then
return M.languages.getByCode(exc.normalize_to)
end
return lang
end
function Util.resolve_context_lang(lang, node_args)
if type(node_args) ~= "table" then return lang end
if node_args.status == M.data.STATUS.INLINE then return lang end
if not (lang.hasType and lang:hasType("etymology-only")) then return lang end
local full = lang.getFull and lang:getFull()
if not full or full:getCode() == lang:getCode() then return lang end
if full.hasAncestor and full:hasAncestor(lang) then return lang end
return full
end
-- Add default values for boolean modifiers (e.g., <unc> becomes <unc:1>)
-- This is needed because Module:parse utilities expects boolean modifiers to have explicit values
function Util.add_boolean_defaults(str, param_mods)
local result = str
for name, spec in pairs(param_mods) do
if spec.type == "boolean" then
-- Replace <name> with <name:1> (but not <name:...> which already has a value)
result = result:gsub("<" .. name .. ">", "<" .. name .. ":1>")
end
end
return result
end
local REQUEST_TEMPLATE_PARAM_MODS = {
rfe = M.parameter_utilities.construct_param_mods {
{ param = { "sort", "y", "m", "fragment", "section" } },
{ param = { "nocat", "box", "noes" }, type = "boolean" },
},
etystub = M.parameter_utilities.construct_param_mods {
{ param = "sort" },
{ param = { "nocat", "nocap", "nodot" }, type = "boolean" },
},
}
function Util.expand_request_template(frame, template_name, param_value, lang_code)
local param_mods = REQUEST_TEMPLATE_PARAM_MODS[template_name]
local with_defaults = Util.add_boolean_defaults(param_value, param_mods)
local parsed = M.parse_utilities.parse_inline_modifiers(with_defaults, {
param_mods = param_mods,
generate_obj = function(text)
if M.yesno(text, false) then
return { is_boolean = true }
end
return { text = text }
end,
})
local template_args = { [1] = lang_code }
for name in pairs(param_mods) do
template_args[name] = parsed[name]
end
if not parsed.is_boolean then
template_args[2] = parsed.text
end
return " " .. frame:expandTemplate({
title = template_name,
args = template_args,
})
end
-- Centralized term formatting: handles suppress_term (-), unknown_term (empty/+), and regular terms
function Util.format_term(term, is_toplevel, opts)
opts = opts or {}
-- suppress_term (-) returns nil
if term.suppress_term then
return nil
end
local lang = term.lang
local exc = Util.get_lang_exception(lang)
if is_toplevel then
local display_text = term.alt or term.title or ""
local sc = term.sc or lang:findBestScript(display_text)
local bold_text = tostring(mw.html.create("strong")
:addClass("selflink")
:wikitext(display_text))
return M.script_utilities.tag_text(bold_text, lang, sc, "term")
end
local link_params = { lang = lang }
link_params.term = not term.unknown_term and term.title or nil
link_params.alt = term.alt
link_params.id = (not term.unknown_term and term.id and term.id ~= "") and term.id or nil
if not (exc and exc.suppress_tr) then
link_params.tr = term.tr
link_params.ts = term.ts
else
link_params.suppress_tr = true
end
link_params.lit = (opts.lit ~= "suppress") and term.lit or nil
if opts.gloss ~= "suppress" then
link_params.gloss = term.gloss
end
link_params.genders = term.genders
if opts.pos ~= "suppress" then
link_params.pos = term.pos
link_params.ng = term.ng
link_params.infl = term.infl
end
if exc and exc.suppress_tr then
link_params.lit = nil
end
if opts.tree_ql ~= "suppress" then
if term.q then
link_params.q = term.q
end
if term.qq then
link_params.qq = term.qq
end
if term.l then
link_params.l = term.l
end
if term.ll then
link_params.ll = term.ll
end
link_params.show_decorations = term.q or term.qq or term.l or term.ll
end
return M.links.full_link(link_params, "term")
end
local __is_content_page_cached
function Util.is_content_page()
if __is_content_page_cached == nil then
__is_content_page_cached = M.pages.is_content_page(mw.title.getCurrentTitle())
end
return __is_content_page_cached
end
local __page_data_cached
function Util.get_page_data()
if not __page_data_cached then
__page_data_cached = M.headword_data.page
end
return __page_data_cached
end
-- Extract base keyword from param (without modifiers)
local function get_keyword_base(param)
if type(param) ~= "string" then return nil end
local base = param:match("^:?([^<]+)") or param:gsub("^:", "")
return base
end
local function is_keyword(param, allow_colon_less)
if type(param) ~= "string" then return false end
local keywords = M.data.keywords
if param:sub(1, 1) == ":" then
local base = get_keyword_base(param)
return keywords[base] ~= nil
end
if allow_colon_less then
local base = get_keyword_base(param)
return keywords[base] ~= nil
end
return false
end
local function get_keyword(param, allow_colon_less)
if type(param) ~= "string" then return nil end
local keywords = M.data.keywords
if param:sub(1, 1) == ":" then
return get_keyword_base(param)
end
if allow_colon_less then
local base = get_keyword_base(param)
if keywords[base] then
return base
end
end
return nil
end
local function normalize_keyword(keyword)
if keyword:sub(1, 1) == ":" then
return keyword
end
return ":" .. keyword
end
-- Resolve keyword (possibly an alias) to its canonical form. Used only at input boundaries
local function get_canonical_keyword(keyword)
if not keyword then return keyword end
return M.data.keyword_canonical[keyword] or keyword
end
local function is_affix_group_keyword(keyword)
local config = keyword and M.data.keywords[keyword]
return config and config.affix_categories or false
end
local function reject_removed_surf_keyword(param)
local base = get_keyword_base(param)
if base == "surf" then
error("The `:surf` keyword has been removed. Use `<surf>` on a formation keyword instead (e.g. `:af<surf>`, `:bor<surf>`).")
end
end
local function copy_keyword_info(source)
local copy = {}
for k, v in pairs(source) do
copy[k] = v
end
return copy
end
local function lowercase_glossary_display(text)
return text:gsub("(%[%[Appendix:Glossary#[^|]+|)([^%]])([^%]]*)%]%]", function(prefix, first, rest)
return prefix .. mw.ustring.lower(first) .. rest .. "]]"
end)
end
local function surf_should_keep_formation_phrase(base)
if not base.phrase then
return false
end
if base.glossary then
return true
end
return not (base.phrase == "from" and (base.text == "From" or base.text == "from"))
end
-- Runtime overrides when <surf> is present on a keyword.
local function get_effective_keyword_info(keyword, modifiers)
local base = M.data.keywords[keyword]
if not base or not modifiers or not modifiers.surf then
return base
end
local effective = copy_keyword_info(base)
local surf_text = "By [[Appendix:Glossary#surface_analysis|surface analysis]],"
local surf_phrase = "by surface analysis,"
effective.new_sentence = true
effective.invisible = "tree"
if surf_should_keep_formation_phrase(base) then
effective.phrase = surf_phrase .. " " .. base.phrase
if base.text then
effective.text = surf_text .. " " .. lowercase_glossary_display(base.text)
else
effective.text = surf_text .. " " .. base.phrase
end
else
effective.text = surf_text
effective.phrase = surf_phrase
end
return effective
end
-- Build text/phrase for nominalization with <g:code> (uses data module for codes only).
local function get_nominalization_label_for_g(code)
if not code or code == "" then return nil end
local codes = M.data.nominalization_g_codes
local adj = codes[code]
if not adj and #code == 2 then
local gender_adj = codes[code:sub(1, 1)]
local number_adj = codes[code:sub(2, 2)]
if gender_adj and number_adj then
adj = gender_adj .. " " .. number_adj
end
end
if not adj then return nil end
local text = adj:gsub("^%l", function(c) return string.upper(c) end) .. " [[Appendix:Glossary#nominalization|nominalization]] of"
local phrase = M.en_utilities.add_indefinite_article(adj .. " [[Appendix:Glossary#nominalization|nominalization]] of", false)
return { text = text, phrase = phrase }
end
local EtymonParser = {}
-- Keyword modifier definitions
EtymonParser.keyword_param_mods = M.parameter_utilities.construct_param_mods {
{ group = "ref" },
{ param = "conj" }, -- conjunction for alternatives: "and", "or", "and/or", etc.
{ param = { "unc", "surf" }, type = "boolean" },
{ param = "text", restrict = { keywords = { "from", "derived" } } },
{ param = "lit", restrict = { affix_group = true } },
{ param = "g", restrict = { keywords = { "nominalization" } } },
{ param = "senseid", restrict = { keywords = { "semantic loan" } } },
}
-- Term modifier definitions
EtymonParser.etymon_param_mods = M.parameter_utilities.construct_param_mods {
{group = {"link", "q", "l", "ref", "infl"}, exclude = {"sc"}},
{param = {"ety", "postype"}},
{param = "unc", type = "boolean"},
{param = "aftype", restrict = {affix_group = true}},
{param = {"bor", "slbor", "lbor"}, type = "boolean", restrict = {affix_group = true}},
}
local function get_clean_param_mods(param_mods)
local clean = {}
for mod_name, mod_def in pairs(param_mods) do
clean[mod_name] = {}
for key, value in pairs(mod_def) do
if key ~= "restrict" then
clean[mod_name][key] = value
end
end
end
return clean
end
function EtymonParser.check_modifier_restrictions(modifiers, current_keyword, param_mods)
for mod_name, mod_value in pairs(modifiers) do
-- Only check restrictions if the modifier has a non-false/nil value
if mod_value then
local mod_def = param_mods[mod_name]
if mod_def and mod_def.restrict then
if mod_def.restrict.affix_group then
if not is_affix_group_keyword(current_keyword) then
local mod_display = mod_value == true and "<" .. mod_name .. ">" or "<" .. mod_name .. ":" .. tostring(mod_value) .. ">"
error("The modifier `" .. mod_display .. "` is only allowed for affix-group keywords (e.g. `:af`, `:blend`, `:univ`).")
end
elseif mod_def.restrict.keywords then
local allowed_keywords = mod_def.restrict.keywords
local is_allowed = false
for _, allowed_keyword in ipairs(allowed_keywords) do
if current_keyword == allowed_keyword then
is_allowed = true
break
end
end
if not is_allowed then
local keyword_list = {}
for _, kw in ipairs(allowed_keywords) do
table.insert(keyword_list, ":" .. kw)
end
local keyword_str = table.concat(keyword_list, #keyword_list == 2 and " or " or ", ")
if #keyword_list > 2 then
-- Replace last comma with "or"
keyword_str = keyword_str:gsub(", ([^,]+)$", " or %1")
end
local mod_display = mod_value == true and "<" .. mod_name .. ">" or "<" .. mod_name .. ":" .. tostring(mod_value) .. ">"
error("The modifier `" .. mod_display .. "` is only allowed for the keyword" .. (#keyword_list > 1 and "s " or " ") .. keyword_str .. ".")
end
end
end
end
end
end
local TERM_RULE_DISALLOW = {
suppress = { field = "suppress_term", label = "suppressed" },
unknown = { field = "unknown_term", label = "unknown" },
family = { field = "is_family", label = "family" },
}
function EtymonParser.check_etymon_limits(count, limits, label, opts)
if not limits then
return
end
opts = opts or {}
local min_etymons = limits.min_etymons
if min_etymons == nil and not opts.skip_default_min then
min_etymons = 1
end
if min_etymons and count < min_etymons then
if min_etymons > 1 then
error("Detected " .. label .. " group with fewer than " .. min_etymons .. " etymons.")
else
error("Detected " .. label .. " with no etymons.")
end
end
if limits.max_etymons and count > limits.max_etymons then
local unit = (limits.max_etymons == 1) and "etymon" or "etymons"
error("Detected " .. label .. " with more than " .. limits.max_etymons .. " " .. unit .. ".")
end
end
function EtymonParser.check_term_rules(etymon_data, entry_lang, rules, label)
label = label or "term"
if rules and rules.disallow then
local disallowed = {}
for _, typ in ipairs(rules.disallow) do
local spec = TERM_RULE_DISALLOW[typ]
if spec and etymon_data[spec.field] then
table.insert(disallowed, spec.label)
end
end
if #disallowed > 0 then
error(label .. " does not support " ..
mw.text.listToText(disallowed, "or") .. " etymons.")
end
end
if etymon_data.is_family then
if rules and rules.family == "disallowed" then
error(label .. " does not support family codes" .. (rules.family_suffix or "."))
elseif not etymon_data.suppress_term then
error("Family codes require suppressed term (use family:-).")
end
end
if rules then
if rules.require_term and (not etymon_data.term or etymon_data.term == "") then
error(label .. " requires a term for each listed form.")
end
if rules.entry_lang then
if Util.get_norm_lang(etymon_data.lang):getFullCode() ~=
Util.get_norm_lang(entry_lang):getFullCode() then
error(label .. " terms must be in the entry language (" ..
entry_lang:getFullCode() .. "), got '" .. etymon_data.lang:getFullCode() .. "'.")
end
end
if rules.ancestor_check then
M.etymology.check_ancestor(entry_lang, etymon_data.lang)
end
elseif etymon_data.is_family and not etymon_data.suppress_term then
error("Family codes require suppressed term (use family:-).")
end
end
function EtymonParser.check_keyword_term(etymon_data, entry_lang, keyword)
local config = M.data.keywords[keyword]
EtymonParser.check_term_rules(etymon_data, entry_lang, config and config.term_rules, "`:" .. keyword .. "`")
end
function EtymonParser.check_supplement_term(etymon_data, entry_lang, supplement_type)
local config = M.data.supplements[supplement_type]
EtymonParser.check_term_rules(etymon_data, entry_lang, config and config.term_rules, "|" .. supplement_type .. "=")
end
-- Parse keyword with modifiers (e.g., ":bor<unc>" or ":bor<ref:{{R:example}}>")
function EtymonParser.parse_keyword_modifiers(param)
if type(param) ~= "string" then return nil, {} end
local base_keyword = get_keyword_base(param)
if not base_keyword then return nil, {} end
local canonical_keyword = get_canonical_keyword(base_keyword)
-- Check if there are any modifiers
if not param:find("<", 1, true) then
return canonical_keyword, {}
end
-- Parse modifiers using the same mechanism as etymon parsing
local rest_with_defaults = Util.add_boolean_defaults(param, EtymonParser.keyword_param_mods)
local function generate_obj(ignored)
return {}
end
local parsed = M.parse_utilities.parse_inline_modifiers(rest_with_defaults:gsub("^:?[^<]+", ""),
{ param_mods = get_clean_param_mods(EtymonParser.keyword_param_mods), generate_obj = generate_obj })
local modifiers = {
unc = parsed.unc or false,
refs = parsed.refs,
text = parsed.text,
lit = parsed.lit,
conj = parsed.conj,
g = parsed.g,
surf = parsed.surf or false,
senseid = parsed.senseid,
}
-- Validate modifiers against restrictions
EtymonParser.check_modifier_restrictions(modifiers, canonical_keyword, EtymonParser.keyword_param_mods)
return canonical_keyword, modifiers
end
local function normalize_keyword_param(keyword_with_mods)
local trimmed = M.string_utilities.trim(keyword_with_mods)
reject_removed_surf_keyword(trimmed:match("^:") and trimmed or (":" .. trimmed))
local base = get_keyword_base(trimmed)
if not base or not M.data.keywords[base] then
error("Invalid keyword '" .. trimmed .. "' in inline etymology")
end
local canonical_base = get_canonical_keyword(base)
local without_colon = trimmed:gsub("^:", "")
local mods_part = without_colon:sub(#base + 1)
local kw_param = normalize_keyword(canonical_base .. mods_part)
EtymonParser.parse_keyword_modifiers(kw_param)
return kw_param
end
local function get_keyword_mod_names()
local names = {}
for mod_name in pairs(EtymonParser.keyword_param_mods) do
names[mod_name] = true
end
return names
end
local function parse_inline_ety_run(ety_string)
local body = ety_string or ""
if body == "" then
error("Empty inline etymology")
end
local keyword_mod_names = get_keyword_mod_names()
local pos = 1
local len = #body
local function parse_err(msg)
error(msg .. " in inline etymology: '" .. body .. "'")
end
local function peek_double()
return body:sub(pos, pos + 1) == "<<"
end
local function mod_name_from_unwrapped(unwrapped)
return unwrapped:match("^<([^:>]+)")
end
local function is_keyword_mod(unwrapped)
local name = mod_name_from_unwrapped(unwrapped)
return name and keyword_mod_names[name] or false
end
local function read_double_bracket()
if not peek_double() then
return nil
end
local start = pos
pos = pos + 2
while pos <= len - 1 do
if body:sub(pos, pos + 1) == ">>" then
local token = body:sub(start, pos + 1)
pos = pos + 2
return token, token:sub(2, -2)
end
pos = pos + 1
end
parse_err("Unmatched <<")
end
local function read_angle_cell()
if body:sub(pos, pos) ~= "<" or peek_double() then
return nil
end
local open = pos
pos = pos + 1
local depth = 1
local i = pos
while i <= len do
local ch = body:sub(i, i)
if ch == "<" then
depth = depth + 1
elseif ch == ">" then
depth = depth - 1
if depth == 0 then
local inner = body:sub(open + 1, i - 1)
pos = i + 1
return inner
end
end
i = i + 1
end
parse_err("Unmatched <")
end
local function read_bare_run()
local start = pos
while pos <= len and body:sub(pos, pos) ~= "<" do
pos = pos + 1
end
return body:sub(start, pos - 1)
end
local function absorb_double_keyword_mods(keyword_str)
while peek_double() do
local saved = pos
local _, unwrapped = read_double_bracket()
if is_keyword_mod(unwrapped) then
keyword_str = keyword_str .. unwrapped
else
pos = saved
break
end
end
return keyword_str
end
local kw_start = pos
while pos <= len and body:sub(pos, pos) ~= "<" do
pos = pos + 1
end
local keyword = body:sub(kw_start, pos - 1)
if keyword:match("^%s*$") then
parse_err("Missing keyword")
end
keyword = absorb_double_keyword_mods(keyword)
local cells = {}
while pos <= len do
if peek_double() then
local _, unwrapped = read_double_bracket()
if is_keyword_mod(unwrapped) then
parse_err("Unexpected keyword modifier " .. unwrapped .. " outside of a keyword")
end
table.insert(cells, "+" .. unwrapped)
elseif body:sub(pos, pos) == "<" then
local inner = read_angle_cell()
if inner ~= "" then
table.insert(cells, inner)
end
else
local bare = read_bare_run()
if bare ~= "" then
if bare:sub(1, 1) ~= ":" then
parse_err("Unexpected bare text '" .. bare .. "' (use :keyword for nested keywords in inline etymology)")
end
if not is_keyword(bare, true) then
parse_err("Invalid keyword '" .. bare .. "' in inline etymology")
end
table.insert(cells, absorb_double_keyword_mods(bare))
end
end
end
return {
keyword = keyword,
cells = cells,
}
end
function EtymonParser.inline_ety_to_pipe(ety_string)
local run = parse_inline_ety_run(ety_string)
if not run.keyword or run.keyword:match("^%s*$") then
return "|"
end
local pipe_parts = { normalize_keyword_param(M.string_utilities.trim(run.keyword)) }
for _, segment in ipairs(run.cells) do
if is_keyword(segment, true) then
table.insert(pipe_parts, normalize_keyword_param(segment))
else
table.insert(pipe_parts, segment)
end
end
return "|" .. table.concat(pipe_parts, "|") .. "|"
end
function EtymonParser.pipe_to_inline_ety(pipe_string)
local cells = {}
for cell in pipe_string:gmatch("([^|]+)") do
if cell ~= "" then
table.insert(cells, cell)
end
end
if #cells == 0 then
return ""
end
local inline_parts = {}
for index, cell in ipairs(cells) do
local base = get_keyword_base(cell)
if base and M.data.keywords[base] then
local without_colon = cell:gsub("^:", "")
local kw_base, mods = without_colon:match("^([^<]+)(.*)$")
local inline_kw = (kw_base or without_colon) .. (mods or ""):gsub("<([^>]+)>", "<<%1>>")
if index > 1 then
inline_kw = ":" .. inline_kw
end
table.insert(inline_parts, inline_kw)
elseif cell:sub(1, 1) == "+" then
local mod = cell:sub(2)
if mod:match("^<.->$") then
mod = mod:sub(2, -2)
end
table.insert(inline_parts, "<<" .. mod .. ">>")
else
table.insert(inline_parts, "<" .. cell .. ">")
end
end
return table.concat(inline_parts, "")
end
function EtymonParser.parse_inline_ety(ety_string, context_lang)
local run = parse_inline_ety_run(ety_string)
local keyword = M.string_utilities.trim(run.keyword)
reject_removed_surf_keyword(":" .. keyword)
if not is_keyword(keyword, true) then
error("Invalid keyword '" .. keyword .. "' in inline etymology <ety:" .. keyword .. "...>")
end
local args = { context_lang:getCode(), normalize_keyword_param(keyword) }
for _, segment in ipairs(run.cells) do
if is_keyword(segment, true) then
table.insert(args, normalize_keyword_param(segment))
else
table.insert(args, segment)
end
end
return args
end
function EtymonParser.parse_etymon(param, context_lang)
if is_keyword(param) then
return nil
end
if type(param) ~= "string" then
return nil
end
local lang, rest
local is_family = false
local before_bracket = param:match("^([^<]*)") or param
local lang_code, rest_match = before_bracket:match("^([a-zA-Z][a-zA-Z0-9._-]*):(.*)$")
if lang_code then
local potential_lang = Util.get_lang(lang_code, true)
if potential_lang then
lang = potential_lang
rest = param:sub(#lang_code + 2)
else
local potential_family = Util.get_family(lang_code)
if potential_family then
lang = potential_family
rest = param:sub(#lang_code + 2)
is_family = true
else
lang = context_lang
rest = param
end
end
else
lang = context_lang
rest = param
end
M.tracking.track_term(rest)
if rest == "" or rest == "+" then
return {
lang = lang,
term = nil,
unknown_term = true,
is_family = is_family,
}
end
if rest == "-" then
return {
lang = lang,
term = nil,
suppress_term = true,
is_family = is_family,
}
end
if not rest:find("<", 1, true) then
return {
lang = lang,
term = M.string_utilities.trim(rest),
is_family = is_family,
}
end
local term_text = rest:match("^([^<]*)") or ""
local is_unknown = (term_text == "" or term_text == "+")
local is_suppress = (term_text == "-")
local function generate_obj(ignored_term)
return { term = (is_unknown or is_suppress) and nil or M.string_utilities.trim(term_text) }
end
local rest_with_defaults = Util.add_boolean_defaults(rest, EtymonParser.etymon_param_mods)
local parsed_obj = M.parse_utilities.parse_inline_modifiers(rest_with_defaults,
{ param_mods = get_clean_param_mods(EtymonParser.etymon_param_mods), generate_obj = generate_obj })
if parsed_obj.id and parsed_obj.id:match("^!") then
parsed_obj.id = parsed_obj.id:sub(2)
parsed_obj.override = true
end
parsed_obj.lang = lang
parsed_obj.is_family = is_family
if is_unknown then
parsed_obj.unknown_term = true
elseif is_suppress then
parsed_obj.suppress_term = true
end
return parsed_obj
end
function EtymonParser.validate(lang, args, id, title, pos, starts_with_lang_code)
-- id is now optional, so only validate if provided
if id then
if mw.ustring.len(id) < 2 then
error("The `id` parameter must have at least two characters.")
end
if id == title or id == Util.get_page_data().pagename then
error("The `id` parameter must not be the same as the page title.")
end
end
local valid_pos = { prefix = true, suffix = true, interfix = true, infix = true, root = true, word = true }
if pos and not valid_pos[pos] then
error("Unknown value provided for `pos`. Valid values: " .. table.concat(require("Module:table").keysToList(valid_pos), ", ") .. ".")
end
local current_keyword = "from"
local current_keyword_explicit = false
local keyword_etymons = {}
local keywords = M.data.keywords
local function checkKeyword()
local config = keywords[current_keyword]
if current_keyword == "from" and not current_keyword_explicit and #keyword_etymons == 0 then
keyword_etymons = {}
return
end
EtymonParser.check_etymon_limits(#keyword_etymons, config, "`:" .. current_keyword .. "`")
keyword_etymons = {}
end
local start_index = starts_with_lang_code and 2 or 1
for i = start_index, #args do
local param = args[i]
if type(param) ~= "string" then
elseif param:sub(1, 1) == ":" and not is_keyword(param) then
reject_removed_surf_keyword(param)
error("Invalid keyword '" .. param .. "'. Did you mean a valid keyword like ':bor', ':inh', etc.?")
elseif is_keyword(param) then
checkKeyword()
current_keyword = get_canonical_keyword(get_keyword(param))
current_keyword_explicit = true
else
local etymon_data = EtymonParser.parse_etymon(param, lang)
if etymon_data then
table.insert(keyword_etymons, param)
EtymonParser.check_keyword_term(etymon_data, lang, current_keyword)
-- Check modifier restrictions
EtymonParser.check_modifier_restrictions(etymon_data, current_keyword, EtymonParser.etymon_param_mods)
-- postype must be "root" or "word"
local VALID_POSTYPES = { root = true, word = true }
if etymon_data.postype and not VALID_POSTYPES[etymon_data.postype] then
error("Invalid <postype:" .. etymon_data.postype .. ">; must be \"root\" or \"word\".")
end
if etymon_data.ety then
local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang)
EtymonParser.validate(etymon_data.lang, inline_args, nil, nil, nil, true)
end
else
table.insert(keyword_etymons, param)
end
end
end
checkKeyword()
end
local DataRetriever = {}
local function format_etymon_id_hint(id_data, idx)
local id = type(id_data) == "table" and id_data.id or id_data
local pos = type(id_data) == "table" and id_data.pos
if id and id ~= "" and id ~= "*" then
return '"' .. id .. '"'
end
if pos and pos ~= "" then
return "unnamed (|pos=" .. pos .. "|)"
end
return "etymon #" .. idx .. " (no |id= on page)"
end
local function etymon_target_page_link(page, norm_lang)
return M.links.full_link({
term = page,
lang = norm_lang,
no_generate_alternants = true,
}, "term")
end
-- Summarize {{etymon}} id slots on a linked page for preview warnings.
local function summarize_available_etymon_ids(ids)
local id_list = {}
local all_idless = true
local target_has_idless = false
local any_pos = false
for i, id_data in ipairs(ids) do
local id = type(id_data) == "table" and id_data.id or id_data
local pos = type(id_data) == "table" and id_data.pos
if id and id ~= "" and id ~= "*" then
all_idless = false
else
target_has_idless = true
end
if pos and pos ~= "" then
any_pos = true
end
table.insert(id_list, format_etymon_id_hint(id_data, i))
end
return {
id_list = id_list,
all_idless = all_idless,
target_has_idless = target_has_idless,
any_pos = any_pos,
count = #ids,
options_text = mw.text.listToText(id_list),
}
end
local function ambiguous_etymon_suggestion(page_link, summary)
if summary.all_idless then
if summary.any_pos then
return " None set `|id=` yet; add a unique `|id=` to each on " .. page_link
.. ", then `<id:identifier>` after the term here. Section order / hints: "
.. summary.options_text .. "."
end
return " None set `|id=` yet; add a unique `|id=` to each {{etymon}} in that section from top to bottom, then `<id:identifier>` after the term here (same value as `|id=`)."
end
return " Specify which one with `<id:identifier>` after the term. Options: " .. summary.options_text .. "."
end
local function warn_ambiguous_etymon_link(page, norm_lang, ids, is_toplevel)
local page_link = etymon_target_page_link(page, norm_lang)
local summary = summarize_available_etymon_ids(ids)
if is_toplevel and summary.target_has_idless then
__state.linked_page_multiple_etymons_idless = true
end
local lang_name = norm_lang:getCanonicalName()
local lead = "Etymology link to " .. page_link .. " is ambiguous (" .. summary.count
.. " {{etymon}} templates for " .. lang_name .. ")."
Util.add_warning(lead .. ambiguous_etymon_suggestion(page_link, summary), true)
end
local function is_mismatched_explicit_id(base_key, cached_args, parent_etymon)
return cached_args == M.data.STATUS.MISSING and not parent_etymon
and #(__state.available_etymon_ids[base_key] or {}) > 0
end
local function maybe_flag_partial_etymology_reference(base_key, etymon_data, cached_args, is_toplevel)
if not is_toplevel or __state.skip_partial_etymology_category then
return
end
if not __state.partial_etymology_targets[base_key] then
return
end
if etymon_data.id and type(cached_args) == "table" then
return
end
__state.linked_page_partial_etymology_sections = true
end
local function is_nonlemma_etymon_template(template_args)
return template_args and M.yesno(template_args.nl, false)
end
local function warn_mismatched_explicit_id(page, norm_lang, base_key, etymon_id)
local page_link = etymon_target_page_link(page, norm_lang)
local summary = summarize_available_etymon_ids(__state.available_etymon_ids[base_key] or {})
local lang_name = norm_lang:getCanonicalName()
local lead = "Etymology link to " .. page_link .. " uses `<id:" .. etymon_id
.. ">`, but no {{etymon}} on that page has `|id=" .. etymon_id .. "|` for " .. lang_name .. "."
Util.add_warning(lead .. " Valid IDs: " .. summary.options_text .. ".", true)
end
-- Given an etymon data, scrape its page and cache the result in the global state object.
function DataRetriever.cache_page_etymons(etymon_page, etymon_title, key, etymon_lang, etymon_id, redirected_from, descendants_is_toplevel)
local content = etymon_title:getContent()
if not content then
__state.cached_etymon_args[key] = M.data.STATUS.REDLINK
return
end
-- Check if the linked page is a redirect. If it is, the template parsing
-- code below will be effectively skipped, and `scrape_page` will be called
-- again on the redirect target (see the bottom of this function)
local lang_section_for_descendants = nil
local redirect_target = etymon_title.redirect_target
if not redirect_target then
content = M.pages.get_section(content, etymon_lang:getFullName(), 2)
if not content then
__state.cached_etymon_args[key] = M.data.STATUS.MISSING
return
end
lang_section_for_descendants = content
end
local etymon_lang_code = etymon_lang:getFullCode()
local lang_page_key = etymon_lang_code .. ":" .. etymon_page
local found_templates_for_lang = {}
local found_ids = {}
local get_node_class = M.template_parser.class_else_type
-- Look for all {{etymon}} templates within the page content using the template parser
-- This way the same page is never parsed more than once
-- Build a map from senseids to their parent etymonids.
local active_etymon_args = nil
local etymology_section_count = 0
local etymology_sections_with_etymon = 0
local current_etymology_has_etymon = false
local current_etymology_has_nonlemma = false
local function finalize_current_etymology_section()
if etymology_section_count == 0 then
return
end
if current_etymology_has_etymon or current_etymology_has_nonlemma then
etymology_sections_with_etymon = etymology_sections_with_etymon + 1
end
current_etymology_has_etymon = false
current_etymology_has_nonlemma = false
end
for node in M.template_parser.parse(content):iterate_nodes() do
local node_class = get_node_class(node)
if node_class == "heading" then
-- A new L2 or etymology section acts as a barrier: an {{etymon}} usage
-- used previously cannot be the parent of any subsequent senseids.
-- Note that we don't have to check for L2s due to the usage of `M.pages.get_section` above.
if node:get_name():find("^Etymology") then
finalize_current_etymology_section()
etymology_section_count = etymology_section_count + 1
active_etymon_args = nil
end
elseif node_class == "template" then
local template_name = node:get_name()
if template_name == "etymon" then
local template_args = node:get_arguments()
-- Check if this etymon is for our language
if template_args[1] == etymon_lang_code then
if is_nonlemma_etymon_template(template_args) then
if etymology_section_count > 0 then
current_etymology_has_nonlemma = true
end
else
if etymology_section_count > 0 then
current_etymology_has_etymon = true
end
table.insert(found_templates_for_lang, template_args)
if template_args.id then
local etymon_key = lang_page_key .. ":" .. template_args.id
__state.cached_etymon_args[etymon_key] = template_args
__state.cached_etymon_pages[etymon_key] = tostring(etymon_page)
table.insert(found_ids, template_args.id)
active_etymon_args = template_args
else
-- Store idless etymon with default key
local etymon_key = lang_page_key .. ":*"
__state.cached_etymon_args[etymon_key] = template_args
__state.cached_etymon_pages[etymon_key] = tostring(etymon_page)
table.insert(found_ids, "*")
active_etymon_args = template_args
end
end
end
elseif active_etymon_args and template_name == "senseid" then
local template_args = node:get_arguments()
-- This should always be true for proper usages of {{senseid}}.
if template_args[1] == etymon_lang_code and template_args[2] then
local sense_id_key = lang_page_key .. ":" .. template_args[2]
__state.senseid_parent_etymon[sense_id_key] = active_etymon_args
__state.cached_etymon_pages[sense_id_key] = tostring(etymon_page)
end
end
end
end
finalize_current_etymology_section()
if lang_section_for_descendants
and etymology_section_count > 1
and etymology_sections_with_etymon > 0
and etymology_sections_with_etymon < etymology_section_count
then
__state.partial_etymology_targets[lang_page_key] = true
end
if descendants_is_toplevel and lang_section_for_descendants and #found_templates_for_lang > 0 then
M.descendants.cache_page_checks({
lang_section = lang_section_for_descendants,
etymon_lang_code = etymon_lang_code,
found_templates_for_lang = found_templates_for_lang,
entry_title = __state.entry_title,
entry_lang_code = __state.entry_lang_code,
entry_lang = __state.entry_lang_code and Util.get_lang(__state.entry_lang_code, true) or nil,
cached_descendants_checks = __state.cached_descendants_checks,
lang_page_key = lang_page_key,
redirected_from = redirected_from,
})
end
local id_data_list = {}
for _, args in ipairs(found_templates_for_lang) do
local id = args.id or "*"
table.insert(id_data_list, { id = id, pos = args.pos })
end
__state.available_etymon_ids[lang_page_key] = id_data_list
if #found_templates_for_lang == 1 then
__state.single_etymons[lang_page_key] = found_templates_for_lang[1]
end
if redirected_from and __state.available_etymon_ids[lang_page_key] then
__state.available_etymon_ids[redirected_from] = __state.available_etymon_ids[redirected_from] or {}
for _, id_data in ipairs(__state.available_etymon_ids[lang_page_key]) do
table.insert(__state.available_etymon_ids[redirected_from], id_data)
end
end
if __state.cached_etymon_args[key] ~= nil or __state.senseid_parent_etymon[key] ~= nil then
-- All done!
return
elseif redirect_target and not redirected_from then
-- Try scraping the redirect.
etymon_page = redirect_target.prefixedText
DataRetriever.cache_page_etymons(etymon_page, redirect_target, lang_page_key .. ":" .. etymon_id, etymon_lang, etymon_id, lang_page_key, descendants_is_toplevel)
__state.cached_etymon_args[key] = __state.cached_etymon_args[etymon_lang_code .. ":" .. etymon_page .. ":" .. etymon_id]
else
__state.cached_etymon_args[key] = M.data.STATUS.MISSING
end
end
local function has_linkable_term(etymon_data)
if etymon_data.is_family or etymon_data.suppress_term or etymon_data.unknown_term then
return false
end
local term = etymon_data.term
if term == nil or term == "" then
return false
end
return M.string_utilities.trim(term) ~= ""
end
local function record_term_id_tracking(etymon_data)
if not has_linkable_term(etymon_data) then
return
end
local term_page = M.links.get_link_page(etymon_data.term, etymon_data.lang)
M.tracking.record_term_id_usage(__state.id_stats, etymon_data, term_page)
end
-- Given an etymon object, scrape its page (if necessary) and return its own etymon arguments as well as the page name.
function DataRetriever.get_etymon_args(etymon_data, is_toplevel)
if not has_linkable_term(etymon_data) then
return M.data.STATUS.MISSING, nil, nil, nil
end
local page = M.links.get_link_page(etymon_data.term, etymon_data.lang)
local norm_lang = Util.get_norm_lang(etymon_data.lang)
local base_key = norm_lang:getFullCode() .. ":" .. page
if etymon_data.id then
local key = base_key .. ":" .. etymon_data.id
local cached_args = __state.cached_etymon_args[key] or __state.senseid_parent_etymon[key]
if cached_args == nil then
local title = mw.title.new(page)
if not title then error('Invalid page title "' .. page .. '" encountered.') end
DataRetriever.cache_page_etymons(page, title, key, norm_lang, etymon_data.id, nil, is_toplevel)
end
cached_args = __state.cached_etymon_args[key] or __state.senseid_parent_etymon[key] -- refresh
-- Get etymon_id from parent if this was resolved via senseid
local parent_etymon = __state.senseid_parent_etymon[key]
local resolved_etymon_id = parent_etymon and parent_etymon.id
local descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = is_toplevel,
base_key = base_key,
lookup = {
explicit_id = etymon_data.id,
parent_etymon = parent_etymon,
},
})
if is_toplevel and descendants_check == nil then
local title = mw.title.new(page)
if title then
DataRetriever.cache_page_etymons(page, title, key, norm_lang, etymon_data.id, nil, true)
descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = true,
base_key = base_key,
lookup = {
explicit_id = etymon_data.id,
parent_etymon = parent_etymon,
},
})
end
end
local mismatched_id = is_mismatched_explicit_id(base_key, cached_args, parent_etymon)
if mismatched_id and is_toplevel then
__state.has_mismatched_id = true
M.tracking.record_mismatched_id_usage(__state.id_stats, norm_lang, page, etymon_data.id)
warn_mismatched_explicit_id(page, norm_lang, base_key, etymon_data.id)
end
maybe_flag_partial_etymology_reference(base_key, etymon_data, cached_args, is_toplevel)
return cached_args, __state.cached_etymon_pages[key], resolved_etymon_id, descendants_check
else
__state.used_idless_etymon = true
if is_toplevel then
__state.toplevel_idless_etymon = true
end
if __state.available_etymon_ids[base_key] == nil then
local title = mw.title.new(page)
if not title then error('Invalid page title "' .. page .. '" encountered.') end
DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, is_toplevel)
end
local ids = __state.available_etymon_ids[base_key] or {}
local count = #ids
-- Try to filter by postype if available and we have multiple candidates
if count > 1 and etymon_data.postype then
local matching_ids = {}
for _, id_data in ipairs(ids) do
if id_data.pos == etymon_data.postype then
table.insert(matching_ids, id_data)
end
end
if #matching_ids == 1 then
local matched_id = matching_ids[1].id
local matched_key = base_key .. ":" .. matched_id
M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "postype")
local descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = is_toplevel,
base_key = base_key,
lookup = { id = matched_id },
})
if is_toplevel and descendants_check == nil then
local title = mw.title.new(page)
if title then
DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, true)
descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = true,
base_key = base_key,
lookup = { id = matched_id },
})
end
end
local matched_args = __state.cached_etymon_args[matched_key]
maybe_flag_partial_etymology_reference(base_key, etymon_data, matched_args, is_toplevel)
return matched_args, __state.cached_etymon_pages[matched_key], nil, descendants_check
end
end
if count == 1 then
local only_id_data = ids[1]
local only_id = (type(only_id_data) == "table" and only_id_data.id) or only_id_data or "*"
M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "single")
local descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = is_toplevel,
base_key = base_key,
lookup = { id_data = only_id_data },
})
if is_toplevel and descendants_check == nil then
local title = mw.title.new(page)
if title then
DataRetriever.cache_page_etymons(page, title, base_key .. ":*", norm_lang, "*", nil, true)
descendants_check = M.descendants.get_lookup_check({
cached_descendants_checks = __state.cached_descendants_checks,
is_toplevel = true,
base_key = base_key,
lookup = { id_data = only_id_data },
})
end
end
local single_args = __state.single_etymons[base_key]
maybe_flag_partial_etymology_reference(base_key, etymon_data, single_args, is_toplevel)
return single_args, __state.cached_etymon_pages[base_key .. ":" .. only_id], nil, descendants_check
elseif count > 1 then
M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "ambiguous")
warn_ambiguous_etymon_link(page, norm_lang, ids, is_toplevel)
maybe_flag_partial_etymology_reference(base_key, etymon_data, M.data.STATUS.AMBIGUOUS, is_toplevel)
return M.data.STATUS.AMBIGUOUS, nil, nil, nil
else
M.tracking.record_idless_resolution(__state.id_stats, norm_lang, page, "missing")
maybe_flag_partial_etymology_reference(base_key, etymon_data, M.data.STATUS.MISSING, is_toplevel)
return M.data.STATUS.MISSING, nil, nil, nil
end
end
end
local function keyword_invisible_in_tree(keyword_info)
if not keyword_info then
return false
end
local inv = keyword_info.invisible
return inv == "all" or inv == true or inv == "tree"
end
-- True when the node has at least one top-level child container visible in the tree.
local function node_has_visible_tree_children(node)
for _, container in ipairs(node.children or {}) do
if not keyword_invisible_in_tree(container.keyword_info) then
return true
end
end
return false
end
-- Count visible term nodes in the tree.
local function get_visible_tree_depth(node, skip_child_rendering)
local max_depth = 1
if skip_child_rendering or not node then
return max_depth
end
for _, container in ipairs(node.children or {}) do
local keyword_info = container.keyword_info
if not keyword_invisible_in_tree(keyword_info) then
local skip_grandchildren = keyword_info and keyword_info.no_child_categories
for _, term in ipairs(container.terms or {}) do
if term.is_duplicate then
if term.original_has_children then
max_depth = math.max(max_depth, 2)
end
else
max_depth = math.max(max_depth, 1 + get_visible_tree_depth(term, skip_grandchildren))
end
end
end
end
return max_depth
end
local function as_param_list(val)
if val == nil then
return {}
end
if type(val) == "table" then
return val
end
if type(val) == "string" and val ~= "" then
return { val }
end
return {}
end
local TreeBuilder = {}
-- Build a unique key for deduplication in the seen table
function TreeBuilder.build_key(lang, title, args)
local norm_lang_code = Util.get_norm_lang(lang):getFullCode()
local is_table = type(args) == "table"
local id = (is_table and args.id) or ""
if title then
return norm_lang_code .. ":" .. M.links.get_link_page(title, lang) .. ":" .. id
end
if is_table and args.status == M.data.STATUS.INLINE then
local content_parts = {}
for i = 1, #args do
content_parts[i] = tostring(args[i])
end
return norm_lang_code .. ":*:" .. id .. "\0" .. table.concat(content_parts, "\0")
end
return norm_lang_code .. ":*:" .. id
end
-- Copy parsed etymon modifiers onto a tree/supplement term node.
function TreeBuilder.apply_etymon_fields(term, etymon_data)
term.id = etymon_data.id
term.gloss = etymon_data.gloss
term.tr = etymon_data.tr
term.ts = etymon_data.ts
term.alt = etymon_data.alt
term.genders = etymon_data.genders
term.pos = etymon_data.pos
term.ng = etymon_data.ng
term.infl = etymon_data.infl
term.refs = etymon_data.refs
term.is_uncertain = etymon_data.unc
term.lit = etymon_data.lit
term.q = etymon_data.q
term.qq = etymon_data.qq
term.l = etymon_data.l
term.ll = etymon_data.ll
term.suppress_term = etymon_data.suppress_term
term.unknown_term = etymon_data.unknown_term
term.is_family = etymon_data.is_family
term.override = etymon_data.override
term.aftype = etymon_data.aftype
term.postype = etymon_data.postype
term.bor = etymon_data.bor
term.lbor = etymon_data.lbor
term.slbor = etymon_data.slbor
end
function TreeBuilder.build_supplement_term(etymon_data, entry_lang, supplement_type)
EtymonParser.check_supplement_term(etymon_data, entry_lang, supplement_type)
local term = {
lang = etymon_data.lang,
title = etymon_data.term,
children = {},
status = M.data.STATUS.OK,
}
TreeBuilder.apply_etymon_fields(term, etymon_data)
return term
end
function TreeBuilder.build_supplement_terms(entry_lang, supplement_type, param_value)
local terms = {}
for _, term_param in ipairs(as_param_list(param_value)) do
if type(term_param) == "string" and term_param ~= "" then
local etymon_data = EtymonParser.parse_etymon(term_param, entry_lang)
if etymon_data then
table.insert(terms, TreeBuilder.build_supplement_term(etymon_data, entry_lang, supplement_type))
end
end
end
return terms
end
-- Attach a |param= supplement defined in etymon_data.supplements (e.g. doublet=).
function TreeBuilder.append_term_supplement(data_tree, entry_lang, supplement_type, param_value)
local config = M.data.supplements[supplement_type]
if not config then
error("Unknown supplement '" .. tostring(supplement_type) .. "'.")
end
local terms = TreeBuilder.build_supplement_terms(entry_lang, supplement_type, param_value)
if #terms == 0 then
return
end
data_tree.supplements = data_tree.supplements or {}
table.insert(data_tree.supplements, {
type = supplement_type,
config = config,
terms = terms,
})
M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, supplement_type, entry_lang, entry_lang, true)
end
function TreeBuilder.build(lang, title, args, seen, depth, stop_recursion)
seen = seen or {}
depth = depth or 0
local is_toplevel = (depth == 0)
if depth > __state.max_depth_reached then
__state.max_depth_reached = depth
end
__state.total_nodes = __state.total_nodes + 1
local lang_code = lang:getCode()
__state.language_count[lang_code] = (__state.language_count[lang_code] or 0) + 1
local current_id = (type(args) == "table" and args.id) or ""
local key = TreeBuilder.build_key(lang, title, args)
local node = { lang = lang, title = title, id = current_id, args = args, children = {}, status = M.data.STATUS.OK }
if type(args) ~= "table" or seen[key] then
node.status = args or M.data.STATUS.MISSING
-- Mark as duplicate if we've seen this node before
if seen[key] then
node.is_duplicate = true
node.duplicate_key = key
local original_node = seen[key]
if type(original_node) == "table" and original_node.children and #original_node.children > 0 then
node.original_has_children = true
end
end
return node
end
node.status = args.status or M.data.STATUS.OK
seen[key] = node
-- If stop_recursion is set, skip parsing children but check for visible children
if stop_recursion then
local keywords = M.data.keywords
local has_visible_children = false
for i = 2, #args do
local param = args[i]
if type(param) == "string" then
local keyword_base = get_keyword_base(param)
if keyword_base and keywords[keyword_base] then
local _, kw_modifiers = EtymonParser.parse_keyword_modifiers(param:sub(1, 1) == ":" and param or (":" .. param))
if not keyword_invisible_in_tree(get_effective_keyword_info(keyword_base, kw_modifiers)) then
has_visible_children = true
break
end
elseif param:sub(1, 1) ~= ":" then
-- It's a term (not a keyword), so there are visible children
has_visible_children = true
break
end
end
end
node.has_visible_children = has_visible_children
return node
end
-- Parse args into keyword containers
local current_keyword = "from"
local current_keyword_modifiers = {}
local current_container = nil
local function ensure_container()
if not current_container or current_container.keyword ~= current_keyword then
local keyword_info = get_effective_keyword_info(current_keyword, current_keyword_modifiers)
current_container = {
keyword = current_keyword,
keyword_info = keyword_info,
keyword_modifiers = current_keyword_modifiers,
terms = {},
}
table.insert(node.children, current_container)
-- Override keyword text/phrase for nominalization with <g:code>
if current_keyword_modifiers.g and current_keyword == "nominalization" then
local labels = get_nominalization_label_for_g(current_keyword_modifiers.g)
if not labels then
local codes = {}
for c in pairs(M.data.nominalization_g_codes) do table.insert(codes, c) end
table.sort(codes)
error("Invalid <g:" .. tostring(current_keyword_modifiers.g) .. ">. Supported codes for nominalization: " .. table.concat(codes, ", "))
end
current_container.keyword_info = copy_keyword_info(keyword_info)
current_container.keyword_info.text = labels.text
current_container.keyword_info.phrase = labels.phrase
end
end
return current_container
end
local parse_context_lang = Util.resolve_context_lang(lang, args)
for i = 2, #args do
local param = args[i]
if is_keyword(param) then
local keyword, modifiers = EtymonParser.parse_keyword_modifiers(param)
if not keyword then
error("Invalid keyword '" .. param .. "'.")
end
current_keyword = keyword
current_keyword_modifiers = modifiers
current_container = nil -- Force new container for new keyword
elseif type(param) == "string" and param:sub(1, 1) == ":" then
reject_removed_surf_keyword(param)
error("Invalid keyword '" .. param .. "'. Did you mean a valid keyword like ':bor', ':inh', etc.?")
elseif type(param) == "string" then
local etymon_data = EtymonParser.parse_etymon(param, parse_context_lang)
if etymon_data then
-- Track keyword usage at top level
M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, current_keyword, lang, etymon_data.lang, is_toplevel)
local term_node = {}
local container
-- Handle suppress_term (-) and unknown_term (empty or +) directly
if etymon_data.suppress_term or etymon_data.unknown_term then
container = ensure_container()
if etymon_data.ety then
local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang)
inline_args.id = etymon_data.id
inline_args.status = M.data.STATUS.INLINE
term_node = TreeBuilder.build(etymon_data.lang, nil, inline_args, seen, depth + 1)
else
term_node = {
lang = etymon_data.lang,
children = {},
status = M.data.STATUS.OK,
}
end
TreeBuilder.apply_etymon_fields(term_node, etymon_data)
else
-- Regular term: fetch arguments from page
record_term_id_tracking(etymon_data)
local etymon_args, page_of, resolved_etymon_id, descendants_check =
DataRetriever.get_etymon_args(etymon_data, is_toplevel)
-- Check for <ety> inline parameter doesn't override the scraped arguments, unless the latter are missing
if etymon_data.ety then
if etymon_args == M.data.STATUS.REDLINK or etymon_args == M.data.STATUS.MISSING then
__state.current_page_has_inline_etymology = true
if is_toplevel then
__state.toplevel_has_inline_etymology = true
end
local inline_args = EtymonParser.parse_inline_ety(etymon_data.ety, etymon_data.lang)
-- Track inline ety keywords too
local inline_keyword = get_keyword(inline_args[2], true)
if inline_keyword and #inline_args >= 3 then
local inline_etymon = EtymonParser.parse_etymon(inline_args[3], etymon_data.lang)
if inline_etymon then
M.tracking.record_keyword_usage(__state.toplevel_keyword_stats, inline_keyword, etymon_data.lang, inline_etymon.lang, is_toplevel)
end
end
inline_args.id = etymon_data.id
inline_args.status = M.data.STATUS.INLINE
etymon_args = inline_args
term_node.page_of = __state.cached_etymon_pages[key] -- term node is on the same page as the parent
else
-- Scraped arguments exist, <ety> is redundant and ignored
__state.current_page_has_redundant_etymology = true
if is_toplevel then
__state.toplevel_redundant_etymology = true
end
end
end
-- Ensure container exists before checking keyword info
container = ensure_container()
-- Check if current keyword has no_child_categories - if so, stop recursion
local keyword_info = container.keyword_info
local should_stop_recursion = (stop_recursion or (keyword_info and keyword_info.no_child_categories))
term_node = TreeBuilder.build(etymon_data.lang, etymon_data.term, etymon_args, seen, depth + 1, should_stop_recursion)
term_node.target_key = Util.get_norm_lang(etymon_data.lang):getFullCode() ..
":" .. M.links.get_link_page(etymon_data.term, etymon_data.lang)
term_node.etymon_id = resolved_etymon_id -- The actual etymon id when resolved via senseid
term_node.page_of = page_of
TreeBuilder.apply_etymon_fields(term_node, etymon_data)
term_node.missing_descendants_header, term_node.missing_descendants_entry =
M.descendants.get_term_sync_flags(current_keyword, term_node.status, descendants_check)
end
table.insert(container.terms, term_node)
end
end
end
return node
end
-- Convert etymology tree to JSON-serializable table
local function tree_to_json(node)
local obj = {
term = node.title,
lang = node.lang:getCode(),
lang_name = node.lang:getCanonicalName(),
id = (node.id and node.id ~= "") and node.id or nil,
status = node.status,
is_uncertain = node.is_uncertain or nil,
is_duplicate = node.is_duplicate or nil,
gloss = node.gloss,
transliteration = node.tr,
transcription = node.ts,
alt = node.alt,
g = node.genders,
pos = node.pos,
ng = node.ng,
infl = node.infl,
children = {},
}
for _, container in ipairs(node.children or {}) do
local keyword_info = container.keyword_info
if keyword_info then
local container_obj = {
keyword = container.keyword,
keyword_label = keyword_info.text,
keyword_abbrev = keyword_info.abbrev,
is_group = keyword_info.is_group or nil,
is_invisible = keyword_info.invisible or nil,
is_uncertain = (container.keyword_modifiers and container.keyword_modifiers.unc) or nil,
terms = {},
}
for _, term in ipairs(container.terms or {}) do
table.insert(container_obj.terms, tree_to_json(term))
end
table.insert(obj.children, container_obj)
end
end
return obj
end
-- Build and return the etymology data tree for a given term.
function export.get_tree(lang, title, args, options)
options = options or {}
__state.entry_title = title
__state.entry_lang_code = lang:getCode()
__state.id_stats = M.tracking.new_id_stats()
__state.skip_partial_etymology_category = options.skip_partial_etymology_category == true
if options.validate then
EtymonParser.validate(lang, args, options.id, title, options.pos, false)
end
local lang_code = lang:getCode()
local start_index = (args[1] == lang_code) and 2 or 1
local tree_args = { [1] = lang_code, id = options.id or args.id }
for i = start_index, #args do
table.insert(tree_args, args[i])
end
__state.cached_etymon_args[lang_code .. ":" .. title .. ":" .. (tree_args.id or "")] = tree_args
local ety_data_tree = TreeBuilder.build(lang, title, tree_args)
if options.json then
return M.JSON.toJSON(tree_to_json(ety_data_tree))
end
return ety_data_tree
end
-- Given a language code, page name and optionally the id= parameter,
-- render the tree and only the etymology tree for the relevant page.
-- Fetches and parses the corresponding {{etymon}} from the requested page,
-- and any further pages needed to render the tree.
-- Parameters can be passed either through the #invoke or as
-- template parameters *through* an #invoke.
function export.render_tree_for_etymon_on_page(frame)
local frame_args = frame.args
local parent_args = frame:getParent().args
local langcode = frame_args[1] or parent_args[1]
local pagename = frame_args[2] or parent_args[2]
local id = frame_args["id"] or parent_args["id"]
local display_title = frame_args["title"] or parent_args["title"]
local parsed_title = mw.title.new(pagename, 0)
local title
if parsed_title.namespace == 0 then
title = M.pages.safe_page_name(parsed_title)
elseif parsed_title.namespace == 118 then
title = "*" .. M.pages.safe_page_name(parsed_title)
else
error("Unsupported namespace for render_tree_for_etymon_on_page: " .. parsed_title.namespace)
end
local lang = Util.get_lang(langcode)
__state.entry_title = title
__state.entry_lang_code = lang:getCode()
__state.id_stats = M.tracking.new_id_stats()
-- Construct etymon_data for DataRetriever.get_args.
local etymon_data = {
lang = lang,
term = title,
id = id
}
local args, pagename = DataRetriever.get_etymon_args(etymon_data, true)
if args == M.data.STATUS.MISSING then
error("The etymon template was not found (language " ..
langcode ..
", title '" ..
title ..
"'" ..
(id and ", ID '" .. id .. "'" or ", no ID given") .. "). Page contents may have changed in the interim.")
end
local tree_title = display_title or title
if lang:stripDiacritics(M.links.remove_links(tree_title)) ~= lang:stripDiacritics(M.links.remove_links(title)) then
M.tracking.track_title_pagename_mismatch(lang)
end
reset_invocation_state()
local ety_data_tree = export.get_tree(lang, tree_title, args, {
validate = true,
id = id,
})
local output = {}
table.insert(output, M.template_styles("Module:etymon/styles.css"))
table.insert(output, M.tree.render({
data_tree = ety_data_tree,
format_term_func = function(term, is_toplevel)
return Util.format_term(term, is_toplevel, {
gloss = "suppress",
pos = "suppress",
lit = "suppress",
tree_ql = "suppress",
})
end,
}))
return table.concat(output)
end
function export.main(frame)
local parent_args = frame:getParent().args
local args = M.parameters.process(parent_args, M.parameters_data.etymon)
local lang = args[1]
local etymon_args = args[2]
local id = args.id
local title = args.title
local text = args.text
local tree = args.tree
local etydate = args.etydate
local doublet = args.doublet
local rfe = args.rfe
local etystub = args.etystub
local is_nonlemma = M.yesno(args.nl, false)
local page_data = Util.get_page_data()
if not title then
title = page_data.pagename
if page_data.namespace == "Reconstruction" then title = "*" .. title end
end
local entry_pagename = page_data.pagename
if page_data.namespace == "Reconstruction" then
entry_pagename = "*" .. entry_pagename
end
if lang:stripDiacritics(M.links.remove_links(title)) ~= lang:stripDiacritics(M.links.remove_links(entry_pagename)) then
M.tracking.track_title_pagename_mismatch(lang)
end
local current_L2 = M.pages.get_current_L2()
if current_L2 then
local norm_lang = Util.get_norm_lang(lang)
local norm_name = norm_lang:getCanonicalName()
if current_L2 ~= norm_name then
local lang_desc = lang:getCode() .. " (" .. lang:getCanonicalName() .. ")"
if norm_lang:getCode() ~= lang:getCode() then
lang_desc = lang_desc .. ", normalized to " .. norm_lang:getCode() .. " (" .. norm_name .. ")"
end
error("Language '" .. lang_desc .. "' does not match the L2 header (" .. current_L2 .. ").")
end
end
reset_invocation_state()
local ety_data_tree = export.get_tree(lang, title, etymon_args, {
validate = true,
pos = args.pos,
id = id,
json = args.json,
skip_partial_etymology_category = is_nonlemma,
})
if args.json then
return ety_data_tree
end
local output = {}
local text_allowlist_mode = M.text_allowed.default_mode or "off"
if text and text_allowlist_mode ~= "off" and not Util.is_text_param_allowed_for_lang(lang) then
local msg = "Etymology texts (parameter <code>text=</code>) are not allowed for " .. lang:getFullName() ..
"; see [[Template:etymon#Text allowlist|Template:etymon § Text allowlist]] for the list of languages that may use the <code>text=</code> parameter."
if text_allowlist_mode == "error" then
error(msg)
else
Util.add_warning(msg, true)
end
end
local lang_exc = Util.get_lang_exception(lang)
if lang_exc and lang_exc.disallow then
local disallow = lang_exc.disallow
local error_text = " for " .. lang:getFullName()
if disallow.ref then
error_text = error_text .. "; see " .. disallow.ref
else
error_text = error_text .. "."
end
if tree and disallow.tree then
error("Etymology trees are not allowed" .. error_text)
end
if text and disallow.text then
error("Etymology texts are not allowed" .. error_text)
end
end
if etydate then
local etydate_param_mods = M.parameter_utilities.construct_param_mods {
{ group = "ref" },
{ param = "refn" },
{ param = "nocap", type = "boolean" },
}
local function generate_etydate_obj(etydate_text)
local etydate_specs = {}
for spec in etydate_text:gmatch("[^,]+") do
table.insert(etydate_specs, mw.text.trim(spec))
end
return { [1] = etydate_specs }
end
local parsed_etydate = M.parse_utilities.parse_inline_modifiers(etydate, { param_mods = etydate_param_mods, generate_obj = generate_etydate_obj })
local etydate_args = {
[1] = parsed_etydate[1],
nocap = parsed_etydate.nocap or false,
}
ety_data_tree.supplements = ety_data_tree.supplements or {}
table.insert(ety_data_tree.supplements, {
type = "etydate",
etydate_text = M.etydate.format_etydate(etydate_args, { omit_refs = true }),
etydate_refs = (parsed_etydate.refs and #parsed_etydate.refs > 0) and parsed_etydate.refs or nil,
})
end
TreeBuilder.append_term_supplement(ety_data_tree, lang, "doublet", doublet)
local has_visible_children = node_has_visible_tree_children(ety_data_tree)
-- Suppress trees for multiword entries and one-step chains
local visible_tree_depth = get_visible_tree_depth(ety_data_tree)
local is_trivial_tree = visible_tree_depth <= 2
local is_multiword = title:find("%s") ~= nil or title:find("_") ~= nil
if tree and (is_multiword or is_trivial_tree) then
tree = false
end
if tree then
table.insert(output, M.template_styles("Module:etymon/styles.css"))
table.insert(output, M.tree.render({
data_tree = ety_data_tree,
format_term_func = function(term, is_toplevel)
return Util.format_term(term, is_toplevel, {
gloss = "suppress",
pos = "suppress",
lit = "suppress",
tree_ql = "suppress",
})
end,
}))
end
local tree_disallowed = lang_exc and lang_exc.disallow and lang_exc.disallow.tree
local ety_tree_json = M.JSON.toJSON(tree_to_json(ety_data_tree))
local anchor = M.anchors.etymonid(lang, id, {
no_tree = args.notree,
title = title,
empty_tree = (not has_visible_children) or tree_disallowed,
ety_tree_json = ety_tree_json,
})
table.insert(output, anchor)
local text_stop_lang_missing = nil
if text then
local max_depth, stop_at_blue_link, stop_at_lang, stop_at_lang_or_bluelink
if text == "++" then
max_depth, stop_at_blue_link = false, false
elseif text == "+" then
max_depth, stop_at_blue_link = 1, false
elseif text == "*" then
max_depth, stop_at_blue_link = false, true
elseif text:match("^:[^*]+%*$") then
-- Stop at a specific language OR first bluelink after it, e.g., ":ota*"
-- If the target language is a redlink, continue to the first bluelink
local lang_code = text:match("^:([^*]+)%*$")
if lang_code and lang_code ~= "" then
local lang_obj = Util.get_lang(lang_code, true)
if lang_obj then
stop_at_lang_or_bluelink = lang_code
else
Util.add_warning('Invalid language code "' .. lang_code .. '" in text parameter. Showing full chain instead.')
max_depth, stop_at_blue_link = false, false
end
else
Util.add_warning('Empty language code in text parameter. Showing full chain instead.')
max_depth, stop_at_blue_link = false, false
end
elseif text:sub(1, 1) == ":" then
-- Stop at a specific language, e.g., ":ar" stops at first Arabic term
local lang_code = text:sub(2)
if lang_code ~= "" then
-- Validate the language code
local lang_obj = Util.get_lang(lang_code, true)
if lang_obj then
stop_at_lang = lang_code
else
Util.add_warning('Invalid language code "' .. lang_code .. '" in text parameter. Showing full chain instead.')
max_depth, stop_at_blue_link = false, false -- default to ++
end
else
Util.add_warning('Empty language code in text parameter. Showing full chain instead.')
max_depth, stop_at_blue_link = false, false -- default to ++
end
else
local num = tonumber(text)
if num and num >= 1 then
max_depth, stop_at_blue_link = num, false
else
error('Invalid text value "' ..
text .. '". Valid values are: "++" (full chain), "+" (first step only), "*" (until first blue link), a number (max steps), ":lang" (stop at language), or ":lang*" (stop at language or first bluelink if redlink)')
end
end
local text_output, text_render_meta = M.text.render({
data_tree = ety_data_tree,
format_term_func = Util.format_term,
lang_matches_stop_code = Util.lang_matches_stop_code,
max_depth = max_depth,
stop_at_blue_link = stop_at_blue_link,
curr_page = page_data.pagename,
nodot = args.nodot,
dot = args.dot,
stop_at_lang = stop_at_lang,
stop_at_lang_or_bluelink = stop_at_lang_or_bluelink,
})
table.insert(output, text_output)
if stop_at_lang and text_render_meta and not text_render_meta.stop_lang_reached then
M.tracking.track_text_stop_lang_missing(lang, stop_at_lang)
text_stop_lang_missing = stop_at_lang
end
end
if rfe then
table.insert(output, Util.expand_request_template(frame, "rfe", rfe, lang:getCode()))
end
if etystub then
table.insert(output, Util.expand_request_template(frame, "etystub", etystub, lang:getCode()))
end
if is_nonlemma then
table.insert(output, " " .. frame:expandTemplate({
title = "nonlemma",
args = {},
}))
end
local categories = {}
if Util.is_content_page() then
M.tracking.track_tree_metrics({
max_depth_reached = __state.max_depth_reached,
total_nodes = __state.total_nodes,
language_count = __state.language_count,
lang = lang,
})
categories = M.categories.build({
data_tree = ety_data_tree,
page_lang = lang,
available_etymon_ids = __state.available_etymon_ids,
senseid_parent_etymon = __state.senseid_parent_etymon,
get_norm_lang_func = Util.get_norm_lang,
lang_exc = lang_exc,
suppress_categories = lang_exc and lang_exc.suppress_categories,
nocat = args.nocat,
tree = tree,
text = text,
exnihilo = args.exnihilo,
toplevel_has_inline_etymology = __state.toplevel_has_inline_etymology,
toplevel_redundant_etymology = __state.toplevel_redundant_etymology,
toplevel_idless_etymon = __state.toplevel_idless_etymon,
has_mismatched_id = __state.has_mismatched_id,
linked_page_multiple_etymons_idless = __state.linked_page_multiple_etymons_idless,
linked_page_partial_etymology_sections = __state.linked_page_partial_etymology_sections,
text_stop_lang_missing = text_stop_lang_missing,
})
M.tracking.track_keywords(__state.toplevel_keyword_stats, lang)
M.tracking.track_page_id(lang, id)
M.tracking.track_ids(__state.id_stats, lang)
end
if #categories > 0 then
table.insert(output, M.categories.format(categories, lang))
end
if __state.warnings then
for i, warning in ipairs(__state.warnings) do
table.insert(output, (i == 1 and "\n" or "") .. warning .. "\n")
end
end
return table.concat(output)
end
return export
83f0fr5rgmxbm610bma1l6w6fdin4eo
Wikikamus:Penyelia/Pengundian/Lynumiss untuk penyelia 4 September 2026
4
141417
375360
375335
2026-09-22T03:18:06Z
PeaceSeekers
3334
/* Keputusan */ Balas
375360
wikitext
text/x-wiki
<big>'''[[Pengguna:Lynumiss|Lynumiss]]'''</big> ([[Perbincangan Pengguna:Lynumiss |Perbualan]] • [[Khas:Sumbangan/Lynumiss|Sumbangan]] • [[Khas:Pusat_pengesahan/Lynumiss|Sejagat]] • <span class="plainlinks">[https://xtools.wmcloud.org/ec/ms.wiktionary.org/Lynumiss Statistik]</span>)
:{| class=wikitable
|-
| '''Diusulkan oleh'''
| [https://ms.wiktionary.org/wiki/Wikikamus:Penyelia/Permohonan#Lynumiss_(Perbincangan_-_Sumbangan) Rombituon]
|-
| '''Jumlah suntingan'''
| 1,297 setakat 03:33, 4 September 2026 (UTC)
|-
| '''Mekanisme'''
|
Waktu pengundian:
* Bermula pada 4 September 2026.
* Berakhir pada <strike>18 September 2026 (2 minggu).</strike> 20 September 2026.
Syarat pengundian:
* Pengguna yang menyokong minimum sebanyak tiga (3) orang.
* Dipersetujui minimum 70% pengundi. Undian 'berkecuali' tidak diambil kira.
* Semua [[Wikikamus:Penyelia/Polisi#Takrifan dan istilah umum|pengguna berdaftar sah]] boleh mengundi.
|}
== Undian ==
<!-- Guna {{Undi|Y}} atau {{Undi|T}} diikuti dengan ~~~~ untuk menandatangan undian-->
# {{Undi|Y}} [[Pengguna:Ultron90|Ultron90]] ([[Perbincangan pengguna:Ultron90|bincang]]) 13:49, 4 September 2026 (UTC)
# {{Undi|Y}} - [[Pengguna:Rulwarih|Rulwarih]] ([[Perbincangan pengguna:Rulwarih|bincang]]) 13:26, 15 September 2026 (UTC)
# {{Undi|Y}} - [[Pengguna:SNN95|SNN95]] ([[Perbincangan pengguna:SNN95|bincang]]) 12:25, 19 September 2026 (UTC)
# {{Undi|Y}} - '''[[Pengguna:Hakimi97|محمد حكيمي]]''' ([[Perbincangan pengguna:Hakimi97|bincang]]) 12:28, 19 September 2026 (UTC)
== Komen ==
== Keputusan ==
Memandangkan undi belum cukup setakat 14 hari, saya panjangkan sedikit tempoh pengundian hingga 20 haribulan. [[Pengguna:PeaceSeekers|PeaceSeekers]] ([[Perbincangan pengguna:PeaceSeekers|bincang]]) 12:26, 19 September 2026 (UTC)
:Undian tamat dengan jumlah persetujuan mencukupi. Pencalonan diterima dan akan dibawa ke Meta bagi urusan lanjut. [[Pengguna:PeaceSeekers|PeaceSeekers]] ([[Perbincangan pengguna:PeaceSeekers|bincang]]) 03:18, 22 September 2026 (UTC)
ek78zcm0kkxnl1w9wh63xphveggaml7
Wikikamus:bdr/lawa
4
145784
375343
2026-09-21T14:06:26Z
Sipatung
10839
Mencipta laman
375343
wikitext
text/x-wiki
==Bahasa {{bahasa|bdr}}==
===Kata sifat===
{{inti|bdr|kata sifat}}
# {{label|1=bdr|2=dialek|3=Sabah}} lawa {{cp|bdr|Lukisan kekanak a '''lawa''' bana.|Lukisan budak itu sangat '''[[cantik]]'''.}}
mmsu7s3eipb4pyfjer1r5e6sjpceffg